We publish working frameworks for experiments in local AI, agent orchestration, and playable evaluation. Each note distinguishes what we know, what we propose, and what still needs to be measured.
Local inference is a system property, not a checkbox. This framework covers the privacy boundary, model and device fit, latency, memory, quality, energy, and graceful fallback.
A model score cannot tell you whether an agent system is reliable. Evaluate the surrounding prompts, tools, state, permissions, control flow, traces, and recovery behavior.
Games make state, actions, objectives, and failure visible. A well-designed playable testbed can expose planning and memory problems that disappear inside a single aggregate score.