Active programs
Questions worth turning into systems.
PumaAI organizes its work into three connected tracks. These are active research directions, not claims of completed products; published methods and results will remain linked here.
Active research track · Private compute
Build a repeatable way to match models and runtimes to personal hardware. The program treats privacy boundaries, usable latency, memory pressure, quality, energy, and fallback behavior as a single deployment problem.
Evaluation framework →
Active research track · Agent systems
Instrument multi-agent workflows so a researcher can reconstruct what happened: which context each agent saw, which tools it used, how state changed, when permissions intervened, and whether recovery actually worked.
Evaluation method →
Active research track · Play
Design small game scenarios that isolate planning, memory, coordination, and adaptation. The goal is an evaluation people can watch, inspect, rerun, and disagree about constructively.
Testbed proposal →
What a project release should contain
When a track produces a public artifact, we intend to publish enough context to evaluate it: the question, system configuration, test cases, measurements, representative traces, failure examples, limitations, and the smallest reproducible implementation we can share.