Research portfolio
More capability.
Clearer evidence.
Keep provides a shared setting for investigating how autonomous systems learn, act, coordinate, and recover. The questions below describe areas of investigation—not claims that each problem has been solved or that the underlying techniques are novel.
01 / Learning
Learning, memory,
and changing evidence
When does retained experience improve performance on genuinely new work? How should an agent handle corrections, withdrawals, and conflicting information without retaining invalid guidance or discarding useful knowledge unnecessarily?
Keep provides source-linked memory and adaptation mechanisms for investigating these questions. Broad learning transfer remains to be established through comparative evaluation.
02 / Skills
Useful and controlled
skill acquisition
How can an agent acquire or reconstruct a procedure, establish that it actually helps, and keep its use within appropriate limits? When should a candidate be promoted, revised, or retired?
Pattern extraction, structural validation, task performance, and permission to execute are different questions. Research must evaluate them separately rather than treat a “validated” label as a complete answer.
03 / Execution
Native execution
and verification
Which combinations of candidate generation, tests, acceptance policies, and operator review improve correctly completed repository work? What additional cost or human effort do those improvements require?
Keep's native coding and acceptance mechanisms provide an implementation base. Generated tests remain proposals to evaluate; passing a test suite is not proof of unrestricted correctness.
04 / Coordination
Coordination, authority,
and recovery
How should interacting operations share resources and authority? What should happen when an effect becomes uncertain, or its original permission or supporting information changes before recovery completes?
Controlled failure cases allow comparison of actual effects, resource obligations, and subsequent useful work. Complete behavior across all execution paths remains an open qualification task.
05 / Human control
Human control, privacy,
and isolation
How can people delegate meaningful work without transferring unnecessary authority or information? Which protections actually hold under a particular arrangement of credentials, processes, tools, and providers?
Both useful task completion and the cost of oversight matter. Configured controls must be distinguished from protections observed in operation.
06 / Evidence
Evidence and
reusable evaluation
What can an evaluator establish about authorization, execution, recovery, and missing observations? Which tools could help examine systems other than Keep?
The aim is to make relevant outcomes and limitations inspectable. Hash-linked records can support integrity checks; they do not, by themselves, prove truth, complete observation, or independent oversight.
The evaluation approach
How progress is evaluated
We distinguish implemented mechanisms, controlled demonstrations, and comparative results. Evaluation should include competent alternatives, held-out work where appropriate, actual outcomes, resource costs, and cases in which the proposed approach fails. A useful negative result can be a reason to simplify the system or change direction.
Resources that expand
the questions we can test
Additional model access, isolated infrastructure, controlled local models, and specialist collaboration can enable broader comparisons. Weight adaptation and accelerator-dependent work require their own implementation and qualification; access to hardware alone is not a demonstrated research result.
An open conversation
Research support
and collaboration
Red Rook AI is exploring support for well-defined research and reusable technical outputs across this portfolio. Potential contributions include research funding, model or compute access, independent evaluation, and specialist collaboration.
Discuss a research opportunity