Selected work

Engineering systems for autonomy, evidence, memory, and real-world effects.

These examples are Cognilode research and internal engineering unless explicitly identified otherwise. They demonstrate technical depth and operating methods; they are not invented customer case studies.

Agent systems · Cognilode R&D · New

Multi-objective evolution for agent teams

We replaced scalar conversation scores, fixed prompt rules and cooldown-driven control with separate uncertain evidence dimensions, challenger populations, causal autonomy attribution and opportunity-cost accounting. Cross-regime historical traces showed that the same prompt family can behave radically differently over time, so the control policy preserves a contextual posterior instead of canonizing one recipe.

  • No hidden-human prediction or single alignment score as routing authority
  • Conversation-level allocation with Pareto and challenger populations
  • Manual prompts separated from machine-caused self-propulsion
  • Cash, public effects, avoided loss and foregone benefit kept distinct
Read the engineering note →
Autonomous systems · Cognilode R&D

Control planes outside the agent

We separate model reasoning from process custody, network authority, resource control, health observation, intervention, and rollback. The goal is to make autonomous work governable without pretending the model itself is an independent source of operational truth.

  • Independent process and resource custody
  • Permission and effect boundaries outside the reasoning loop
  • Human intervention and recovery paths
  • Provider-neutral execution substrate
Read the architecture →
Evidence · Cognilode R&D

Observability beyond model traces

For consequential autonomous work, API traces are only one plane. We reconstruct what the system saw, decided, delegated, spent, changed, delivered, and how humans altered the trajectory.

  • Messages, tool calls, delegation, and host effects joined by identity
  • Human interventions retained as causal events
  • Cost and outcome evidence kept distinct from mere activity
  • Failure states preserved rather than rewritten as success
Read the field note →
Memory & context · Cognilode R&D

Long-horizon memory without infinite context

Our memory work treats retrieval, summarization, procedural learning, provenance, revision, and deletion as one lifecycle. The useful test is not whether a fact can be retrieved once, but whether accumulated experience improves later work without overwhelming the agent.

  • Hierarchical summaries and role-specific continuity
  • Semantic compression and retention budgets
  • Procedural memory separated from raw trajectories
  • Historical corrections preserved with provenance
Read the tree-first memory implementation →
Computer use · Cognilode R&D

Browser automation that can survive real interfaces

We combine DOM-first mechanics with visual fallback, authenticated browser custody, exact action identity, bounded retries, and downstream evidence. This is designed for applications and operational workflows where visual-only automation is too slow or fragile.

  • DOM mechanics before expensive visual reasoning
  • Visual/VLM fallback for non-DOM controls and inspection
  • Authenticated session isolation and custody
  • Exactly-once semantic effects and provider readback
Engineering operations · Cognilode R&D

Software changes with explicit acceptance boundaries

We treat “the agent finished” and “the result is acceptable” as separate states. Repository work carries source identity, tests where useful, review evidence, delivery state, and known limitations rather than collapsing everything into a green status.

  • Exact source and artifact identity
  • Tests and review kept separate from delivery
  • Known failure and rollback coordinates retained
  • Reusable public interfaces preferred over bespoke glue
Apply the same discipline

Have a system where the hard part is making AI actually work?

Bring the blocked outcome, existing evidence, and operational constraints. We can determine whether the next step is architecture, implementation, evaluation, data work, or something simpler.

Discuss a project