Product, technical & vision brief · July 2026

An agent finishing a task and the task actually being done are two different things.

Waldo is the agent that cares for you: one user-owned intelligence carrying the context a person should not have to rebuild—priorities, commitments, boundaries, corrections, capacity, relationships, and outcome history—across the models, tools, and surfaces they use. Kennel is its first home on the Mac.

One agent · many presences Model-agnostic orchestration Suggest before execute Health as life context

Why we arrived here

Powerful agents still leave the coordination and consequences to the person.

At Atlan, Shivansh built more than 30 production agent instances and kept seeing expert users remember why each session existed, move context, catch blockers, and judge the result themselves. Suyash was running a design studio while training for an Ironman; his work tools and health tools were useful, but the coordination still lived in his head.

Learning 01

Health is foundational context

Energy, recovery, load, and capacity change what a realistic plan looks like. They belong in the agent's understanding of the person, but they do not define the product boundary.

Learning 02

Agent users already feel the pain

People using Codex and similar tools must remember why sessions exist, move context, catch waiting decisions, and judge whether the result was useful.

Learning 03

Completion is not an outcome

A green check, stopped process, commit, or final message is evidence about an agent session. It is not proof that the person's actual goal is complete.

Learning 04

The first surface should stay present

Kennel runs beside the tools people already use. It can hold attention, decisions, outcomes, and re-entry without asking the user to adopt another isolated workflow.

What changed: we began with health because it is personal, longitudinal context. Building it taught us that insight without action becomes another dashboard. We kept health as permissioned capacity context and changed the wedge to the place where the coordination problem is already visible: agent-native work.

One ordinary agent day

The session can stop while the human problem remains open.

Kennel makes the hidden handoff between machine activity and human responsibility visible.

1 · DelegateWork starts in CodexThe user gives an agent a goal, context, and constraints.
2 · ObserveKennel stays presentThe selected session, conversation, and processing state remain attributable.
3 · Agent stopsThe provider says “done”That closes or pauses an Agent Session. It does not prove the intended result.
4 · ReviewEvidence meets intentThe user sees what changed, what is missing, and which judgment still belongs to them.
5 · ReconcileConfirm, correct, deferThe outcome can be accepted, reopened, transferred, or consciously released.
6 · ContinueWaldo carries the lessonFuture work can inherit the correction and outcome—not merely another transcript.

The point: the user should not need to reconstruct yesterday before deciding what deserves attention today. Waldo preserves the re-entry point, the evidence, and the consequence—not just the transcript.

One agent · many presences

There is one Waldo for the person. Every surface is a different way to reach the same agent.

Models, interfaces, and devices will change. The person's identity, continuity, permission policy, corrections, and right to judge closure should remain durable.

Kennel

The work presence

A native Mac home for agent sessions, exceptions, evidence, decisions, open loops, and orchestration insight. It stays beside the tools people already use.

Waldo mobile

The life presence

The place where the person owns and corrects their priorities, commitments, boundaries, health context, relationships, routines, and capacity.

Waldo Core

The continuity and permission layer

The durable runtime, memory, policy, routing, evidence, and learning machinery that lets many models and surfaces behave as one governed personal agent.

One personal identity across changing surfaces

mindmap
  root((Waldo<br/>one agent for one person))
    Context
      Priorities and commitments
      Boundaries and corrections
      Capacity and health
      Relationships and routines
    Software presences
      Kennel on Mac
      Waldo on mobile
      Browser and messaging
    Intelligence fabric
      Coding agents
      General and specialist models
      Tools and skills
    Durable core
      Memory and provenance
      Permission policy
      Outcome history
      Open Loop continuity
    Future bodies
      Wearables and home objects
      Vehicles and robots
      One identity and policy
          

Kennel

Know what your agents are doing, what they produced, and what still needs you.

Kennel turns scattered agent activity into a legible control surface: what happened, what needs judgment, what evidence exists, what remains open, and what the system is learning about how your agents work.

1. ConnectBring Codex and later other agents through explicit provider boundaries.
2. ObserveKeep sessions, conversation, processing, decisions, and artifacts attributable.
3. ContinueReturn to the exact task, context, and point of interruption.
4. AskBring waiting decisions, missing evidence, and consequential judgment back to the person.
5. VerifyAccept, correct, reject, defer, reopen, or consciously release a claimed outcome.
6. UnderstandReveal agent habits: repeated stalls, useful steering moves, strong workflows, and weak handoffs.
7. OrchestrateBrief the next agent and choose the model, skill, and workflow that fit the outcome.
Supervision

Truthful state

Every session, decision, artifact, and provider event retains provenance so the person can understand what happened and where to return.

Judgment

Reviewable outcomes

What the work was meant to achieve, what evidence exists, and which human judgment is still missing.

Insight

Habits and orchestration

Where an agent repeatedly gets stuck, which corrections matter, which workflows succeed, and how that history should change future orchestration.

What Waldo owns

Waldo rents model intelligence, but owns the relationship between a person and their agents.

The bet is not that one model wins forever. The durable layer is where changing models become useful for one specific person without taking ownership away from them.

Own

Context

Priorities, commitments, boundaries, health and capacity, relationships, routines, corrections, and current Open Loops.

Own

Workflow

Briefs, conversations, handoffs, approvals, evidence review, delivery, re-entry, and closure.

Own

Trust

Explain, preview, approve, edit, undo where possible, audit, correct memory, export, revoke, and delete.

Rent and route

Models

Frontier, open, specialist, local, and future models enter through adapters and compete for the work they are best suited to perform.

The product triangle

Product, consumer trust, and technical depth have to compound together.

Product

Brief, Chat, Handoff, Patrol

Simple surfaces for understanding the day, asking for help, handing off work, supervising exceptions, and returning to what remains open.

Consumer

Taste, habit, emotion, safety

The agent must feel calm and specific enough that a person can let it closer to work, health, relationships, and consequential choices.

Technical

Runtime, router, memory, evals

The harness makes every suggestion scoped, attributable, permissioned, recoverable, model-routed, measured, private, and economically viable.

Brief

What matters now

A calm synthesis of capacity, commitments, consequences, and the decisions that deserve attention.

Chat

Think with Waldo

A conversation grounded in the person's current life and work context, not a blank thread.

Handoff

Delegate with control

A reviewable plan that states the goal, context, tools, limits, evidence contract, and permission required.

Patrol

Supervise exceptions

Continuous awareness of agent work that stays quiet until a decision, risk, or unresolved consequence needs the person.

Why all three matter: product without harness depth becomes another dashboard. Harness depth without consumer trust becomes another developer tool. A warm character without useful decisions becomes theatre.

The learning loop

Waldo should learn from steering and verified outcomes, not from activity volume.

Token counts, session duration, commits, and tool calls can describe activity. They cannot tell us whether the work mattered or whether the person is finished.

Observed

Evidence

Provider events, artifacts, changed files, decisions requested, plans, corrections, and user responses.

Inferred

Candidate pattern

A possible habit, recurring blocker, preferred steering move, or unfinished commitment. It remains an inference.

Confirmed

User-owned truth

The person accepts, edits, rejects, defers, or releases the candidate. Correction is part of the product.

Applied

Better future work

Waldo briefs the next agent, protects a boundary, proposes a follow-up, or chooses a better workflow.

The agent that cares for you: Waldo is not trying to maximize session completion. It carries the person's commitments, capacity, boundaries, and consequences long enough to help the real outcome move.

Behavioral evidence without scoring

Agent traces can reveal useful patterns, but the product must belong to the person being interpreted.

Studying adjacent behavioral-evidence systems strengthened our belief that plans, corrections, tool choices, and outcomes can teach a personal agent how someone works. It also clarified what Waldo should not become.

What we carry forward

Evidence-linked, longitudinal help

Patterns should point back to their evidence, accumulate across time, express uncertainty, and help the person re-enter work without reconstructing everything.

What we reject

Opaque scoring or external judgment

Waldo does not turn agent activity into a builder score, productivity grade, admissions signal, or irreversible personality claim. The user can inspect, correct, reject, or release every important interpretation.

The design consequence: behavioral evidence should help the person understand and steer their own agents. It should never become an opaque score produced for someone else. The useful unit is an evidence-linked pattern the user can inspect, correct, and apply.

Architecture

A local, durable projection of agent work feeds a separate personal continuity layer.

Provider activity is admitted through narrow adapters. It becomes durable evidence, not automatic authority over the user's outcome, memory, or future actions.

Provider adapters

Official provider interfaces enter through explicit adapters. Version, provenance, consent, size, and deduplication gates decide what can affect durable state.

Kennel local runtime

An append-only SQLite event ledger, deterministic replay, normalized events, and one public projection make the current Mac state explainable and recoverable.

Independent truth machines

Agent Session, Outcome Verification, and Open Loop remain separate contracts joined by identifiers and evidence. A provider event cannot silently close a human loop.

Waldo personal context

Priorities, commitments, boundaries, corrections, outcome history, and carefully scoped health context give future work the person-level continuity missing from isolated sessions.

Orchestration

With consent, Waldo can brief the next session, route a decision, preserve an open loop, propose re-entry, and choose the provider or workflow best suited to the task.

Trust and privacy

No blanket home-directory crawl, ambient screenshots, global input capture, or invented authority. Sensitive health values remain in their governed data plane; agent context is minimized and purpose-bound.

The agent platform

Every serious personal agent must solve eight problems beyond calling a model.

These are the platform responsibilities Waldo keeps first-party even when providers, transports, tools, and interfaces change.

Context

What enters intelligence

Layered, purpose-bound compilation with just-in-time tools, selected personal context, evidence, and progressive compaction—never the person's whole life in every prompt.

Action

What can change the world

Typed tools, reviewable skills, explicit blast radius, per-purpose permissions, previews, approvals, and recoverable failure.

Memory

What survives

Typed, provenance-bearing, inspectable, correctable memory written through a gate so a model cannot silently rewrite personal truth.

Initiation

What wakes the agent

KAIROS lets user intent, schedules, events, unresolved consequences, and changing context wake Waldo. Its tick-and-decide gate keeps the agent quiet when nothing deserves attention.

Verification

What proves it worked

Independent evidence, replayable traces, user judgment, and later outcomes challenge the agent's completion claim. Generation never grades itself.

Life intelligence

What makes help realistic

Capacity, health, calendar, commitments, relationships, routines, and boundaries shape what a good plan means for this person now.

Delivery

How help reaches the person

Channel-native cards, conversation, desktop presence, and selective notification ordered by consequence and timing—not engagement.

Threading

How thought stays coherent

Topic-persistent threads, exact re-entry points, shared identity across surfaces, and continuity that survives dates, model changes, and handoffs.

Tools, skills, and transports are different: a tool is a typed capability; a skill is a reviewable way of using capabilities; MCP and provider APIs transport them. None of these grants authority. Waldo keeps context, permission, memory, evidence, and outcome policy first-party.

Technical depth

The harness is the contract between a probabilistic model and the real world.

Runtime, routing, evaluation, memory, privacy, permissions, and cost are product decisions. They determine whether an always-present agent can be useful without becoming careless, expensive, or impossible to trust.

Principle 01

Deterministic seams

Models can propose. Code owns authentication, validation, permissions, idempotency, state transitions, and audit. A prompt is never a security boundary.

Principle 02

One writer per truth

Provider activity, outcome evidence, personal memory, and human closure have explicit owners. A projection or cache cannot silently become authority.

Principle 03

Memory never authorizes

A remembered approval is information about the past. Every new action still needs a current, purpose-bound grant.

Principle 04

Contract-first adapters

Models, providers, tools, and surfaces enter through typed capabilities. The surrounding personal-agent contract stays stable when any one of them changes.

Principle 05

Durability before autonomy

Long work needs journals, replay, idempotency, bounded retries, and explicit failure states before it earns broader authority.

Principle 06

Required context, not maximal context

Compile only what the declared outcome requires. Personal context remains permissioned, purpose-bound, correctable, and removable.

The durable personal-agent loop

flowchart TB
  subgraph UNDERSTAND["1 · Understand"]
    direction LR
    T["Trigger or user intent"] --> S["Attributable Agent Session"]
    S --> C["Context compiler<br/>purpose + permitted context"]
  end
  subgraph ACT["2 · Act safely"]
    direction LR
    M["Selected provider model"] --> D["Typed tool dispatcher"]
    D --> G{"Permission + policy gates"}
    G -->|allowed| E["Evidence + durable journal"]
    G -->|needs judgment| H["Return decision to human"]
    H --> E
  end
  subgraph LEARN["3 · Learn with the user"]
    direction LR
    O["Outcome Verification"] --> L["Open Loop disposition"]
    L --> R["Correctable memory + next brief"]
  end
  C --> M
  E --> O
  R -. "next useful action" .-> C
          
The compiled prompt — REASONS anatomy

Context is compiled from typed layers, not written as one giant prompt. REASONS is the anatomy: a provider-shaped working brief assembled for the intended outcome using only the personal context, tools, evidence, and safeguards that purpose requires.

RRequirementsThe trigger, intended outcome, constraints, and definition of done.
EEntitiesThe people, artifacts, commitments, and permitted personal context involved.
AApproachThe selected workflow or skill and why it fits this task.
SStructureAvailable tools, channel, evidence contract, and output shape.
OOperationsThe ordered steps, prior evidence, and explicit handoffs.
NNormsVoice, user preferences, correction history, and interaction boundaries.
SSafeguardsPermission, privacy, safety, and failure rules enforced outside the prompt.

Why compile instead of append: a personal agent may know a great deal, but useful context is not maximal context. Compilation controls relevance, privacy, latency, cost, and the chance that old information distorts the current job.

Five-layer memory — user-owned continuity

The five-layer design separates short-lived task context from durable personal truth. The important idea is not a particular database: every layer has a purpose, provenance, lifecycle, correction path, and authority limit.

Memory tiers · from working context to user-owned archive

flowchart TB
  T0["Layer 0 · Working context<br/>volatile, task-bounded, rebuilt"]
  T1["Layer 1 · Typed personal memory<br/>facts, events, discoveries, preferences, advice"]
  T2["Layer 2 · Episodes and evidence<br/>attributable sessions, corrections, outcomes"]
  T3["Layer 3 · Skills and procedures<br/>reviewable ways of working"]
  T4["Layer 4 · Archive and export<br/>history, deletion, recovery, portability"]
  T4 --> T3 --> T2 --> T1 --> T0
  T0 -. "new evidence, never direct truth" .-> T2
  T2 -. "candidate claim" .-> T1
  T1 -. "user correction or release" .-> T2
              
Scribe

Staged, evidence-linked memory

Observations enter a memory inbox. A governed write path promotes them only with provenance, scope, confidence, correction, reversibility, expiry, and deletion.

Reflection

Patterns without silent truth

Background reflection can connect episodes, surface recurring patterns, decay stale confidence, and propose memory changes. The person keeps the right to inspect, correct, release, export, or delete them.

Time and retrieval

History without rewriting it

Memory records when something was true and when Waldo learned it. Retrieval fuses relevance, recency, confidence, and the current purpose instead of treating every old fact equally.

Security — authority is enforced at every seam

Security is a sequence of independent refusals. Seeing evidence, inferring a habit, remembering a preference, or receiving provider completion never grants permission to act.

Defense in depth · enduring contract

flowchart LR
  I["Verified identity"] --> P["Fresh purpose-bound permission"]
  P --> A["Capability allowlist"]
  A --> T["Untrusted-input and taint checks"]
  T --> Z["Schema validation + sanitization"]
  Z --> X["Human approval for consequential action"]
  X --> V["Evidence and output verification"]
  V --> J["Attributable journal + audit"]
  J --> R["Revocation, correction, deletion"]
              
Memory

“The user approved this before.”

Historical information only.

Authority

“This action is allowed now.”

Requires a current grant for this purpose and capability.

Evidence

“Here is what actually happened.”

Recorded without silently closing the human loop.

The security principle: identity, permission, capability, input trust, approval, evidence, and audit are separate gates. Passing one never implies another, and a prompt is never a security boundary.

Owned orchestration intelligence

Models get better for everyone. Waldo gets better at helping one person.

Waldo does not need to replace foundation models. It can become the intelligence that decides when to stay quiet, what context is required, which model or tool fits, what permission is needed, how success should be judged, and what the next agent should inherit.

The user-owned learning flywheel

flowchart TB
  C["Consented context<br/>intent + commitments + capacity + memory"] --> P["Waldo brief or proposal"]
  P --> U["User steers<br/>approve, edit, reject, defer, correct"]
  U --> A["Agent or tool acts"]
  A --> E["Outcome evidence"]
  E --> J["Human judgment<br/>close, reopen, transfer, release"]
  J --> M["Correctable memory + Open Loops"]
  M --> C
  E --> R["Routing and workflow insight"]
  R --> P
          
What compounds

Intended, permitted, corrected

The durable history is not raw prompt volume. It is what the person intended, allowed, changed, verified, and consciously left open.

What improves

Timing, routing, and workflow

Outcome history teaches Waldo when not to interrupt, which model or skill fits, which evidence matters, and where a human judgment belongs.

What stays owned

The person's continuity

Models and surfaces can change without forcing the user to surrender their memory, permissions, preferences, or accumulated ways of working.

From model access to a Waldo intelligence gateway

flowchart LR
  I["Intent + permitted context"] --> W["Waldo Intelligence Gateway"]
  W --> Q{"Quality, privacy,<br/>latency, cost, tools"}
  Q --> F["Frontier reasoning"]
  Q --> O["Open or local model"]
  Q --> S["Specialist model or skill"]
  F --> V["Evidence + outcome review"]
  O --> V
  S --> V
  V --> W
  W --> N["Better next orchestration"]
          
01

Route broadly

Use replaceable model and tool adapters rather than binding the person's agent identity to one provider.

02

Trace meaningfully

Capture intent, route, permission, evidence, correction, cost, and outcome—not surveillance for its own sake.

03

Learn policy

Improve task classification, action timing, context selection, workflow choice, and route quality from consented outcomes.

04

Specialize later

Where outside models remain weak, Waldo can develop specialized orchestration intelligence while continuing to use the best external capability.

Economics is part of the product: skip first, route second, escalate last. An always-present agent stays viable by resolving routine cases deterministically, choosing the cheapest sufficient capability, and spending frontier intelligence only where the outcome justifies it.

The trust ladder

Authority is earned per action, not granted globally.

Waldo expands from understanding to action through visible usefulness, task-specific permission, evidence, reversibility, and repeated user-confirmed success.

Stage 1

Observe

Read permitted context, explain what matters, show uncertainty, and do nothing by default.

Stage 2

Suggest

Offer a concrete next move, preparation brief, recovery adjustment, re-entry point, or agent workflow.

Stage 3

Approve

Preview the plan and blast radius, request a purpose-bound grant, and let the person edit, defer, or refuse.

Stage 4

Automate selectively

Only reversible, bounded behaviors graduate after repeated success, clear audit, revocation, and a reliable exception path.

Continuous does not mean noisy: Waldo can wake often and act rarely. It should preserve exact re-entry points and interrupt only when timing, consequence, and the need for human judgment justify attention.

The people building Waldo

Systems depth, consumer taste, and a long habit of making things.

The team met through building: Shivansh and Ashish became friends at school over iOS jailbreaking; years later Shivansh met Suyash in the university Computer Center and showed him how to build a website by describing it to an AI tool.

Shivansh Fulper

Founder · engineering and architecture

At Atlan, he built more than 30 production agent instances and learned where capable agents still leave purpose, context, blockers, and outcome judgment to people. His earlier work spans OpenFn/C4GT, Project EKA data-curation infrastructure, native apps, and independent model implementation.

Suyash Pingale

Founder · product, experience and brand

Through SAPIEN, product work, and the lived collision of a design studio with Ironman training, he brings the consumer discipline agent infrastructure normally lacks: making permissions, uncertainty, attention, and control feel legible.

Ashish Tembhekar

Founding Engineer

A former AI engineer who shares core technical execution with Shivansh. His role is deliberately stated as Founding Engineer rather than cofounder unless the team's formal structure changes.

The product triangle: engineering makes the system dependable; product taste makes truth and control understandable; the consumer relationship makes the technology worth keeping around.

Long horizon

Software establishes the agent. Physical bodies become additional surfaces for the same identity.

Most physical AI work begins with an industrial arm, warehouse system, vehicle, or humanoid. We are interested in objects an ordinary person would actually welcome into their home.

Body

A physical surface

A desk object, wearable, home device, or eventual robot through which the personal agent can sense, communicate, and act.

Continuity

The same Waldo

Each body should not become another disconnected assistant. The person's context, corrections, relationships, and outcome history travel with the agent.

Permission

One user-owned policy

Home, wearable, vehicle, and robot surfaces act under explicit permissions and boundaries that the person can inspect, change, or revoke.

Why software first: the continuity and permission layer must be trustworthy before a personal agent is given sensors, movement, or physical-world authority. HTX Studio is an inspiration for the future craft of desirable consumer forms; it is not a partner, dependency, or current Waldo hardware program.

The company we are building

One agent that remains on the person's side as intelligence spreads everywhere.

Kennel begins with the coordination problem already visible in agent-native work. Waldo carries the same identity, context, permission policy, corrections, and outcome history across life, software, and eventually physical bodies—so every new model makes the person's agent more capable without making their life more fragmented.

See Waldo

Public anchors

Product and provider references

Waldo

Product

heywaldo.in

Codex

Provider contracts

Hooks · App Server

Cloudflare

Durable runtime substrate

Durable Objects