Why this team is building Waldo
Shivansh and Suyash introduce the founders and the personal experience behind the company.
Open founder video ↗Product, technical & vision brief · August 2026
Agents can already produce code, research, plans, and documents. But when the output arrives, the responsibility does not disappear. You still have to decide whether to trust it, what happens next, who it affects, and what is still unfinished.
That work lands in the middle of the life you already have: your calendar, commitments, relationships, health, and whatever capacity you have that day. Each tool sees only one piece.
We are building Waldo so you do not have to hold all of that together alone. Waldo understands the context you choose to share, coordinates the agents and apps working for you, keeps track of what actually became true, and returns when the judgment is yours.
The problem we kept seeing
Shivansh saw this firsthand at Atlan, where he says he built and operated more than 30 production agent instances. Even expert users still had to carry the purpose of each session, move context between tools, catch waiting decisions, and work out whether the original problem was actually solved.
For Suyash, the same pressure showed up while running a design studio and training for an Ironman, with work, commitments, and health spread across tools that never understood how those things affected one another.
Every agent result becomes one more thing to read, trust, decide on, or remember. A finished run does not automatically mean the real problem is solved.
Those decisions land alongside messages, commitments, relationships, routines, health, and everything else that did not quite get finished.
Waldo is being built to understand the person, coordinate the work, and carry the outcome beyond whichever task or tool produced it.
A useful external signal: YC recently described moving from a simple internal agent loop to more than 50 Hermes agents serving individual employees as personal assistants, then building QM because the fleet became difficult to manage. We take that as a sign of what comes next: once agents become useful to each person, someone has to keep identity, context, permissions, and unfinished work coherent across them. Waldo carries that problem back to the individual. Many specialist agents may work for you; Waldo remains the continuing relationship across work and life.
The belief underneath Waldo: the more work machines can carry out, the more important it becomes to protect the person who must judge it and live with what happens next.
What Waldo is
Waldo carries the context you choose to share: why the work matters, what you have promised, where your boundaries are, which decisions belong to you, and what happened the last time.
With your permission, Waldo keeps the wider understanding of your work and life in view—your calendar, messages, commitments, routines, and health—so it can help you plan the day, follow through, and know what can wait.
You can change models, tools, and devices without rebuilding that relationship from zero. The intelligence underneath Waldo can change. Your history, corrections, permissions, and unfinished work should still belong to you.
Long term: one user-owned Waldo across work, life, devices, and eventually physical forms. Specialist agents, models, and surfaces can change. Waldo keeps your context, permissions, and outcomes coherent so you remain the author of what happens on your behalf.
One Waldo · many places
These are not separate assistants. Mobile, Kennel, and the agent harness are different parts of one Waldo relationship.
Ask, capture, catch up, plan the day, and handle a quick decision wherever you are. Waldo keeps the work connected to the personal context you choose to share.
When agent work needs careful judgment, Kennel gives you the space to see what happened, understand the evidence, and decide what happens next.
Behind both, the harness carries permitted context to the right agent or tool, remembers what it was allowed to do, recovers when something fails, and brings back what changed and what remains.
The broader goal: make agentic help useful without asking people to become agent operators. The surface can change; Waldo should still know where you left off.
Kennel
We are starting on the Mac with people who already use coding agents because this pressure is visible there today. Instead of opening every session and reading every update, you can see what changed, where an agent is stuck, which decision genuinely needs you, and what can wait.
When a decision needs more space than a phone can give it, Kennel is where Waldo brings the work, evidence, and agent state together. The session is still there when you need the detail. It is no longer the only way to understand the work.
Every session, decision, artifact, and provider event retains provenance so the person can understand what happened and where to return.
What the work was meant to achieve, what evidence exists, and which human judgment is still missing.
Where work repeatedly stalls, which corrections matter, which workflows reach acceptance, and what the person confirms should change future orchestration.
The experience we are building
This is the target cross-surface experience, not a claim of current end-to-end integration. The technical contracts matter underneath it. The experience should remain simple.
Across surfaces: a piece of work might begin as a quick request on your phone. Waldo can send the right parts to specialist agents, keep routine progress out of your way, and open Kennel at the exact decision that needs closer inspection. Once you decide, the work can continue in the background and return to your phone as a result, receipt, or clear next step. You should not have to reconstruct the story at any point.
Current product truth · internal evidence · August 2026
Since May 2026, we have built foundations in Kennel on Mac, Waldo on mobile, and the agent harness underneath them. We use these foundations internally. We do not have external users or revenue yet.
These statements summarize dated internal acceptance records and repository snapshots. They are not independently inspectable from this page; supporting artifacts can be shared with reviewers.
Controlled internal acceptance with Codex covers bounded session discovery, conversation history, processing state, continuing the same task, and archive cleanup.
The mobile app provides foundations for conversation, briefings, permissions, and personal context. The harness provides foundations for durable runs, governed actions, recovery, and delivery.
Production connectors, cross-app search, workflows, artifacts, multi-channel delivery, and one accepted piece of work that remains understandable across surfaces are not yet a finished experience.
The next proof: one real piece of work that begins with intent, moves through agents and tools, asks for the right human decision, shows what changed, reaches acceptance, survives a restart, and remains understandable from another Waldo surface.
Watch and read
These are the short companion artifacts for reviewers who want the people, product, and company narrative before going deeper into the system.
Shivansh and Suyash introduce the founders and the personal experience behind the company.
Open founder video ↗See the Mac surface for understanding agent activity, evidence, consequential judgments, accepted outcomes, and what remains open across sessions.
Open product video ↗The product wedge, founder story, market, business model, and long-term physical-AI direction.
What remains yours
That promise has to live in the product model, not only in the language. Context, outcomes, permissions, and history should remain yours even as the intelligence underneath Waldo changes.
Priorities, commitments, boundaries, health and capacity, relationships, routines, corrections, and current Open Loops.
Outcome intent, Work Units, briefs, conversations, handoffs, judgment, evidence, delivery, acceptance, re-entry, and Open Loop disposition.
Explain, preview, approve, edit, undo where possible, audit, correct memory, export, revoke, and delete.
Frontier, open, specialist, local, and future models enter through adapters and compete for the work they are best suited to perform.
The product triangle
Simple surfaces for understanding the day, asking for help, handing off work, supervising exceptions, and returning to what remains open.
The agent must feel calm and specific enough that a person can let it closer to work, health, relationships, and consequential choices.
The harness makes every suggestion scoped, attributable, permissioned, recoverable, model-routed, measured, private, and economically viable.
A calm synthesis of capacity, commitments, consequences, and the decisions that deserve attention.
A conversation grounded in the person's current life and work context, not a blank thread.
A reviewable plan that states the goal, context, tools, limits, evidence contract, and permission required.
Continuous awareness of agent work that stays quiet until a decision, risk, or unresolved consequence needs the person.
Why all three matter: product without harness depth becomes another dashboard. Harness depth without consumer trust becomes another developer tool. A warm character without useful decisions becomes theatre.
The learning loop
Token counts, session duration, commits, and tool calls can describe activity. They cannot tell us whether the work mattered or whether the person is finished.
Provider events, artifacts, changed files, decisions requested, plans, corrections, and user responses.
A possible habit, recurring blocker, preferred steering move, or unfinished commitment. It remains an inference.
The person accepts, edits, rejects, defers, or releases the candidate. Correction is part of the product.
Waldo briefs the next agent, protects a boundary, proposes a follow-up, or chooses a better workflow.
The agent that cares for you: Waldo is not trying to maximize session completion. It carries the person's commitments, capacity, boundaries, and consequences long enough to help the real outcome move.
Behavioral evidence without scoring
Studying adjacent behavioral-evidence systems strengthened our belief that plans, corrections, tool choices, and outcomes can teach a personal agent how someone works. It also clarified what Waldo should not become.
Patterns should point back to their evidence, accumulate across time, express uncertainty, and help the person re-enter work without reconstructing everything.
Waldo does not turn agent activity into a builder score, productivity grade, admissions signal, or irreversible personality claim. The user can inspect, correct, reject, or release every important interpretation.
The design consequence: behavioral evidence should help the person understand and steer their own agents. It should never become an opaque score produced for someone else. The useful unit is an evidence-linked pattern the user can inspect, correct, and apply.
Target architecture and authority
The target system is one product, not one giant database. Cross-surface Outcome state belongs to the governed Backend; local Mac and mobile effects retain local authority; provider activity enters as attributable evidence, never automatic authority over memory, acceptance, or future action.
Outcome, Work Unit, Actor, Artifact, Evidence, Verification, Judgment Request, Authority Grant, Acceptance, Open Loop, Schedule, Presence, and Context Claim will use versioned identifiers and conformance fixtures across TypeScript and Swift.
The governed cloud plane is designed to own shared Outcome state, durable runs, Work Unit orchestration, schedules, cloud connectors, policy, delivery, recovery, artifact versions, evidence events, and cross-surface presence.
The desktop plane is designed to be authoritative for local provider observation, local action grants, and its durable local projection. It will mediate and record the person's Needs You judgments, acceptance, re-entry, and Operator Mode inspection while syncing minimized events and receipts rather than indiscriminate local transcripts.
The mobile plane is designed to be authoritative for device consent, capture, protected local context, notifications, and offline actions. It will mediate and record lightweight user judgment without becoming the judgment authority. The target privacy invariant is that raw health and other sensitive device data must remain local by default; only purpose-bound claims or snapshots may leave the device.
Models, agents, people, deterministic workflows, connected services, and future machines will advertise typed capabilities, authority requirements, evidence contracts, cost, latency, cancellation, and recovery behavior.
Agent Session, Outcome, Evidence, Verification, Acceptance, Open Loop, and Context Claim must remain separate. The target system must prohibit blanket home-directory crawls, ambient screenshots, global input capture, raw cross-device sync, and authority inferred from memory.
The agent platform
These are the platform responsibilities Waldo keeps first-party even when providers, transports, tools, and interfaces change.
Layered, purpose-bound compilation with just-in-time tools, selected personal context, evidence, and progressive compaction—never the person's whole life in every prompt.
Typed tools, reviewable skills, explicit blast radius, per-purpose permissions, previews, approvals, and recoverable failure.
Typed, provenance-bearing, inspectable, correctable memory written through a gate so a model cannot silently rewrite personal truth.
KAIROS lets user intent, schedules, events, unresolved consequences, and changing context wake Waldo. Its tick-and-decide gate keeps the agent quiet when nothing deserves attention.
Versioned artifacts, receipts, independent verification, exact Judgment Requests, authorized acceptance, and explicit Open Loop disposition challenge every actor's completion claim. Generation never grades itself.
Capacity, health, calendar, commitments, relationships, routines, and boundaries shape what a good plan means for this person now.
Channel-native cards, conversation, desktop presence, and selective notification ordered by consequence and timing—not engagement.
Shared Outcome identifiers, exact re-entry points, cross-surface presence, durable dispositions, and continuity that survives dates, model changes, handoffs, restarts, and deliberate rest.
Tools, skills, and transports are different: a tool is a typed capability; a skill is a reviewable way of using capabilities; MCP and provider APIs transport them. None of these grants authority. Waldo keeps context, permission, memory, evidence, and outcome policy first-party.
Technical depth
Runtime, routing, evaluation, memory, privacy, permissions, and cost are product decisions. They determine whether an always-present agent can be useful without becoming careless, expensive, or impossible to trust.
Models can propose. Code owns authentication, validation, permissions, idempotency, state transitions, and audit. A prompt is never a security boundary.
Provider activity, outcome evidence, personal memory, and human closure have explicit owners. A projection or cache cannot silently become authority.
A remembered approval is information about the past. Every new action still needs a current, purpose-bound grant.
Models, providers, tools, and surfaces enter through typed capabilities. The surrounding personal-agent contract stays stable when any one of them changes.
Long work needs journals, replay, idempotency, bounded retries, and explicit failure states before it earns broader authority.
Compile only what the declared outcome requires. Personal context remains permissioned, purpose-bound, correctable, and removable.
The durable personal-agent loop
flowchart TB
subgraph UNDERSTAND["1 · Understand"]
direction LR
T["Trigger or user intent"] --> O["Outcome<br/>intent + constraints + acceptance policy"]
O --> W["Work Unit plan<br/>agents + people + workflows"]
W --> C["Context compiler<br/>purpose + permitted context"]
end
subgraph ACT["2 · Act safely"]
direction LR
M["Selected actor or provider"] --> D["Typed capability dispatcher"]
D --> G{"Permission + policy gates"}
G -->|allowed| E["Artifact, effect, receipt<br/>+ durable evidence"]
G -->|needs judgment| H["Needs You<br/>exact Judgment Request"]
H -->|approved grant| E
end
subgraph LEARN["3 · Learn with the user"]
direction LR
V["Independent Verification"] --> A["Authorized Acceptance"]
A --> L["Open Loop disposition<br/>close, defer, reopen, release"]
L --> R["Correctable memory + exact re-entry"]
end
C --> M
E --> V
R -. "next useful action" .-> C
Context is compiled from typed layers, not written as one giant prompt. REASONS is the anatomy: a provider-shaped working brief assembled for the intended outcome using only the personal context, tools, evidence, and safeguards that purpose requires.
Why compile instead of append: a personal agent may know a great deal, but useful context is not maximal context. Compilation controls relevance, privacy, latency, cost, and the chance that old information distorts the current job.
The five-layer design separates short-lived task context from durable personal truth. The important idea is not a particular database: every layer has a purpose, provenance, lifecycle, correction path, and authority limit.
Memory tiers · from working context to user-owned archive
flowchart TB
T0["Layer 0 · Working context<br/>volatile, task-bounded, rebuilt"]
T1["Layer 1 · Typed personal memory<br/>facts, events, discoveries, preferences, advice"]
T2["Layer 2 · Episodes and evidence<br/>attributable sessions, corrections, outcomes"]
T3["Layer 3 · Skills and procedures<br/>reviewable ways of working"]
T4["Layer 4 · Archive and export<br/>history, deletion, recovery, portability"]
T4 --> T3 --> T2 --> T1 --> T0
T0 -. "new evidence, never direct truth" .-> T2
T2 -. "candidate claim" .-> T1
T1 -. "user correction or release" .-> T2
Observations enter a memory inbox. A governed write path promotes them only with provenance, scope, confidence, correction, reversibility, expiry, and deletion.
Background reflection can connect episodes, surface recurring patterns, decay stale confidence, and propose memory changes. The person keeps the right to inspect, correct, release, export, or delete them.
Memory records when something was true and when Waldo learned it. Retrieval fuses relevance, recency, confidence, and the current purpose instead of treating every old fact equally.
Security is a sequence of independent refusals. Seeing evidence, inferring a habit, remembering a preference, or receiving provider completion never grants permission to act.
Defense in depth · enduring contract
flowchart LR
I["Verified identity"] --> P["Fresh purpose-bound permission"]
P --> A["Capability allowlist"]
A --> T["Untrusted-input and taint checks"]
T --> Z["Schema validation + sanitization"]
Z --> X["Human approval for consequential action"]
X --> V["Evidence and output verification"]
V --> J["Attributable journal + audit"]
J --> R["Revocation, correction, deletion"]
Historical information only.
Requires a current grant for this purpose and capability.
Recorded without silently closing the human loop.
The security principle: identity, permission, capability, input trust, approval, evidence, and audit are separate gates. Passing one never implies another, and a prompt is never a security boundary.
Owned orchestration intelligence
Waldo does not need to replace foundation models. It can become the intelligence that decides when to stay quiet, what context is required, which model or tool fits, what permission is needed, how success should be judged, and what the next agent should inherit.
The user-owned learning flywheel
flowchart TB
C["Consented context<br/>intent + commitments + capacity + memory"] --> P["Waldo brief or proposal"]
P --> U["User steers<br/>approve, edit, reject, defer, correct"]
U --> A["Agent or tool acts"]
A --> E["Outcome evidence"]
E --> J["Human judgment<br/>close, reopen, transfer, release"]
J --> M["Correctable memory + Open Loops"]
M --> C
E --> R["Routing and workflow insight"]
R --> P
The durable history is not raw prompt volume. It is what the person intended, allowed, changed, verified, and consciously left open.
Outcome history teaches Waldo when not to interrupt, which model or skill fits, which evidence matters, and where a human judgment belongs.
Models and surfaces can change without forcing the user to surrender their memory, permissions, preferences, or accumulated ways of working.
From model access to a Waldo intelligence gateway
flowchart LR
I["Intent + permitted context"] --> W["Waldo Intelligence Gateway"]
W --> Q{"Quality, privacy,<br/>latency, cost, tools"}
Q --> F["Frontier reasoning"]
Q --> O["Open or local model"]
Q --> S["Specialist model or skill"]
F --> V["Evidence + outcome review"]
O --> V
S --> V
V --> W
W --> N["Better next orchestration"]
Use replaceable model and tool adapters rather than binding the person's agent identity to one provider.
Capture intent, route, permission, evidence, correction, cost, and outcome—not surveillance for its own sake.
Improve task classification, action timing, context selection, workflow choice, and route quality from consented outcomes.
Where outside models remain weak, Waldo can develop specialized orchestration intelligence while continuing to use the best external capability.
Economics is part of the product: skip first, route second, escalate last. An always-present agent stays viable by resolving routine cases deterministically, choosing the cheapest sufficient capability, and spending frontier intelligence only where the outcome justifies it.
Trust and ownership
Model providers will keep getting better. We want that intelligence to compete for the work it does best. The lasting part should be the relationship you own: your context, corrections, permissions, outcomes, and the ability to inspect, edit, export, revoke, or delete them.
Read permitted context, explain what matters, show uncertainty, and do nothing by default.
Offer a concrete next move, preparation brief, recovery adjustment, re-entry point, or agent workflow.
Preview the plan and blast radius, request a purpose-bound grant, and let the person edit, defer, or refuse.
Only reversible, bounded behaviors graduate after repeated success, clear audit, revocation, and a reliable exception path.
Memory is not permission: something Waldo learned yesterday does not authorize it to act today. Consequential actions should remain specific, visible, and reversible wherever possible.
Care and attention
Waldo should not make you supervise more software, monitor more behavior, or stay permanently available. It should carry routine responsibility quietly and interrupt you only when the timing, consequence, or authority genuinely belongs to you.
Filter drafts, collapse duplicates, resolve reversible cases within policy, batch non-urgent choices, and reserve interruption for decisions whose consequence or authority belongs to the person.
Every Outcome can define sufficient evidence, acceptable quality, time and cost limits, and valid dispositions such as accept, defer, transfer, reopen, or consciously release.
Explicit boundaries, calendar load, rest, and permissioned health context can shape timing and plans. Waldo must not infer laziness, morality, personality, or commitment from behavioral or body traces.
Record what became true, what is waiting, what was released, and the exact next re-entry point so rest does not require keeping every obligation alive in working memory.
The product test: does Waldo leave you with fewer things that must stay alive in your head? Sometimes the right outcome is done. Sometimes it is deferred. Sometimes it no longer deserves to be carried.
The people building Waldo
As the founders tell it, Shivansh and Ashish became friends at school over iOS jailbreaking. Years later, Shivansh met Suyash in the Computer Center at IIITDM Jabalpur and showed him how to build a website by describing it to an AI coding tool. Waldo is the first company the three are building together. The experience and work history below come from the founders and their supporting records.
Founder · engineering and architecture
His work at Atlan showed Shivansh exactly where capable agents still leave intent, context, follow-up, and judgment to the person. His earlier work spans Project EKA's data pipeline, OpenFn/C4GT, native apps, and independent model implementation.
Founder · product, experience and brand
Suyash brings the part technical agent products often miss: what makes powerful software understandable and worth keeping around. Through SAPIEN and his product and brand work, he has learned how to turn complex systems into experiences people can trust.
Founding Engineer
Before Waldo, Ashish spent nine months working as an AI engineer. He built much of Waldo's first app and health-data pipeline, validated Health Connect on real Android hardware, and now works across native iOS, Supabase, and agent infrastructure.
The product triangle: engineering makes the system dependable; product taste makes truth and control understandable; the consumer relationship makes the technology worth keeping around.
The physical world
We do not think personal agents will remain inside chat windows forever. Over time, the same Waldo could meet you through a desk object, wearable, home device, vehicle, or small robot instead of giving every object a separate assistant with its own memory and agenda.
Maintenance, field service, inspections, contractor coordination, equipment repair, and acceptance can combine AI preparation with accountable human execution and real-world evidence.
Sensors, wearables, home devices, equipment, and vehicles enter through capability manifests that declare identity, location, telemetry, required authority, failure behavior, and observable completion.
Every physical effect must declare safety class, preconditions, permitted action, live state, abort path, human handoff, reversibility or recovery, telemetry, evidence, and acceptance authority.
Why software first: memory, permission, evidence, interruption, and recovery must work before a personal agent is trusted with sensors, movement, or physical authority. This is not a current Waldo hardware program. The form may change. The person it works for should not.
Waldo is the company
Waldo carries the context you choose to share across work and life, coordinates the agents and tools working for you, brings you in when your judgment matters, and keeps hold of what remains until you decide what happens next.
Kennel is Waldo's first home on the Mac and its first market wedge—not the company. Over time, the same Waldo can meet you through mobile, messaging, voice, connected services, and eventually physical forms. The surface changes; the relationship, memory, permissions, and loyalty to the person do not.
Meet WaldoPublic anchors