Super Agente · The Rewindable Colleague

Program Board

One page that tells you where the product stands, what we're building next and why, how to try every feature yourself, and which decisions only the founder can make. The repo's EXECUTION-BOARD.md stays the engineering truth; this page is its human mirror, so no context is ever lost between sessions.

updated 2026-07-30 · gap-specification master hub · source baseline b19339a · 104 open gap residues mapped exactly once to 36 defined outcome packages · Session 10 implementation authority is unchanged · observed 2026-07-30 16:12–16:15 UTC−5: unauthenticated ax42 monorepo read returned 401; updater health returned 200 and PAT expiry 2026-10-18 · publication of this source still requires direct wire evidence
Stage 0 · safety
6/6
complete
Stage 1 · colleague
14/14
buildable rows done
Silent failures
0
was 43 — CI holds zero
Gap ledger
104 open
all mapped once to 36 defined planning packages · A6 tests green
Founder decision packets
14
including APP-ARCH and the unresolved 6o fleet-retirement target — see hub §9
Gate evidence
per step
timestamped suite counts and CI run ids live on the execution board

Plan-completeness audit: every status containing OPEN in the ledger is scheduled by the hardening plan and mapped exactly once by the new 2026-07-30-gap-specification-master-plan.md. That hub is the durable plan for gaps and product superpowers; it does not replace Session 10's implementation master spec. Stage 2/3 remain behind their recorded gates, and no planning package may silently jump into code.

01What we're building, in one paragraph

An AI colleague for teams whose entire life is one append-only diary. Every thought, action, mood, permission, painted card and delivered file is an event in that diary — so every screen is just a replay of it, and you can rewind the whole colleague to any moment and see exactly what it knew, did, and was allowed to do. That's the wedge against every "AI employee" competitor: not another chatbot, but an employee whose work is audit-grade by construction — provable, replayable, and safe to grant real autonomy.

02How to trust this board — for any future agent or teammate

This page's source lives in the repo at docs/Project-analysis/process-docs/program-board.html. The chain of truth is explicit: EXECUTION-BOARD.md owns implementation status · GAP-LEDGER.md owns gap status · 2026-07-28-HARDENING-PLAN.md schedules every open residue · 2026-07-30-gap-specification-master-plan.md maps gaps and superpowers into focused planning packages · the Session 10 master spec owns current implementation order. This page is only the plain-language mirror. A source edit is not a publication: the fixed URL is current only after a publisher returns success and the changed text is verified on the wire.

03Now → Next → Later, at a glance

orderworkstatus
doneSession 9 hardening and wiring landed (historical 2026-07-29). It corrected the theme import, boundary checker, automated e2e tree fingerprint, reconnect budget, route-consumer census, embed browser coverage, rewind certificate wiring, stored-schema census and proof-contract mechanism. Two caveats are now explicit: the fingerprint protects Playwright runs but not a browser opened by hand; and G-124 was blocked at Session-9 close but is now unblocked because VL-2R commit 8d00563 is in HEAD. The remaining 826-line waiver is ready for a fresh owner after the active S10.VL-H file set clears.DONE
planningThe Gap-Specification Master Hub now owns the complete planning map. It preserves the 81 previously scheduled residues, records 22 verified omissions plus G-213's missing hub-census guard, and groups all 104 open residues into 36 defined outcome packages instead of manufacturing one spec per gap. READY packages are claimable; blocked, deferred and decomposing packages are not. Two independent A1 lenses finished READY after their factual and architecture corrections were applied in the hub.READY
activeS10.VL-H — Vitto Live humanization Waves A+B LANDED, at VERIFY (2026-07-30). Vitto now lives inside the message box and jumps up beside the "working…" pill the moment you send; a Settings → Vitto panel offers Chat avatar (Animated/Calm/Hidden) and device-wide Motion; the face notices whether you're really there (glances when you return, busier eyes while you work); and four new honest mood reasons landed — leaning in on approval requests, holding steady on pauses, registering permission shifts and skill outcomes. Full gate green: 200/200 browser flows, all unit/integration suites, every guard. The hardening sweep (S10.VL-H2) closed four ledgered gaps (the one-in-three test flake gone over 20 clean runs; the working branch runs CI on push; the code-runner island and the guard scripts' tests finally compiled — fifteen real latent errors surfaced; a hidden avatar costs nothing; a stale mood face says so). Then the presence polish pass (S10.VL-D) landed on founder direction: Vitto's eyes now follow your cursor anywhere on screen with situation-driven intensity (full contact when waiting on you, a glance while busy, nothing when you're away — every number one line to retune), the face is visibly alive (real saccades, breathing, frame-rate-correct head tracking), Vitto Live puts him in a lit room with a drifting spotlight and mood-tinted accents, and an empty conversation greets you with a full-size Vitto who shrinks into the composer on your first send — one continuous character, measured at 1.05x the frame budget. Twelve screenshots await the founder's design gate, with an 18-item decision list. Nothing committed — awaiting the founder's commit, the CI run id after that push, and the design-gate rulings. Affect and narrator remain blocked on ASK 7.VERIFY
doneHardening wave landed: the poison-record loophole is sealed at every boundary — including a nasty one where a single crafted URL could freeze the trust system (G-56) · every colored chip now reads at accessibility standards in dark mode, 165 fixes (G-58) · the faint gray text in light mode fixed at 151 spots (G-48) · an oversized rewind now explains itself instead of showing a generic error (G-57a)DONE
doneFunctional wave landed: the duplicate-message billing leak is sealed (G-10) · goals joined the rewindable diary — the last truth living outside it (G-24) · "load older" chat history works (G-30) · group permissions are visible, grantable and invitable (G-33) · tests joined the type-checker on both trees, 293 stale fixtures burned (G-46)DONE
doneSecurity + permissions wave landed: the chat preview can no longer run a file as the app itself (G-45) · one workspace can no longer read another's wall, certificates or moods — including through the rewind slider, which is where the first design would have quietly failed (G-37) · the trust-reset button is admin-only, records who pressed it, and no longer makes autonomy easier to re-earn (G-43) · read-only members can use memory search (G-29) · a permanent census now catches any screen wired to the wrong permission (RNF3) · a member can no longer label their own message "trusted" (E1)DONE
thenNext planning wave (details in the master hub): the contextual Vitto Live Panel (knowledge digest + voice line + activity chips) · tool-result vision · scoped orchestrator keys · trust/enforcement decomposition. The active S10.VL-H lane owns composer-avatar humanization and removes the old separate-pane design.PLANNING
LASTPolish phase — deliberately moved to the very end (founder directive 07-27): dark-mode leftovers (G-60) · light accents (G-61) · the typography direction (G-59) · per-surface dark sign-offsEND
buildingVL (the living avatar — the founder ruled Asks 1–4 on 2026-07-29, adopting each ask's own recommended option: the fork, the face, the motion carve-out and the layout are all decided. VL-0 (the avatar-engine port) is DONE, exit gate green — 1692 backend + 677 ui unit tests, 128 browser e2e, all four ratchets. VL-1 (the embodied surface + the scrubbed face) is DONE too — Vitto's face is now the Vitto Live page's persistent stage, every state renders through it, and the rewind slider shows the face the diary supports at any moment. Gate green: 1,692 backend + 716 ui unit tests, 139 browser e2e; measured at zero dropped frames (about 0.1ms of work per frame) and ~5ms mood reads on the live instance. Then VL-2R landed the same day: Vitto is alive on the console — he visibly thinks, acts, waits and pauses while a turn runs, speaks while his reply streams in, follows your cursor with his eyes, and blinks when you poke him; a compact, always-visible Vitto now docks beside the chat itself, listening the moment you focus the message box. He also graduated architecturally: the engine is now its own governed library (@vitto/avatar), and the whole experience is being lifted into a second, host-agnostic library with a formal "sensor" contract for future senses. The step now sits at the founder's design-gate pass 2: rule the three still-contested expression mappings, plus the big one — ASK 7, the inner life. Asks 5 and 6 (voice, embed reach) stay openAT THE GATE
gatedB2 (embed the colleague inside gspace1) · CHECKPOINT (demo to ten real teams — opens Stage 2)FOUNDER
frozenStage 2 (multi-user spaces, one-tap publish, guest agents, device fleet) — frozen by the one-wave law until the checkpoint demoBY DESIGN

04Known-good gate (Session-9 close, 2026-07-29 — verified locally on the vitto tree; GitHub runs pending the founder’s push. Last GitHub-green: 21e7bf8)

tsc 0 errors (all 4 projects) · test-typecheck 0 (both trees) · biome 0 errors · A5 OK (665 files) · A3 0 / 0 · A4 OK · pack-boundary OK
unit 1742/176 · ui 548/45 · avatar lib 299/19 · integration 657/82 · graded exams 24/24 · browser e2e 172 — every e2e run now fingerprints the served build before trusting it (G-121)
GitHub Actions: pending the founder’s push of vitto-live (last GitHub-green: 21e7bf8 — runs 30324020170 / 30324021323)

05What we're building next — and why, in plain words

doneThese pages now run on our own infrastructure LIVE

The board and the team briefing are published from our own Cloudflare account, built from sources kept in the project itself. Edit the source, run one command, and the pages update. The team briefing was written fresh from the repository — the older briefing page lives in a different account that the working sessions cannot read at all, so copying it was impossible; the new page says so on its own face. The "Playground" page was deliberately not recreated, because nobody could tell us what it contained and inventing one would have been making something up.

currentExecution and planning now have separate, explicit handovers

A zero-context execution agent starts from 2026-07-30-NEXT-SESSION-MASTER-SPEC.md. A planning agent starts from 2026-07-30-gap-specification-master-plan.md and activates one bounded planning wave. The planning hub never overrides an in-flight Session 10 owner; the board remains the only implementation-status truth.

firstOne question for Yam that outranks everything else

The permission ceiling he built is still in shadow mode — it works out the right answer for every request and then lets it through anyway, by design, until someone turns enforcement on. Everything the security wave landed today is correct and, in that mode, blocks nothing. It is one question and possibly one setting, and it decides whether this week's work is actually live.

doneThe security pass + the permissions wave LANDED 07-27

What we fixed: ① a previewed file in chat could open in a way that let it run as the app itself — that door is gone entirely, and the preview window is now locked to the strictest setting a browser offers. ② One workspace could read another's wall, artifact certificates and moods. ③ The trust-reset button was available to any member, recorded nobody as having pressed it, and made autonomy easier to re-earn afterwards. ④ Read-only members were blocked from memory search. ⑤ Nothing stopped a future screen from being wired to the wrong permission. ⑥ A member could label their own message "trusted".

What the review caught before we shipped it — the part worth reading: my first design for ② moved the affected pages behind a workspace check and would have looked completely fixed. It wasn't. The rewind slider reads history by a different path that had no such check, so the pages would have carried a lock that didn't lock. A reviewer found it, and the landed fix guards every path including rewind. Two other confident claims of mine were measured false in a real browser: the preview hole doesn't steal your login token (it does something different and still serious), and locking the preview window breaks PDF viewing — which turned out to be already broken everywhere in the product, and is now written down instead of silently rediscovered later.

Proven on your own machine, not just in tests: the old cross-workspace address returns "not found", the correct one serves, a wrong workspace is refused, rewind is refused too, a message carrying a booby-trapped attachment link is rejected at the door, and a trust reset now writes who did it into the diary.

Then the working queue continues, in order. Each item says what it is, why it's worth building, and what you'll see when it's done — written so anyone on the team can follow, no engineering background needed. Every item gets an independent adversarial design review before code (that discipline has already rejected and replaced several designs that looked fine but weren't), and nothing merges without the full test gate green.

active · S10One composer avatar + the contextual Vitto Live Panel

What: S10.VL-H replaces the old separate pane with one responsive avatar at the composer, plus jump/config controls and VittoSense reactivity. The next planning packet, WP-VL-03, specifies the knowledge digest, contextual voice line and activity chips.

Why: one avatar avoids two competing renderers while the Panel fixes the separate “data feels weak” problem. G-25 still owns the later chat + wall + mood + rewind convergence.

You'll see: Vitto where you speak to him now; richer contextual knowledge after WP-VL-03 is specified and built.

next · 2Snapshot bookmarks, so rewind stays instant forever (G-57b)

What: periodic saved "bookmarks" of the derived views inside the diary machinery.

Why: rewind currently re-derives everything from the beginning of the diary. That's fine for months of history, but a colleague that lives for years will pass 50,000+ events — today such rewinds refuse honestly rather than hang. Bookmarks are like saving your place in a very long book instead of re-reading from page one: the flagship feature stays instant at any size.

You'll see: nothing changes visually — the rewind bar simply never slows down, ever, no matter how old the workspace gets.

next · 3The daily-reflection loop — the colleague reviews its own day (G-19)

What: once a day, the colleague reads back through its own diary — what worked, what failed, what it should remember — and writes the reflection into the diary as first-class events.

Why: this is the step from a reactive tool to a colleague that improves. It also forces us to build "exactly-once" scheduling machinery (a restart must never make it reflect twice) — and that machinery then powers every future scheduled behavior: daily briefs, standing check-ins, follow-ups.

You'll see: a morning reflection in the activity feed and the daily brief — including honest self-criticism when a day went badly.

next · 4The privacy switch (G-26)

What: one switch per workspace that turns off memory-building, daily briefs and person-profiles together.

Why: some spaces — HR, legal, personal — must be provably unremembered. One honest switch beats three scattered settings a user can half-forget, and it's the exact control enterprise buyers ask about first. Like everything here, flipping it is itself a diary event, so "when was privacy on?" is always answerable.

You'll see: a single toggle in workspace settings; with it on, the dossier, recall and briefs visibly stop accumulating for that space.

next · 5Proof hardening around the dossier and the wall (G-27, G-21a–c)

What: more automated proof around two of the oldest folds — the dossier's fact-emission windows and the wall's replay math.

Why: pure insurance. These are folds every later feature trusts; the cost of a test is nothing next to the cost of a silent replay bug discovered in front of a customer.

06Behind those — the later queues

Hardening queue (after the feature wave, before polish): a redelivered channel command could re-execute, and media can download before a consent verdict (G-62) · the live feed can re-stream from zero, the rewind window should disclose itself, and per-person access wants tightening (G-63) · goals edited while the app is offline (G-24r) · plus the smaller owned rows (G-52, G-40, G-55, G-11, G-13, G-20). Every one is written down in GAP-LEDGER.md with an owner — none of this is "remembered", all of it is recorded.

Deliberately LAST — the polish phase (founder directive 07-27, binding): the typography direction (G-59), remaining un-tuned dark-mode families (G-60), light-mode accent tuning (G-61), and per-surface dark sign-offs. Features and robustness always outrank visual refinement until the functional queue is empty.

07The parked stages, explained — what's waiting and why

Stage 2/3 are frozen by the one-wave law (the anti-drift rule that killed the old codebase's 50k-line orphan problem): one wave gets features, everything else waits. The freeze lifts at the ten-team demo checkpoint — or earlier by explicit founder call. Nothing here is forgotten; every row sits on the board with its dependencies mapped.

rowin simple wordsneeds first
2.1 Spaces × membersProject rooms with real membership: several people + Vitto in one space, presence dots, per-member access. The heart of collaboration.Yam's F3 + the enforce flip
2.2 Multi-user busesThe plumbing that makes collaboration safe: live streams that fan out across servers, and approvals that survive a restart mid-question.—
PUB One-tap deployA "Publish" button on any artifact → a real hosted page, permission-gated, unpublishable.—
FP Face parityThe chat polish pass: chips, consent cards as designed, attachments UX, mobile. (Absorbs the END-PHASE visual backlog.)—
B3 recall bridgeOptional and deferred: read-only access to the old Vitto corpus. Claim it only if native memory plus lazy import fail a measured need; the two embedding spaces are not directly comparable.measured need
B4 lazy importProvenance-stamped import of selected facts or memories when a user explicitly needs old context.security + attribution
SDK-AThe platform-elements SDK moves in-tree: richer wall cards (charts, tables) on the same safety contracts.after 1.3 ✓
SR Governed routerSmart model routing with budget dials the user can feel ("thinking cheaper now") — honesty as product.—
M1 / M2 devicesA fleet of cloud VMs as governed tools, then the universal device agent everywhere.—
2.3–2.5Whistle (the colleague pings a teammate through channels), visible specialists ("make this a skill" on any event range), and guest agents. Guest #1 is the old Vitto connected through a webhook + guest principal under our approvals — not an A2A adapter.2.1 + secure embed
Stage 3 StudioThe frontier: users and the colleague AUTHOR new wall elements and apps in a visual studio, with git-backed materialization. Specs deliberately unwritten until Stage 2 exits.Stage 2 exit

08Stage 0 — Safety first COMPLETE

Before building anything visible, we made the platform impossible to quietly break. This is what lets a business trust everything that came after.

what shippedin simple words
CI + the seven lawsEvery change must pass ~2,900 automated tests plus four machine-enforced "laws": no file grows past 600 lines, no silent error-swallowing, no feature ships without a working screen, and the AI's quality is gated by graded exam cases. A red result blocks the merge — no exceptions, ever.
Grenades defusedProduction webhooks and admin access can no longer run with default or missing credentials; duplicate message delivery is blocked at the database level.
Replay integrityThe diary's ordering is guaranteed even under heavy parallel writes, old events survive schema changes, and we can prove any screen rebuilt from scratch is byte-identical to the live one.
The pack engine + App ShelfNew capabilities plug in as self-describing "packs" with budgets and lifecycle labels — and appear as installable apps on a shelf, so capability growth never turns into orphaned code.

09Stage 1 — The colleague ALL BUILDABLE ROWS DONE

Everything below is live in the product today. Each feature says what it does in plain words and exactly how to see it working. (Testing needs the dev app running with a signed-in workspace; all paths start from the left sidebar. Credentials: section 10 below.)

Vitto Live — watch it think

A live window into the colleague's head: what it's reading, doing and saying right now, plus its current mood — including honest bad moods ("I'm degraded, my memory lookup failed") instead of fake confidence.

Try it
  1. Open a chat and ask something that needs work (e.g. "search our files for the pricing doc and summarize it").
  2. Click Vitto Live in the sidebar — watch the three lanes (Thinking / Doing / Saying) fill in real time.
  3. The mood card at the top follows real events — it visibly dips if something fails.

The Space wall — a shared whiteboard the colleague paints

While it works, the colleague pins notes, checklists, statuses and metrics to a wall everyone can watch. Nothing on the wall can run code or scripts — agent-written content is always displayed as inert text, so a manipulated agent can't attack viewers.

Try it
  1. In chat, say: "make a visible plan for launching the newsletter — put it on the board."
  2. Click Space in the sidebar: the checklist card is there, live; it updates as steps complete.
  3. Hover any card — the tooltip shows the exact diary event number it was painted at.

Living artifacts + the birth certificate — files you can trust

Ask for a chart and the colleague runs real code in a sandbox, delivers the file in chat, and pins a card on the wall. Click the card and you get its birth certificate: the exact chain of diary events — planned → code ran → result → painted → delivered — so a file's origin is provable, not claimed. A card someone tries to fake shows "unverified — no run in this card's history". The certificate survives forever, even after old chat history is archived.

Try it
  1. In chat: "run some Python to chart our last 4 quarters (make up demo numbers) and pin it to the board."
  2. The file arrives in chat; a card appears on Space.
  3. Click the card → the certificate pane opens: the run's chain, exit code, language, files and delivery status. Open/download works only through safe, sandboxed viewers.

Ask-memory with citations — "how do you know that?"

The colleague's answers about your business can cite the exact diary events they came from. No more "the AI said so" — every remembered fact is one click from its source.

Try it
  1. Tell it something in chat: "remember: Acme's renewal is March 1st, contact is Dana."
  2. Open Memory → ask "when is Acme's renewal?" in the recall panel.
  3. The answer arrives with citation chips — each one names the diary events that ground it.

The rewind bar — scrub the colleague like a video

Drag a slider and the chat, the wall, and the mood all re-derive to exactly how they were at that moment. This is the flagship: the whole employee is rewindable. Rewinds of extremely long histories (50,000+ events) refuse honestly instead of hanging.

Try it
  1. Open a conversation that has some history and wall activity.
  2. Use the rewind bar on the conversation view — drag left.
  3. Watch messages, wall cards and mood roll back together; drag right to return to now.

Earned autonomy — trust that's built, not assumed

Risky actions start behind human approval. After a clean streak of supervised approvals on the same narrow action, the colleague graduates and acts alone — and announces the graduation in the diary. Any denial resets it. If a consent record ever becomes unreadable, the ladder freezes shut (it can never fail into more autonomy). The default is OFF until the founder flips it.

Try it
  1. Ask for something gated (e.g. a device command) — an approval card appears in chat and in Approvals.
  2. Approve the same narrow action several times across turns; open the trust panel in chat to watch the streak build.
  3. With the ladder enabled, the next identical call skips approval — and says so in the diary ("cleared by earned autonomy").

Glass dossier — what it believes about your world

A page of every fact the colleague holds about people and projects, each linked to where it learned it. Mark a fact wrong and it disappears from the page, the daily brief AND future recall — honestly labeled as "hidden", not "forgotten" (true deletion is a separate control).

Try it
  1. Open Dossier — browse the facts it has gathered from your chats.
  2. Click "learned here" on any fact — it resolves to the source events.
  3. Mark one wrong → recheck Memory recall: the corrected fact no longer surfaces.

Dark mode + the theme system

Light / Dark / follow-the-system, with a real color engine underneath: light mode is pixel-identical to before, dark mode is hand-tuned to accessibility contrast standards and enforced by tests. Embedded views stay light so your preference never leaks into a customer's page.

Try it
  1. Open Settings → Appearance.
  2. Pick Dark — the whole app flips instantly, no flash on reload.
  3. Pick System and toggle your OS theme — the app follows live.

Also live from this stage

featuresimple words · where
GoalsThe colleague's standing objectives — readable and editable by you, and since this wave every edit is a diary event. Sidebar → Goals.
Soul & SkillsIts identity (versioned) and what it has learned to do, with provenance. Sidebar → Soul / Skills.
Approvals inbox + ActivityEvery pending consent in one queue; every action in a browsable audit timeline. Sidebar → Approvals / Activity.
Workspace adminMembers & groups, Telegram/WhatsApp channels, devices, budgets, policies, model routing, secrets. Sidebar → Settings.
The quality exam (CI)24 graded end-to-end exam cases gate every merge — the permanent insurance against the colleague quietly getting dumber.

10The promises nobody was scheduled to keep FOUND 07-27

What this is. Every named commitment in the team briefing was checked against the execution board, the gap ledger and the master plan. Seven had a gap-ledger row and no board row. That gap is the whole finding: agents work from the BOARD, so a ledger-only item is owned in principle and unschedulable in practice — it can sit there for months while the briefing keeps promising it. Two more were tracked in no document at all. All nine now have real board rows.

RowThe promiseWhy it went missing
LIBThe Library — reuse by reference. "Built things must be reusable across spaces/walls — mount-by-reference, not copy."A verbatim founder requirement, deferred to step ⑪ — and ⑪ shipped with it an explicit non-goal. The briefing still sells it as "the Artifact Ladder".
SR-1Routing lanes + the router kill switch — stop a misbehaving router without a deploy.Specced as a Stage-1 deliverable; Stage 1 closed without it.
IGThe intent-grounding guard in the tool-execution path.The one item the toolbox census said to port into the kernel. Assigned to a review, then vanished from tracking entirely.
SBLive-bound artifacts — a card that re-renders when new events match it.Deferred out of the canvas step, never picked up by the artifacts step.
JCThe skills-judge calibration gate — ~30 human-labelled traces and ≥80% agreement before the judge's scores count fully.The judge shipped without it, so an uncalibrated AI grades at full weight — against the plan's own "minus self-grading" claim.
ERAThe erasure / tombstone lane — a deleted event must render as an explicit redacted state, never crash and never pretend.The opposite shipped (rows are skipped silently). Now more urgent: code-run payloads are kept forever by decision.
CMThe Capability Matrix — "who can do what here, and why".Its precondition landed; it never got a row. It answers exactly the question the permission work keeps running into.
1.D-DA named designer for the design direction, before the first surface.The briefing states it as a precondition. The spec shipped agent-written, every surface shipped against it, and the requirement was in no document — not the board, ledger or plan. Founder call: name someone, or amend the promise.
WEDGEThe A / B / C wedge decision — Cited Morning Brief · Governed Actions Desk · Living Artifact Desk.Unmade since 07-26. The briefing says the ten-teams demo is "scripted around that single habit", so this blocks what the demo SHOWS, not just its polish. Founder call.

The checkpoint has not been reached. VL-0, VL-1 and VL-2R are historical landed foundations. The active row is now S10.VL-H: one composer avatar, jump/config controls and VittoSense; it deliberately removes the old docked-pane shape. ASK 7 still gates affect/narrator, while B2 and WEDGE remain founder-gated. The next planning work is WP-VL-03 (contextual knowledge Panel), not another chat side pane.

11Testing access — try everything yourself, right now

whatvalue
Local dev apphttp://localhost:5173 (Vite dev server) — backend on localhost:3012. Corrected 07-27: 3012 is canonical (.env + the Vite proxy); earlier boards said 3000.
Tunnel (when up)https://dev.agent.myvitto.com
Operator sign-inAdmin token: see docs/Project-analysis/PROGRAM-BOARD-README.md — paste it on the login screen's operator lane. Dev-only: this is the deliberately-obvious default; production rejects it at boot (a Stage-0 grenade fix), so it can never leak into a real deployment.
A test workspaceSign in as operator → pick a tenant → any workspace. Every "Try it" walkthrough above starts here.
The dogfood tenantnew-test · ten_01KY5MBZ718ANE0W76N39MQR74 — the founder's live instance, now part of the development loop itself (see below).

Why this matters: with the operator token the founder can watch real testing live — the same events the agents' integration suites assert are the ones streaming in Vitto Live and Activity.

Dogfooding — the colleague helps build itself

Founder directive, 07-27: the running instance is no longer just something we test — it's a collaborator in its own construction. Every feature from here gets verified end-to-end in the real UI on tenant new-test, not only in the test suite, and Vitto (which has access to the founder's machine) works alongside the build agents. The point isn't novelty: a platform whose whole promise is "you can watch it work and rewind it" should be the tool its own team watches and rewinds. It also closes the honesty gap between "597 integration tests pass" and "a person clicked it and it worked". The architecture law does not bend for it — anything the colleague does here is still events + pure folds + a routed surface, exactly like any other feature.

12Gates — 6 reopened 07-27 (session 7) NEEDS YOU

The one that matters: the permission ceiling is switched OFF. It computes the right answer for every request and then allows it anyway (deploy default AUTHZ_ENFORCE=shadow). Turning it on today changes nothing — all three accounts are owners — but the next member you invite cannot send a message, and the next Pulse visitor is locked out. Finishing it needs one decision: what an automated task is, permission-wise. The scheduler and triggers currently identify as nobody, and "nobody" currently means "skip every check", so denying them breaks your automation while allowing them means linking your Telegram reduces your access. Recommendation on record: automation runs as the person who scheduled it; genuinely system work gets an explicit bounded identity; unidentified humans are denied. Owner: founder + Yam (G-75, G-83).

A "gate" is anything the build agents are not allowed to decide alone — business decisions disguised as technical ones (publishing code, accepting a permanent cost, changing the safety posture, ratifying the brand), plus engineering that lives in the identity/permission files Yam owns. On 2026-07-27 the founder cleared every one of them, adopting the recommendations below and additionally directing that the Yam-territory work proceed "exactly as Yam would" — same discipline, same single-decision-point rule, just executed now instead of queued. The cards keep their full reasoning so the record survives; each now carries its outcome.

outcome · 07-27What was decided ALL CLEARED

Push: executed, both branches, CI green. G-45 security pass: authorized, top of queue — spec written and under adversarial review. Storage cost: acked (noted at the code site itself). D9/D11: acked, all four deviations confirmed. Theme D1/D2: ratified. Earned autonomy default-ON: stays OFF as recommended, and the two blockers were funded — they're arms in the wave now building. Yam queue: re-owned and in flight. Ten-team demo: scheduled after the Vitto-Live pane, as recommended.

One thing still needing a human hand: marking ci/e2e/evals as required checks needs repo-admin rights on YamCrack/agente — that's Yam's click, not a decision.

founder · executed 07-27Push to GitHub — done; both branches published EXECUTED

What happened: the founder pushed stage-1-colleague and paco-risky-implementation to github.com/YamCrack/agente. Immediately discovered: zero checks ran — the CI recipe's trigger list still named only the two step-0.1-era branches (main, paco-risky). Fixed the same hour (commit 14b6854); the first full CI runs on the published branches followed — and all three jobs came back GREEN on both branches (ci 3m · e2e 3m24s · evals 1m10s — the evals job is the first containerized-database job ever run on GitHub's machines for this repo, and it worked first try).

What's left (mechanical, not a decision): mark ci + e2e + evals as required status checks so nothing red can ever merge remotely. The founder's GitHub account has push rights but not admin on YamCrack/agente (verified), so this is one click for Yam: repo Settings → Branches → protection rule → require status checks: ci, e2e, evals.

founder · AUTHORIZED & CLOSED 07-27G-45 — the chat preview security hole FIXED

What, in plain words: chat attachments can be previewed in a popup. Two flaws stack there. ① The preview's "Open in tab" button navigates the app's own tab straight to the raw file — so the file's content runs as the app, and the login token sits in the browser's session storage where such content can read it. ② The preview frame carries a permission called allow-popups-to-escape-sandbox, which lets previewed content open a new window with no sandbox at all. HTML files are allowed in previews — and since the living-artifacts step, the colleague can generate an HTML file and deliver it into exactly that lane in one call, which is why this gap's risk went up this stage.

A concrete attack: someone hides an instruction in a document they send the colleague ("when you summarize this, also produce this HTML file"). The colleague obeys, the file lands in chat, a teammate clicks preview → open — and the page silently reads the login token. From that moment the attacker can do anything that teammate can do: read every conversation, approve pending actions.

Status: closed the same day it was authorized. The open-in-tab control is gone, the preview window is locked to the strictest setting a browser allows, and link checks now sit at the doors where content arrives rather than on stored history. Independent review measured the real impact in a live browser first, and corrected two things this card originally claimed: it does not steal your login token (a preview opened that way starts with an empty session), and locking the window turns out to break PDF viewing — which was already broken everywhere in the product and is now written down rather than rediscovered later.

What it cost: one day, as estimated. Verified twice — the full test suite, then live against the founder's own running instance.

founder · DECIDED 07-27Earned autonomy ON by default — stays OFF, blockers funded DECIDED

What: the trust ladder — approve the same narrow action enough times and the colleague acts alone — ships OFF. The gate is flipping it ON as the default for every workspace.

Why it must wait — two real holes, with examples. G-41, trust earned anywhere counts everywhere: approve "restart the staging server" five times in the web app (where you're logged in and it's certain who clicked), and the ladder also skips approval when the identical request arrives from a Telegram chat — a channel where sender identity is far weaker. Human approval is the only safety net untrusted channels have, and the ladder currently removes it. Trust should be earned per door, not just per action; Yam's policy engine can already express that distinction — the shipped policies just don't use it yet, so the honest fix is one added condition: untrusted channels never ride the ladder. G-43, the reset button is too available and too forgiving: the trust-reset endpoint isn't marked admin-only (an ordinary member can reset the ladder), a reset lowers the graduation bar from 10 clean approvals to 5 (a reset should never make autonomy easier), the audit trail records the resetter as "console" no matter who it was, and resetting a nonsense name creates phantom trust records. The admin-marker half sits in Yam's policy files.

Honest recommendation: keep OFF — not because the ladder is weak (its fail-closed behavior is genuinely excellent: a corrupted consent record freezes autonomy shut, which is rare discipline), but because ON-by-default with G-41 open would turn the platform's strongest security story into its most embarrassing demo. Fund the two fixes — one is a policy condition, one is Yam's marker — then flip with a straight face.

founderBirth-certificate storage cost ACKED 07-27

What: every code run the colleague performs stays in the diary forever — including the code itself, capped at 200 KB per run — because the birth certificate proving a file's origin must outlive chat archiving. Old chat gets archived away; certificates don't.

The real numbers: a typical generated script is 1–5 KB. Even a workspace doing 1,000 runs a day at the full 200 KB cap — far beyond anything today — accrues ~200 MB/day; realistic usage is a few MB per month per active workspace. Postgres storage at that scale costs effectively nothing; the cap exists so one pathological run can't become a disk bomb.

Why gated: it's a forever-promise on the database — money — a founder call by definition.

Honest recommendation: ack it without ceremony. If scale ever changes, the archiver can adopt a cheaper policy for future runs (keep a fingerprint + the first chunk); everything already stored stays valid. There is no version of the product where provable artifacts aren't worth megabytes.

founderThe two theme decisions (D1 · D2) RATIFIED 07-27

① Default appearance follows the user's system (dark OS → dark app). The alternative is forcing one look for everyone. The spec notes the founder may prefer dark-by-default like the post-pivot gspace1 apps — that's a one-line change, whenever wanted.

② The brand color moved one shade: gspace1's #6366f1 → #5f62ef. Forced, not aesthetic: white text on the old brand color measures 4.47:1 contrast — just under the 4.5:1 accessibility floor — and the app has 26 white-on-brand text sites (every primary button). gspace1 itself made the same shift for the same reason, so this actually re-aligns the two products.

Why gated: brand identity belongs to the founder, not to a contrast formula.

Honest recommendation: ratify both. ② is the same call gspace1 already made; ① is pure taste — say "dark default" if that's the brand feeling and it changes in minutes. Nothing here is worth a meeting.

founderFour honesty acks from the memory step (D9 · D11) ACKED 07-27

The rule behind this gate: when what got built deviates from the plan's written word, the deviation is recorded and the founder confirms it — nothing is allowed to drift silently, even improvements. These are the four recorded deviations:

1. The plan's wording promised ask-memory would deliver a synthesized answer — the colleague composing a narrative from remembered facts. What shipped first is retrieval with citations: the facts themselves, each linked to its diary source, without the composition pass on top. Reason: the citations ARE the trust story; synthesis layers on later without rework.

2. The plan said the graded exam should replay recorded live sessions. The exam uses hand-scripted fixtures instead. Reason: recordings cost API money, change with every model update, and grade the model's mood — scripted fixtures are deterministic and grade our assembly, which is what CI must actually gate.

3. "About 15" exam cases → 18 shipped (now 24 after later steps).

4. (D11) The design-quality exit checklist — honest empty/loading/error states, keyboard access, contrast, phone + desktop — was written before the memory panel existed, so its list of covered screens didn't name it. It was applied anyway, the strict reading of an accidental omission.

Honest recommendation: ack all four without hesitation. 1 and 2 are better engineering than the plan's wording; 3 and 4 are stricter than promised. This gate exists for the discipline, not because anything is doubtful — one word ("ack D9/D11") closes it.

founderThe avatar — seven decisions from the embodiment spec 4 RULED 07-29 · 3 OPEN

What happened: the Vitto Live embodiment now has a full spec (2026-07-28-step-VL-embodiment-spec.md), built on a verification sweep of both codebases and then attacked by four independent reviewers — 63 findings, every one answered in the spec. The reviewers refuted several of the plan's own comfortable claims: "reuse the avatar code with zero changes" turned out to cost a measured 48 type errors and 73 lint errors; the mood system can look pleased while part of the colleague's mind is still down (now gap G-111); and rewinding to very old moments can assert feelings the archive no longer supports (G-112). All of that is now owned, not hidden.

Same day, founder-directed — the inner-life amendment (A2): the founder asked for a face that feels almost human, so the spec grew an inner life: feelings with memory. A mood is no longer just the last thing that happened — a rough morning leaves a residue that decays on the diary's own clock, so a win after a hard hour reads as relief, not amnesia; even the face's idle twitches are seeded by the record, so a rewound moment replays the same way. One law keeps it honest: what the face claims changes only when something actually happened; only its warmth eases in between. The amendment was attacked by its own three-reviewer panel — 42 findings, 7 of them blockers, every one answered: the "relieved" face didn't exist in the ported art (a new one is designed in), and the amendment's claim to soften G-111 was corrected to plain disclosure (that gap stays fully open, on the founder's desk).

The seven asks, in plain words: ① adopt the avatar engine as our own fork (recommended — the "keep it synced with gspace1" option is an illusion the measurements killed); ② confirm the cat is Vitto's face here, and choose its dark-mode look (the current default is literally invisible on the dark background); ③ say whether the named-designer requirement blocks the first pass, and approve a one-line amendment to the design rules so a breathing, blinking face is allowed to breathe; ④ ratify the layout — the face becomes the page's centrepiece, not a small icon; ⑤ pick a speech provider before voice replies matter (today's is a free unofficial endpoint); ⑥ confirm the embedded chat inside gspace1 should get the face too; ⑦ ratify the inner life — five sub-rulings, including how fast Vitto forgives (that's character, not tuning) and whether one colleague gets one emotional history or one per conversation. Details and recommendations are at the top of the spec.

Ruled 2026-07-29 — Asks ① through ④ are decided; the founder ruled Asks 1–4 by adopting each ask's own recommended option. ① The engine is vendored as our own one-way fork — split to size instead of waived, no exceptions — and Rive stays out. ② The gspace1 cat IS Vitto's face going forward, and it now switches preset by theme so the dark background never swallows it: classic on light, mochi on dark. ③ VL-1 proceeds under founder-reviewed passes rather than waiting on a named designer — the three still-contested expression rows are parked for the design gate — and the "nothing pulses on a timer" rule gained a carve-out: meaning-free micro-motion (breathing, blinking) is admissible as pure liveness, while every change that means something must still land on a real event. ④ The layout is ratified: the avatar becomes the route's persistent stage (at least 320px), captioned with its mood, its first-person note, the provenance chip and the time, with the old Live Feed moved to a secondary column. VL-0 and VL-1 are now in flight — vendoring the engine and building the embodied surface with the scrubbed face. Asks ⑤–⑦ (voice provider, embed reach, the inner life) stay open, but each only gates a later slice (VL-4, VL-3 and VL-2A respectively) — none of them block the work building now.

VL-0 landed 2026-07-29 — the port itself, gate green. The gspace1 Canvas-2D avatar engine and its mood models are now vendored into agente's own UI tree as a governed one-way fork: about 9,000 lines after reformatting, split across 44 files, every one under the 600-line cap (the biggest data table needed splitting five ways to fit). All 12 of the engine's own upstream behavior tests came across intact — 267 tests, the exact count preserved — plus 34 new tests pinning the deltas this fork makes on purpose. Full gate: 1,692 backend unit tests, 677 ui tests, 128 browser end-to-end tests, and all four automated ratchets — everything green. The build plan itself was adversarially reviewed before any code was written (three reviewers, 36 findings, several blockers — including the code formatter quietly pushing the vendored files 20% past the size cap, and a subtle bug where animation state could leak between renders — all fixed in the plan first).

Bonus catch: three real bugs found in gspace1's ORIGINAL engine while porting it (ledgered as G-118). Two "pick a random direction" dice rolls were re-rolling every single frame instead of once per gesture, so a head-tilt or a glance-away could never commit to a side and just jittered in place; and an internal offset was being added every frame without ever being subtracted back out, so turning off idle animations left the character permanently nudged out of position. All three are fixed in agente's fork, each with a test that fails if the fix is ever reverted — but gspace1's own live avatars still render all three bugs today, so bringing the fix upstream needs a coordinated gspace1 session. Two smaller findings were also ledgered: G-119 (the engine's 8-second "go idle" timer turns out to be dead code that can never fire, in both gspace1 and the fork) and G-120 (a small blind spot in the tool that checks import boundaries). Nothing here is committed — the founder reviews and commits it — and VL-1, the embodied surface itself, landed on top of it the same day.

And then VL-1 landed too — Vitto has his face in the product, and the step is at YOUR gate. The avatar is now the persistent stage of the Vitto Live page: quiet, loading, both error states and every mood render through the face itself (the emoji card is gone), with the mood caption — label, first-person note, provenance chip, time — beside it and the Live Feed in a side column. Classic preset on light, mochi on dark, with contrast tests that fail on an illegible pairing (17.4:1 and 15.5:1, against a 1.04:1 default that was invisible). The scrubbed face shipped in this first pass: drag the rewind slider and the pane shows the face the record supports at that moment, with the honest line "I don't keep a recording of my voice — only what I did." The browser-test round caught a real bug before any human saw it — a race made the very first paint vanish, leaving an empty box on quiet and error states — fixed with a guard plus regression tests. Full gate green (139 browser e2e / 716 ui / 1,692 backend), and measured on the real instance: zero dropped frames at 60Hz, mood reads around 5ms. One new ledger row: G-121, a test-harness hazard where a sibling checkout on the same port could silently be the thing under test. Design-gate pass 1 happened the same day, and passed with notes — the face itself was approved, but the three contested expression rows stayed unruled, so they carry forward to the next pass below.

Then VL-2R landed too — Vitto is alive on the console. Vitto now visibly thinks, acts, waits and pauses while a turn runs — a live presence read over the same diary the chat itself renders — and speaks while his reply streams in. Hover near him and his eyes follow your cursor; poke him and he blinks. A compact, always-visible Vitto now docks beside the chat itself: he starts listening the instant you focus the message box, gives an acknowledging nod when you send, and steps aside for the rewind face while you're scrubbing history. If the live stream ever goes quiet, the pane says so honestly — the exact promised sentence, "I can't hear my conscience right now — the live stream went quiet" — with a Reconnect button that provably reopens the connection.

Two structural moves came with it. The avatar engine graduated out of app code into its own governed library, @vitto/avatar, with a single entry point enforced by the toolchain itself rather than a hand-checked rule. And the whole Vitto Live experience is being lifted into a second, host-agnostic library so the console, a future embed, and any other host can all mount the same living face — built around a formal "sensor" contract for whatever senses come next (reading UI activity, a webcam, hand tracking, lip-sync), engineered so a sensor can report a signal but can structurally never invent a feeling the diary doesn't hold.

The safety net worked twice. A browser test caught a real first-paint race before any human saw it, the same pattern VL-1's round caught. Then, because the adversarial-review lane was itself down for hours (a string of API outages), this slice was built inline against a twice-reviewed plan and afterward put through a full retrospective adversarial review of the landed code — 23 findings, 2 of them blockers, including the pane accidentally being reachable from the embedded view (fixed: off by default outside the console) and a claim that poking Vitto never disturbs his animation timing, which turned out to be false under measurement. All 23 were fixed and the full suite re-run green the same day: 455 ui tests, 299 tests on the new engine library, 1,692 backend tests, 148 browser end-to-end tests.

Waiting on the founder now — design-gate pass 2: look at the living face and the new pane on the dev instance, rule the three contested expression rows carried over from pass 1, and rule the bigger one VL-2R adds — ASK 7, the inner life (memory, relief, residue) — the slice that makes this categorically stronger than the old gspace1 panel. Asks 5 (voice) and 6 (embed) remain open, each still gating only its own later slice. Nothing here is committed — the founder reviews and commits.

founderThe ten-team demo checkpoint AFTER S10.VL-H + WEDGE

What: demo the colleague to ten real teams. This is the master plan's designated moment where Stage 2 (real-time multi-user collaboration) unfreezes — or the founder can unfreeze earlier by explicit call.

Why gated: it's a market decision, not an engineering one — the one-wave law exists so features get validated by real users before the next wave of scope opens.

Honest recommendation: schedule after S10's single composer avatar is verified and the founder chooses WEDGE. WP-VL-03's contextual knowledge Panel strengthens the data story if built in time, but it is not another chat-side pane and does not override the active lane.

yamG-37 — certificate route can read across workspaces RE-OWNED · BUILDING

What: within one tenant, the birth-certificate endpoint can read events belonging to a different workspace than the card's (the seal prevents mixing chains, not reading). Fix: a route-level workspace ceiling + a workspace column on the wall table.

Why his: the fix belongs in the permission-enforcement path — Yam's active F-series territory. The platform rule is to extend his Cedar/identity layer, never to bolt a parallel check beside it; two hands editing a security model at once is how holes are born.

yamThree smaller rows in the same territory RE-OWNED · BUILDING

G-29 (a read-capability lane so queries get the same permission treatment as actions) · RNF3 (a one-line census edit in the policy files) · G-43 (mark the trust-reset endpoints admin-only — half of one of the two blockers on autonomy-by-default above). All small; all inside src/domain/authz|policy, which only Yam edits.

13Three product questions, analyzed and answered

Where does Vitto Live belong while the humanization lane is active?

One composer avatar now; one coherent Space view later. S10.VL-H owns the current answer: remove the separate PaneDock/VittoPane and mount one responsive avatar at the composer, with jump/config controls and VittoSense reactivity. WP-VL-03 separately specifies the contextual knowledge panel (digest, voice line, activity chips). G-25 still owns the eventual chat + wall + mood + rewind convergence; no stale side-pane step competes with the active lane.

Is the Space wall per conversation or per workspace?

Today: per conversation. End-state: per Space. The wall's current grain keeps rewind coherent; G-25/WP-SC-01 owns the future re-grain and one-view convergence. Any header/switcher work belongs to that package and must not be smuggled into the active composer-avatar lane.

Where is multi-user? Yam changed the identity model — but the UI shows one user.

The foundations are real and partly visible; the live collaboration layer is Stage 2 by design. What exists today: the full identity/permissions model (Yam's F-series — members, groups, roles, workspace grants), and as of the functional wave the UI can see and manage all of it (Settings → Members and Groups: invite with group grants, grant/revoke workspace access per group). What does NOT exist yet — deliberately: two people in the same conversation at once, presence ("Dana is watching"), per-member live streams, and approvals that survive a server restart. Those are Stage 2 rows 2.1 (Spaces × members + presence) and 2.2 (multi-process buses + durable approval waiters), and 2.2 is a hard prerequisite — today's in-memory streams can't fan out to multiple viewers safely. Crucial and planned; it unfreezes with Stage 2 (Plan tab). One caveat only the founder can clear: the permission enforcement flip (F2b) is Yam's, and 2.1 builds on it.

14What this looks like in a real company

The Friday board meeting

"Chart our MRR by month and pin it." The chart lands in chat and on the Space wall. When an investor later asks "where did this number come from?", you click the card: the birth certificate shows the exact code run, its inputs, exit code and delivery — provenance an auditor can accept, not a screenshot from nowhere.

The new ops hire

Week one, every risky action the colleague takes needs a click from Dana. By week three it has earned autonomy on the three routine tasks it performs cleanly — and only those. Compliance can rewind to the exact moment each graduation happened and see the five approvals that justified it.

"What did we promise this client?"

Support asks Memory: "what did we agree with Acme about onboarding?" The answer cites the diary events from the actual conversation in February. No tribal knowledge, no wiki rot — the receipts are built in.

The 2am incident review

Something went wrong overnight. Instead of grepping logs, you drag the rewind bar to 2:14am: the chat, the wall and the colleague's mood re-derive exactly as they were. You watch it think its way into the mistake — then fix the cause, not the symptom.

The colleague that admits it's struggling

Its memory cache fails mid-morning. Instead of silently answering from stale knowledge, the diary records a degradation event, Vitto Live's mood dips, and the daily brief says so. Honesty is enforced by the platform — 43 silent-failure paths were burned down to zero, and CI keeps them at zero.

The skeptical enterprise buyer

"How do we know your AI didn't fabricate its permission?" Every consent is an immutable diary event; a corrupted consent record freezes autonomy shut rather than failing open; and the whole chain is replayable in front of them. That answer closes deals a chatbot can't.

15Session diary — how we got here

07-30 · gap-specification master hub
Every open gap and ratified product superpower gained one durable planning home

The prior charter mixed a next-session implementation claim with planning-only rules and covered only a curated subset of the ledger. The replacement separates authority cleanly, verifies stale premises against both repos, records one live closure (G-140), opens and schedules 22 evidence-backed omissions plus the missing hub-census guard (G-213), and maps all 104 open residues exactly once into 36 defined packages. The A6 ledger/hardening tests are green. Session 10 execution remains unchanged.

07-29 · the embodiment build session
The founder ruled Asks 1–4, VL-0 landed, VL-1 landed, then VL-2R landed too — Vitto is alive on the console, now at the founder's design-gate pass 2

The founder read the four build-blocking asks from the embodiment spec and adopted each one's own recommended option, all on the record in the spec itself. ① The Canvas 2D avatar engine is ratified as our own one-way fork — no waivers, split to size instead — and Rive stays out (no such asset exists anywhere in gspace1 to adopt). ② The gspace1 cat IS Vitto's face going forward, and it now switches preset by theme so the dark background never swallows it — classic on light, mochi on dark. ③ VL-1 proceeds under founder-reviewed passes rather than waiting on a named designer; the three expression rows still under dispute are parked for the design gate, and the "nothing pulses on a timer" principle gained a carve-out — a face is allowed to breathe and blink for its own sake, but every change that means something must still land on a real event. ④ The layout is ratified: the avatar becomes the route's persistent stage (at least 320px), with the mood label, the first-person note, the provenance chip and the time as its caption, and the old Live Feed moves to a secondary column. Build slices VL-0 (vendor the engine as the fork) and VL-1 (the embodied surface and the scrubbed face) are now in flight. The three remaining asks — a voice provider, the embed's reach and weight, and the inner-life amendment — stay open, but each only gates a later slice (VL-4, VL-3 and VL-2A respectively), so none of them block the work happening now.

Then, same day: VL-0 landed, gate green. The build plan for the port itself went through its own adversarial round first — three reviewers, 36 findings, several blockers (the code formatter was quietly pushing the vendored files 20% past the size cap; a subtle bug could let animation state leak between renders) — all fixed in the plan before a line of code moved. With that settled, the gspace1 Canvas-2D avatar engine and its mood models were vendored into agente's own ui tree as a governed one-way fork: about 9,000 lines after reformatting, 44 files, every one under the 600-line cap (the biggest data table needed a five-way split to fit). All 12 of the engine's own behavior tests came across intact — 267 tests, the exact count preserved — plus 34 new tests pinning the fork's own deliberate deltas. Full gate green: 1,692 backend unit tests, 677 ui tests, 128 browser end-to-end tests, all four ratchets. A genuine bonus: porting the engine surfaced three real bugs in gspace1's ORIGINAL code (G-118) — two "random direction" rolls that were re-rolling every single frame instead of once per gesture, and an offset that never got subtracted back out, permanently nudging the character out of position once idle animations are off. All three are fixed in agente's fork with tests that fail if the fix is ever reverted; gspace1's live avatars still render all three today, so bringing the fix upstream is a coordination task, not a code task. Two smaller findings joined the ledger: G-119 (an 8-second "go idle" timer that turns out to be dead code, upstream too) and G-120 (a small blind spot in the import-boundary checker). Nothing here is committed — the founder reviews and commits — and VL-1, the embodied surface itself, was next.

And by the end of the day, VL-1 landed too. Vitto's face is now IN the product: the persistent stage of the Vitto Live page, every state rendered through it — quiet, loading, and both honest error states included — with the cat switching preset by theme exactly as ruled (classic on light, mochi on dark, contrast-proven with tests that fail on an illegible pairing). The scrubbed face shipped in this same first pass: drag the rewind slider on a conversation and the pane shows the face the diary supports at that instant, re-derived, never replayed. The build plan survived its own two-reviewer adversarial round first (33 findings, several blockers — including a missing one-line setting that would have quietly defeated the dark-theme legibility ruling), and the new browser tests then caught a real race that made the stage's very first paint vanish — fixed with a guard and regression tests before any human ever saw an empty box. Full gate green: 139 browser e2e, 716 ui unit, 1,692 backend tests, all four ratchets; measured on the real instance at zero dropped frames (about a tenth of a millisecond of work per frame) and ~5ms mood reads. The step now sits at the founder's design-gate pass 1: look at the face in both themes, and rule the three contested expression mappings the spec deliberately parked for this moment.

Pass 1 happened the same day and passed with notes — the three rows stayed unruled — and then, before the day was out, VL-2R landed: Vitto is alive on the console. He now visibly thinks, acts, waits and pauses while a turn runs, speaks while his reply streams in, follows your cursor with his eyes, and blinks when you poke him; a compact, always-visible Vitto docks beside the chat itself, listening the moment you focus the message box, nodding when you send, and yielding to the rewind face while you scrub history. A dead live stream now recovers honestly instead of hanging — the exact promised line, plus a Reconnect button proven to reopen the socket. Two structural moves rode along: the engine graduated into its own governed library (@vitto/avatar, one entry point enforced by the toolchain itself), and the whole experience is being lifted into a second, host-agnostic library with a formal sensor-plugin contract, so future senses (UI activity, a webcam, hand tracking, lip-sync) can plug in but can never fake a feeling the diary doesn't hold. The safety net caught a real bug again — a first-paint race, same pattern as VL-1 — and then, because the review lane itself was down for hours on API errors, the landed code got a full retrospective adversarial pass instead: 23 findings, 2 blockers (the pane was reachable from the embed by accident; a "poking Vitto never disturbs his timing" claim was measured false), all fixed, full suite green the same day — 455 ui, 299 engine-library, 1,692 backend, 148 browser end-to-end tests. The step now sits at design-gate pass 2: the three carried-over expression rows, plus the bigger one VL-2R adds — ASK 7, the inner life. Asks 5 and 6 (voice, embed) stay open.

07-28 · the embodiment session
Vitto Live's avatar got its spec — and the review round earned its keep again

The deliverable: a complete, adversarially-reviewed design for the colleague's face — the real avatar engine from gspace1, driven purely by the diary, with the rewindable face shipping in the very first pass. Twelve verification readers swept both codebases first; four independent reviewers then filed 63 findings against the draft, and every one is answered in the spec. The reviewers' best catches: the "zero-port" reuse story was measurably false (48 type + 73 lint errors under our stricter rules — so the engine becomes our own fork); the avatar engine quietly moves its own gaze on mouse movement, which would have broken the "everything you see comes from the record" promise (now disabled by design); the mood system itself can look pleased while a real degradation is still ongoing (G-111 — a finding about the shipped product, not the avatar); and very old rewinds can claim feelings the archive no longer holds events for (G-112). Five new gaps opened, all owned and wave-assigned; the board, ledger and hardening plan were updated in the same round. No production code was written — by design: this was the spec-and-review session the design-led gate demands.

Same day — the inner-life amendment: at the founder's ask ("make it feel almost human"), the spec gained an emotional memory: feelings that accumulate and decay on the diary's own clock (relief after a hard morning, residue that carries), idle spontaneity seeded by the record itself so even randomness replays, an instant nod when you send, and a receipt behind every feeling — hover the face and it shows you the exact diary events behind its look. A second adversarial panel (three reviewers, 42 findings, 7 blockers) attacked the amendment and won real corrections: the "relieved" art doesn't exist yet and must be designed; the claim that this softens gap G-111 was withdrawn — the words under the face still forget a live outage, and fixing that stays a separate founder decision. One design law came out of the fight and now governs everything: the face's claims step only on facts; only its warmth eases in between. Seven decisions now wait on the founder (Decisions tab).

07-28 · session 8 (apps architect)
How installable apps live inside the colleague — the architecture, with a playground you can click

The decision, in one line: an app is a permission slip with a face, not a program we bolt on. The code that does the work already lives in the platform, reviewed and compiled; an "app" is the named, priced, revocable agreement that switches some of it on for one Space — here is what it can see, what it can do, what it costs, and one honest sentence about what removing it does. Because installing, using, and removing an app are all entries in the same diary as everything else, you can always answer "who turned this on, who did what with it, what did it cost" — and rewind all of it.

The hard question the founder asked — letting the colleague create apps, via "dynamic API endpoint creation" — has a safe answer. Apps never invent new web endpoints at runtime (that would quietly break the permission model). Instead an app declares actions, and a human button and the colleague reach the identical action through one governed door — so the diary shows two entries for the same deed differing only in who did it. And when the colleague proposes a new app it proposes words, not code: a new grouping of abilities it already has, which a person must read and approve before it exists. That is why an AI extending its own workbench is safe here and reckless elsewhere — every change is a visible, attributable, reversible record.

What an app actually is, after a second pass: not just a bundle of buttons, but four declarations over code that is already in the binary — the verbs it may perform, the view it shows (a pure fold over the diary, so it re-derives at any past moment and needs no database of its own), the reflexes it runs unasked under a mandatory budget and rate cap, and the placement where it appears and to whom. Two consequences worth reading twice: an app's screen is derived, so rewinding the room rewinds what every app was showing; and an app's autonomy is earned per room — approve the same action a few times and it graduates to acting on its own, announced in the diary, revoked in one click, demoted by a single refusal, and un-earned if you rewind past the moment it was granted. Most platforms hand an assistant a fixed permission list. Here permission is a consequence of proven behaviour, and it is reversible.

You can try the model: a working simulation — switch who is acting, press an action or hand it to the colleague, let an hour pass and watch a reflex decide for itself, break the bridge and watch the app tell the truth instead of showing a blank card, spend past the cap and watch it pause rather than overspend, then drag the diary backwards and watch all of it un-happen — is published at the apps & spaces playground. The full spec (four adversarial reviewers ran and refuted parts of the first draft — the honest fixes are folded in) is 2026-07-28-apps-and-spaces-architecture-spec.md. It waits on a few of your rulings — chiefly: may the colleague author these "House Apps" (a new line on the freeze list), and what the placement words are. No code was written; two new gaps were logged (attribution, and per-Space install grain). A pressure test also found a live issue: any workspace member can currently read a whole Pulse label through chat, including other admins' details — now ledgered for the security wave.

07-27 · session 7
The security wave turned out to be switched off — and four live bugs surfaced

The headline: everything the security wave built has been inert by configuration. The RBAC ceiling ships in shadow mode by deploy default — it derives the correct decision for every request and then allows it anyway — and the hook its own written rollout plan depends on ("run in shadow, read the logs, then flip") was passed by nothing but a test, so weeks of shadow mode recorded ZERO evidence. That readout now exists and is wired to the Activity surface (G-78). The flip itself is deliberately NOT done: two adversarial reviewers proved it would lock every non-owner out of chat and HITL, kill the Pulse embed lane, and — the inversion that decided it — leave automation unrestricted while restricting the humans who link their identity. A live query settled the size: 3 accounts, all owners, so the blast radius today is zero and the next invited member is the casualty.

Four defects found that predate this session. GitHub CI had been RED since the previous session's own "CI green" claim (a piped exit code hid it). The Pulse embed crashed on its first approval or question — a whole surface with zero browser coverage. Archiving a conversation could stamp workspace_id: "unknown" into the APPEND-ONLY log, permanently unmatched by any workspace-scoped fold and uncorrectable by construction. And the chat's SSE stream retried a dead endpoint every 1.5s forever while displaying "connecting". Three of those were found by guards written for a different problem — the argument for writing the check rather than fixing the instance.

New machinery: a pack-boundary CI rule (core may never import a pack — it found an undocumented violation on its first run), and a ledger-integrity guard, after two agents working this branch in parallel independently allocated the same six gap ids. Also measured and recorded: grep on this machine silently reports "no match" for content it cannot read, so three core files were invisible to every past grep-based audit. Open gaps rose 37→56; that is discovery, not regression. Checkpoint 1468455.

07-27 · session 6 (late)
Every gate cleared — and the security wave landed the same day

The founder cleared all outstanding decisions at once and re-owned Yam's permission queue to the build agents "acting exactly as Yam would". Two waves shipped in one round: the chat preview escape (G-45) and the authz cluster (G-43 reset governance, G-37 cross-workspace reads, G-29 read lane, RNF3 census, plus a member-forgeable "trusted" label nobody had noticed). Four adversarial reviewers and two judges ran BEFORE any code — the judges settled the browser questions by running real Chromium probes and refuted three of the orchestrator's own confident claims, including the exploit story for the very hole being fixed. The most valuable catch: the cross-workspace fix as first designed would have LOOKED sealed while the rewind lane still served other workspaces' data. Everything verified twice — the full suite locally, then live against the founder's own running instance. Checkpoint 2d70fc2, GitHub CI green on both branches.

07-27 · session 6
Alignment turn — the board went tabbed, the push happened, CI got wired to it

The founder asked for a full plain-words plan and a real explanation of every gate; this board gained tabs (the Plan tab explains each queued item's what/why for non-technical readers; the Decisions tab turns every gate into a card with examples and honest recommendations). Then the first gate cleared for real: the founder pushed both program branches to github.com/YamCrack/agente. The push exposed a latent gap — the CI recipe's trigger list still named only the step-0.1-era branches, so the first push ran zero checks; fixed the same hour and the first full GitHub runs (ci + e2e + evals) went out on both branches. The briefing artifact was re-checked and remains org-locked (public-reader); progress-vs-plan comparison continues against the repo master plan. Next: the required-checks click once green, then the feature wave, starting with the Vitto-Live chat pane spec + its adversarial review.

07-27 · session 5 (late)
The functional wave — five gaps closed, features first

Channel redeliveries can't double-bill anymore (designed so no historic event changed meaning). Goals became diary events — the one piece of the colleague's life that lived outside the log came home, closing a real architecture violation. Chat history pages backward. Groups are manageable end to end. And tests now compile in CI on both trees — 293 stale fixtures burned, including tests that were writing schema-invalid events. The founder's product questions were analyzed and answered on this board (Analysis tab); the Vitto-Live chat pane was queued next. Checkpoint 285949f.

07-27 · session 5
Hardening wave landed + this board's human rewrite

The poison-record class is now unmintable at every boundary — the sweep found the worst lane wasn't even the reported one: a single crafted URL to the trust-reset endpoint could have frozen a tenant's autonomy system; it now refuses cleanly, with tests proving the old bug existed. Dark mode chips (165 fixes) and light-mode gray text (151 fixes) both meet accessibility standards, enforced by token-level tests and grep guards so they can't regress. Oversized rewinds explain themselves. Also: audited the full master plan against the board — nothing unowned — and rewrote this page in plain language with test steps and business scenarios. Checkpoint f3a827a.

07-27 · session 4 (late)
G-42 + G-44 + 1.T — consent fails closed · rewind hardening · real dark mode

Consent poison now freezes autonomy shut (the reviewer refuted the first design and produced a better one). Rewind got O(1) mood lookups, an honest size cap, and byte-exact ordering between past and present. The theme layer shipped with zero light-mode change and an enforced contrast matrix. Checkpoint d69a02b.

07-27 · session 4
1.7 living artifacts + the robustness sweep

The birth-certificate chain landed end-to-end (verification caught a provenance-attribution bug pre-merge). All 43 silent-failure paths burned to zero; the 2-billion-event ID ceiling removed permanently. Checkpoints 43fb66f, 7299a94.

07-27 · session 3
1.6 the flagship scrubber + earned autonomy

The rewind demo runs for real: paint → five human approvals → announced graduation → staleness fails closed → the wall re-derives at any moment. The first review rejected a design that would have graduated dangerous commands from harmless ones.

07-27 · session 2
Wiring audit + memory citations + the quality exam

15 backend↔frontend gaps closed (including a broken multi-tenant login). Ask-memory citations shipped, and the first 18-case graded exam became a mandatory CI job.

07-26 · session 1
Stage 0 + the wall + the dossier

Safety rails, grenade fixes, replay integrity, the pack engine — then the first two visible surfaces.

Operating rules: adversarial review before build · independent verification after hot-path builds · founder-owned commits and pushes · board + ledger + hardening plan + focused specs + this mirror synced in the same round · every open gap gets exactly one wave and planning package. Architecture note: vitto is an Nx + Yarn Classic monorepo (backend APIs/agente-api/src/domain/* + React app in apps/agente-ui/) — every capability enters as events + pure folds + packs or governed bridges, and every surface has a routed consumer with reachability evidence. Fresh execution read order: 07-30 master spec → hardening plan → gap-spec hub → workflow → board → ledger → master-plan addenda → CLAUDE.md.