← All work

Case study 04 · Gabi OS

Gabi OS

A personal AI operating system for reliable work across assistants.

AI assistants are individually powerful, but real work breaks down between sessions: context fragments, status goes stale, the human becomes the transport layer, and “done” becomes a claim instead of an evidenced state. Gabi OS is an internal product I built and continue to evolve to make that workflow explicit, bounded and verifiable.

AI systems · Product engineering · Internal product

The problem

Where AI-assisted work breaks down

Different assistants do not naturally share dependable state. Each new session starts by rediscovering what the last one knew, so long prompts and repeated explanations use up context and attention before any useful work begins.

When information has to move between tools, the person in the middle becomes the orchestration layer: copying instructions in, pasting results out, and remembering what actually happened.

Agents also need edges. Without explicit scope, permissions and stop conditions, a small task can turn into open-ended discovery or unintended change. And a model saying a task is complete is not evidence that it is. Missing or conflicting evidence has to stay visible rather than quietly becoming “done”.

The system

How a piece of work moves through it

Gabi OS architecture, from intent to canonical state. The same steps are described in the text below.
  1. Human and chat control planeDecides what matters and approves consequential actions.
  2. Context compilerAssembles the smallest sufficient context from canonical state.
  3. Bounded Task EnvelopeFixes scope, exclusions, acceptance criteria, required evidence, limits and permissions.
  4. Approved workerCarries out only the bounded task, on an approved host.
  5. Execution receiptReturns structured evidence of what was actually done.
  6. ReconciliationChecks the receipt against canonical state and the acceptance criteria.
  7. Canonical stateRecords the accepted outcome, or keeps the gap visible. The next context package starts here.
  • Canonical state feeds Control TowerThe human-readable view of current work and its integrity.
  • Shared authenticated context (MCP) is shared with AssistantsDifferent assistants retrieve the same durable context instead of relying on copy and paste.

Architecture

The parts that make it dependable

One canonical state

Gabi OS keeps one canonical operational state instead of a separate tracker for each assistant. Workstreams, tasks and execution runs live in Supabase with source references and explicit lifecycle states, so any assistant can recover active work from the same record rather than from conversation history.

  • DONE
  • ACTIVE
  • PARTIAL
  • WAITING
  • BLOCKED
  • UNKNOWN
  • RECONCILIATION_REQUIRED

Missing evidence is not silently converted into DONE.

Bounded Task Envelopes

Before anything executes, work is compiled into a bounded Task Envelope. It carries the relevant canonical state, the exact scope and exclusions, acceptance criteria, the evidence required, budgets and limits where relevant, stop conditions and permissions. Its job is to stop a small task from turning into uncontrolled discovery or mutation.

Context compilation

Instead of handing every model the entire history, Gabi OS assembles the smallest context package that is sufficient for the task. Large logs and raw output stay behind evidence pointers rather than being injected into every prompt.

Execution receipts and reconciliation

Workers return structured receipts, not prose. Receipts are reconciled deterministically against canonical state, and completion depends on accepted evidence and read-back. A successful run does not by itself complete a whole workstream, and an ambiguous result returns RECONCILIATION_REQUIRED instead of a guess.

Human and permission boundaries

Some actions deliberately stop and ask. The execution loop halts before consequential permissions rather than automating through them.

Automation should remove unnecessary human relay, not remove necessary human judgement.

Shared context through MCP

An authenticated Model Context Protocol endpoint gives different assistants access to the same durable context. Requests are authenticated, authorised and audited, so information no longer has to be carried between tools by hand.

Control Tower

The Control Tower is the human-readable operational surface for Now, Plan, Sessions and execution, Tasks, History, and Integrity and reconciliation, projected from canonical state. It is part of the v1 operating architecture and is still evolving.

What I built

  • AI-system architecture and the orchestration model
  • Product and frontend engineering in TypeScript and Next.js
  • Supabase-backed canonical state and lifecycle model
  • Authenticated MCP shared-context integration
  • Deterministic workflow, governance and reconciliation logic
  • Testing and acceptance gates

Verified evidence

Accepted, with a clear boundary

The v1 operating architecture reached accepted canonical completion in October 2026, after passing its acceptance gates.

AI Operating System v1 acceptance

tests passed
1,115
tests skipped
3
TypeScript
Pass
scoped lint
Pass
production build
Pass

Gabi OS is an internal system under active development. The core AI Operating System v1 architecture and implementation reached accepted verification, while source-integrity, scaling and continuing operational work remain active. This evidence describes verified architecture and implementation; it is not a claim that every component is deployed or live.

What it demonstrates

AI is most useful when the system around it knows what is true, what is allowed, and what still needs proof.

This is not a collection of prompts or an AI demo. It is AI designed as part of a software system: deterministic code handles deterministic work, AI is used where semantic reasoning earns its place, state is explicit, permissions have boundaries, outputs require verification, provenance matters, and people stay in the loop where judgement is required.

Practical AI systems

Building an AI workflow that needs to hold up outside the demo?

Let’s talk about internal workflows, agents, knowledge systems or human-review architecture, designed around real constraints. No magic, and no promise of full autonomy.

Practical AI systems · Frontend / product engineering

Discuss an AI workflow