AI Pilot Ledger

A governance layer for enterprise AI investment

Enterprises aren't short on AI pilots — they're short on proof. AI Pilot Ledger is a lightweight, accountability-first dashboard concept built to close that gap: one screen where every AI initiative carries a defined metric, a named owner, and a visible trajectory, from kickoff to kill-or-scale decision.

This case study walks through the research that shaped it, four design iterations built through applied AI collaboration, and the reasoning behind each change — from a working prototype to a decision-ready tool.

Overview

Build Specs (How I Built It)

Stack — React (JSX), component-based architecture

State Management — React Hooks (useState, useEffect, useCallback)

Data Layer — Persistent key-value storage, scoped per session

Design System — Custom dark-mode UI tokens; Archivo (display), Inter (body), JetBrains Mono (data)

Logic Layer — Rules-based status engine — pilot status is computed from progress-to-target ratio and time-to-deadline, never manually set

Data Visualization — Custom SVG line chart, rendered from milestone-logged time-series data

AI Collaboration — Iteratively built with Claude Sonnet 5 & Opus 4.8 (Anthropic) across four design cycles, each shaped by targeted, stakeholder-level critique. Iteration MethodRapid prototyping with incremental, scoped revisions — Version History — v1–v4, each preserved as a discrete artifact to document design reasoning over time

Enterprise AI investment has outpaced enterprise AI proof. Industry research paints a consistent picture of where that gap comes from:

Context

Fatigue compounds the problem. After a few stalled initiatives, leadership attention quietly withdraws — and the next legitimate pilot inherits skepticism it didn't earn.

Most pilots never reach production. A large majority of enterprise AI projects stall before they scale — not because the models underperform, but because the surrounding organization isn't ready to absorb them.

ROI is assumed, not measured. Projects are frequently greenlit on projected returns that are never revisited after launch — success criteria are defined loosely, if at all, before a pilot begins.

Ownership is diffuse. When a pilot stalls, there's often no single person accountable for the outcome — just a rotating cast of stakeholders watching different metrics.

The common thread isn't a technology failure. It's a tracking failure — enterprises are testing agents and automation faster than they're building the discipline to evaluate them.

Opportunity

The research pointed to a clear gap: most organizations had AI dashboards for usage, but nothing that enforced accountability before a pilot ever launched.

I saw an opportunity to design something that was:

  • All-inclusive — one ledger for every AI pilot across departments, instead of scattered spreadsheets and status decks.

  • Accountable — every entry tied to a named, responsible owner — not a team, not a vendor, a person.

  • Informative at a glance — priority, progress, cost, and status visible in a single view, built for how a decision-maker actually scans a room full of competing initiatives.

That opportunity became AI Pilot Ledger — designed and iterated through four working versions, each responding to the kind of scrutiny a real VP would apply before trusting a tool with their AI agent pilot portfolio.

Showcase

AI Pilot Ledger is a fully interactive prototype: pilots are logged with a required baseline, target, deadline, and named accountable owner. Status is never manually set — it's computed automatically from progress pace and time remaining. Each pilot opens into a detail view with a live progress-over-time chart, milestone history, and a "last updated" freshness indicator that surfaces silent, neglected pilots before they quietly become budget write-offs.

Built as a self-contained React application with persistent storage, sortable by priority and urgency, and structured around one operating principle: no pilot enters the ledger without a number, a name, and a date attached to it.

AI Pilot Ledger

Design & Build

Version 1 - Here’s where it started

The first version proved the core mechanic: a pilot ledger where nothing gets tracked without a baseline, a target, and a deadline.

Status — on track, at risk, stalled, hit — was computed automatically rather than self-reported, removing the temptation to quietly mark a struggling pilot "green."

It was intentionally minimal: the goal was to validate that forcing the metric was enough to change how a pilot got tracked.

Version 2 — Seeing the opportunity for better output

My critique on v1 was that once the mechanic worked, the framing didn't for the progress metric: Mixing raw units — hours, dollars, percentages — in a single progress column made pilots impossible to compare side by side.

I also thought to myself, “aren’t companies struggling with seeing ROI in AI investments right now?” But that would be difficult to determine, so how can I capture the "ROI intent" without it as a headline metric? After reflection, I decided to place ultra descriptive captions underneath the Agent project header to specifically capture value intent. (Most pilots currently don't have clean, attributable return figures this early.)

v2 normalized progress into a single percentage, split cost into its own column, and replaced ROI with a plain-language objective — the actual business problem each pilot was meant to solve.

Version 3 — Standing in the decision-maker's shoes

Looking at v2 the way a VP would, two things didn't hold up. Once I would double-click into an agent-build project, the smaller pop-up’s bare "log value" input had no value explanation — why would a leader be manually saving a number with no context? And a flat list gave no sense of what mattered most among five competing pilots—I decided to add prioritization markers.

Ultimately, v3 reframed data entry as milestone logging — a value paired with a one-line "what happened" — and fed that history directly into a live progress-over-time chart. A priority tier was added and the list was reordered so the pilots demanding attention surfaced first, automatically.

Version 4 — Where accountability closed the loop

My critique on v3 was the remaining gap of ownership—who would be accountable? A dashboard can show a stalled pilot in red without ever answering the next question: who do I call?

v4 added a required, named accountable owner to every pilot — no longer a team label, but a person and their role — alongside a "last updated" freshness signal that flags pilots gone quiet even when the deadline math still looks fine. Visual refinements followed: tightened priority badges, corrected spacing, and header copy rewritten to read as a decision tool rather than a case-study demo.

Takeaway

Each iteration answered a question a real decision-maker would actually ask: What are we measuring? Can I compare these fairly? What's actually going on? Who's responsible? That's the design discipline this project demonstrates — not just building an AI-powered tool, but interrogating it the way its end user would, until the product held up under that scrutiny.

For enterprises evaluating AI initiatives, the implication is direct: the tooling gap isn't more dashboards — it's dashboards that refuse to let a pilot exist without a number, a name, and a deadline attached to it. That single constraint is what turns AI experimentation into something a leadership team can actually govern.

For hiring managers, this project is a demonstration of applied AI product thinking: identifying a real, research-backed business problem, prototyping fast, and iterating based on the specific, practical objections a stakeholder would raise — not just polishing pixels, but sharpening judgment.

Next
Next

Executive Briefing Center