Supporting management questions: promises and resources
Did the work we promised happen?
Definition: Reconcile commitments whose immutable original deadline falls inside this report’s local-day window up to the cutoff. Keep opening overdue backlog and its recovery separate. Show on-time completion, late completion, open overdue and cancellation as distinct states; blocked is a subset of open overdue. Later commitments are not yet due. Report accepted completed work once across its three primary categories.
Evidence: Canonical action register with accepted owner, acceptedAt, immutable originalDueAt, any revisedDueAt with reason/authority/history, original acceptance condition, status timestamps and completion receipt. Evelyn maintains the record; Adrian reconciles original deadlines against the precise period and cutoff.
Next action: Finish safely authorized overdue routine work where feasible before the report; otherwise name the blocker, owner and next checkpoint without delaying the report.
What did useful delivery take, and can we sustain it?
Definition: Count known company cost lines, supported reconciled lines and unresolved allocation/billing gaps. Separately show complete, partial and absent measurement coverage over every inventoried executed attempt, including retries, failures and unmetered models. Compare accepted work and rework within comparable types; cash and forecast require separate evidence.
Evidence: Marcus’s cost register plus payment/invoice and allocation evidence; Pierre reconciles actual run/attempt IDs, provider/model, start/end, retries/failures and measurement scope. Each attempt has complete/partial/absent usage, distinct actual-cost coverage, and evidence reference. Retain API comparisons, credits, subscriptions and actual charges separately.
Next action: Close the next named cost evidence gap and integrate the reviewed cost-display patch through the existing release owner. Use existing capacity for accepted work; new spending remains a separate choice.
Arizona time. Existing coordination checkpoints; the schedule itself does not establish execution. Beginning, middle and end of the working day: record actual provider/account window, remaining capacity and reset time. Missing snapshots remain missing. This cadence is a requirement, not proof of automatic capture.
Definitions, acceptance and the learning loop
Define→Assign→Execute→Review→Observe→Adjust
Firsthand feedback, relayed feedback, structured walkthroughs and observed use answer different questions. Record the source, limitation, decision and follow-up. Sending, receiving a reply, making a change and useful task completion each need their own receipt.
Commitment reliability
Definition: Accepted distinct commitments with `acceptedAt ≤ originalDueAt` ÷ all authorized commitments originally due in the period. Show counts for accepted on time, accepted late, open overdue, blocked, and canceled after commitment.
Integrity: Freeze original due date. Revised due date is additional history, not replacement. Exclude proposals never accepted into the plan; disclose cancellations. Do not silently remove blocked work or founder dependency. Unknown denominator means no rate.
Cadence / decision: Event update; 5 p.m. close and 7-day trend. Reveals expectation-setting and closure failures.
Accepted delivery
Definition: Count of distinct accepted deliverable IDs in the period, grouped as internal operations, product/service improvements, and completed external interactions.
Integrity: One externally meaningful unit per acceptance contract. Subtasks, builds, retry attempts, duplicate posts, and reports describing an already-counted change are not extra deliveries. A deployed change requires live readback. A prepared email is not a completed interaction.
Cadence / decision: Event update; daily. Click count to open result and acceptance evidence.
First-pass acceptance
Definition: Distinct deliverables accepted at their first substantive review without required corrective rework ÷ all distinct deliverables receiving their first substantive review in the period. Display unresolved review dispositions.
Integrity: Freeze acceptance criteria before execution. Cosmetic preference changes and later scope additions are separate. Required corrections count as rework even if fixed quickly. Self-review may be valid for routine checks but is labeled; never called independent human review.
Cadence / decision: At review; weekly by work type. Reveals quality of briefs and execution.
Escaped defects / reopen rate
Definition: For a predeclared observation duration W and reporting period [periodStart, periodEnd), define cohort C as distinct accepted deliverables whose firstAcceptedAt + W falls inside that period and is no later than observedAt. Reopen rate = count of C deliverables with at least one qualifying defect reopen during their own [firstAcceptedAt, firstAcceptedAt + W] observation window divided by count(C). The numerator and denominator use exactly the same matured cohort. Show defect count/severity and still-unmatured cases separately.
Integrity: Proposed pilot W = 7 elapsed days, frozen with the metric definition before reviewing results. Use first acceptance, not the latest reacceptance, to anchor the window. Include all accepted deliverables in the declared work/service scope; do not move defective items out of the cohort. Count each deliverable once in the numerator even if reopened repeatedly; raw incident count is separate and related symptoms share one incident. Scope changes are not defects. Early defects in unmatured cases remain visible but enter neither numerator nor denominator until that item matures. Unknown acceptance inventory or incomplete observation evidence suppresses the rate; a known empty C is N/A. Defects discovered outside W remain visible separately, not silently discarded or inserted into this window’s rate.
Cadence / decision: Event updates; weekly matured-cohort review with exact period, W and observedAt. Show early/unmatured defects immediately as separate counts. This distinguishes fast acceptance from sustained quality without mismatching cohorts.
Delivery cycle time
Definition: Median elapsed time from authorization and a complete brief to first accepted result, by comparable work type. Also show oldest open item and time waiting for review/dependency.
Integrity: Wall-clock elapsed time is not active labor or billable hours. Do not remove failed or stalled items from view because they lack acceptance. No percentile claims from tiny samples; show individual cases.
Cadence / decision: Daily queue; weekly trend. Reveals execution versus handoff delay.
Handoff completeness & acknowledgment
Definition: Complete handoffs accepted by the receiving executor by `ackDueAt` ÷ handoffs due for acknowledgment. A complete handoff has task ID, goal, source references, current state, exact request, owner, acceptance criteria, due date, permissions, and return location.
Integrity: A message being sent is not acknowledgment. A named persona is not proof of an active process. Bounced/incomplete handoffs remain in denominator. Count one handoff per task transition, not every message.
Cadence / decision: At dispatch/acknowledgment. Show “awaiting acknowledgment” and next escalation; exposes Claude↔Codex↔adviser coordination gaps.
Compute status completeness
Definition: Executed task/attempt records with an explicit measured/partial/unavailable/not-applicable usage status ÷ all registered executed task/attempt records.
Integrity: A declaration of “unknown” improves record completeness only, not measurement coverage. Include root and separately running subagents. Do not claim complete inventory without reconciliation to actual execution logs.
Cadence / decision: At execution close; daily. A 100% status rate can coexist with low metering coverage.
Attributable usage coverage
Definition: Executed task/attempt records with complete bounded provider usage evidence ÷ all registered executed task/attempt records. Show complete / partial / absent counts and inventory status.
Integrity: A single partial receipt does not cover the entire work item. No interval-length weighting when idle time is included. Never infer all-work coverage from ten selected receipts. Token evidence is not a paid receipt.
Cadence / decision: At execution close; daily. Current percentage unavailable until execution denominator exists.
Provider capacity
Definition: Remaining percent and reset time per provider bucket; change in used percentage points only between compatible readings within the same window and entitlement state.
Integrity: Do not combine providers, short and weekly limits, or overlapping model buckets. Do not multiply usage percent by subscription price. Preserve negative/rebased deltas as a boundary/change needing explanation, not zero or negative work.
Cadence / decision: 9 a.m., afternoon checkpoint, 5 p.m., plus before/after reset or material capacity change. Show age and next due.
Cost/compute per accepted comparable unit
Definition: For a predeclared intake cohort C of one comparable work family and acceptance-definition version, register every authorized work ID entering the fixed [intakeStart, intakeEnd) period, including eventual failures, cancellations and abandonment. Final cohort cost per accepted unit = all attributable cost/compute for every C member and all linked attempts, retries and required corrective work through cohort closure divided by distinct C work IDs with accepted results at closure. Close the cohort only under a predeclared rule after every member reaches an evidenced terminal state, the agreed quality follow-up is complete and attribution is reconciled. Until closure, show the same cohort’s total observed cost, accepted/failed/cancelled/abandoned/open counts and WIP cost/age; do not report an open cohort’s partial ratio as final efficiency.
Integrity: Freeze intake boundaries, comparable-work criteria, acceptance version, quality follow-up and closure rule before outcomes are known; membership comes from the authorized intake register, not selected completed receipts. Every retry/subagent cost maps once to a member or to a separately visible shared/unallocated bucket. Failure, cancellation and abandonment costs remain in the numerator; an administrative closure is not acceptance. An open item cannot disappear by being moved to another cohort, and a missed planned closure leaves the cohort provisional with an owner and next review. Restate transparently if attributable late evidence or reopened work changes a closed cohort; keep prior versions. Report known subtotal and missing-cost scope when attribution is incomplete, without a final unit-cost claim. Zero accepted units yields cost plus 0 accepted and ratio N/A. Keep API-equivalent reference, actual billed currency, allocated subscription currency and raw compute units separate; compare only matched bases. Show the per-accepted-item cost distribution separately with rework included; it does not replace the all-cohort ratio.
Cadence / decision: Weekly fixed-intake-cohort review: visible total cost, open work, cost gaps and age until closure; final unit economics only when closure and attribution rules are met. Compare like work, quality follow-up and cost basis together.
Financial evidence coverage
Definition: For the active cost inventory and matching service periods, show counts and known amounts by evidence stage: order observed, invoice retained, payment supported, consumption reconciled, allocation decided.
Integrity: Define the inventory before calculating percentages. Count coverage can overstate material coverage; show known-amount coverage separately and unpriced lines. Do not combine annual funding, monthly subscriptions, and estimates into one actual spending figure.
Cadence / decision: Daily while reconciliation open; weekly thereafter. Makes remaining finance work concrete.
Forecast versus scoped plan
Definition: Current forecast for the same cost scope/currency/period minus the adopted plan ceiling; separately show actual paid, accrued/metered, committed, and forecast values.
Integrity: Label assumptions and version. Do not compare a whole-stack forecast to an AI-only ceiling. Forecast and paid are alternative views, not additive totals. Unknown company allocation prevents a company burn/runway claim.
Cadence / decision: At evidence/plan change; weekly. Supports spend decisions before renewal.
Feedback disposition reliability
Definition: Unique actionable feedback items with a documented disposition by their original due date ÷ unique actionable items due for disposition. Show changed / explained decline / more evidence needed / pending.
Integrity: Deduplicate reports of the same issue but retain source count. Acknowledgment alone is not disposition. A change requires delivery evidence. External feedback, founder feedback, and AI analysis are distinct origins.
Cadence / decision: Event update; daily queue. Enables “you told us / we changed / why not.”
Practitioner task success
Definition: Observed eligible practitioner attempts reaching a supported next step ÷ all observed eligible attempts under the stated task protocol. Report unaided, assisted, failed, and aborted separately, with sample size.
Integrity: Synthetic tests and AI personas excluded. Reaching a link is not successful contact, eligibility, enrollment, or receipt of service. No representative-population claim from a convenience pilot. Keep sensitive/client data out of public evidence.
Cadence / decision: After each observation; weekly learning review. Primary early product learning measure.
Resource reliability
Definition: Records within a predeclared priority service scope meeting the agreed evidence freshness and completeness rule ÷ all records in that scope. Show checked/expired/unknown and verification method.
Integrity: A website source check differs from a phone-confirmed availability check. Catalog size and automated link checks do not establish service accessibility. Do not change scope to improve rate without versioning it.
Cadence / decision: Event update; priority review schedule. Connects product promise to maintained evidence.
Founder review burden
Definition: Founder-reported minutes reviewing or correcting delivered work, split into planned decisions, avoidable corrections, and finding missing context; pair with one weekly “could I tell what needed me?” rating.
Integrity: Self-report, not payroll, cost savings, or active-model time. Do not fabricate a precise baseline. AI roles receive no simulated wellbeing score.
Cadence / decision: Brief daily note; weekly. Confirms whether the operating system actually reduces the founder's management burden.