Questions inbox
Work pauses durably while a question is open and resumes when you answer.
Composer Initiative Tools: List/Count and Create from Brief
agent:openai · Composer Initiative Tools: List/Count and Create from Brief · 8/19/2026, 6:55:30 PM
Every work item is complete, but the brief assessment could not verify any of its 2 criteria — it was not shown the full content of the delivered work. Is this initiative done?
The core runtime tools and server-side behavior appear to be in place: list_initiatives is registered and returns an authoritative list and count under server-authoritative scope (#517); create_initiative_from_brief is a confirm-only, canonical mutation with idempotency, structured failures, warnings, and a link to the new initiative (#518), and the client surfaces provenance, grounded counts (including zero), links, and warnings (workspaces #210). Server authorization is preserved and never taken from model input, and mutations are idempotent and confirmed, not inferred. Beyond Wave‑1, the Wave‑2 in‑initiative reads (planning status, work items, waves, questions, deferred scope, brief revisions) and Wave‑3 answer mutations (continuation decisions and initiative questions) are also registered and integrated (#521–#526, #533), with confirm/verbatim UX wired in the client (#222). What is not shown end‑to‑end is the Composer’s natural‑language triggering of list/count: the brief’s user‑visible behavior requires that asking “How many initiatives are there?” or “List the initiatives.” causes an invocation of list_initiatives and returns the exact current count based on live data. The repository evidence proves the tool exists, is registered, and its results are rendered, but it does not demonstrate the model/provider policy reliably invoking the tool from those NL prompts. Given the bounded evidence and omitted test contents, this acceptance remains partial pending an explicit prompt→tool invocation proof. Everything else in scope is covered by delivered capabilities and client wiring. Every criterion came back "cannot confirm" rather than "not done", and the delivered work exceeded the evidence budget, so the diff content of 1 completed item was withheld from the assessment (their changed-file lists were shown, but not their content) — so this is an unverified verdict, not an adverse one. Check the delivered work yourself before waiving anything: waiving records each criterion as a gap you accepted, and if the work is in fact complete that is a false record. Criteria the assessment could not confirm: - NL-to-tool behavior for count/list prompts (Composer invokes list_initiatives on questions like “How many…?”) — The brief requires that asking the Composer causes it to invoke list_initiatives and return the exact count. The repo shows the tool is registered and UI renders groundedCount, but provided evidence does not conclusively show prompt→tool invocation policy/tests (content for #517 tests was omitted). - Empty initiative collection returns a count of zero (authoritatively) — Workspaces #210 renders groundedCount: 0 distinctly; list_initiatives is designed to carry total counts. End-to-end prompt→tool tests for the empty case were not visible in the bounded evidence.