Initiative Brief — Make the Composer agent real: tool-calling over platform data
COMPLETEDIntent— the specification (6,186 chars); the plan below decomposes it into work items. Click to expand.
# Initiative Brief — Make the Composer agent real: tool-calling over platform data ## Desired outcome The workspace Composer answers real questions about the operator's own platform data — grounded in what the system actually holds, showing what it consulted, and safe by construction — instead of echoing the message back. Asking "why does this work item have RFC in the title?" returns an answer read from that initiative's work item, not a guess and not an echo. ## Motivation The Composer is the platform's primary conversational surface and the one users reach for first. Today every message returns `Acknowledged: <your text>`, because the runtime is running its deterministic offline provider. That is not an unfinished feature so much as an unfinished last mile. The foundation shipped: a server-side runtime with a bounded context window, a swappable provider adapter, a capability registry with per-action authorization, a confirmation flow for mutations with durable audit receipts, and redacted telemetry. Two things are missing, and neither is small: 1. The only provider that can emit tool calls is the **mock**. The real provider (`agent_text`) is text-only by design, so even switched on it cannot look anything up. 2. The capability registry ships **one stubbed read** and **one mutate placeholder**. There is nothing real for a provider to call. The result is a surface that looks finished and answers nothing. Closing this converts a demo into the platform's actual front door. ## Current state - The `conversation_runtime` feature flag is **enabled** in the production database. - `CONVERSATION_RUNTIME_PROVIDER=agent_text` is now staged in OutShine's prod `.env` but **not yet active** — it takes effect on the next restart/deploy. - No `agent_configs` entry keyed `conversation_runtime` exists in Outsold. Until it does, a send under `agent_text` raises `RuntimeProviderError` and **fails outright** — strictly worse than the mock's echo. The Outsold bridge URL and service token are both configured. - `AgentTextRuntimeProvider` is explicitly text-only: "it does not emit tool calls (the mock covers tool orchestration)". - `build_default_registry()` returns `workspace.read_summary` (a stubbed read) and `workspace.request_action` (a mutate placeholder that is never executed). - Bounds already enforced: 20 messages / 12,000 characters of context per turn, and a hard ceiling of 3 provider↔tool round trips per message. - The client/server conversation contract mismatches were fixed server-side (outshine #456, #459); send and history now work end to end. - The confirmation and receipt paths exist in code but have never been exercised end to end from the client. ## Product principles - Answer from the platform's own data, never from model memory. An ungrounded answer about an operator's initiatives is worse than no answer. - When the assistant cannot answer, it says so plainly and names what it would need. - Every mutation is explicitly confirmed by the user and leaves a durable receipt visible in the conversation. - The user can see what the assistant consulted to produce an answer. - A misconfigured backend produces an honest, actionable message — never a silent echo and never an opaque failure. ## Architectural principles - Provider credentials stay server-side and resolve through the Outsold `agent_configs` registry. No provider key belongs in a workspace app or in OutShine's own `.env`. - The server adapts to the client contract, not the reverse (the direction established in #394/#416/#417/#428 and applied again in #456/#459). - Context is always bounded. No unbounded transcript ever reaches a provider. - Capabilities authorize per-action against the caller's real permissions. Being in a conversation grants nothing. - The provider stays swappable behind the adapter; the runtime never depends on one vendor's tool-call shape. ## Constraints - Single production box with 7.6 GiB RAM shared by every service. Builds and deploys serialize; a careless build has frozen this box before. - CI must stay fully offline. The mock remains the default so no test ever reaches a provider. - Per-turn cost and latency must be bounded and observable before this is enabled for real traffic. - Message content must never enter logs; telemetry stays counts and low-cardinality labels. - The workspace Composer's existing UX contract (send, history, receipts) must not regress. ## Invariants - The default provider is the offline mock. - No provider credential reaches the client or the logs. - Every tool call is authorized against the caller's actual permissions at call time. - Every mutation produces a durable, user-visible receipt. - Context handed to a provider is always bounded by both message count and character count. ## Expected behavior - Asking about an initiative, work item, or release returns an answer read from that record, including its actual text. - The answer shows what the assistant consulted, so the user can verify it. - Asking the assistant to change something produces an explicit confirmation prompt, and after confirming, a receipt that survives a reload. - While the assistant is doing multi-step work, the Composer shows that it is working rather than appearing hung. - If the provider is misconfigured or unavailable, the Composer says so plainly and the message is preserved. ## Open questions - Which provider and model for the `conversation_runtime` agent config — and what is the acceptable per-turn cost? - Which read capabilities land first? Proposed order: initiatives and work items (the reported use case), then workspace directory, then releases and worker health. - Should Orchestrator reads go through the canonical `/orchestrator/*` HTTP surface or a service-to-service call with the existing workspaces↔orchestrator service JWT? - Who owns the cost ceiling and monitors it once real traffic starts? - Does the confirmation flow need a UX pass before mutating capabilities are exposed, given its client paths have never run end to end? - Should tool-call support be added to `AgentTextRuntimeProvider` or land as a third provider, leaving the text-only bridge untouched for other callers?
Repository scope
targets the work spans — independent of where the brief was authoredPrimary repository: outshine
- outshineactive · primary
- orchestratoractive
- workspacesactive
- hubactive
brief revision 4 assessed as delivered
assessed with 82% confidenceEnd-to-end, the Composer can now answer questions about initiatives, work items, releases, and worker/queue health by invoking read capabilities against the Orchestrator’s canonical HTTP surface and returning the records’ actual text with citations. The UI renders consulted sources, shows a working indicator during multi‑step turns, preserves provenance across reloads, and surfaces explicit provider errors while preserving the user’s message. Provider selection is resolved from Outsold’s agent_configs (one place), with the executing service holding credentials; OpenAI and Claude tool-calling transports exist behind a vendor-agnostic seam. Defaults keep CI fully offline; telemetry is bounded and content-free. Mutations/receipts are explicitly deferred for this wave, so their absence is not unmet scope.
Plan waves — 2 approved waves
planning closed- Wave 1foundation-firstAPPROVED6 work items
Make the workspace Composer actually answer questions from platform data by enabling tool-calling in the server runtime, adding real read capabilities (initiatives and work items) with citations and per-action authorization, integrating provider configuration via Console’s agent_configs, hardening failure modes and telemetry, and updating the Composer UI to surface sources and in-progress state — all with mock as default and CI fully offline. Mutations/receipts and broader read surface are deferred to a follow-on wave after UX and cost/rollout decisions.
Deferred scope
- Expose at least one real mutate capability with confirmation and durable receipts end-to-end — Mutation UX has never run end-to-end in the client and may need a UX pass; exposing mutations without this risks user confusion and irreversible changes. (unblocks: After read-only E2E verification in staging and operator confirmation that the confirmation/receipt UX is acceptable.) · retracted
- Additional read capabilities (workspace directory, releases, worker health) — Focus the first wave on the primary use case (initiatives and work items) to reduce risk; extend surface subsequently. (unblocks: After initial read capabilities are in production and telemetry shows acceptable latency/cost.) · retracted
- Production enablement with selected provider/model and cost guardrails — Requires explicit provider/model selection and a cost ceiling owner; enabling without these is risky under the single-box constraint. (unblocks: Operator decision recorded; staging verification complete per runbook; telemetry dashboards in place.) · retracted
Continuation checkpoint: Merge and validate sequences 0–5; demonstrate in staging that the Composer answers initiative/work-item questions with citations using the tool-capable provider adapter (still testable offline). Operator selects provider/model and cost ceiling and confirms data access path for Orchestrator-owned records; decide whether to proceed with mutations/receipt UX. · decided: Answered retrospectively at initiative close. Provider/model: two vendors supported via the conversation_runtime agent config — OpenAI gpt-5 (shipped) and Claude claude-opus-5 (outshine#484). Orchestrator reads go through the canonical /orchestrator/* surface. No mutate capability was exposed; the confirmation/receipt UX is unexercised and is filed as RM-0007. Cost ownership for a paid provider is filed as RM-0008; the workspace directory read as RM-0006. - Wave 2foundation-firstAPPROVED6 work items
Wave 2 enables grounded Q&A over Orchestrator releases and worker/queue health via the canonical /orchestrator/* surface, fixes the broken work-item read path from wave 1, and adds a Claude (claude-opus-5) tool-calling transport selectable solely via agent_configs. Reads only; no mutations. Default provider remains the offline mock and CI stays fully offline.
Deferred scope
- Workspace directory read capability — Deferred from wave 2 of the Composer tool-calling initiative: the source of truth is undecided — OutShine's own membership tables versus a canonical platform directory service. Choosing wrongly means rework and a cross-service mismatch, so the decision gates the work rather than the other way round. (unblocks: Operator confirms the canonical source for directory data and the fields a directory read should return; a reference implementation or API surface is identified.) · filed → RM-0006
- Expose one or more mutate capabilities with confirmation and durable receipts end-to-end — Deferred through both waves of the Composer tool-calling initiative. The confirmation and receipt paths exist in code but have never been exercised end to end from the client, so exposing a mutation would put an unproven, irreversible path in front of users. (unblocks: Staging validation of the confirmation/receipt UX with signoff that it is acceptable, and receipt persistence verified to survive a reload.) · filed → RM-0007
- Production enablement of a paid provider with cost guardrails — Deferred through both waves. Per-turn cost ceiling ownership and monitoring were never assigned, and the platform runs on a single box, so cutting over to a paid provider without a named owner and working dashboards risks unbounded spend nobody is watching. (unblocks: Operator assigns a cost owner, verifies dashboards and alerts, and approves the production cutover steps in the runbook.) · filed → RM-0008
Continuation checkpoint: Eligible for wave 3 when: (1) OutShine bugfix merged; (2) Orchestrator releases endpoints live; (3) Releases and worker-health read capabilities passing CI with offline mocks; (4) Claude transport merged with offline tests; and (5) runbook/docs updated. Optional operator decision: confirm canonical source for workspace directory and assign cost owner for production enablement. · decided: Answered retrospectively at initiative close. Provider/model: two vendors supported via the conversation_runtime agent config — OpenAI gpt-5 (shipped) and Claude claude-opus-5 (outshine#484). Orchestrator reads go through the canonical /orchestrator/* surface. No mutate capability was exposed; the confirmation/receipt UX is unexercised and is filed as RM-0007. Cost ownership for a paid provider is filed as RM-0008; the workspace directory read as RM-0006.
Work items
| # | Title | Repository | State | Issue | PR / CI | Review | |
|---|---|---|---|---|---|---|---|
| w2·0 ⛓ | Fix Orchestrator work-item read client path and add regression tests | outshine | COMPLETED | #481 | #486merged · CI passed | c1: approve | |
| w1·0 | Conversation runtime: harden failure modes and add bounded, content-free telemetry | outshine | COMPLETED | #460 | #464merged · CI passed | c1: approve | |
| w2·1 | Add canonical endpoints: GET /orchestrator/releases, /releases/{id}, and the single work-item read | orchestrator | COMPLETED | #453 | #456merged · CI passed | c1: approve | |
| w1·1 ⛓ | Add Console agent_configs client and provider resolution for conversation runtime | outshine | COMPLETED | #461 | #465merged · CI passed | c1: approve | |
| w2·2 ⛓ | Add release read capabilities (list/detail) via canonical Orchestrator surface with citations | outshine | COMPLETED | #482 | #492merged · CI passed | c1: approve | |
| w1·2 ⛓ | Introduce a tool-call-capable provider adapter (separate from the text-only provider) | outshine | COMPLETED | #462 | #466merged · CI passed | c1: approve | |
| w1·3 ⛓ | Add read capabilities for initiatives and work items with citations and authorization | outshine | COMPLETED | #463 | #467merged · CI passed | c1: approve | |
| w2·3 ⛓ | Add worker/queue health read capability using existing /orchestrator/workers endpoints | outshine | COMPLETED | #483 | #493merged · CI passed | c1: approve | |
| w1·4 ⛓ | Composer UI: show consulted sources, working state, and explicit provider error surfaces | workspaces | COMPLETED | #156 | #157merged · CI passed | c1: approve | |
| w2·4 | Add Claude (claude-opus-5) tool-calling transport and provider selection via agent_configs | outshine | COMPLETED | #484 | #487merged · CI passed | c1: approve | |
| w2·5 ⛓ | Runbook/docs: enabling Claude and new reads (releases, worker health) with cost/observability guardrails | hub | COMPLETED | #67 | #68merged · no CI in this repo | c1: approve | |
| w1·5 ⛓ | Runbook: enable conversation runtime with agent_configs and observability/cost guardrails | hub | COMPLETED | #65 | #66merged · no CI in this repo | c1: approve |
Release candidate
ELIGIBLEAll work complete — merged, reviewed, and unblocked. Release and deployment remain manual.
aed178afee · review approvebf0f8381d8 · review approve9bd9dffcb2 · review approve2497b2af0b · review approve96058a8734 · review approve81401c7f83 · review approvec0b6e2d74a · review approve16a4dce927 · review approve44bc912954 · review approve4aa7262cc0 · review approve3fd2659ce2 · review approvea752f0ab80 · review approveRelease planning
Releases are cut in the Release planning section. Associating this initiative requires an admin role.
Associating an initiative to a release only records the link — it never changes the initiative’s own state or work. Releases are optional.
Questions
agent:claude · Fix Orchestrator work-item read client path and add regression tests · 8/15/2026, 6:40:17 PM
The wave-1 `orchestrator.read_work_item` capability currently requests `GET /orchestrator/work-items/{work_item_id}` on the canonical Orchestrator surface and 404s; `GET /orchestrator/initiatives/{initiative_id}` works. What is the exact correct canonical path template for reading a single work item? Candidates I'm weighing: (a) a different flat resource token, e.g. `/orchestrator/work_items/{id}` or `/orchestrator/issues/{id}`; (b) nested under its initiative, e.g. `/orchestrator/initiatives/{initiative_id}/work-items/{work_item_id}` (if so, the capability needs an `initiative_id` argument too); or (c) something else. Please give the literal path template so the fix and its regression tests assert the real canonical path rather than a guess.
Answer: (c) — none of the candidates exist today. The canonical SHAPE is (b), but the member route itself is missing and must be added. Do not build around it; there is no acceptable interim. ## The literal path template GET /orchestrator/initiatives/{initiativeId}/work-items/{workItemId} Assert exactly that in the regression tests. It returns 404 right now because the handler does not exist. It must be added in Xyence/orchestrator at: apps/web/src/app/orchestrator/initiatives/[id]/work-items/[workItemId]/route.ts ## Evidence (probed live against the running orchestrator with a valid admin token) /orchestrator/work-items/{id} 404 /orchestrator/work_items/{id} 404 /orchestrator/issues/{id} 404 /orchestrator/initiatives/{initiativeId}/work-items 200 /orchestrator/initiatives/{initiativeId}/work-items/{workItemId} 404 /orchestrator/initiatives/{initiativeId}/work-items/{workItemId}/execution 200 Under `apps/web/src/app/orchestrator/initiatives/[id]/work-items/` there are exactly three route.ts files: the collection, `[workItemId]/execution`, and `[workItemId]/actions/[actionId]`. The `[workItemId]` segment already exists and has children; only its own handler is absent. This is a hole in an otherwise-complete hierarchy, not a deliberate flat-vs-nested decision. ## Why nested rather than a flat resource token 1. The member segment is already established by its two existing children. 2. The Workspaces UI addresses work items identically (`apps/orchestrator/src/app/(workspace)/initiatives/[id]/work-items/[workItemId]/page.tsx`), so the read path and the citation URL agree. 3. A flat route would introduce a second addressing scheme for a single resource. ## Two further corrections this work item must make 1. The citation URL is broken the same way. `_shape_work_item` in `app/services/conversation_runtime/orchestrator_reads.py` builds `_link_for(base_url, f'/orchestrator/work-items/{work_item_id}')`. That path exists in neither the API nor the Workspaces UI. It must use the same nested template, which means it needs the initiative id to construct. 2. The capability needs an `initiative_id` argument alongside `work_item_id`. Note that it cannot be discovered from the sibling read: `GET /orchestrator/initiatives/{id}` returns `{initiative, brief, workItemCount}` — a count, not the items. Treat `initiative_id` as required input to `orchestrator.read_work_item`. ## Explicitly rejected Do NOT fetch `GET /orchestrator/initiatives/{initiativeId}/work-items` and select the member client-side. It works (the list carries `id` and the full `body`), but it is an interim workaround, it is O(n) over the collection, and the list entries omit `initiative_id`, which `_shape_work_item` needs for the parent-initiative citation. We want the correct endpoint, not a workaround. ## Cross-repo dependency This work item is scoped to Xyence/outshine, but the endpoint above lands in Xyence/orchestrator. The OutShine fix cannot pass its regression tests until that route exists. Coordinate with the wave-2 Orchestrator work item that is already adding `GET /orchestrator/releases` and `GET /orchestrator/releases/{id}` — the new work-item member route belongs with it, following the same auth, response-shape, and testing conventions as the neighbouring canonical handlers.
agent:openai · 8/15/2026, 1:09:36 PM
Do we expose any mutate capability in the next wave, and does the confirmation flow need a UX pass first before we allow real mutations?
Two operator decisions are required to complete the rollout: selecting the vendor/model and cost ceiling for the conversation runtime (and who owns monitoring), and confirming the data access path for initiative/work-item reads when those are Orchestrator-owned. These affect irreversible cost exposure and integration posture and cannot be inferred from repositories.
Answer: No. Reads only next wave, and yes — the confirmation flow needs to be exercised first.
agent:openai · 8/15/2026, 1:09:36 PM
Confirm the data access path for reading initiatives/work items when they are Orchestrator-owned: use the canonical HTTP surface or an existing service-to-service call with the workspaces↔orchestrator JWT.
Two operator decisions are required to complete the rollout: selecting the vendor/model and cost ceiling for the conversation runtime (and who owns monitoring), and confirming the data access path for initiative/work-item reads when those are Orchestrator-owned. These affect irreversible cost exposure and integration posture and cannot be inferred from repositories.
Answer: Use the canonical HTTP surface.
agent:openai · 8/15/2026, 1:09:36 PM
Select the vendor/model for the conversation_runtime provider and the acceptable per-turn cost ceiling (who owns and monitors it).
Two operator decisions are required to complete the rollout: selecting the vendor/model and cost ceiling for the conversation runtime (and who owns monitoring), and confirming the data access path for initiative/work-item reads when those are Orchestrator-owned. These affect irreversible cost exposure and integration posture and cannot be inferred from repositories.
Answer: claude-opus-5
Timeline
No events yet.