Follow-up to PR #798 (#795 final). PR #798 closed the total-timeout /
hard-cap issue, but left two residual code-path inconsistencies that
#802 captures:
- Submit path and refresh path used different trace-polling code:
submitMessage called `pollRunnerTrace` + `waitForAgentResult` (a
dual function pair from pre-#798), while hydrate() never subscribed
to the running trace at all. After a page refresh, the user saw
the cloud-api "running" state but no live trace events.
- The submit fail branch only dispatched a local "Code Agent 超时"
message; it never called `persistConversation`, so the failure was
invisible after F5 — the cloud-api conversation would show
"running" forever from the user's perspective.
This PR collapses both into a single `subscribeToTrace` entry point
and wires the refresh path into the same active subscription:
- `state/runner-trace.ts` — replace `waitForAgentResult` +
`pollRunnerTrace` with a single `subscribeToTrace(config)` that
handles both the result endpoint and the trace endpoint in one
`for(;;)` loop. Calls `onActivity` on EVERY successful poll (not
just on new events) so the per-poll inactivity window in
`fetchJson` does not fire on a healthy polling loop. Terminal
status -> `onComplete`; 5xx -> keep polling, no activity; 4xx ->
`onInfrastructureError`; `signal.aborted` -> silent return. The
hard-cap / 4x / `TRACE_HARD_CAP_ATTEMPTS` machinery is gone (was
already removed by #798).
- `state/trace-reattach.ts` — new `useTraceReattach` hook that
subscribes to the running trace on hydrate (same `subscribeToTrace`
call as submitMessage). Reuses an existing agent message slot for
the traceId if one exists; otherwise creates a placeholder. Persists
on both onComplete and onInfrastructureError. Path unification
achieved: one entry, two callers.
- `state/workbench.ts`:
- `submitMessage` calls `subscribeToTrace` directly with a
single-line `onActivity: () => updateActivity()` plus
`onComplete` / `onInfrastructureError` callbacks. The
`onInfrastructureError` branch now ALSO calls
`persistConversation`, persisting the fail state to cloud-api.
- The re-attach useEffect is replaced with a one-liner
`useTraceReattach({...})` call (workbench.ts stays under the
400-line guard).
- `scripts/fetchJson-inactivity.test.ts` — third case tightened:
200ms-cadence activity for 10s with `timeoutMs: 3000`; the
request must live the full 10s and abort only after activity
stops. Asserts `elapsed >= 10_000`.
- `docs/reference/spec-v02-hwlab-cloud-web.md` — the "Code Agent
Timeout Model" section is renamed "Code Agent Timeout Model 与
Path Unification" and now also pins: (a) `subscribeToTrace` is
the only entry; `waitForAgentResult` / `pollRunnerTrace` /
`TRACE_HARD_CAP_ATTEMPTS` MUST NOT exist; (b) `useTraceReattach`
MUST be wired in; (c) `persistConversation` is required in both
onComplete and onInfrastructureError; (d) `state.workspace.activeTraceId`
changes MUST re-attach. Includes a distilled history of
#775 / #777 / #791 / #795 / #797 / #798 / #802.
Verification
- `bun run scripts/tsc-check.ts` → strict React TSX, 0 explicit any
- `bun run check` → 5 pass / 0 fail (3 inactivity + 2 dist-contract)
- `bun run build` → 9 dist files verified fresh; `app.js`
238460 → 240990 B
- `state/trace-reattach.ts` is now wired in; `state/runner-trace.ts`
exports only `subscribeToTrace` (verified by tsc + grep)
- `workbench.ts` line count 400 (under the 400-line guard)
- New 200ms-cadence 10s test passes; old 100ms-cadence 2s test
and total-timeout 200ms test still pass.
Refs: HWLAB #802, #795, #777, #791, #798
Co-authored-by: Codex Agent <codex@hwlab.local>
Follow-up to PR #797: drop every wall-clock or attempt-count ceiling from
`state/runner-trace.ts` so the Web UI's Code Agent inactivity-timeout
contract is total-cap-free. The previous PR kept a
`max(totalTimeoutMs * 4, totalTimeoutMs + 60s)` outer-loop safety net
plus a `TRACE_HARD_CAP_ATTEMPTS = 120` poll-count cap, which would
re-introduce the same regression class that #795 was filed against: a
long-running operation that keeps emitting trace events could still be
killed by the outer ceiling even with an active `activityRef`. The
"Code Agent Timeout Model" section of
`docs/reference/spec-v02-hwlab-cloud-web.md` now pins the contract:
- The sole abort signal is per-poll `fetchJson` inactivity-timeout,
driven by `activityRef`.
- The outer `waitForAgentResult` / `pollRunnerTrace` loops are
`for (;;)`. They only exit when the upstream returns a terminal
status, the request fails with a non-5xx, or the inactivity window
fires inside `fetchJson`.
- The per-poll `getAgentChatResult` call passes the full
`totalTimeoutMs`; the inactivity window is **not** shrunk by
wall-clock elapsed time.
- `TRACE_HARD_CAP_ATTEMPTS` and the 4x / +60s outer cap are deleted.
- Runaway protection is the caller's responsibility: Web users have
cancel / steer / close-tab; CLI has `--timeout-ms`. The browser does
not introduce an implicit cap.
Changes
- `web/hwlab-cloud-web/src/state/runner-trace.ts`:
- Drop `const hardCapMs = Math.max(totalTimeoutMs * 4, totalTimeoutMs + 60_000)`
and the `if (Date.now() - startedAt >= hardCapMs) break;` checks in
both `waitForAgentResult` and `pollRunnerTrace`.
- Replace `while (Date.now() - startedAt < hardCapMs)` with
`for (;;)`. Drop the now-unused `startedAt` locals and the
`let attempt = 0; attempt > TRACE_HARD_CAP_ATTEMPTS` guard.
- Pass `totalTimeoutMs` directly to `api.getAgentChatResult` instead
of `totalTimeoutMs - (Date.now() - startedAt)` so the per-poll
inactivity window is stable.
- Delete `const TRACE_HARD_CAP_ATTEMPTS = 120;` (no longer referenced).
- `web/hwlab-cloud-web/scripts/fetchJson-inactivity.test.ts`:
- Replace the "1.5s cadence 6.5s" case with a stricter "200ms cadence
10s" case: `timeoutMs=3000` inactivity window, continuous 200ms
activity for 10s, then stop. Assert `elapsed >= 10_000` and that
the abort is the inactivity-timeout error. This pins the
`elapsed >= 10000` invariant in the spec.
- `docs/reference/spec-v02-hwlab-cloud-web.md`:
- Strengthen the existing bullet in "在系统中的职责划分" to
explicitly forbid total-timeout, hard cap, `codeAgentTimeoutMs * N`,
poll-count cap, and wall-clock-shrunk per-poll windows.
- Add a dedicated "Code Agent Timeout Model" section after "## 内部架构"
that distils the contract, lists the three invariants, and records
the history of #775 / #777 / #791 / #795 leading to the final state.
Verification
- `bun run scripts/tsc-check.ts` → strict React TSX, 0 explicit any
- `bun run check` → 5 pass / 0 fail (3 inactivity + 2 dist-contract)
- `bun run build` → 9 dist files verified fresh; `app.js` 241802 → 238460 B
(delta reflects the deleted hardCap / TRACE_HARD_CAP_ATTEMPTS code)
- New 200ms-cadence 10s inactivity test runs for ~13s and asserts the
request lives the full window — proving the outer cap is gone.
Refs: HWLAB #795 (final), #777, #791, #775
Co-authored-by: Codex <codex@local>
The React migration (#756) and the 7-round #775 patch left the
`fetchJson` inactivity-timeout surface half-wired: `client.ts` and the
`api.*` helpers accepted an `ActivityRef`, but `state/workbench.ts`
`submitMessage` and `state/runner-trace.ts` `waitForAgentResult` /
`pollRunnerTrace` never built or forwarded one. The Web UI was therefore
falling back to a **total** timeout: a long-running Code Agent operation
(e.g. `bootsharp` that finished in ~9.7s but the assistant message
chain took ~37s) was killed at `codeAgentTimeoutMs` even while trace
events kept flowing, producing "Code Agent 在超时内未返回可显示正文".
User-facing symptom in #795 is exactly this: the 4th turn on
`D601-F103-V2 bootshar[p]` timed out in the Web UI while the
backend completed the work and emitted `agentrun:result:completed`.
Changes
- `state/workbench.ts` adds a `useRef<number>`-backed `lastActivityAtRef`,
`updateActivity`, `buildActivityRef` and exposes `recordActivity` on the
store. `submitMessage` now bumps activity on the initial submit, on
every trace snapshot, and threads a single `submitActivityRef` through
`api.sendAgentMessage` / `api.steerAgentMessage`, `pollRunnerTrace`,
and `waitForAgentResult`.
- `state/runner-trace.ts` `waitForAgentResult` and `pollRunnerTrace`
accept an `ActivityRefSource`, pass it to `api.getAgentChatResult` so
each per-poll `fetchJson` uses the inactivity-timeout branch, and
`pollRunnerTrace` mutates the ref's `lastActivityAt` / `lastEventLabel`
/ `waitingFor` whenever a fresh snapshot arrives. The outer
`while`-loop hard ceiling becomes `max(totalTimeoutMs * 4, totalTimeoutMs + 60s)`
as a safety net for runaway operations; the per-request
inactivity-timeout (driven by `activityRef`) is the primary abort
signal, matching the pre-React `app-device-pod.ts:fetchJson` contract.
- `components/command-bar/CommandBar.tsx` adds an `onTyping` prop and
fires it from the textarea `onChange` so user-typing resets the
inactivity clock.
- `App.tsx` adds a stable `noteActivity` (`useCallback`) wired to
`store.recordActivity` and passes it as `onTyping` to `<CommandBar>`.
Test
- `scripts/fetchJson-inactivity.test.ts` adds three cases:
1. `fetchJson` without `activityRef` honours the total window.
2. `fetchJson` with 100ms-cadence activity survives the configured
window and aborts ~1-2s after activity stops with the
"无新活动" error (the HWLAB #795 regression assertion).
3. `fetchJson` with 1.5s-cadence activity (the actual
`pollRunnerTrace` interval) survives past the 5s window and only
aborts after activity stops at 6.5s.
Verification
- `bun run scripts/tsc-check.ts` passes (strict React TSX, 0 explicit any)
- `bun run check` 5 pass / 0 fail (3 new + 2 dist-contract)
- `bun run build` 9 dist files verified fresh; `app.js` 234676 → 238740 B
- orphan-listener check (per HWLAB #748): no `el.X.addEventListener`
in `*.tsx` whose X is not in the `el` literal.
Refs
- HWLAB #795 (this issue)
- HWLAB #775 round 1 #777 added the `fetchJson` `activityRef` plumbing
in `client.ts` but stopped short of the workbench / runner-trace layer
- HWLAB #756 React migration baseline
Co-authored-by: Codex <codex@local>
- Replace monolithic app.ts / app-conversation.ts / app-device-pod.ts /
app-skills.ts / app-helpers.ts / app-session-tabs.ts / auth.ts and the
364-line index.html with a Vite-built React + TypeScript SPA.
- New src/ module tree: main.tsx, App.tsx, components/{layout,auth,
conversation,command-bar,device-pod,skills,settings,help},
hooks/, services/api, services/auth, state, types, logic/ (pure
helpers from the old tree), styles/.
- Pure-logic modules (composer-policy, code-agent-facts, code-agent-
status, live-status, message-markdown, app-trace, app-session-tabs,
runtime) move to src/logic/ with their tests preserved; tests still
auto-discover.
- Build pipeline switches from Bun.build to Vite: vite.config.ts
inlines everything into a single app.js + app-assets/ for the dist;
scripts/dist-contract.ts now runs vite build and asserts the dist
against a fresh build instead of comparing dist to source.
- scripts/check.ts now verifies the React entry, the mount-only
index.html, Vite dist, the composer / session / device pod / skills
/ settings / help contracts, the React state shape, and the new
test file locations. It also refuses to run if any legacy app-*.ts
or auth.ts survives at the top level.
- New scripts/frontend-guard.ts enforces the 400-line migration target,
refuses the legacy el map and document.getElementById markers,
refuses to keep app-*.ts / auth.ts at the top level, and verifies
the required module directory layout.
- Update scripts/src/dev-cloud-workbench-smoke-lib.mjs to read the new
src/services/, src/logic/ paths and refresh cloudWebModuleSourceFiles.
- index.html is now a mount shell only (no login shell, no workbench
shell, no device pod sidebar markup). React renders the full UI from
src/main.tsx, satisfying the issue-756 migration goal of catching
unclosed tags at the TypeScript/JSX compile step instead of relying
on browser DOM healing.
- web:check (now: check + frontend-guard + tsc-check + bun test) is
green: 51 pass / 0 fail, Vite build produces dist/app.js (379 kB)
+ dist/index.html + dist/app-assets/. Vite build is reproducible;
dist freshness is verified against a fresh Vite build.
- Grandfathered pure-logic files (> 600 lines): src/logic/app-trace.ts
(1367), src/logic/live-status.ts (1006), src/logic/code-agent-facts.ts
(548), src/logic/code-agent-status.ts (572). Reason: each is a tested
pure-logic classifier with a public contract consumed by both the
CLI renderer and the Web renderer; splitting now would fork the
public function shape across two PRs and the migration target is the
React entry + module split, not a pure-logic refactor.
Refs #756. Closes#756 once merged + deployed to hwlab-v02.
Co-authored-by: Codex <codex@local>