enhancement
major
minor
major
minor
major
#29649
TL Views: a job's final state is overwritten by a stale "running" snapshot, so <job-status> shows a completed job as running
Problem
A <start-job> whose body reports frequently (here ChunkedScriptJobBody with update-interval="200" and a jobMessage(...) per item) sometimes ends with its <job-status> stuck on "Läuft": the phase of the completion active, the elapsed time counting on, the cancel button offered — for minutes, although the job has ended and the command waiting for it has already continued.
Observed in the TL Development app (Trac sync of a 6-ticket component, snapshot build 42). The client received the correct terminal state (status: completed, result: "Probe lnw", cancelable: false, finishedAt: 1790344546737) at 31542 ms and, 9 ms later, a running state (result: null, cancelable: true, finishedAt: null) that overwrote it. A second run of the same sync showed the completed state correctly — the defect is intermittent.
Cause
JobRunner builds every snapshot under _lock but delivers it (deliver(state) → _target.set(state)) after the lock is released. Two threads deliver: the worker (report(), finish()) and the scheduler thread running the throttling flush(). The interleaving
- a report is held back, a flush() is scheduled;
- flush() runs on the scheduler thread, takes a RUNNING snapshot under the lock, releases the lock;
- the worker's body returns, finish() takes the COMPLETED snapshot and delivers it;
- the scheduler thread now delivers its RUNNING snapshot, which becomes the channel's value for good.
The _status != RUNNING guard in flush() does not help, since the status was still RUNNING when the snapshot was taken. The same race exists between two report() deliveries (an older message overwriting a newer one), which is harmless only because the next report corrects it — nothing corrects the terminal one.
Expected
Snapshots reach the channel in the order they were taken; in particular, once the terminal snapshot has been delivered, no earlier snapshot may be written afterwards.
Lösung
JobRunner serialisiert das Erstellen eines Snapshots und seine Auslieferung: ein eigener Monitor wird in report(), flush() und finish() vom Erstellen des Snapshots (weiterhin unter _lock) bis einschließlich des Schreibens in den Kanal gehalten. Snapshots erreichen den Kanal damit in der Reihenfolge, in der sie erstellt wurden; ist der End-Snapshot ausgeliefert, findet ein späterer Flush den Job beendet vor und liefert nichts mehr aus. cancel() nimmt weiterhin nur _lock, ein Abbruch aus dem Request wartet also nie auf eine Auslieferung. Ein Regressionstest erzwingt die beschriebene Verschränkung deterministisch.