Why isn't this task running?
Agent Queue used to answer that in a log line nobody could query. Making it answerable turned readiness from a scan into one indexed column, and every kind of waiting into the same kind of row.
Twelve agents are running on the box. Forty-odd tasks are in the queue. One of
them — the one you actually care about — has been sitting there for an hour,
and the dashboard says DEFINED next to it, which tells you nothing.
For most of Agent Queue’s life the honest answer to why isn’t this running
was: read the daemon log and hope. A method called _log_scheduler_blockers
wrote a heuristic sentence to disk every cycle. It was often right. It was
never queryable, it disagreed with the dashboard, and no other part of the
system — not the CLI, not the supervisor agent that is supposed to notice
stuck work — could see it.
This post is about what it took to make that question answerable, because the answer turned out not to be a better log line. It was three changes to the substrate underneath.
The scan that could only say yes
The old model had one implicit edge type. task_dependencies was a pair of
columns, (task_id, depends_on_task_id), and it meant exactly one thing:
this blocks that. Everything else got overloaded onto it. Plan subtasks used
parent_task_id and a dependency edge, with a special case in the
promotion cascade that treated an IN_PROGRESS parent as satisfied. There was
no way to say “run when all of X’s children are done”, no way to say “run only
if X failed”, and no way to record that this task was discovered while working
on that one.
Worse than the thinness was how readiness was computed. Every five seconds the
cascade re-derived dependency satisfaction for every DEFINED and BLOCKED
task by walking their dependency lists. Nothing was persisted, so nothing was
indexable, and — the part that actually bit — the scan only ever ran in the
positive direction. A task that had already been promoted to READY and then
gained a blocking edge was never demoted. The scheduler would happily pick it
up and run it against a prerequisite that had not happened.
Edges that mean something
The first change was to give edges a type. task_dependencies gained a
dep_type column, and four of the types block:
dep_type |
Reads as | Satisfied when |
|---|---|---|
blocks |
this task runs after the dep | the dep is COMPLETED |
parent-child |
this task is a child of a container | the container is anything but DEFINED — it has been released |
waits-for |
this task fans in over a container’s children | every child of that container is COMPLETED |
conditional-blocks |
this task runs only if the dep failed | the dep failed terminally — BLOCKED, or FAILED with the retry budget spent |
Four more types — discovered-from, related, duplicates, supersedes —
never block anything. They are provenance, and they exist so that the graph can
record why a task is there without that fact accidentally becoming a
scheduling constraint.
waits-for is the one worth dwelling on, because it is dynamic. A finalise or
review task waits on a container; workers that discover follow-up work
mid-flight attach it to that same container with a parent-child edge. The
finalise task, which may already have been unblocked, is re-blocked by
children that did not exist when it was written. That is not a special case in
the cascade — it falls out of evaluating the same predicate again.
conditional-blocks needed a disposal rule to avoid the obvious rot. Once the
dependency reaches COMPLETED, a contingency edge on it can never be
satisfied, so a cascade step auto-closes any dependent whose only remaining
unsatisfied blockers are conditional edges on completed deps. They close
COMPLETED with a no-op outcome and a task.skipped_conditional event.
Contingency tasks never accumulate in the queue.
One column, recomputed in the transaction
The second change is the one that made the question cheap to ask.
tasks.is_blocked is a single boolean column, and it is a pure projection —
never authored, only derived:
is_blocked(t) = 1iff any blocking edge from t is unsatisfied, or any gate attached to t is not resolved.
It is graph blockedness only. The transient reasons — no idle agent, the workspace is locked, the budget is spent, the provider is cooling down — are deliberately not in the row. They change every tick, and a column that changes every tick is a write amplifier, not an index.
The recompute runs inside the caller’s transaction, not on a timer. Every mutation that could change some row’s projection does read, write and recompute in one block, so a reader can never observe the mutation without its projection. The set of columns that can trigger one is small and explicit:
# ``tasks`` columns the predicate reads. A write to any of them can change
# some row's projection, so every mutating query method must recompute when
# its value set touches one — not just on ``status``: the
# ``conditional-blocks`` clause reads ``retry_count >= max_retries`` to tell a
# transient failure from a terminal one, so bumping a retry counter alone can
# flip a contingency task.
PROJECTION_INPUT_COLUMNS = frozenset({
"status", "retry_count", "max_retries", "branch_name", "repo_id",
})The algorithm is small. Seed with the tasks the mutation touched, expand to
their direct dependents and to any waits-for waiter over a container they are
a child of, then evaluate the predicate for the whole affected set in one
set-based UPDATE. Within a single transaction every other task’s status is
fixed, so the predicate is a pure function and each task needs exactly one
evaluation. If the transaction itself moves several statuses — a bulk graph
creation, a conditional auto-close, an admin skip cascade — re-seed with the
rows whose status changed and repeat to a fixpoint. Every row that flips emits
task.blocked or task.unblocked after commit.
What this bought is worth stating plainly. Promotion stopped being a scan and
became DEFINED AND is_blocked = 0 — one indexed query. The scheduler learned
to skip READY AND is_blocked = 1, which closed the demotion hole. And the
ready frontier — the set of tasks that would be picked next — became something
you can select rather than something you have to simulate.
Everything that waits is a row
The third change was to stop encoding why a task is waiting in its status.
The old model had a bespoke status for each external condition:
AWAITING_APPROVAL, AWAITING_PLAN_APPROVAL, WAITING_INPUT, each with its
own polling loop and its own Discord buttons. Every new condition — CI is
green, a timer expired, another project’s task finished — would have needed
another status, another poller and another piece of UI.
A gate replaces all of them. It is a first-class wait record, with a type, a question, an optional deadline, and a join table attaching it to the tasks it blocks. One gate can block many tasks; one task can wait on many gates.
| Gate type | Resolved when |
|---|---|
human |
a person runs gate resolve — the only path |
timer |
the clock passes the target |
pr-merged |
the PR is merged and its work reached the default branch |
ci-run |
gh run reports success |
event |
a matching event fires, with the persisted event table as the restart-safe backstop |
task |
a named task reaches COMPLETED, across projects if need be |
An open gate blocks its waiters exactly like a blocking dependency — the
is_blocked predicate simply includes an EXISTS over the join — and
resolving one triggers the same in-transaction recompute as completing a
dependency. A deterministic sweep runs each cycle: it resolves satisfied timer
and task gates with pure SQL and a clock, polls the GitHub-backed ones on a
60-second throttle, re-checks event gates against events persisted since the
gate was created, and marks anything past its deadline as expired.
Expired continues to block. A timed-out approval that silently self-approved
would be a much more interesting bug than a stuck queue, so the sweep fires
gate.expired for escalation and leaves the task where it is.
Merged is not the same as on the default branch
The pr-merged row in that table carries a scar, and it is the best argument
in the whole design for making waiting explicit.
A pr-merged gate exists to tell a dependent task “your prerequisite has
shipped”. Resolving it the obvious way — ask GitHub whether the PR is merged —
is wrong, because a PR whose base is a feature branch merges without putting
a single line on the default branch.
The fix is that the sweep asks a merged PR a second question: if the PR’s base
is not the project’s default branch, the gate stays open until origin/<base>
is an ancestor of origin/<default>. An unanswerable question — no gh, no
auth, no network — resolves the gate as before, because unknowable must not
wedge every dependent shut forever.
The merge side grew the same kind of second question. integration.merge_ci_policy
decides what a non-green status-check rollup does before a merge: nothing, a
logged verdict, or a refusal. It ships as warn. The gate exists because
GitHub’s does not — main carries no required status check, and gh pr merge
merged 29 of the last 30 PRs red. And a green rollup only proves the head
passed against the base as it was when the run started, so two PRs each green
on a stale base can put a red main together;
integration.merge_require_up_to_date folds that into the same verdict, and
ships as true.
The answer, as a typed list
With edges typed, blockedness projected and waiting made into rows, “why isn’t
this running” becomes a query with a stable shape. aq task explain returns
an ordered list of {code, detail, ref} — a closed enum, a human sentence, and
the id of the thing to go look at:
aq task explain --task-id vivid-falcon.8 --json{
"schema_version": 1,
"data": {
"task_id": "vivid-falcon.8",
"status": "DEFINED",
"is_blocked": true,
"reasons": [
{
"code": "blocked_dependency",
"detail": "blocked by blocks dep 'vivid-falcon.3' (Information architecture and copy deck) status=IN_PROGRESS",
"ref": "vivid-falcon.3"
},
{
"code": "blocked_gate",
"detail": "gate 'g-4f21' (human: Approve the copy deck) status=open",
"ref": "g-4f21"
}
]
}
}Graph reasons come first, because they are the persistent story: a hold:*
label, then unsatisfied edges, then unresolved gates. Capacity reasons come
second, read from the last scheduler tick’s cached snapshot —
no_idle_agent, no_compatible_agent, workspace_locked, budget_exhausted,
rate_limited, project_paused. Recovery states get their own codes rather
than an empty answer: paused_backoff names the time a cooldown ends,
paused_manually says which command resumes it, and blocked_terminal says
that a hard failure will not auto-recover and needs a restart or a reopen.
The important property is not the vocabulary. It is that the same reason builder feeds the log line, the CLI, the dashboard and the supervisor agent, so none of them can disagree about why a task is waiting. That was the actual failure before — four surfaces, four heuristics, four answers.
What generalises
Three things carried over from this that I would do again on any queue with a graph under it.
Persist the projection, not the reason. is_blocked is one column, derived
in the same transaction as the mutation that could change it, and it turned the
hot path from a scan into an index lookup. The transient reasons stayed out of
the row precisely because they are transient.
Make every kind of waiting the same kind of row. Human approval, a timer, a PR, a CI run, another project’s task — one table, one blocking rule, one sweep. The alternative is a new status, a new poller and a new button per condition, and they diverge immediately.
Assume the operator will ask why, and make the answer a query. A log line is a description of the system for whoever is reading the log. A typed reason list is an interface, and once it exists the dashboard and the supervisor agent get the same answer the human gets, for free.