Why isn't this task running?

Agent Queue used to answer that in a log line nobody could query. Making it answerable turned readiness from a scan into one indexed column, and every kind of waiting into the same kind of row.

Twelve agents are running on the box. Forty-odd tasks are in the queue. One of them — the one you actually care about — has been sitting there for an hour, and the dashboard says DEFINED next to it, which tells you nothing.

For most of Agent Queue’s life the honest answer to why isn’t this running was: read the daemon log and hope. A method called _log_scheduler_blockers wrote a heuristic sentence to disk every cycle. It was often right. It was never queryable, it disagreed with the dashboard, and no other part of the system — not the CLI, not the supervisor agent that is supposed to notice stuck work — could see it.

This post is about what it took to make that question answerable, because the answer turned out not to be a better log line. It was three changes to the substrate underneath.

The scan that could only say yes

The old model had one implicit edge type. task_dependencies was a pair of columns, (task_id, depends_on_task_id), and it meant exactly one thing: this blocks that. Everything else got overloaded onto it. Plan subtasks used parent_task_id and a dependency edge, with a special case in the promotion cascade that treated an IN_PROGRESS parent as satisfied. There was no way to say “run when all of X’s children are done”, no way to say “run only if X failed”, and no way to record that this task was discovered while working on that one.

Worse than the thinness was how readiness was computed. Every five seconds the cascade re-derived dependency satisfaction for every DEFINED and BLOCKED task by walking their dependency lists. Nothing was persisted, so nothing was indexable, and — the part that actually bit — the scan only ever ran in the positive direction. A task that had already been promoted to READY and then gained a blocking edge was never demoted. The scheduler would happily pick it up and run it against a prerequisite that had not happened.

Edges that mean something

The first change was to give edges a type. task_dependencies gained a dep_type column, and four of the types block:

dep_type Reads as Satisfied when
blocks this task runs after the dep the dep is COMPLETED
parent-child this task is a child of a container the container is anything but DEFINED — it has been released
waits-for this task fans in over a container’s children every child of that container is COMPLETED
conditional-blocks this task runs only if the dep failed the dep failed terminally — BLOCKED, or FAILED with the retry budget spent

Four more types — discovered-from, related, duplicates, supersedes — never block anything. They are provenance, and they exist so that the graph can record why a task is there without that fact accidentally becoming a scheduling constraint.

waits-for is the one worth dwelling on, because it is dynamic. A finalise or review task waits on a container; workers that discover follow-up work mid-flight attach it to that same container with a parent-child edge. The finalise task, which may already have been unblocked, is re-blocked by children that did not exist when it was written. That is not a special case in the cascade — it falls out of evaluating the same predicate again.

conditional-blocks needed a disposal rule to avoid the obvious rot. Once the dependency reaches COMPLETED, a contingency edge on it can never be satisfied, so a cascade step auto-closes any dependent whose only remaining unsatisfied blockers are conditional edges on completed deps. They close COMPLETED with a no-op outcome and a task.skipped_conditional event. Contingency tasks never accumulate in the queue.

One column, recomputed in the transaction

The second change is the one that made the question cheap to ask. tasks.is_blocked is a single boolean column, and it is a pure projection — never authored, only derived:

is_blocked(t) = 1 iff any blocking edge from t is unsatisfied, or any gate attached to t is not resolved.

It is graph blockedness only. The transient reasons — no idle agent, the workspace is locked, the budget is spent, the provider is cooling down — are deliberately not in the row. They change every tick, and a column that changes every tick is a write amplifier, not an index.

The recompute runs inside the caller’s transaction, not on a timer. Every mutation that could change some row’s projection does read, write and recompute in one block, so a reader can never observe the mutation without its projection. The set of columns that can trigger one is small and explicit:

python
# ``tasks`` columns the predicate reads.  A write to any of them can change
# some row's projection, so every mutating query method must recompute when
# its value set touches one — not just on ``status``: the
# ``conditional-blocks`` clause reads ``retry_count >= max_retries`` to tell a
# transient failure from a terminal one, so bumping a retry counter alone can
# flip a contingency task.
PROJECTION_INPUT_COLUMNS = frozenset({
    "status", "retry_count", "max_retries", "branch_name", "repo_id",
})

The algorithm is small. Seed with the tasks the mutation touched, expand to their direct dependents and to any waits-for waiter over a container they are a child of, then evaluate the predicate for the whole affected set in one set-based UPDATE. Within a single transaction every other task’s status is fixed, so the predicate is a pure function and each task needs exactly one evaluation. If the transaction itself moves several statuses — a bulk graph creation, a conditional auto-close, an admin skip cascade — re-seed with the rows whose status changed and repeat to a fixpoint. Every row that flips emits task.blocked or task.unblocked after commit.

What this bought is worth stating plainly. Promotion stopped being a scan and became DEFINED AND is_blocked = 0 — one indexed query. The scheduler learned to skip READY AND is_blocked = 1, which closed the demotion hole. And the ready frontier — the set of tasks that would be picked next — became something you can select rather than something you have to simulate.

Everything that waits is a row

The third change was to stop encoding why a task is waiting in its status.

The old model had a bespoke status for each external condition: AWAITING_APPROVAL, AWAITING_PLAN_APPROVAL, WAITING_INPUT, each with its own polling loop and its own Discord buttons. Every new condition — CI is green, a timer expired, another project’s task finished — would have needed another status, another poller and another piece of UI.

A gate replaces all of them. It is a first-class wait record, with a type, a question, an optional deadline, and a join table attaching it to the tasks it blocks. One gate can block many tasks; one task can wait on many gates.

Gate type Resolved when
human a person runs gate resolve — the only path
timer the clock passes the target
pr-merged the PR is merged and its work reached the default branch
ci-run gh run reports success
event a matching event fires, with the persisted event table as the restart-safe backstop
task a named task reaches COMPLETED, across projects if need be

An open gate blocks its waiters exactly like a blocking dependency — the is_blocked predicate simply includes an EXISTS over the join — and resolving one triggers the same in-transaction recompute as completing a dependency. A deterministic sweep runs each cycle: it resolves satisfied timer and task gates with pure SQL and a clock, polls the GitHub-backed ones on a 60-second throttle, re-checks event gates against events persisted since the gate was created, and marks anything past its deadline as expired.

Expired continues to block. A timed-out approval that silently self-approved would be a much more interesting bug than a stuck queue, so the sweep fires gate.expired for escalation and leaves the task where it is.

Merged is not the same as on the default branch

The pr-merged row in that table carries a scar, and it is the best argument in the whole design for making waiting explicit.

A pr-merged gate exists to tell a dependent task “your prerequisite has shipped”. Resolving it the obvious way — ask GitHub whether the PR is merged — is wrong, because a PR whose base is a feature branch merges without putting a single line on the default branch.

The fix is that the sweep asks a merged PR a second question: if the PR’s base is not the project’s default branch, the gate stays open until origin/<base> is an ancestor of origin/<default>. An unanswerable question — no gh, no auth, no network — resolves the gate as before, because unknowable must not wedge every dependent shut forever.

The merge side grew the same kind of second question. integration.merge_ci_policy decides what a non-green status-check rollup does before a merge: nothing, a logged verdict, or a refusal. It ships as warn. The gate exists because GitHub’s does not — main carries no required status check, and gh pr merge merged 29 of the last 30 PRs red. And a green rollup only proves the head passed against the base as it was when the run started, so two PRs each green on a stale base can put a red main together; integration.merge_require_up_to_date folds that into the same verdict, and ships as true.

The answer, as a typed list

With edges typed, blockedness projected and waiting made into rows, “why isn’t this running” becomes a query with a stable shape. aq task explain returns an ordered list of {code, detail, ref} — a closed enum, a human sentence, and the id of the thing to go look at:

bash
aq task explain --task-id vivid-falcon.8 --json
json
{
  "schema_version": 1,
  "data": {
    "task_id": "vivid-falcon.8",
    "status": "DEFINED",
    "is_blocked": true,
    "reasons": [
      {
        "code": "blocked_dependency",
        "detail": "blocked by blocks dep 'vivid-falcon.3' (Information architecture and copy deck) status=IN_PROGRESS",
        "ref": "vivid-falcon.3"
      },
      {
        "code": "blocked_gate",
        "detail": "gate 'g-4f21' (human: Approve the copy deck) status=open",
        "ref": "g-4f21"
      }
    ]
  }
}
aq task explainTwo graph reasons on one task: an unsatisfied blocking edge and an unresolved human gate. Every reason carries the id of the thing to go and look at, which is what makes the answer actionable rather than merely true.

Graph reasons come first, because they are the persistent story: a hold:* label, then unsatisfied edges, then unresolved gates. Capacity reasons come second, read from the last scheduler tick’s cached snapshot — no_idle_agent, no_compatible_agent, workspace_locked, budget_exhausted, rate_limited, project_paused. Recovery states get their own codes rather than an empty answer: paused_backoff names the time a cooldown ends, paused_manually says which command resumes it, and blocked_terminal says that a hard failure will not auto-recover and needs a restart or a reopen.

The important property is not the vocabulary. It is that the same reason builder feeds the log line, the CLI, the dashboard and the supervisor agent, so none of them can disagree about why a task is waiting. That was the actual failure before — four surfaces, four heuristics, four answers.

What generalises

Three things carried over from this that I would do again on any queue with a graph under it.

Persist the projection, not the reason. is_blocked is one column, derived in the same transaction as the mutation that could change it, and it turned the hot path from a scan into an index lookup. The transient reasons stayed out of the row precisely because they are transient.

Make every kind of waiting the same kind of row. Human approval, a timer, a PR, a CI run, another project’s task — one table, one blocking rule, one sweep. The alternative is a new status, a new poller and a new button per condition, and they diverge immediately.

Assume the operator will ask why, and make the answer a query. A log line is a description of the system for whoever is reading the log. A typed reason list is an interface, and once it exists the dashboard and the supervisor agent get the same answer the human gets, for free.