Skip to content
HomeConsoleGet started

Common symptoms when running agents on runstate, what usually causes them, and where to look.

Start with the error code (every SDK error has one) and the error table. For a stuck run, rs.diagnostics.run(runId) and the console’s Resources and Tasks pages show who holds what and what is waiting.

Calls fail with UnavailableError after about 30 seconds

Section titled “Calls fail with UnavailableError after about 30 seconds”

The SDK can’t reach the API. Most often RUNSTATE_BASE_URL isn’t set, so the client is trying http://localhost:8080. Set it to https://api.getrunstate.com. Each attempt times out after requestTimeoutMs (10 s) and is retried twice, which is where the 30 seconds come from.

ConfigError: missing apiKey or missing spaceId

Section titled “ConfigError: missing apiKey or missing spaceId”

RUNSTATE_API_KEY or RUNSTATE_SPACE_ID isn’t set in the environment of the process that constructs the client, and wasn’t passed as an option.

The key lacks a permission. Taking quota units, reserving or settling budget, and creating quotas, budgets or work limits all need resource_config. See Configuration & auth. FORBIDDEN also means the key isn’t allowed for this space.

Each space allows 600 requests per minute by default, and idle workers poll: consume() with the default pollMs of 250 makes up to 240 requests a minute per worker while the queue is empty. Empty polls aren’t billed, but they count toward the request rate. Raise pollMs for idle workers, or raise the space’s requestsPerMinute up to your plan’s limit (per-space limits).

  • Different run. Workers only receive tasks submitted under the exact run they attach to (not its parent or children). Check that rs.scope(runId) uses the id the producer submitted to.
  • Task group members are delivered on the group’s own run: use (await group.status()).scopeId.
  • Different queue name. Queue names are per space and case-sensitive.
  • Admission. Tasks submitted with requirements or workLimit are still visible to receive()/consume(), which bypass admission. If you use admission, workers should call admit().

The queue doesn’t exist in this space yet. Call rs.mailboxes.ensure(name) before submitting, consuming or creating a task group on it. The same applies to pools, quotas, budgets and work limits.

  • No worker is consuming that queue under that run.
  • With admission: a required pool or quota is short, the work limit is full, or a quota has a cooldown. admit() returns null until everything is available; check Resources in the console.
  • The run was cancelled or its deadline passed. rs.scope(runId).status() shows effective: 'INACTIVE'.

LeaseLostError, or a lease’s signal aborts, during long work

Section titled “LeaseLostError, or a lease’s signal aborts, during long work”

The SDK couldn’t renew in time, so it gave the lease up. Common causes:

  • The event loop was blocked. Renewal runs on timers in your process. CPU-heavy synchronous work in Node.js, or blocking calls inside an async def in Python, stop renewals. Move heavy work off the event loop (worker threads, asyncio.to_thread) or use a longer leaseSeconds.
  • The lease is too short for your network or pauses. Try 60 to 120 seconds for long model calls.
  • The run was cancelled. Renewals are refused after cancellation, so every held lease in the run is lost on its next renewal. That is the intended stop signal.

CLAIM_HELD but nobody is working on the key

Section titled “CLAIM_HELD but nobody is working on the key”

The previous owner probably crashed and its lease hasn’t expired yet; it will within leaseSeconds of its last renewal. rs.diagnostics.run(runId) lists holders and their earliest lease expiry, and expiredRecent shows leases that just expired. Use wait: true to queue behind the owner instead of failing.

  • Leases from crashed agents return only when they expire.
  • Pool units taken through admit() stay held until the admission lease expires, even after the task completes. Use a leaseSeconds close to the real work time. See admission.
  • The organization may be at its “agents working at once” limit (CONCURRENCY_LIMITED); acquisitions wait through it.

The work is being retried. Either the handler throws (check onError / on_error), or the work outlasts its lease: receive() and admit() don’t renew, so a delivery held past leaseSeconds is handed to another worker. After 5 attempts the task becomes FAILED.

A task with that key already exists in the run with different input. Inputs are compared by content fingerprint, so a timestamp or random id inside the input makes every retry “different”. Keep volatile values out of the input, or use a different key.

submit() returns an old result instead of doing the work again

Section titled “submit() returns an old result instead of doing the work again”

Task keys are permanent within a run: submitting an existing key with the same input returns the existing task, even if it finished long ago. Use a new key or a new run to redo work.

Only your wait ended. The task is still running on the server. Call result() again later, or re-attach with rs.scope(runId).task(taskId).

A worker keeps polling after the run’s deadline

Section titled “A worker keeps polling after the run’s deadline”

After a deadline passes, requests fail with DEADLINE_EXCEEDED, which consume() treats as a temporary error and keeps polling through. Stop the worker with its signal / stop_event when your run ends, or cancel the run (cancellation stops consume()).

quota.take() waits much longer than the window

Section titled “quota.take() waits much longer than the window”
  • A cooldown is active: check cooldownUntil in quota.status().
  • The queue is strict FIFO, so a waiter asking for many units holds back smaller requests behind it.
  • Waiters left by crashed processes stay queued until their time-to-live and can be granted units first.

The space has more queued and in-progress messages than its maxBacklog (1,000 by default). Add workers, or raise maxBacklog (per-space limits).

An N_ACCEPTED group finalized right away with threshold_unreachable

Section titled “An N_ACCEPTED group finalized right away with threshold_unreachable”

Groups are re-evaluated in the background. If the group has fewer members that could still be accepted than threshold (including right after creation, before you’ve submitted enough), it fails. Submit at least threshold members immediately after creating the group. See task groups.

RuntimeWarning: coroutine 'Delivery.complete' was never awaited

Section titled “RuntimeWarning: coroutine 'Delivery.complete' was never awaited”

The handler is a plain def, but delivery methods are coroutines. Make the handler async def and await delivery.complete(...). This applies to the blocking Runstate client too.

The blocking Runstate client hangs inside a handler

Section titled “The blocking Runstate client hangs inside a handler”

An async def handler runs on the blocking client’s own event loop, so calling a blocking method from it waits on itself. Inside async def handlers, await the async objects you were given; call blocking methods only from plain def handlers or outside handlers. See the blocking client.

Compute the HMAC over the raw request body exactly as received, prefixed by the x-kernel-timestamp value and a dot, using the signing secret from destination creation. Re-serializing parsed JSON changes the bytes. See Webhooks.