Overview
Findings are structured health observations. Each one includes:- Severity — how urgent the issue is
- Evidence — the specific metrics that triggered the finding
- Probable cause — the most likely reasons this is happening
- Confirm steps — commands to verify the root cause
- Safe fix — what to do about it
REDIS_DOWN
Severity: CRITICAL
- Redis process is down
- Network or firewall blocking the connection
- Wrong
REDIS_URL
NO_WORKERS
Severity: CRITICAL
- All workers have crashed or been stopped
CELERY_BROKER_URLdoesn’t match the running broker- Workers are running but the broker is unreachable from their side
WORKERS_MISSING
Severity: CRITICAL
worker_offline_grace_seconds (default: 90 s). The agent reports missing_workers in the payload, and the backend turns any missing_workers > 0 into a CRITICAL alert.
How it works: the long-running agent (kanari run) tracks a high-water mark of the alive worker count. When the current count falls below that baseline for longer than the grace period, a WORKERS_MISSING finding is raised. Detection is count-based, not name-based — see Known limitations below.
Payload fields:
expected_workers— the high-water baseline countmissing_workers— how many workers are currently absent
- One or more worker processes crashed (OOM, unhandled exception)
- Workers were stopped for a deployment but not yet restarted
- Network partition between workers and the broker
WORKER_OFFLINE
Severity: CRITICAL
- Worker process crashed (check OOM, unhandled exception)
- Worker was stopped and not restarted
- Network partition between worker and broker
STUCK_TASK
Severity: HIGH
max_task_runtime_seconds (default: 30 minutes). The worker slot occupied by this task is blocked.
Probable causes:
- Task is deadlocked waiting on an external resource (database lock, HTTP timeout)
- Infinite loop in task code
- Missing
time_limiton a task that can legitimately hang
QUEUE_BACKLOG_*
Severity: HIGH (critical queues) / MEDIUM (regular queues)
max_queue_size. Tasks are accumulating faster than workers can process them.
The finding ID includes the queue name: QUEUE_BACKLOG_EMAILS, QUEUE_BACKLOG_CELERY, etc.
Probable causes:
- Insufficient worker capacity for current load
- Spike in task production (traffic burst, batch job)
- Worker(s) offline or processing very slowly
- All workers busy with a different queue
QUEUE_SLA_BREACH_*
Severity: HIGH
max_wait_time_seconds. Requires KanariStampPlugin for accurate measurement.
The finding ID includes the queue name: QUEUE_SLA_BREACH_EMAILS, etc.
Probable causes:
- Workers not consuming from this specific queue (routing misconfiguration)
- All worker slots full with other work
- Worker process crashed silently
LATENCY_UNAVAILABLE
Severity: MEDIUM
QUEUE_SLA_BREACH detection is blind.
Cause:
Celery + Redis doesn’t add timestamps to queue messages by default. task_send_sent_event=True does not help — it emits events to the Celery event stream, not to the queue messages that Kanari reads.
Fix:
Add one line to your Celery app:
HIGH_SATURATION
Severity: MEDIUM
saturation_pct, active, concurrency
Probable causes:
- Sustained high load without enough worker capacity
- Long-running tasks consuming all slots
worker_prefetch_multiplier > 1causing tasks to be reserved but not executing
Known limitations: worker-offline detection
Worker-offline detection is count-based and lives in the agent’s memory. Known gaps:- Baseline resets on agent restart. The high-water mark is in-memory. After a restart the baseline starts from the currently-alive workers, so a worker that died while the agent was down is not detected until the fleet is seen complete again. Persisting to disk is a deliberate non-goal — it does not resolve the crash-vs-scaledown ambiguity during downtime.
- Count-based, not name-based. We alert on how many workers are missing, not which one. This is intentional: it stays robust when worker names are ephemeral (Kubernetes, autoscaling), where name-tracking would produce constant churn.
- Crash vs. intentional scale-down is indistinguishable without an external
signal. Default is fail-loud (keep alerting); set
worker_auto_resolve_secondsto re-baseline automatically after a window. - One-shot
kanari auditcannot detect missing workers. It has no cross-cycle history, so its baseline equals the current alive count. Only the long-running agent (kanari run) accumulates a baseline.