Skip to main content

Overview

Findings are structured health observations. Each one includes:
  • Severity — how urgent the issue is
  • Evidence — the specific metrics that triggered the finding
  • Probable cause — the most likely reasons this is happening
  • Confirm steps — commands to verify the root cause
  • Safe fix — what to do about it
Findings drive the system status and exit code:

REDIS_DOWN

Severity: CRITICAL
Kanari cannot connect to Redis. No queue metrics are available. Probable causes:
  • Redis process is down
  • Network or firewall blocking the connection
  • Wrong REDIS_URL
Confirm:
Fix:

NO_WORKERS

Severity: CRITICAL
No Celery workers are responding. Tasks will queue indefinitely with no processing capacity. Probable causes:
  • All workers have crashed or been stopped
  • CELERY_BROKER_URL doesn’t match the running broker
  • Workers are running but the broker is unreachable from their side
Confirm:
Fix:

WORKERS_MISSING

Severity: CRITICAL
The number of alive workers has dropped below the observed peak (baseline) and the deficit has persisted longer than worker_offline_grace_seconds (default: 90 s). The agent reports missing_workers in the payload, and the backend turns any missing_workers > 0 into a CRITICAL alert. How it works: the long-running agent (kanari run) tracks a high-water mark of the alive worker count. When the current count falls below that baseline for longer than the grace period, a WORKERS_MISSING finding is raised. Detection is count-based, not name-based — see Known limitations below. Payload fields:
  • expected_workers — the high-water baseline count
  • missing_workers — how many workers are currently absent
Probable causes:
  • One or more worker processes crashed (OOM, unhandled exception)
  • Workers were stopped for a deployment but not yet restarted
  • Network partition between workers and the broker
Confirm:
Fix:
Config knobs:

WORKER_OFFLINE

Severity: CRITICAL
A specific worker that was previously registered is not responding to pings. Probable causes:
  • Worker process crashed (check OOM, unhandled exception)
  • Worker was stopped and not restarted
  • Network partition between worker and broker
Confirm:
Fix:

STUCK_TASK

Severity: HIGH
A task has been running longer than max_task_runtime_seconds (default: 30 minutes). The worker slot occupied by this task is blocked. Probable causes:
  • Task is deadlocked waiting on an external resource (database lock, HTTP timeout)
  • Infinite loop in task code
  • Missing time_limit on a task that can legitimately hang
Confirm:
Fix:

QUEUE_BACKLOG_*

Severity: HIGH (critical queues) / MEDIUM (regular queues)
A queue’s depth has exceeded max_queue_size. Tasks are accumulating faster than workers can process them. The finding ID includes the queue name: QUEUE_BACKLOG_EMAILS, QUEUE_BACKLOG_CELERY, etc. Probable causes:
  • Insufficient worker capacity for current load
  • Spike in task production (traffic burst, batch job)
  • Worker(s) offline or processing very slowly
  • All workers busy with a different queue
Confirm:
Fix:

QUEUE_SLA_BREACH_*

Severity: HIGH
The oldest task in a queue has been waiting longer than max_wait_time_seconds. Requires KanariStampPlugin for accurate measurement. The finding ID includes the queue name: QUEUE_SLA_BREACH_EMAILS, etc. Probable causes:
  • Workers not consuming from this specific queue (routing misconfiguration)
  • All worker slots full with other work
  • Worker process crashed silently
Confirm:
Fix:

LATENCY_UNAVAILABLE

Severity: MEDIUM
Queues have tasks pending but no timestamps are present in the messages. Kanari cannot measure how long tasks have been waiting, so QUEUE_SLA_BREACH detection is blind. Cause: Celery + Redis doesn’t add timestamps to queue messages by default. task_send_sent_event=True does not help — it emits events to the Celery event stream, not to the queue messages that Kanari reads. Fix: Add one line to your Celery app:
See the Latency Tracking guide for full instructions.

HIGH_SATURATION

Severity: MEDIUM
Workers are above 80% utilization. New tasks will queue instead of executing immediately. If load continues, queues will start growing. Evidence fields: saturation_pct, active, concurrency Probable causes:
  • Sustained high load without enough worker capacity
  • Long-running tasks consuming all slots
  • worker_prefetch_multiplier > 1 causing tasks to be reserved but not executing
Fix:

Known limitations: worker-offline detection

Worker-offline detection is count-based and lives in the agent’s memory. Known gaps:
  • Baseline resets on agent restart. The high-water mark is in-memory. After a restart the baseline starts from the currently-alive workers, so a worker that died while the agent was down is not detected until the fleet is seen complete again. Persisting to disk is a deliberate non-goal — it does not resolve the crash-vs-scaledown ambiguity during downtime.
  • Count-based, not name-based. We alert on how many workers are missing, not which one. This is intentional: it stays robust when worker names are ephemeral (Kubernetes, autoscaling), where name-tracking would produce constant churn.
  • Crash vs. intentional scale-down is indistinguishable without an external signal. Default is fail-loud (keep alerting); set worker_auto_resolve_seconds to re-baseline automatically after a window.
  • One-shot kanari audit cannot detect missing workers. It has no cross-cycle history, so its baseline equals the current alive count. Only the long-running agent (kanari run) accumulates a baseline.
Feedback and discussion welcome — these tradeoffs are deliberately conservative.