> ## Documentation Index
> Fetch the complete documentation index at: https://getkanari.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Findings Reference

> Every finding Kanari can raise, with causes and fixes

## Overview

Findings are structured health observations. Each one includes:

* **Severity** — how urgent the issue is
* **Evidence** — the specific metrics that triggered the finding
* **Probable cause** — the most likely reasons this is happening
* **Confirm steps** — commands to verify the root cause
* **Safe fix** — what to do about it

Findings drive the system status and exit code:

| Highest severity | System status | Exit code |
| - | - | - |
| CRITICAL | CRITICAL | 2 |
| HIGH | DEGRADED | 2 |
| MEDIUM or LOW | WARNINGS | 1 |
| No findings | OK | 0 |

***

## REDIS\_DOWN

<Info>Severity: **CRITICAL**</Info>

Kanari cannot connect to Redis. No queue metrics are available.

**Probable causes:**

* Redis process is down
* Network or firewall blocking the connection
* Wrong `REDIS_URL`

**Confirm:**

```bash theme={null}
redis-cli -u $REDIS_URL ping
# Expected: PONG
```

**Fix:**

```bash theme={null}
# Check Redis status
systemctl status redis

# Verify URL
echo $REDIS_URL

# Restart if needed
systemctl restart redis
```

***

## NO\_WORKERS

<Info>Severity: **CRITICAL**</Info>

No Celery workers are responding. Tasks will queue indefinitely with no processing capacity.

**Probable causes:**

* All workers have crashed or been stopped
* `CELERY_BROKER_URL` doesn't match the running broker
* Workers are running but the broker is unreachable from their side

**Confirm:**

```bash theme={null}
celery -A your_app inspect ping
```

**Fix:**

```bash theme={null}
# Start a worker
celery -A your_app worker --loglevel=info

# Check broker URL matches workers
echo $CELERY_BROKER_URL
```

***

## WORKERS\_MISSING

<Info>Severity: **CRITICAL**</Info>

The number of alive workers has dropped below the observed peak (baseline) and the deficit has persisted longer than `worker_offline_grace_seconds` (default: 90 s). The agent reports `missing_workers` in the payload, and the backend turns any `missing_workers > 0` into a CRITICAL alert.

**How it works:** the long-running agent (`kanari run`) tracks a high-water mark of the alive worker count. When the current count falls below that baseline for longer than the grace period, a `WORKERS_MISSING` finding is raised. Detection is count-based, not name-based — see [Known limitations](#known-limitations-worker-offline-detection) below.

**Payload fields:**

* `expected_workers` — the high-water baseline count
* `missing_workers` — how many workers are currently absent

**Probable causes:**

* One or more worker processes crashed (OOM, unhandled exception)
* Workers were stopped for a deployment but not yet restarted
* Network partition between workers and the broker

**Confirm:**

```bash theme={null}
celery -A your_app inspect ping
celery -A your_app inspect active_queues
```

**Fix:**

```bash theme={null}
# Check for OOM kill
dmesg | grep -i "killed process"
journalctl -u celery --since "1 hour ago" | grep -i "error\|killed"

# Restart workers
celery -A your_app worker --loglevel=info
```

**Config knobs:**

```yaml theme={null}
thresholds:
  worker_offline_grace_seconds: 90   # how long a drop must persist before alerting
  worker_auto_resolve_seconds: null  # null = alert until workers return; set e.g. 1800 to auto-resolve
```

***

## WORKER\_OFFLINE

<Info>Severity: **CRITICAL**</Info>

A specific worker that was previously registered is not responding to pings.

**Probable causes:**

* Worker process crashed (check OOM, unhandled exception)
* Worker was stopped and not restarted
* Network partition between worker and broker

**Confirm:**

```bash theme={null}
celery -A your_app inspect ping -d worker-name@hostname
```

**Fix:**

```bash theme={null}
# Check for OOM kill
dmesg | grep -i "killed process"
journalctl -u celery --since "1 hour ago" | grep -i "error\|killed"

# Restart the specific worker
celery -A your_app worker --hostname=worker-name@hostname --loglevel=info
```

***

## STUCK\_TASK

<Info>Severity: **HIGH**</Info>

A task has been running longer than `max_task_runtime_seconds` (default: 30 minutes). The worker slot occupied by this task is blocked.

**Probable causes:**

* Task is deadlocked waiting on an external resource (database lock, HTTP timeout)
* Infinite loop in task code
* Missing `time_limit` on a task that can legitimately hang

**Confirm:**

```bash theme={null}
# See what's active on the stuck worker
celery -A your_app inspect active -d stuck-worker@hostname

# Check task logs for the task ID shown in evidence
```

**Fix:**

```bash theme={null}
# Add a timeout to prevent future occurrences
@app.task(time_limit=300, soft_time_limit=270)
def your_task():
    ...

# Revoke the currently stuck task
celery -A your_app control revoke <task_id> --terminate
```

***

## QUEUE\_BACKLOG\_\*

<Info>Severity: **HIGH** (critical queues) / **MEDIUM** (regular queues)</Info>

A queue's depth has exceeded `max_queue_size`. Tasks are accumulating faster than workers can process them.

The finding ID includes the queue name: `QUEUE_BACKLOG_EMAILS`, `QUEUE_BACKLOG_CELERY`, etc.

**Probable causes:**

* Insufficient worker capacity for current load
* Spike in task production (traffic burst, batch job)
* Worker(s) offline or processing very slowly
* All workers busy with a different queue

**Confirm:**

```bash theme={null}
redis-cli LLEN queue-name
celery -A your_app inspect active
```

**Fix:**

```bash theme={null}
# Scale workers immediately
celery -A your_app worker --concurrency=8 --loglevel=info &

# Or increase concurrency on existing workers (via autoscale)
celery -A your_app worker --autoscale=10,2
```

***

## QUEUE\_SLA\_BREACH\_\*

<Info>Severity: **HIGH**</Info>

The oldest task in a queue has been waiting longer than `max_wait_time_seconds`. Requires [KanariStampPlugin](/docs/guides/latency-tracking) for accurate measurement.

The finding ID includes the queue name: `QUEUE_SLA_BREACH_EMAILS`, etc.

**Probable causes:**

* Workers not consuming from this specific queue (routing misconfiguration)
* All worker slots full with other work
* Worker process crashed silently

**Confirm:**

```bash theme={null}
# Check which queues workers are subscribed to
celery -A your_app inspect active_queues

# Inspect the oldest raw message
redis-cli LINDEX queue-name -1 | python3 -m json.tool
```

**Fix:**

```bash theme={null}
# Verify workers are subscribed to the right queues
celery -A your_app worker -Q emails,celery --loglevel=info

# Scale workers for this specific queue
celery -A your_app worker -Q emails --concurrency=4 --loglevel=info
```

***

## LATENCY\_UNAVAILABLE

<Info>Severity: **MEDIUM**</Info>

Queues have tasks pending but no timestamps are present in the messages. Kanari cannot measure how long tasks have been waiting, so `QUEUE_SLA_BREACH` detection is blind.

**Cause:**
Celery + Redis doesn't add timestamps to queue messages by default. `task_send_sent_event=True` does not help — it emits events to the Celery event stream, not to the queue messages that Kanari reads.

**Fix:**

Add one line to your Celery app:

```python theme={null}
from kanari_agent.stamps import KanariStampPlugin
KanariStampPlugin.install(app)
```

See the [Latency Tracking guide](/docs/guides/latency-tracking) for full instructions.

***

## HIGH\_SATURATION

<Info>Severity: **MEDIUM**</Info>

Workers are above 80% utilization. New tasks will queue instead of executing immediately. If load continues, queues will start growing.

**Evidence fields:** `saturation_pct`, `active`, `concurrency`

**Probable causes:**

* Sustained high load without enough worker capacity
* Long-running tasks consuming all slots
* `worker_prefetch_multiplier > 1` causing tasks to be reserved but not executing

**Fix:**

```bash theme={null}
# Add more workers
celery -A your_app worker --concurrency=8 &

# Or reduce prefetch to free up slots faster
# In celery config:
worker_prefetch_multiplier = 1
```

***

## Known limitations: worker-offline detection

Worker-offline detection is count-based and lives in the agent's memory. Known gaps:

* **Baseline resets on agent restart.** The high-water mark is in-memory. After a
  restart the baseline starts from the currently-alive workers, so a worker that
  died while the agent was down is not detected until the fleet is seen complete
  again. Persisting to disk is a deliberate non-goal — it does not resolve the
  crash-vs-scaledown ambiguity during downtime.
* **Count-based, not name-based.** We alert on *how many* workers are missing, not
  *which* one. This is intentional: it stays robust when worker names are ephemeral
  (Kubernetes, autoscaling), where name-tracking would produce constant churn.
* **Crash vs. intentional scale-down is indistinguishable** without an external
  signal. Default is fail-loud (keep alerting); set `worker_auto_resolve_seconds`
  to re-baseline automatically after a window.
* **One-shot `kanari audit` cannot detect missing workers.** It has no cross-cycle
  history, so its baseline equals the current alive count. Only the long-running
  agent (`kanari run`) accumulates a baseline.

Feedback and discussion welcome — these tradeoffs are deliberately conservative.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.