> ## Documentation Index
> Fetch the complete documentation index at: https://docs.monk.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Watcher and recovery

> What the orchestrator recovers by itself, what Watcher adds, and why Watcher's fixes need your approval

Two separate things happen when part of your app fails. The **orchestrator** recovers managed workloads within its policies. **Watcher**, if you set it up, notices problems, writes a diagnosis and tells you. They are different features, and only the orchestrator acts without asking.

## Orchestrator recovery

The orchestrator (`monkd`) on your cluster restarts and reschedules the workloads it manages, within its configured policies, retry limits and the capacity available on your nodes. That happens whether or not Watcher is installed.

Some limits:

* Recovery doesn't go beyond its policies. When a workload keeps crashing because of a bug or a bad setting, restarting it doesn't fix the cause.
* A failed health check on its own doesn't always mean a restart.
* Managed services created as entities, such as a cloud database, recover according to their provider's own behavior, not Monk's.

## Watcher

Watcher is a monitoring stack that runs on your cluster, in the `system/watcher` group on the node with the system tag. It has two parts:

* **watcher-agent** polls the health of nodes and workloads and detects crashes and threshold breaches.
* **watcher-ai** takes those alerts, adds context from logs and status, and writes an assessment with a recommended next step.

Watcher reports. It doesn't change your deployment. **Any fix it proposes needs a person to approve it**, and the fix is then carried out through your coding agent and Monk, with the usual plan review in the local dashboard.

### What it watches, and the defaults

| Setting | Default |
| - | - |
| Container restarts before a crash alert | 3 within 5 minutes |
| Consecutive failed health checks | 3 |
| Node CPU | 70% for 5 minutes |
| Node memory | 80% for 5 minutes |
| Node disk | 85%, 2 checks in a row |
| Workload CPU | 70% for 5 minutes |
| Workload memory | 80% for 5 minutes |
| Workload disk | 90%, 3 checks in a row |
| Poll interval | 15 seconds |
| Log lines collected per alert | 100 |
| Alert context kept for | 24 hours |
| Skip the local development node | On |

You can change any of these at setup, or later by running setup again.

### Slack notifications

Slack is Watcher's notification channel. Alerts go to a Slack Incoming Webhook, which Monk asks for through a credential form in the [local dashboard](/getting-started/local-dashboard). By default only the AI-refined alerts are sent, which keeps the channel quieter. Watcher still runs without Slack, but nothing is pushed to you.

Slack is one-way. You can't chat with Monk or approve changes from Slack.

An alert can include a **Fix with Monk** link. It opens your coding agent with the alert's context, so you can ask for a fix. Whatever your coding agent proposes then goes through the normal plan and approval.

### Setting it up

```
/monk set up watcher
```

Setup creates a cluster service token and a scoped API key and stores them as secrets on the cluster. You review the plan in the local dashboard before anything is deployed. Watcher needs an active cluster. Monk can also report whether Watcher is running and remove it, which asks for approval.

Watcher's monitoring and diagnosis use watcher credits from your plan. See [Pricing](https://monk.io/pricing).

## What neither of them does

* Neither one scales your app or resizes machines. Ask Monk to do that. See [Scale and resize](/guides/scale-and-resize).
* Neither one rolls back a release. See [Deployments](/concepts/deployments).
* Neither one repairs data.

## Related

<CardGroup cols={2}>
  <Card title="Watcher and alerts" icon="bell" href="/guides/watcher-and-alerts">
    Set up Watcher and Slack
  </Card>

  <Card title="Logs and debugging" icon="magnifying-glass" href="/guides/logs-and-debugging">
    Status and logs on demand
  </Card>

  <Card title="Deployments" icon="rocket" href="/concepts/deployments">
    Redeploys and rollbacks
  </Card>

  <Card title="Security" icon="shield" href="/concepts/security">
    Approvals and credentials
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.