Skip to main content
Two separate things happen when part of your app fails. The orchestrator recovers managed workloads within its policies. Watcher, if you set it up, notices problems, writes a diagnosis and tells you. They are different features, and only the orchestrator acts without asking.

Orchestrator recovery

The orchestrator (monkd) on your cluster restarts and reschedules the workloads it manages, within its configured policies, retry limits and the capacity available on your nodes. That happens whether or not Watcher is installed. Some limits:
  • Recovery doesn’t go beyond its policies. When a workload keeps crashing because of a bug or a bad setting, restarting it doesn’t fix the cause.
  • A failed health check on its own doesn’t always mean a restart.
  • Managed services created as entities, such as a cloud database, recover according to their provider’s own behavior, not Monk’s.

Watcher

Watcher is a monitoring stack that runs on your cluster, in the system/watcher group on the node with the system tag. It has two parts:
  • watcher-agent polls the health of nodes and workloads and detects crashes and threshold breaches.
  • watcher-ai takes those alerts, adds context from logs and status, and writes an assessment with a recommended next step.
Watcher reports. It doesn’t change your deployment. Any fix it proposes needs a person to approve it, and the fix is then carried out through your coding agent and Monk, with the usual plan review in the local dashboard.

What it watches, and the defaults

You can change any of these at setup, or later by running setup again.

Slack notifications

Slack is Watcher’s notification channel. Alerts go to a Slack Incoming Webhook, which Monk asks for through a credential form in the local dashboard. By default only the AI-refined alerts are sent, which keeps the channel quieter. Watcher still runs without Slack, but nothing is pushed to you. Slack is one-way. You can’t chat with Monk or approve changes from Slack. An alert can include a Fix with Monk link. It opens your coding agent with the alert’s context, so you can ask for a fix. Whatever your coding agent proposes then goes through the normal plan and approval.

Setting it up

Setup creates a cluster service token and a scoped API key and stores them as secrets on the cluster. You review the plan in the local dashboard before anything is deployed. Watcher needs an active cluster. Monk can also report whether Watcher is running and remove it, which asks for approval. Watcher’s monitoring and diagnosis use watcher credits from your plan. See Pricing.

What neither of them does

  • Neither one scales your app or resizes machines. Ask Monk to do that. See Scale and resize.
  • Neither one rolls back a release. See Deployments.
  • Neither one repairs data.

Watcher and alerts

Set up Watcher and Slack

Logs and debugging

Status and logs on demand

Deployments

Redeploys and rollbacks

Security

Approvals and credentials