# Operations and monitoring

> How ZoPanel watches your server: service watchdog, alerts by e-mail, Telegram or webhook, uptime and usage history, the Tuning advisor, doctor and logs.

Source: https://zopanel.net/docs/operations  
Updated: 2026-10-07

ZoPanel watches the server for you. A watchdog restarts services that stop or hang, uptime checks find websites that stop answering, and alerts reach you by e-mail, Telegram or webhook. When something goes wrong, `zopanel ctl doctor` tells you what is broken and how to fix it.

## Service watchdog

Every minute the ZoPanel agent checks the installed services. It only watches units that are installed and **enabled** (set to start on boot):

- A service that is **not running** is started again.
- A service that runs but **does not answer** is restarted after its health probe fails **3 times in a row**. This catches services that hang or run out of file descriptors.

| Service | Health probe |
| --- | --- |
| nginx | TCP connection to 127.0.0.1:80 |
| mariadb | Unix socket `/run/mysqld/mysqld.sock` |
| postfix | SMTP greeting (`220`) on 127.0.0.1:25 |
| dovecot | IMAP greeting (`* OK`) on 127.0.0.1:143 |
| rspamd | TCP 127.0.0.1:11332 |
| redis-server | TCP 127.0.0.1:6379 |
| pdns | Its API on 127.0.0.1:8081 |
| zopanel | The panel's `/healthz` check |
| pure-ftpd, postgresql, mongod, docker, fail2ban, cron, ssh, zp-webmail | Running state only |

Safety rules:

- **At most 3 restarts per service per hour.** After that the watchdog gives up and leaves the service to you, and you get a "Service down" alert. It watches the service again once it recovers.
- **nginx is only restarted if `nginx -t` passes.** If the configuration is invalid, you get an alert with the error instead.
- **Package upgrades come first.** The watchdog pauses while `apt` or `dpkg` is running.
- systemd restarts crashed daemons within seconds (ZoPanel adds a drop-in for this, at most 5 times in 10 minutes). The watchdog covers what systemd cannot see, such as a service that hangs.

### Held services

If you stop a service on the **Services** page, the watchdog leaves it stopped. It stays stopped after an agent restart too. The service is watched again once you start it, from the panel or with `systemctl start`. To keep a service off for good, turn off **Start on boot**.

Every watchdog action (restarted, failed, gave up, recovered) is recorded in the **Activity Log**.

## Alerts

Set up alerts in **Settings → Alerts**. Turn on one or more channels:

| Channel | What you enter |
| --- | --- |
| Telegram | Bot token (create a bot with @BotFather) and chat ID (from @userinfobot) |
| Email (SMTP) | SMTP host, port (587 by default; 465 uses TLS from the start), username, password, From, To (comma separated) |
| Webhook (Discord / Slack) | Type (Discord, Slack or JSON) and an `https://` URL |

Choose the events under **Events**, set the thresholds, click **Save**, then click **Send test**. The bot token, SMTP password and webhook URL are never shown again after you save them.

| Event | When it is sent |
| --- | --- |
| Website down | A website fails two checks in a row, and again when it recovers |
| Service down or restarted | The watchdog restarted a service, could not restart it, or gave up |
| High CPU usage / High memory usage | The average over the last minute reaches the CPU % or RAM % threshold (90% by default) |
| Disk almost full | The root filesystem reaches the Disk % threshold (90% by default), or the forecast says it will reach 90% within 14 days |
| SSL expiring soon / SSL renewal failed | A certificate expires within 14 days, or a renewal fails |
| Backup failed | A backup fails, or accounts have no recent automatic backup |
| Mail queue growing | 300 or more messages wait in the mail queue |
| Server IP on a blocklist | Checked once a day |
| Mail sending limit reached | An account reaches its hourly e-mail limit |
| Panel update failed or rolled back | An update failed, was rolled back, or could not be checked |
| Domain expiring soon | 30, 7 and 1 days before a domain's registration expires |
| Fleet server unreachable | A server on the **Servers** page stops answering |
| Login from a new IP, Failed panel logins | Sign-in events |
| Malware detected, Deployment failed, Update available | As named |

All events are on by default except **Failed panel logins** and **Update available**. The same alert is sent at most once every 30 minutes.

Customers and resellers can get their own alerts (website down, SSL, backups, mail limit, malware, disk quota) by e-mail or Telegram in **My account → Notifications**.

For billing systems and automation, **Settings → Hooks** sends signed JSON webhooks when accounts, websites, databases, mail or certificates change. These are separate from alerts.

## Uptime and usage history

**Uptime.** Every minute ZoPanel requests each active website through the local nginx. A 5xx answer or no answer counts as down. A site is marked down after two failures in a row. The website's **Domains** tab shows uptime for 24 hours, 7 days and 30 days, with a bar per day for the last 90 days.

Uptime checks never keep a sleeping site awake: sites whose PHP is idle-stopped are not checked, and a check does not count as a visit. While customers' memory is short, checks pause; you get the memory alert instead.

**Usage history.** Every 5 minutes ZoPanel records each account's CPU (percent of one core) and memory. The graphs show averages and peaks over 24 hours, 7 days, 30 days or 1 year.

- Administrators: **Accounts**, then **Usage history** in an account's menu.
- Customers and resellers: **Resources**.

## The Tuning advisor

**Tuning** (admin menu) gives suggestions based on measurements of this server, computed locally:

- PHP-FPM: whether the PHP workers allowed in total fit in RAM;
- the MariaDB buffer pool size (about 20% of RAM suits a shared hosting server);
- missing swap;
- OPcache turned off for a PHP version;
- memory compression, and the kernel command to get it (see [Performance](/docs/performance));
- nginx `worker_processes` versus CPU cores, and high CPU load;
- WordPress sites without the page cache;
- disk use, with a forecast once it has 5 days of data.

Some suggestions have an **Apply** button: MariaDB buffer size (MariaDB restarts, so databases are unavailable for a few seconds), swap file, nginx workers (reloaded without downtime) and page cache for WordPress sites.

## zopanel ctl doctor

Run this first when something looks wrong. It checks the server and changes nothing:

```bash
zopanel ctl doctor
```

It checks that `zopanel-agent`, `zopanel`, `nginx` and `mariadb` are active, that the agent and the panel health check (port 8888 by default) answer, the panel database integrity, at least 10% free space on `/`, `/home` and `/var`, more than 200 MB of available memory, a panel certificate valid for more than 14 days, NTP sync, the last update's result and that the activity log has not been changed.

Each `[FAIL]` line shows the command or panel page that fixes it. The command exits with status 1 when it finds a problem, so you can use it in scripts.

## zopanel ctl support-bundle

```bash
zopanel ctl support-bundle
```

This writes `/root/zopanel-support-YYYYMMDD-HHMMSS.tar.gz`, readable by root only. It contains the doctor report, software versions, failed units, memory and disk usage, the last 500 log lines of `zopanel`, `zopanel-agent`, nginx, MariaDB, Postfix, Dovecot and `/var/log/nginx/error.log`, and the panel configuration with passwords, secrets, tokens and keys removed.

## Logs

| What | Where |
| --- | --- |
| Panel and agent | `journalctl -u zopanel -u zopanel-agent -n 200` |
| Failed panel logins (used by fail2ban) | `/var/log/zopanel/auth.log` |
| Website access and error logs | `/var/log/zopanel/sites/<domain>.access.log`, `<domain>.error.log`; also on the website's **Logs** tab |
| PHP errors of an account | `~/logs/php<version>_errors.log` in the account's home |
| PHP-FPM pool | `/var/log/zopanel/php/<user>-<version>.log` |
| nginx | `/var/log/nginx/error.log` |
| Actions in the panel | **Activity Log** |

On a website's **Logs** tab, **Diagnostics** (**Run diagnostics**) reads its recent errors and suggests fixes. See [Troubleshooting](/docs/troubleshooting) for common problems.
