Operations and monitoring
How ZoPanel watches your server: service watchdog, alerts by e-mail, Telegram or webhook, uptime and usage history, the Tuning advisor, doctor and logs.
ZoPanel watches the server for you. A watchdog restarts services that stop or hang, uptime checks find websites that stop answering, and alerts reach you by e-mail, Telegram or webhook. When something goes wrong, zopanel ctl doctor tells you what is broken and how to fix it.
Service watchdog
Every minute the ZoPanel agent checks the installed services. It only watches units that are installed and enabled (set to start on boot):
- A service that is not running is started again.
- A service that runs but does not answer is restarted after its health probe fails 3 times in a row. This catches services that hang or run out of file descriptors.
| Service | Health probe |
|---|---|
| nginx | TCP connection to 127.0.0.1:80 |
| mariadb | Unix socket /run/mysqld/mysqld.sock |
| postfix | SMTP greeting (220) on 127.0.0.1:25 |
| dovecot | IMAP greeting (* OK) on 127.0.0.1:143 |
| rspamd | TCP 127.0.0.1:11332 |
| redis-server | TCP 127.0.0.1:6379 |
| pdns | Its API on 127.0.0.1:8081 |
| zopanel | The panel's /healthz check |
| pure-ftpd, postgresql, mongod, docker, fail2ban, cron, ssh, zp-webmail | Running state only |
Safety rules:
- At most 3 restarts per service per hour. After that the watchdog gives up and leaves the service to you, and you get a "Service down" alert. It watches the service again once it recovers.
- nginx is only restarted if
nginx -tpasses. If the configuration is invalid, you get an alert with the error instead. - Package upgrades come first. The watchdog pauses while
aptordpkgis running. - systemd restarts crashed daemons within seconds (ZoPanel adds a drop-in for this, at most 5 times in 10 minutes). The watchdog covers what systemd cannot see, such as a service that hangs.
Held services
If you stop a service on the Services page, the watchdog leaves it stopped. It stays stopped after an agent restart too. The service is watched again once you start it, from the panel or with systemctl start. To keep a service off for good, turn off Start on boot.
Every watchdog action (restarted, failed, gave up, recovered) is recorded in the Activity Log.
Alerts
Set up alerts in Settings → Alerts. Turn on one or more channels:
| Channel | What you enter |
|---|---|
| Telegram | Bot token (create a bot with @BotFather) and chat ID (from @userinfobot) |
| Email (SMTP) | SMTP host, port (587 by default; 465 uses TLS from the start), username, password, From, To (comma separated) |
| Webhook (Discord / Slack) | Type (Discord, Slack or JSON) and an https:// URL |
Choose the events under Events, set the thresholds, click Save, then click Send test. The bot token, SMTP password and webhook URL are never shown again after you save them.
| Event | When it is sent |
|---|---|
| Website down | A website fails two checks in a row, and again when it recovers |
| Service down or restarted | The watchdog restarted a service, could not restart it, or gave up |
| High CPU usage / High memory usage | The average over the last minute reaches the CPU % or RAM % threshold (90% by default) |
| Disk almost full | The root filesystem reaches the Disk % threshold (90% by default), or the forecast says it will reach 90% within 14 days |
| SSL expiring soon / SSL renewal failed | A certificate expires within 14 days, or a renewal fails |
| Backup failed | A backup fails, or accounts have no recent automatic backup |
| Mail queue growing | 300 or more messages wait in the mail queue |
| Server IP on a blocklist | Checked once a day |
| Mail sending limit reached | An account reaches its hourly e-mail limit |
| Panel update failed or rolled back | An update failed, was rolled back, or could not be checked |
| Domain expiring soon | 30, 7 and 1 days before a domain's registration expires |
| Fleet server unreachable | A server on the Servers page stops answering |
| Login from a new IP, Failed panel logins | Sign-in events |
| Malware detected, Deployment failed, Update available | As named |
All events are on by default except Failed panel logins and Update available. The same alert is sent at most once every 30 minutes.
Customers and resellers can get their own alerts (website down, SSL, backups, mail limit, malware, disk quota) by e-mail or Telegram in My account → Notifications.
For billing systems and automation, Settings → Hooks sends signed JSON webhooks when accounts, websites, databases, mail or certificates change. These are separate from alerts.
Uptime and usage history
Uptime. Every minute ZoPanel requests each active website through the local nginx. A 5xx answer or no answer counts as down. A site is marked down after two failures in a row. The website's Domains tab shows uptime for 24 hours, 7 days and 30 days, with a bar per day for the last 90 days.
Uptime checks never keep a sleeping site awake: sites whose PHP is idle-stopped are not checked, and a check does not count as a visit. While customers' memory is short, checks pause; you get the memory alert instead.
Usage history. Every 5 minutes ZoPanel records each account's CPU (percent of one core) and memory. The graphs show averages and peaks over 24 hours, 7 days, 30 days or 1 year.
- Administrators: Accounts, then Usage history in an account's menu.
- Customers and resellers: Resources.
The Tuning advisor
Tuning (admin menu) gives suggestions based on measurements of this server, computed locally:
- PHP-FPM: whether the PHP workers allowed in total fit in RAM;
- the MariaDB buffer pool size (about 20% of RAM suits a shared hosting server);
- missing swap;
- OPcache turned off for a PHP version;
- memory compression, and the kernel command to get it (see Performance);
- nginx
worker_processesversus CPU cores, and high CPU load; - WordPress sites without the page cache;
- disk use, with a forecast once it has 5 days of data.
Some suggestions have an Apply button: MariaDB buffer size (MariaDB restarts, so databases are unavailable for a few seconds), swap file, nginx workers (reloaded without downtime) and page cache for WordPress sites.
zopanel ctl doctor
Run this first when something looks wrong. It checks the server and changes nothing:
zopanel ctl doctor
It checks that zopanel-agent, zopanel, nginx and mariadb are active, that the agent and the panel health check (port 8888 by default) answer, the panel database integrity, at least 10% free space on /, /home and /var, more than 200 MB of available memory, a panel certificate valid for more than 14 days, NTP sync, the last update's result and that the activity log has not been changed.
Each [FAIL] line shows the command or panel page that fixes it. The command exits with status 1 when it finds a problem, so you can use it in scripts.
zopanel ctl support-bundle
zopanel ctl support-bundle
This writes /root/zopanel-support-YYYYMMDD-HHMMSS.tar.gz, readable by root only. It contains the doctor report, software versions, failed units, memory and disk usage, the last 500 log lines of zopanel, zopanel-agent, nginx, MariaDB, Postfix, Dovecot and /var/log/nginx/error.log, and the panel configuration with passwords, secrets, tokens and keys removed.
Logs
| What | Where |
|---|---|
| Panel and agent | journalctl -u zopanel -u zopanel-agent -n 200 |
| Failed panel logins (used by fail2ban) | /var/log/zopanel/auth.log |
| Website access and error logs | /var/log/zopanel/sites/<domain>.access.log, <domain>.error.log; also on the website's Logs tab |
| PHP errors of an account | ~/logs/php<version>_errors.log in the account's home |
| PHP-FPM pool | /var/log/zopanel/php/<user>-<version>.log |
| nginx | /var/log/nginx/error.log |
| Actions in the panel | Activity Log |
On a website's Logs tab, Diagnostics (Run diagnostics) reads its recent errors and suggests fixes. See Troubleshooting for common problems.