DocsOperations and monitoring

Operations and monitoring

How ZoPanel watches your server: service watchdog, alerts by e-mail, Telegram or webhook, uptime and usage history, the Tuning advisor, doctor and logs.

ZoPanel watches the server for you. A watchdog restarts services that stop or hang, uptime checks find websites that stop answering, and alerts reach you by e-mail, Telegram or webhook. When something goes wrong, zopanel ctl doctor tells you what is broken and how to fix it.

Service watchdog

Every minute the ZoPanel agent checks the installed services. It only watches units that are installed and enabled (set to start on boot):

  • A service that is not running is started again.
  • A service that runs but does not answer is restarted after its health probe fails 3 times in a row. This catches services that hang or run out of file descriptors.
Service Health probe
nginx TCP connection to 127.0.0.1:80
mariadb Unix socket /run/mysqld/mysqld.sock
postfix SMTP greeting (220) on 127.0.0.1:25
dovecot IMAP greeting (* OK) on 127.0.0.1:143
rspamd TCP 127.0.0.1:11332
redis-server TCP 127.0.0.1:6379
pdns Its API on 127.0.0.1:8081
zopanel The panel's /healthz check
pure-ftpd, postgresql, mongod, docker, fail2ban, cron, ssh, zp-webmail Running state only

Safety rules:

  • At most 3 restarts per service per hour. After that the watchdog gives up and leaves the service to you, and you get a "Service down" alert. It watches the service again once it recovers.
  • nginx is only restarted if nginx -t passes. If the configuration is invalid, you get an alert with the error instead.
  • Package upgrades come first. The watchdog pauses while apt or dpkg is running.
  • systemd restarts crashed daemons within seconds (ZoPanel adds a drop-in for this, at most 5 times in 10 minutes). The watchdog covers what systemd cannot see, such as a service that hangs.

Held services

If you stop a service on the Services page, the watchdog leaves it stopped. It stays stopped after an agent restart too. The service is watched again once you start it, from the panel or with systemctl start. To keep a service off for good, turn off Start on boot.

Every watchdog action (restarted, failed, gave up, recovered) is recorded in the Activity Log.

Alerts

Set up alerts in Settings → Alerts. Turn on one or more channels:

Channel What you enter
Telegram Bot token (create a bot with @BotFather) and chat ID (from @userinfobot)
Email (SMTP) SMTP host, port (587 by default; 465 uses TLS from the start), username, password, From, To (comma separated)
Webhook (Discord / Slack) Type (Discord, Slack or JSON) and an https:// URL

Choose the events under Events, set the thresholds, click Save, then click Send test. The bot token, SMTP password and webhook URL are never shown again after you save them.

Event When it is sent
Website down A website fails two checks in a row, and again when it recovers
Service down or restarted The watchdog restarted a service, could not restart it, or gave up
High CPU usage / High memory usage The average over the last minute reaches the CPU % or RAM % threshold (90% by default)
Disk almost full The root filesystem reaches the Disk % threshold (90% by default), or the forecast says it will reach 90% within 14 days
SSL expiring soon / SSL renewal failed A certificate expires within 14 days, or a renewal fails
Backup failed A backup fails, or accounts have no recent automatic backup
Mail queue growing 300 or more messages wait in the mail queue
Server IP on a blocklist Checked once a day
Mail sending limit reached An account reaches its hourly e-mail limit
Panel update failed or rolled back An update failed, was rolled back, or could not be checked
Domain expiring soon 30, 7 and 1 days before a domain's registration expires
Fleet server unreachable A server on the Servers page stops answering
Login from a new IP, Failed panel logins Sign-in events
Malware detected, Deployment failed, Update available As named

All events are on by default except Failed panel logins and Update available. The same alert is sent at most once every 30 minutes.

Customers and resellers can get their own alerts (website down, SSL, backups, mail limit, malware, disk quota) by e-mail or Telegram in My account → Notifications.

For billing systems and automation, Settings → Hooks sends signed JSON webhooks when accounts, websites, databases, mail or certificates change. These are separate from alerts.

Uptime and usage history

Uptime. Every minute ZoPanel requests each active website through the local nginx. A 5xx answer or no answer counts as down. A site is marked down after two failures in a row. The website's Domains tab shows uptime for 24 hours, 7 days and 30 days, with a bar per day for the last 90 days.

Uptime checks never keep a sleeping site awake: sites whose PHP is idle-stopped are not checked, and a check does not count as a visit. While customers' memory is short, checks pause; you get the memory alert instead.

Usage history. Every 5 minutes ZoPanel records each account's CPU (percent of one core) and memory. The graphs show averages and peaks over 24 hours, 7 days, 30 days or 1 year.

  • Administrators: Accounts, then Usage history in an account's menu.
  • Customers and resellers: Resources.

The Tuning advisor

Tuning (admin menu) gives suggestions based on measurements of this server, computed locally:

  • PHP-FPM: whether the PHP workers allowed in total fit in RAM;
  • the MariaDB buffer pool size (about 20% of RAM suits a shared hosting server);
  • missing swap;
  • OPcache turned off for a PHP version;
  • memory compression, and the kernel command to get it (see Performance);
  • nginx worker_processes versus CPU cores, and high CPU load;
  • WordPress sites without the page cache;
  • disk use, with a forecast once it has 5 days of data.

Some suggestions have an Apply button: MariaDB buffer size (MariaDB restarts, so databases are unavailable for a few seconds), swap file, nginx workers (reloaded without downtime) and page cache for WordPress sites.

zopanel ctl doctor

Run this first when something looks wrong. It checks the server and changes nothing:

zopanel ctl doctor

It checks that zopanel-agent, zopanel, nginx and mariadb are active, that the agent and the panel health check (port 8888 by default) answer, the panel database integrity, at least 10% free space on /, /home and /var, more than 200 MB of available memory, a panel certificate valid for more than 14 days, NTP sync, the last update's result and that the activity log has not been changed.

Each [FAIL] line shows the command or panel page that fixes it. The command exits with status 1 when it finds a problem, so you can use it in scripts.

zopanel ctl support-bundle

zopanel ctl support-bundle

This writes /root/zopanel-support-YYYYMMDD-HHMMSS.tar.gz, readable by root only. It contains the doctor report, software versions, failed units, memory and disk usage, the last 500 log lines of zopanel, zopanel-agent, nginx, MariaDB, Postfix, Dovecot and /var/log/nginx/error.log, and the panel configuration with passwords, secrets, tokens and keys removed.

Logs

What Where
Panel and agent journalctl -u zopanel -u zopanel-agent -n 200
Failed panel logins (used by fail2ban) /var/log/zopanel/auth.log
Website access and error logs /var/log/zopanel/sites/<domain>.access.log, <domain>.error.log; also on the website's Logs tab
PHP errors of an account ~/logs/php<version>_errors.log in the account's home
PHP-FPM pool /var/log/zopanel/php/<user>-<version>.log
nginx /var/log/nginx/error.log
Actions in the panel Activity Log

On a website's Logs tab, Diagnostics (Run diagnostics) reads its recent errors and suggests fixes. See Troubleshooting for common problems.


← Fleet and multiple servers Performance and capacity →