Skip to content
← Back to case studies

Real-time observability and alerting

Deployment of Zabbix, Nagios, and Grafana with instant Telegram bot alerts for critical routers and servers. Before: no unified visibility, reactive failure detection dependent on user reports.

Key metrics

Key facts

  • Scope: core routers and servers in production (200+ hosts).
  • Alerts: notifications in under 60 s via Telegram.
  • Goal: operational visibility, lower MTTR, and fast response.

Stack

  • Zabbix
  • Nagios
  • Grafana
  • Telegram Bot
Sanitized observability diagram
  1. Routers/servers
  2. Zabbix/Nagios
  3. Grafana
  4. Telegram alerts

Logical view without hosts, IPs, internal thresholds, or sensitive service names.

Context

The ISP's operation relied on manual failure detection. Alerts arrived through customer calls, not monitoring. There was no unified dashboard and no response-time metrics. The on-call team needed real-time visibility and actionable alerts.

Actions

Anonymized global Zabbix view with problems, hosts, and monitoring metrics
Anonymized Zabbix: global view of problems, hosts, and operational status.
Anonymized Nagios summary with host groups, services, and check status
Anonymized Nagios: status summary for critical hosts and services.
Anonymized Grafana dashboard with traffic, CPU, memory, and storage metrics
Anonymized Grafana: operational metrics for traffic, resources, and availability.

Results