02 Observability & Automation

Server Monitoring & Observability System

Real-time telemetry pipeline capturing host health, container metrics, and network latency with automated webhook alerts to Discord prior to client-visible degradation.

Role
Systems & Automation Engineer
Technologies
Node.js · Redis · WebSocket · Bash Scripting · Discord Webhook API · Systemd
Key Result
Sub-Second Metric Sampling Pipeline · Proactive Incident Alerting Before Escalations · Automated Node Heartbeat Tracking
Server Monitoring & Observability System preview
Observability & Automation Systems & Automation Engineer

Context & Overview

In high-concurrency game hosting and virtual server environments, operational incidents often manifest abruptly — a runaway Java thread or saturated disk I/O can degrade latency for hundreds of active players within seconds. Heavy commercial monitoring agents introduced undesirable memory footprints on budget VPS slices. This necessitated a bespoke, lightweight telemetry collector and real-time dispatcher engineered specifically for host health and rapid discord-centric alert delivery.

My Role & Responsibilities

  • Designed the architecture for distributed telemetry polling across multiple host nodes.
  • Developed low-overhead shell collectors querying /proc and kernel counters directly to prevent agent CPU hogging.
  • Built a centralized aggregation service using Node.js and Redis pub/sub.
  • Created intelligent webhook alert pipelines with debounce logic to prevent notification storms during transient spikes.
  • Configured automated systemd service isolation and auto-restart policies.

The Problem & Operational Challenges

  • Resource Overhead: Heavy monitoring suites (Prometheus/Grafana node_exporter stacks) consumed up to 300MB RAM per host, which was impractical for lightweight compute nodes.
  • Alert Fatigue: Raw threshold alerting caused rapid alert spam whenever memory spiked briefly during world saves or startup sequences.
  • Delayed Response Times: Operational engineers were frequently notified of outages by customer complaints rather than proactive telemetry alarms.

Architectural Approach & Engineering Decisions

  1. Lightweight Shell & Proc Collectors: Replaced heavy metric agents with lean POSIX shell scripts that parse /proc/stat, /proc/meminfo, and disk queues, outputting compact JSON payloads every 2 seconds.
  2. Redis In-Memory Buffering & Debouncing: Centralized metric ingest writes directly to Redis timeseries rings. Alert evaluations require a metric to exceed safety thresholds across 3 consecutive cycles before dispatching an alarm, filtering out 98% of transient spikes.
  3. Structured Webhook Integration: Formatted rich embeds directed to private Discord operations channels, featuring color-coded severity states (Warning, Critical, Resolved), affected host tags, and recommended triage commands.

System Architecture

[ Host Node 01 ]                [ Host Node 02 ]
  │ Direct /proc parsing          │ Direct /proc parsing
  ▼ (Every 2s via Systemd timer)  ▼
[ Local Agent Daemon ]          [ Local Agent Daemon ]
        │                               │
        └──────────────┬────────────────┘
                       ▼ (Secure WebSocket / HTTPS)
              [ Central Ingest API ]
                       │
                       ▼
              [ Redis Cache / Buffer ]
                       │
             [ Alert Evaluation Engine ]
             (Threshold & Debounce Checks)
                       │
        ┌──────────────┴──────────────┐
        ▼                             ▼
 [ Operations Dashboard ]   [ Discord Webhook Pipeline ]
 (Real-time WS Broadcast)    (Actionable Alert Dispatch)

Measurable Outcomes & Limitations

  • Memory Efficiency: Telemetry collection daemon uses less than 15MB RAM per bare-metal host, representing over 90% resource reduction compared to off-the-shelf agents.
  • Preemptive Resolution: Reduced average time-to-detect (TTD) from several minutes down to under 10 seconds, allowing engineers to intervene before players experienced disconnections.
  • Zero Alert Storms: Implemented exponential backoff and grouping logic, eliminating false-positive notification floods.
  • Honest Limitations: Relies on WebSocket connectivity to the central collector; during total upstream network loss on a host, heartbeats detect silence rather than receiving root-cause diagnostics until link restoration.

Technologies Used

  • Telemetry Core: Node.js, POSIX Shell (Bash), Linux Procfs
  • Data Layer: Redis (Pub/Sub & Time-window caching)
  • Transport: WebSocket (ws), HTTPS, TLS Encryption
  • Notification: Discord Webhook API (Rich Embed formatting)
  • Process Management: Systemd Units, Systemd Timers