Skip to main content
DevOps & InfrastructureEngineered by Shah Meer

WatchTower

Built a self-hosted, asynchronous uptime and API performance monitoring engine leveraging Celery workers, Redis brokers, FastAPI, and Next.js for high-frequency telemetry.

WatchTower architecture and interface preview

99.99%

Uptime Tracking

Granular second-by-second availability

Async Celery

Worker Concurrency

Non-blocking event loop execution

Up to 10s

Metrics Frequency

Configurable ping intervals per endpoint

100%

Self-Hosted

Zero reliance on expensive third-party SaaS

The Problem Statement

Operational Constraints & Friction

Monolithic commercial performance suites are cost-prohibitive, opaque, and overly complex for decentralized endpoint health, latency tracking, and API uptime diagnostics.

Engineered Solution

Architectural Execution

Built a self-hosted, asynchronous uptime and API performance monitoring engine leveraging Celery workers, Redis brokers, FastAPI, and Next.js for high-frequency telemetry.

WatchTower — System Architecture

High-Level Topology & Data Flow

Real-time Pings with Multi-Region Worker Queues
Ingestion Layer

Next.js 15 Real-time Dashboard & Telemetry Visualizer

Client requests, sensors & telemetry

Gateway / Proxy

FastAPI REST Core & WebSocket Bridge

Core Service Cluster
Celery Beat Periodic Scheduler
Celery Distributed Worker Swarm
HTTP/TCP Probe Telemetry Agent
⚡ Queue: Redis Message Broker (Celery Tasks)
Data & State Tier
PostgreSQL Time-Series Metric Store
Redis Health Cache

ACID Persistence & Read/Write separation

Technical Specifications & Internals

Internal Topology & Data Mechanics

Celery Beat schedules asynchronous HTTP ping tasks through a Redis broker to Celery worker pools. Results are aggregated, buffered, and committed to PostgreSQL, while a Next.js frontend fetches telemetry via FastAPI REST endpoints.

Technologies Used
Next.jsFastAPICeleryRedisPostgreSQLTailwind CSSTypeScript

Engineering Obstacles Overcome

Managing concurrent background network workers smoothly without blocking main application routing loops or overwhelming the relational database with high-frequency check logs.

Key Capabilities & Implementations
  • Asynchronous distributed uptime checks
  • Interactive real-time telemetry dashboard
  • Histogram latency metrics & percentile reporting
  • Multi-target endpoint health scoring
  • Configurable ping frequencies & failure alerting

Planned Enhancements & Roadmap

+ Slack & Webhook Alerts+ SSL Certificate Expiry Tracking+ Multi-Region Worker Distribution

Project Delivery & Impact

Engineered a lightweight, non-blocking monitoring microservice suite that safely schedules high-frequency telemetry requests across distributed targets.