Nordlicht Logistics GmbH
Freight forwarding and warehouse operations · Hamburg, Germany · 420 employees
Warehouse and customs HTTPS endpoints placed on Cloud Gateway routing with DNS failover monitoring across two origin nodes.
- RPO
- 24 hours (daily backups; WAL archive retained 7 days)
- RTO
- 4 hours for database recovery; 1 hour for compute replacement
- SLA
- Growth tier: automated DNS failover probes; email support within 1 business day
Situation
Nordlicht runs a warehouse-management platform that must remain available across three shifts. Prior to engagement, alerting was limited to a hosting-provider dashboard, PostgreSQL backups were untested, and a disk-full event on a Saturday night halted goods-out for 3 hours 40 minutes. The internal IT team of four could not staff nights or weekends.
Environment
- Region and placement
- eu-central-1 (Frankfurt) with a warm standby in eu-west-1
- Hosts
- 8 production VMs (application, PostgreSQL primary/replica, Redis, jump host, reporting)
- Operating systems and runtime
- Ubuntu 22.04 LTS; PostgreSQL 15; Redis 7
- Data stores
- 2.4 TB warehouse and customs data; 380 GB nightly change volume
- Connectivity and access
- Site-to-site IPsec from the Hamburg DC to the cloud VPC; no public database endpoints
Agreed scope
- 24/7 uptime monitoring on all eight nodes and the VPN tunnels
- Daily encrypted backups of PostgreSQL with quarterly restore tests
- Query and index review of the goods-out and inventory schemas
- Security baseline: SSH key-only access, fail2ban, unattended-upgrades, UFW
Work performed
- Deployed Prometheus exporters and Alertmanager with PagerDuty routing to Veltis on-call.
- Implemented pgBackRest to object storage with AES-256, 14-day retention, and a documented restore runbook.
- Added missing indexes on shipment_events and stock_moves; reduced p95 goods-out query time from 1.8 s to 210 ms.
- Closed inbound 5432/6379 from the public internet; restricted admin access to the jump host and a named engineer group.
- Ran a full restore drill on 12 March 2026; recovered the primary to a staging host in 47 minutes.
Measured results
Unplanned downtime (rolling 12 months)
11 minutes
Goods-out query p95
210 ms (from 1.8 s)
Last restore drill
47 minutes, verified
Open P1 incidents
0
Service tier: Cloud Gateway & Route Manager. Commercial inquiries: contact operations.