Skip to main content
Retail•Observability

Observability at Scale for a Global Retailer

Global Athletic Apparel Retailer
Global
10,000+ employees
2022–2024
observability engagement

The Challenge

Prometheus solves metrics collection well and long-term storage badly. Its local time-series database is built for weeks of retention on a single node, which left the retailer facing three problems simultaneously. Retention: seasonal retail analysis depends on year-over-year comparison, and a fortnight of history cannot answer whether this peak trading period behaved like the last one. Availability: a single Prometheus is a single point of failure, yet running redundant replicas produces two divergent datasets rather than one authoritative view. Locality: with infrastructure spread across regions, each Prometheus sees only its own slice, so any platform-wide question means querying every instance and reconciling the answers by hand. The retailer needed one query surface across all of it, plus dashboards and alerting that on-call engineers would actually trust at three in the morning.

Solution Overview

PlatOps built and operated the retailer's monitoring and observability platform from 2022 to 2024, using Prometheus for metrics collection, Thanos for long-term storage, high availability, and global query, and Grafana for dashboards and alerting.

The Results

Prometheus-based metrics collection across the platform
Thanos providing long-term storage, high availability, and a global query view
Grafana dashboards and alerting operated for the retailer

"PlatOps built and ran the retailer's Prometheus, Thanos, and Grafana observability platform from 2022 to 2024."

P
PlatOps engagement lead
Global Athletic Apparel Retailer

Key Takeaways

  • Prometheus is a collection and alerting engine, not a long-term metrics store; pairing it with Thanos separates those concerns cleanly.
  • Thanos deduplicates redundant Prometheus replicas, so high availability stops producing two conflicting answers to the same query.
  • A global query view matters more than raw retention — the value is asking one question across every region, not merely keeping data longer.
  • Seasonal businesses need year-over-year metric retention, which default Prometheus retention cannot support.
  • Object-storage backing makes multi-year retention cheap enough that retention becomes a product decision rather than a cost ceiling.

Key Outcome

2022–2024
observability engagement

Technologies Used

PrometheusThanosGrafanaKubernetesAlertmanager

Want Similar Results?

Let's discuss how we can help your organization achieve its goals.

Get Free Assessment

Ready to Write Your Success Story?

Join the organizations that have transformed their security and infrastructure with PlatOps.