SK // the layer beneath
SURFACE · ORBIT 0 m
scroll to descend ▾
devops · sre · platform engineering

SOURAV
KUNDUthe layer beneath

Five years keeping production boring for six real-money gaming brands and a global streaming platform. You didn't notice anything. That was the job.

5+ yrs production 99.99% SLA held $20k/mo recovered 50+ domains owned
layer 01 · dns.resolve

Fifty domains answer to me.

At Momentum I own the domain estate end to end: 50+ apex domains, zones, records, certificates, and the drift between them. I built the self-service platform that provisions Cloudflare zone, DNS, WAF and Workers in one shot, and the health monitor that pages me about breakage before a customer ever notices it.

Cloudflare DNSRoute 53 ACME / certs50+ zones
layer 02 · edge.terminate

I hold the front door.

Six real-money gaming brands attract constant abuse, so I tune the WAF rules and rate limits against real attack traffic, wrote the Cloudflare Workers that steer users to healthy regions mid-incident, and put Zero Trust in front of everything internal. My favourite security events are the ones nobody ever hears about.

WAFWorkers Zero TrustDDoS shielding
layer 03 · compute.schedule

Four years of production Kubernetes.

At Warner Bros. Discovery I ran EKS end to end (Terraform provisioning, upgrades, disaster recovery) and tuned HPA on requests/sec and queue depth until 5xx bursts stopped happening. I cut deploys from 60 minutes to under 10, and my non-prod hibernation handed back ~$20k a month without a single developer noticing.

EKSHelm · ArgoCD HPA on custom metrics−90% deploy time
layer 04 · data.persist

I rehearse the bad day.

PostgreSQL, Redis and Kafka carry the state I protect. I automated etcd, EBS and S3 snapshots and actually restored from them in DR drills, because a backup you've never restored is just a wish. Migrations run inside my deploy path, not beside it, so schema changes ship like code.

PostgreSQLRedis KafkaDR: rehearsed
layer 05 · telemetry.emit

Five years carrying the pager.

24/7 on-call at WBD; the incident.io rotations at Momentum are mine too. I build the dashboards people actually open mid-incident: Datadog, Prometheus, Grafana, plus an OpenTelemetry rollout wired browser → SSR → API. That's how my 99.99% stayed a number instead of a slogan.

DatadogPrometheus · Grafana OpenTelemetrySLO / error budgets
layer 09 · on-call · bottom of the stack

Still answering the page.

You didn't notice anything.
That was the job.

▴ resurface