Senior DevOps Engineer

Predictable.
Scalable.
Boring.

Senior DevOps Engineer with 5+ years running high-traffic infrastructure. I take systems that are expensive, fragile or slow to deploy and make them predictable, scalable and fast — whether on-prem, hybrid, or pure cloud (k3s, EKS, AWS).

Right now: a programmatic ad exchange across three datacenters, edge tuned past 1M req/sec. Before that, 200+ services on managed AWS.

Bangalore, India · Full-time · Remote-friendly

121 bare-metal servers3 datacenters, all in Ansible
1.0M+ req/sTested ceiling; ~494k in production
200+ services on AWSMulti-region DR, RTO under 10 min
$240–360K/yr avoidedMoved the exchange off its CDN

Aug 2025 – Present

DevOps Engineer → Senior DevOps Engineer

The Cool Company (formerly Insticator) · programmatic ad exchange

~120 bare-metal servers across 3 datacenters, 3 k3s clusters and the AWS estate alongside them.

Jul 2021 – Aug 2025

Associate → DevOps → Senior DevOps Engineer

Conga · enterprise revenue lifecycle SaaS

200+ microservices on managed AWS. CI/CD for 8+ teams, multi-region DR, and the FinOps work.

Case Studies

Experience Log

Selected case studies.

High-Volume Ad Exchange (SSP) — 2 datacenters • ~494k req/s edge • Lead

Moving the exchange off Cloudflare

HAProxyRoute 53AnsibleTLSWAF

Key Impact

$240–360K/yr contract avoided
2.7–3.5 ms backend response
0% 5xx at steady state

The Challenge

Every ad request went through Cloudflare's proxy. For a lot of users it replaced their IPv4 address with IPv6, and that address ended up in the bid request. Bidders use the IP for location and fraud checks, and they were pricing IPv6 traffic 20 to 30% below IPv4, so the substitution was suppressing bid values across the exchange. Cloudflare's proposed fix was an Enterprise contract at $240–360K a year.

The Goal

Rebuild what Cloudflare was doing (TLS, routing, rate limiting, filtering) on our own HAProxy edge across two datacenters and two CPU architectures, and keep rollback down to a single DNS change.

The Solution

Built two proof-of-concept HAProxy boxes from bare metal, then load-tested one ARM server to find its real ceiling before moving any traffic.

Set up Route 53 with latency routing per exchange and health-checked automatic failover, later pinning all 50 states and DC with geo records after AWS re-mapped between 4 and 54% of US traffic four times in a single week.

Replaced Cloudflare's WAF and rate limiting: per-IP ceilings at roughly 2.4x observed legitimate peak, and a 429 cap above 5,000 rps per source.

Kept a Cloudflare-proxied box as a warm backup for the whole cutover, and moved one endpoint and one region at a time, checking against authoritative nameservers rather than cached resolvers. Staging it that way is what surfaced a WAF rule ported across from Cloudflare too broadly — 403s on beacon traffic for two EU hostnames — in time to scope it the same day.

Diagnosed a load imbalance on cutover day within the hour: connections stuck in CLOSE_WAIT pushed one box to 98.7% CPU while two others sat at 40%.

Trade-offs

Running the edge ourselves means TLS, patching, capacity and on-call are ours. That is what it costs to have the bid path under our own control instead of a vendor's.

The WAF rules are ours to maintain now. A vendor ships rule updates; we decide ours from our own traffic.

DNS TTL went up in stages to cut Route 53 cost by about 54%, so any change, including a rollback, propagates more slowly.

The Outcome

Bidders see real client IPs again. Backend response stayed at 2.7–3.5 ms, steady-state 5xx stayed at 0%, and three old HAProxy servers were retired once traffic was off them.

Systems

Kernel and OS tuning

An audit across 14 layers and every host in three regions rejected four proposed optimisations on measurement. These are some that survived it.

amd_pstate

Worker nodes were dropping into low-power states between traffic bursts and paying to climb back out on every spike. Pinned the CPU floor to the CPPC nominal value, with a selectable driver per host group, proved on one canary before the fleet.

hard-stop-after

On 80-core edge boxes a plain HAProxy reload left old and new workers overlapping and cost a one-to-two minute dip with 5xx. A hard-stop guard ended it. The package upgrade path was also triggering reloads on its own and bypassing the role’s gate.

99-sysctl.conf

The procps symlink was missing, so every sysctl setting on the fleet silently reverted on reboot. Everything anyone had tuned was only true until the next restart. Found while bringing up the third region, along with two more unit-ordering bugs of the same shape.

index-type shmem

Aerospike never returns primary-index memory to the OS, so free-memory alerting on those clusters was reporting a problem that did not exist. Rewrote the alerts to match what the process actually does.

Post-Mortems

Incident investigations

Six from the incident record, picked because the first answer was wrong in each one.

Jul 2026

A VPN that was down for seven days before anyone noticed

What I found

Vector had grown to 3.6 GB on a 4 GB instance with no swap and taken the primary AWS VPN down with it. Failover then did exactly what it was built to do: the VIP moved to the backup, traffic kept flowing, and nothing degraded enough to alert. The outage lasted a week because the recovery worked.

What changed

Memory bounded on the gateway, and the gap went straight into the status page catalog. A successful failover is indistinguishable from health unless something alerts on running degraded.

May 2026

504s on 42 to 64% of requests, and nothing had changed

What I found

A managed OpenSearch domain was failing most requests. The load balancer target group was still pinned to ENI addresses that no longer existed, left behind by a volume-shrink blue/green that had otherwise completed cleanly.

What changed

Target group repaired, then the pressure underneath it: deleting 1,556 stale indices took the shard count from the 4,000 cap down to 918, and heap from 79-85% down to 43-50%.

Aug 2026

A certificate rotation that only broke the clients nobody tests with

What I found

A wildcard rotation went out with the leaf but no intermediate. Browsers papered over it by fetching the missing certificate through AIA, so every manual check looked fine. Go and Node clients do not do that, and they failed. The verification step had used curl -k, which skips the exact check that would have caught it.

What changed

Rotation reverted, and a strict-verification test written that fails the way a real client fails rather than the way a browser does.

May 2026

Daily replication lag that turned out not to be ours

What I found

Recurring MirrorMaker2 lag pages, arriving at idle traffic levels, which ruled out load. Not the NIC, not Aerospike XDR, not WireGuard. It was packet loss on the public transatlantic transit underneath, on a path that was asymmetric: one carrier outbound, a different one on the way back.

What changed

This is what the cross-DC mtr exporter was built for. Fixed-IP ping probes could not survive a reroute or say which hop was losing; per-hop probing with carrier attribution can.

Sep 2026

A compactor that filled its own disk

What I found

A compaction concurrency change I had made filled a 600 Gi volume. Expanding it live to 900 Gi fixed the immediate problem and exposed a second one: replica pressure on three control-plane root disks, traced to the filesystem trim cron running against Longhorn.

What changed

Scratch sizing documented per compaction group, so the next person raising concurrency knows what it costs in disk before they raise it.

Sep 2026

An alert reporting reboots that never happened

What I found

Leaseweb’s NTP served bad time in three episodes. Clocks stepped forward by anywhere from 4 seconds to 8 minutes, a VPN address flapped twice because its health check read wall-clock time, and an alert reported reboots across the fleet. It was measuring boot time, which moves when the clock steps.

What changed

Monotonic uptime proved no host had rebooted. Then multi-source chrony replaced single-source timesyncd on 105 hosts across three regions, one host at a time, with no downtime.

Things I've Built

Open Source

Sairo object browser

Sairo

v3.5 · Apache-2.0 · Open Source

Object storage, beautifully engineered.

Most object-storage UIs crawl to a halt past a few thousand objects. Sairo stays instant at 15.5 million — folder navigation is a constant-time index lookup that lands at the network floor.

A self-hosted, S3-compatible storage platform — far more than a browser. It breaks down storage cost across 13 providers with optimization recommendations, and exposes an AI-queryable analytics layer over MCP. Built full-stack (FastAPI · React · Go CLI · SQLite FTS5) and proven on a live 269 TB / 15.5M-object deployment.

15.5M · 269TB
Live deployment
191,231×
Indexed vs naive listing, 2M objects
13
Providers, cost intel
26
AI tools (MCP)

Constant-time at any scale

Folder navigation is an index lookup — 0.002ms on 2M-object buckets (191,231× faster than the naive path) — plus SQLite FTS5 full-text search across millions of keys.

Adaptive delta crawler

Picks up new objects in ~30s vs a 7-minute full re-crawl, auto-parallelizing across sub-prefixes (16 workers / 12 buckets).

Cost intelligence

Per-folder cost across 13 providers with live AWS pricing; cold-data, duplicate & lifecycle-gap detection with tiering-savings recommendations.

AI storage intelligence

A 26-tool MCP server (+4 guided workflows) exposes the whole analytics layer to Claude Desktop, Cursor & any MCP client — hardened with 75 security tests.

Memory-bounded uploads

Direct browser→S3 multipart upload for files of any size (114 MB/s), with just-in-time part signing — server RAM stays ~100MB regardless of file size.

Per-user S3-key scoping

Every call uses the logged-in user’s own S3 credentials, so IAM scopes what they can see. Plus RBAC, 2FA, OAuth/LDAP, encrypted creds & an audit log.

Single container · zero external dependencies · Helm chart · Go CLI · 394 end-to-end tests · cosign-signed multi-arch images + SBOM

Python · FastAPIReact · ViteGo CLISQLite FTS5MCPDocker · Helmboto3

Dependencies

Tech Stack

The tools and technologies I work with daily.

Cloud

AWSAzureEKSRoute 53S3CloudFrontCloudWatchIAM

Bare Metal & OS

LinuxDebian/UbuntuIPMI/BMCsysctl & kernelsystemd

Infrastructure as Code

AnsibleTerraformHelmRancher Fleet

Kubernetes

k3sRancherLonghornCiliumArgo RolloutsIngressk3d

Edge & Network

HAProxyNginxkeepalivedWireGuardiptablesmtrBBR

Observability

PrometheusThanosGrafanaAlertmanagerPagerDutypromtoolOpenTelemetry

Logging & Search

OpenSearchELK StackVectorFluent Bit

Streaming

KafkaMirrorMaker2Kafka ConnectSchema RegistryKafbatNATSIceberg

Data & Analytics

TrinoAerospikePostgreSQLRedisDruidAirflowSupersetPolars

Security & Identity

OpenBaoSSH CAOkta SSOSAML/OIDCKubernetes RBACCheckmarxfail2ban

CI/CD

GitHub ActionsJenkinsSpinnakerCircleCIDockerBlue/Green

Languages & Tooling

PythonGoBashGroovyPowerShellSQLFastAPIPlaywrightMCPClaude Agent SDK

Handshake

Let's Connect

Currently open to new opportunities. Have a question or just want to say hi?

Or download my resume