DevOps Architect
- Developed an AI-powered DevOps Assistant (FastAPI, LangGraph, MCP servers) automating infrastructure and application troubleshooting, reducing engineering effort for incident diagnosis and improving MTTR.
- Architected and deployed an enterprise LLM Gateway (LiteLLM) centralising AI access across teams with governance, usage analytics, cost controls, and policy enforcement, eliminating direct provider integrations.
- Drove the AI platform strategy as part of the core AI adoption team, defining engineering standards, governance models, success metrics, and adoption frameworks across engineering teams.
- Led the adoption of AIOps-driven observability, replacing static threshold-based alerting with context-aware anomaly detection, reducing alert fatigue and accelerating incident response.
- Delivered ~$240K/yr in infra savings through Spot adoption, platform modernisation, observability consolidation, and resource right-sizing.
- Enabled multi-region platform architecture, supporting business expansion into new geographies while improving application availability, resilience, and user latency.
- Led three cross-functional Platform Engineering teams across multiple product suites, driving hiring, budget planning, engineering excellence, and operational maturity.
- Directed migration of 70+ microservices from ECS to EKS - P99 latency down 60%, P75 by ~30%, P50 by ~12%, near-zero errors at peak via KEDA autoscaling.
- Implemented a secure software supply chain: SBOM generation, KMS-backed Cosign image signing, and Kyverno policy enforcement, ensuring only trusted images are deployed.
- Defined the enterprise observability strategy: a unified stack with Grafana, Prometheus/Mimir, Tempo, Loki, Faro, OpenTelemetry, and Fluent Bit for monitoring, logging, and tracing.
- Built and operated scalable data processing platforms on Amazon EMR and Elasticsearch, enabling reliable, continuous data ingestion and analytics workloads.
- Led the evaluation, commercial negotiation, and rollout of Kloudfuse, replacing Datadog to cut observability stack cost by 35% while enhancing capability.
- Executed zero-downtime AWS account migrations at enterprise scale, leveraging Terraform to automate cross-account infrastructure migration with no production disruption.
- Established a self-service developer platform using Backstage, reducing DevOps dependencies and accelerating service onboarding across distributed engineering teams.
- Designed and operationalised the Business Continuity and Disaster Recovery (BCP/DR) strategy for mission-critical production systems, strengthening organisational resilience.
- Championed FinOps adoption by implementing Kubecost and AWS Cost Explorer with cost attribution dashboards, improving infrastructure cost transparency and accountability.