Chandra Sekhara Reddy Ummadi
AI + DevOps Engineer
13+ years engineering AWS cloud infrastructure that scales — now building that same discipline into AI-powered workflows and agentic systems.
About
I'm an infrastructure engineer with 14+ years of experience designing and operating cloud systems that stay reliable under real production load. My work spans AWS platform engineering, Kubernetes, and building Infrastructure-as-Code that other engineers can actually reuse — not just scripts that work once.
Currently a Senior/Staff Infrastructure Engineer at NTT DATA, where I focus on Terraform module architecture, CI/CD pipeline design with GitHub Actions and Jenkins, and AWS infrastructure spanning EC2, EKS, RDS, and IAM. Before that, I spent 5 years at JPMorgan Chaseas an SRE, where I learned what "highly available" actually means when real money depends on it.
I'm currently extending that same infrastructure discipline into AI-powered workflows — building agentic tooling and DevOps automation that combines 14 years of hands-on operations knowledge with modern AI engineering practices, rather than treating them as separate disciplines.
Outside of infrastructure, I spend time hiking and exploring off-the-beaten-path places, and I'm a parent navigating the entirely different kind of complexity that comes with that.
Experience
2022 — Present
Digital Engineering Sr Staff Engineer · NTT DATA
Redesigned 12 CI/CD deployment pipelines using GitHub Actions and Terraform, cutting deployment time by 40% and enabling on-demand releases. Built a reusable GitHub Actions workflow library (15+ workflows) adopted by 8 teams, reducing pipeline maintenance by 60%. Created 25+ Terraform modules published to a private registry, enabling 3x faster environment provisioning. Led migration of 200+ VMs to AWS using AWS MGN + DMS, achieving 99.9% data integrity and a 40% infrastructure cost reduction. Designed a multi-account VPC architecture (Inspection, Shared Services, Workloads) with Transit Gateway and Route 53 Resolver across 20+ VPCs. Deployed a Prometheus/Grafana + Loki monitoring stack on EKS (50+ services), defining 200+ alerts with 95% noise reduction.
- AWS
- Terraform
- GitHub Actions
- EKS
- Prometheus
- Grafana
- Transit Gateway
2017 — 2022
Senior Systems Administrator (SRE) · JPMorgan Chase
Supported Site Reliability Engineering initiatives by defining operational best practices, improving system reliability, and enhancing production stability. Automated 15+ manual runbooks (patch management, cert rotation) using Python, saving 40 hours/month of ops time. Implemented monitoring tools to proactively identify and mitigate performance issues. Managed production incidents for Tier-1 applications, reducing Mean Time to Recovery (MTTR) from 45 minutes to 12 minutes through standardized runbooks and automated diagnostics. Performed root cause analysis and implemented preventive measures to improve platform reliability. Served as primary on-call for Tier-1 services.
- Python Automation
- SRE
- Incident Management
- Monitoring
- Runbooks
2012 — 2017
Server Support Specialist · IBM India Pvt Ltd
Managed virtualized server environments across multiple data centers. Migrated business applications to cloud platforms, automated server provisioning, and streamlined patch management to reduce vulnerabilities and downtime.
- Linux
- AIX
- Virtualization
- Automation
Projects
Reusable Terraform Module Library
A production-grade module repository covering VPC, EC2, EKS, RDS, Lambda, IAM, Load Balancer, and 10+ other AWS services — built for reuse across multiple environments (dev/staging/prod) rather than one-off scripts.
- Terraform
- AWS
- IaC
AWS Application Migration (MGN) — SAP Lift-and-Shift
Led a 44-instance SAP EC2 migration between VPCs using AWS Application Migration Service, including a custom CloudFormation template for networking, HA NAT gateways, and SAP-specific security groups.
- AWS MGN
- CloudFormation
- VPC
- SAP
Claude Code Skills Library for Infrastructure Engineering
Built a personal library of reusable Claude Code skills encoding IaC standards, cloud security review patterns, and a custom stdlib-only Python cross-stack asset scanner with HTML report output.
- Python
- AI Tooling
- DevOps Automation
Contact
Open to conversations about infrastructure, AI-powered DevOps, or roles where both matter. Reach out directly: