Hi, I'm Jeevanantham
Site Reliability Engineer @Cisco Systems
Site Reliability Engineer with 6+ years ensuring 99.99% uptime for Cisco Webex Meetings across 50+ Kubernetes clusters and 8 global data centers serving 100M+ users. Drove 40% MTTD reduction and 35% MTTR reduction through automation and continuous monitoring. Expertise in Kubernetes, CI/CD, IaC, and cloud platforms (AWS, Azure).

Reliability, at a Glance
A day in the life — the tooling, pipelines, and clusters I keep healthy in production.
CI/CD Pipeline
webex-meetings · mainCluster Fleet · 8 DCs
50+ healthyDelivery Pipeline
Every commit ships through an automated, observable path — build to verified in production, with progressive delivery and instant rollback.
Source
git push
Build
compile + lint
Test
unit / integration
Artifact
image + scan
Deploy
ArgoCD → K8s
Verify
SLO smoke check
Rolling
Zero-downtime, surge + maxUnavailable tuned
Blue / Green
Instant switch, instant rollback
Canary
Progressive traffic shift w/ auto-analysis
Reliability Scorecard
Engineering for four nines — measured with SLOs, error budgets, and the DORA metrics that map directly to delivery performance.
Service Availability
rolling 30-day SLO
18% burned this window
Deployment Frequency
Multiple / day
EliteLead Time for Changes
−60%
faster releasesChange Failure Rate
<5%
95% clean deploysMean Time to Restore
−35%
faster recoveryMy Toolchain
The stack I use every day to ship, scale, and observe cloud-native systems — Kubernetes, CI/CD, IaC, and full-stack observability.
Cloud
Containers & Orchestration
CI/CD
Infrastructure as Code
Observability
Languages
Other
Deployment History
Every role, shipped with measurable impact — a rollout log of my SRE, DevOps, and software engineering journey.
Adecco India (Jan 2024 – May 2026) · Tekgence India (May 2026 – Present)
- Maintained 99.99% uptime for Webex Meetings by managing 50+ K8s clusters across 8 DCs serving 100M+ users
- Cut infrastructure costs by 25% through HPA scale-down stabilization and time-based cronjob scaling
- Enabled zero-downtime deployments using PodDisruptionBudgets during cluster maintenance
- Managed Helm charts for 50+ microservices with HPA, PDB, resource quotas, and sidecar patterns
- Resolved 95% of pod failures within SLA by debugging CrashLoopBackOff, OOMKilled, ImagePullBackOff issues
- Saved 15+ hours weekly by developing 20+ Python/Bash automation scripts for operations
- Slashed deployment time by 60% with Python Kubernetes Client tools for parallel execution across clusters
- Built Grafana dashboard propagation tool for consistent deployment across 8+ data centers
- Developed cluster auth automation with HashiCorp Vault integration for secure token management
- Lowered MTTD by 40% by designing 25+ Grafana dashboards for availability, error rates, latency metrics
- Decreased MTTR by 35% through automated alerting and runbook execution
- Conducted RCA for 100+ incidents, trimming repeat incidents by 60%
- Led AIOps initiatives integrating ML models, improving anomaly detection by 40%
- Published 'Copilot Chat History Search' VS Code extension on Marketplace
- Delivered 10+ enterprise Java applications with AWS SDK for cloud integrations
- Achieved 95% deployment success rate with Jenkins and Docker-based CI/CD pipelines
- Designed Kafka clusters processing 500K+ messages daily on Kubernetes and Helm charts
- Built PDF Q&A chatbot using LangChain/LLM with 85% query accuracy
- Integrated Hugging Face models, cutting document processing time by 50%
Associate Software Engineer
Torry Harris Integration Solutions- Shortened provisioning time by 60% using Terraform for Azure infrastructure
- Containerized 15+ Kafka applications using Docker for consistent deployments
- Developed 10+ UiPath RPA workflows, automating 200+ hours monthly
Featured Projects
Tools and applications I've shipped — presented the way I work with them: as version-controlled repositories.
Cloud Instance Manager
AWS EC2 management platform built with Spring Boot featuring JWT authentication, RBAC, instance lifecycle control, security-group management, and multi-region support.
Copilot Chat History Search
VS Code extension for searching GitHub Copilot conversations, published on the VS Code Marketplace for enhanced developer productivity.
PDF Q&A Chatbot
LangChain-based AI chatbot with vector search for document analysis. Implements chunking, embeddings, and semantic search for natural-language Q&A.
SLO Dashboards
Grafana SLO/SLI dashboards with error-budget burn-rate alerts, availability tracking, and custom PromQL queries integrated with Prometheus.
Foundation Layers
The base image my career is built on — academic layers provisioned before the stack went live.
B.E. Computer Science
2015 – 2019Velammal Engineering College, Anna University
Higher Secondary Education
2014 – 2015Adhiyaman Matric Hr. Sec. School
Get In Touch
Have a project idea, an opportunity, or just want to say hi? Feel free to reach out.