Available for opportunities

Hi, I'm Jeevanantham

Site Reliability Engineer @Cisco Systems

Site Reliability Engineer with 6+ years ensuring 99.99% uptime for Cisco Webex Meetings across 50+ Kubernetes clusters and 8 global data centers serving 100M+ users. Drove 40% MTTD reduction and 35% MTTR reduction through automation and continuous monitoring. Expertise in Kubernetes, CI/CD, IaC, and cloud platforms (AWS, Azure).

webex-sre · production operational
Jeevanantham P
Jeevanantham P
Site Reliability Engineer
Chennai, India
99.99%Availability
Error Budget82%
MTTD
-40%
MTTR
-35%
requests / sec100M+ users
6+
Years Experience
99.99%
Uptime Delivered
50+
K8s Clusters
100M+
Users Served
Live Ops

Reliability, at a Glance

A day in the life — the tooling, pipelines, and clusters I keep healthy in production.

zsh — jeeva@webex-sre
jeeva@webex-sre:~$ whoami
Site Reliability Engineer @ Cisco Webex Meetings
jeeva@webex-sre:~$ kubectl get pods -A --field-selector=status.phase=Running | wc -l
2418 pods healthy across 50+ clusters
jeeva@webex-sre:~$ cat skills.yaml
k8s: [Helm, ArgoCD, HPA, PDB]
observability:[Prometheus, Grafana, ELK]
automation: [Python, Bash, Terraform, Vault]
jeeva@webex-sre:~$ ./contact.sh --open-to-work
✔ Available for SRE / DevOps roles
jeeva@webex-sre:~$

CI/CD Pipeline

webex-meetings · main
Build
Test
Package
Deploy
Monitor

Cluster Fleet · 8 DCs

50+ healthy
DFW
IAD
LON
FRA
SIN
SYD
AMS
SJC
CI/CD

Delivery Pipeline

Every commit ships through an automated, observable path — build to verified in production, with progressive delivery and instant rollback.

pipeline · main passing
Source logo

Source

git push

Build logo

Build

compile + lint

Test

unit / integration

Artifact logo

Artifact

image + scan

Deploy logo

Deploy

ArgoCD → K8s

Verify

SLO smoke check

Rolling

Zero-downtime, surge + maxUnavailable tuned

Blue / Green

Instant switch, instant rollback

Canary

Progressive traffic shift w/ auto-analysis

SLOs & DORA

Reliability Scorecard

Engineering for four nines — measured with SLOs, error budgets, and the DORA metrics that map directly to delivery performance.

99.99%Availability

Service Availability

rolling 30-day SLO

Error budget82% left

18% burned this window

Deployment Frequency

Multiple / day

Elite

Lead Time for Changes

−60%

faster releases

Change Failure Rate

<5%

95% clean deploys

Mean Time to Restore

−35%

faster recovery
Technical Expertise

My Toolchain

The stack I use every day to ship, scale, and observe cloud-native systems — Kubernetes, CI/CD, IaC, and full-stack observability.

Kubernetes logoKubernetes
Docker logoDocker
Helm logoHelm
ArgoCD logoArgoCD
Jenkins logoJenkins
GitLab CI logoGitLab CI
GitHub Actions logoGitHub Actions
Terraform logoTerraform
Ansible logoAnsible
Prometheus logoPrometheus
Grafana logoGrafana
Elasticsearch logoElasticsearch
Kibana logoKibana
Git logoGit
AWS logoAWS
Azure logoAzure
Vault logoVault
Kafka logoKafka
Python logoPython
LangChain logoLangChain
Kubernetes logoKubernetes
Docker logoDocker
Helm logoHelm
ArgoCD logoArgoCD
Jenkins logoJenkins
GitLab CI logoGitLab CI
GitHub Actions logoGitHub Actions
Terraform logoTerraform
Ansible logoAnsible
Prometheus logoPrometheus
Grafana logoGrafana
Elasticsearch logoElasticsearch
Kibana logoKibana
Git logoGit
AWS logoAWS
Azure logoAzure
Vault logoVault
Kafka logoKafka
Python logoPython
LangChain logoLangChain
☁️

Cloud

AWS logoAWSEC2 logoEC2EKS logoEKSVPCIAMAzure DevOps logoAzure DevOpsAKS logoAKSAzure OpenAI logoAzure OpenAI
🐳

Containers & Orchestration

Docker logoDockerKubernetes logoKubernetesArgoCD logoArgoCDHelm logoHelmHPAVPAPDB
🔄

CI/CD

Jenkins logoJenkinsGitLab CI/CD logoGitLab CI/CDGitHub Actions logoGitHub Actions
🏗️

Infrastructure as Code

Terraform logoTerraformAnsible logoAnsible
📊

Observability

Prometheus logoPrometheusGrafana logoGrafanaELK Stack logoELK StackAppDynamicsThousandEyes
💻

Languages

Python logoPythonBash logoBashJava logoJava
🔧

Other

HashiCorp Vault logoHashiCorp VaultKafka logoKafkaLangChain logoLangChainMCPGit logoGit
Career

Deployment History

Every role, shipped with measurable impact — a rollout log of my SRE, DevOps, and software engineering journey.

CS

Site Reliability Engineer

Cisco Systems (Contract)
Jan 2024 – Present Chennai, India

Adecco India (Jan 2024 – May 2026) · Tekgence India (May 2026 – Present)

99.99%50+100M+25%
  • Maintained 99.99% uptime for Webex Meetings by managing 50+ K8s clusters across 8 DCs serving 100M+ users
  • Cut infrastructure costs by 25% through HPA scale-down stabilization and time-based cronjob scaling
  • Enabled zero-downtime deployments using PodDisruptionBudgets during cluster maintenance
  • Managed Helm charts for 50+ microservices with HPA, PDB, resource quotas, and sidecar patterns
  • Resolved 95% of pod failures within SLA by debugging CrashLoopBackOff, OOMKilled, ImagePullBackOff issues
  • Saved 15+ hours weekly by developing 20+ Python/Bash automation scripts for operations
  • Slashed deployment time by 60% with Python Kubernetes Client tools for parallel execution across clusters
  • Built Grafana dashboard propagation tool for consistent deployment across 8+ data centers
  • Developed cluster auth automation with HashiCorp Vault integration for secure token management
  • Lowered MTTD by 40% by designing 25+ Grafana dashboards for availability, error rates, latency metrics
  • Decreased MTTR by 35% through automated alerting and runbook execution
  • Conducted RCA for 100+ incidents, trimming repeat incidents by 60%
  • Led AIOps initiatives integrating ML models, improving anomaly detection by 40%
  • Published 'Copilot Chat History Search' VS Code extension on Marketplace
TH
Sep 2021 – Jan 2024 Bangalore, India
10+95%500K+85%
  • Delivered 10+ enterprise Java applications with AWS SDK for cloud integrations
  • Achieved 95% deployment success rate with Jenkins and Docker-based CI/CD pipelines
  • Designed Kafka clusters processing 500K+ messages daily on Kubernetes and Helm charts
  • Built PDF Q&A chatbot using LangChain/LLM with 85% query accuracy
  • Integrated Hugging Face models, cutting document processing time by 50%
TH

Associate Software Engineer

Torry Harris Integration Solutions
Aug 2019 – Sep 2021 Bangalore, India
60%15+10+200+ hours
  • Shortened provisioning time by 60% using Terraform for Azure infrastructure
  • Containerized 15+ Kafka applications using Docker for consistent deployments
  • Developed 10+ UiPath RPA workflows, automating 200+ hours monthly
Provisioning

Foundation Layers

The base image my career is built on — academic layers provisioned before the stack went live.

LAYER 02compsci:degree provisioned

B.E. Computer Science

2015 – 2019

Velammal Engineering College, Anna University

AlgorithmsOperating SystemsNetworksDBMSOOP
Chennai, IndiaGrade: 74/100
LAYER 01science:hsc provisioned

Higher Secondary Education

2014 – 2015

Adhiyaman Matric Hr. Sec. School

MathematicsPhysicsComputer Science
Uthangarai, IndiaGrade: 94/100
Contact

Get In Touch

Have a project idea, an opportunity, or just want to say hi? Feel free to reach out.

© 2026 Jeevanantham P. All rights reserved.

Built with Astro & React