
Calix is Hiring Site Reliability Engineer (SRE) | GCP | Kubernetes | ArgoCD | Grafana | Linux | Python | GitOps
Calix is a leading cloud and software platform company helping Communication Service Providers (CSPs) simplify operations, automate infrastructure, and deliver exceptional subscriber experiences. Through its cloud-first, AI-powered Calix One platform, Calix enables broadband providers to modernize operations, accelerate innovation, reduce costs, and build reliable digital experiences for millions of users worldwide.
As the broadband industry continues to evolve, Calix is driving the next generation of cloud-native networking, automation, observability, and AI-powered operations. Join a team building scalable cloud platforms that empower communication providers across the globe.
---
Job Overview
Calix is looking for a Site Reliability Engineer (SRE I) to join its Engineering team in Bangalore. This role focuses on maintaining highly available, scalable, secure, and resilient production services running on Google Cloud Platform (GCP).
As an SRE, you'll bridge the gap between Software Development and Operations by leveraging DevOps principles, GitOps deployment practices, Kubernetes, Linux administration, cloud infrastructure, observability, networking, and AIOps automation.
You'll work closely with developers, platform engineers, and operations teams to automate deployments, improve production reliability, troubleshoot infrastructure issues, and support mission-critical cloud-native applications.
---
Key Responsibilities
GitOps & Continuous Deployment
- Deploy, manage, and roll back containerized applications using ArgoCD.
- Maintain GitOps workflows for reliable application delivery.
- Support CI/CD deployment pipelines and release automation.
- Ensure consistent deployments across production environments.
Site Reliability Engineering
- Maintain highly available and scalable production services on Google Cloud Platform.
- Monitor production workloads and ensure platform reliability.
- Participate in production incident management and on-call rotations.
- Improve platform resilience through automation and operational excellence.
Infrastructure & Kubernetes Operations
- Deploy, manage, and troubleshoot workloads running on Google Kubernetes Engine (GKE).
- Monitor Kubernetes clusters, pods, services, and networking components.
- Support container lifecycle management and workload optimization.
- Investigate infrastructure bottlenecks affecting production services.
Linux System Administration
- Troubleshoot Linux operating system performance issues.
- Investigate CPU utilization, memory leaks, storage constraints, and kernel-level issues.
- Manage Linux processes, file systems, permissions, networking, and operating system performance.
- Utilize system diagnostic tools for infrastructure troubleshooting.
Networking & Cloud Infrastructure
- Troubleshoot networking issues across cloud environments.
- Diagnose DNS, HTTP/HTTPS, SSL/TLS, TCP/IP, routing, subnetting, and Kubernetes networking problems.
- Investigate connectivity issues between cloud VPCs and Kubernetes clusters.
- Support cloud networking and service communication.
Monitoring, Observability & AIOps
- Monitor production environments using Grafana Labs ecosystem.
- Analyze metrics, logs, traces, and alerts.
- Utilize AI-powered operations (AIOps) platforms for event correlation and anomaly detection.
- Improve monitoring dashboards, alert quality, and operational visibility.
Incident Management
- Respond to production alerts and service incidents.
- Perform root cause analysis (RCA).
- Restore production services with minimal downtime.
- Collaborate with engineering teams during incident resolution.
Automation
- Develop automation scripts using Python, Bash, or Go.
- Automate operational tasks and repetitive infrastructure activities.
- Improve deployment efficiency and platform reliability through automation.
---
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related discipline.
- Strong understanding of Site Reliability Engineering (SRE) principles.
- Knowledge of Google Cloud Platform (GCP).
- Hands-on experience with Kubernetes and containerized environments.
- Experience using Git version control.
- Familiarity with ArgoCD GitOps workflows.
- Strong Linux administration skills.
- Good troubleshooting and analytical abilities.
- Excellent communication and collaboration skills.
---
Required Technical Skills
- Site Reliability Engineering (SRE)
- Google Cloud Platform (GCP)
- Google Kubernetes Engine (GKE)
- Kubernetes
- Docker
- ArgoCD
- GitOps
- CI/CD
- Git
- Linux Administration
- Ubuntu
- Debian Linux
- TCP/IP
- DNS
- HTTP
- HTTPS
- SSL/TLS
- gRPC
- Networking
- OSI Model
- CIDR
- Subnetting
- Routing
- Kubernetes Networking
- Ingress
- Services
- CNI
- Grafana
- Prometheus
- Grafana Loki
- Tempo
- Mimir
- Observability
- Monitoring
- Logging
- Alerting
- Incident Management
- Root Cause Analysis (RCA)
- Python
- Bash
- Go
- AIOps
- Automation
- Cloud Infrastructure
---
Preferred Skills
Candidates with experience in the following areas will have an added advantage:
- Machine Learning for Operations (MLOps)
- Log-based Machine Learning Models
- Event Correlation
- Anomaly Detection
- Cloud Security
- Infrastructure Automation
- Performance Tuning
- Production Support
- Cloud Networking
---
Technologies You'll Work With
- Google Cloud Platform (GCP)
- Google Kubernetes Engine (GKE)
- Kubernetes
- ArgoCD
- Grafana
- Prometheus
- Loki
- Tempo
- Mimir
- Git
- Python
- Bash
- Go
- Linux
---
Why Join Calix?
- Work on enterprise-scale cloud-native infrastructure.
- Build highly available production platforms serving millions of users.
- Gain experience with Google Cloud Platform and Kubernetes.
- Work with modern GitOps and cloud-native deployment practices.
- Learn AI-powered operations and observability technologies.
- Collaborate with experienced SRE, Cloud, and Platform Engineering teams.
- Excellent opportunities for technical growth and career development.
---
Benefits
- Full-Time Employment
- Flexible Hybrid Work Model
- Enterprise Cloud Infrastructure Exposure
- Modern Kubernetes & GitOps Stack
- Learning & Career Development
- Collaborative Engineering Culture
- Exposure to AI-powered Operations (AIOps)
---
Job Details
- Role: Site Reliability Engineer (SRE I)
- Employment Type: Full-time
- Experience Level: Entry Level
- Location: Bangalore, Karnataka, India
- Work Mode: Hybrid (20 days per quarter from Bangalore office)
- Job Requisition ID: R-11783
---
Site Reliability Engineer Jobs India, SRE Jobs Bangalore, Google Cloud Platform Jobs, GCP Engineer Jobs, Kubernetes Engineer Jobs, ArgoCD Jobs, GitOps Engineer Jobs, DevOps Engineer Jobs Bangalore, Linux Administrator Jobs, Cloud Infrastructure Jobs, Grafana Jobs, Prometheus Jobs, Loki Jobs, Observability Engineer Jobs, Cloud Operations Engineer Jobs, Production Support Engineer Jobs, Python Automation Jobs, Google Kubernetes Engine Jobs, AI Operations Jobs, Cloud Native Engineer Jobs.
---
Application Process
Interested candidates can apply using the Apply Now button below.
Job Overview
| Role | Calix is Hiring Site Reliability Engineer (SRE) | GCP | Kubernetes | ArgoCD | Grafana | Linux | Python | GitOps |
| Company | Calix |
| Job Type | full-time |
| Experience Level | Entry Level |
| Street Address | Rmz Infinity Tower-2, B, Old Madras Rd, Sadanandanagar, Bennigana Halli, Bengaluru |
| City | Bangalore |
| State / Region | Karnataka |
| Postal Code | 560001 |
| Country | India |
| Salary Range | INR 12,00,000 - 22,00,000 /Year |
Required Skills
Related Jobs
Ready to Apply?
Don't miss out on this opportunity. Apply now and take the next step in your career.
Apply Now