Get 10x More Interviews Without Applying Manually | OptimHire

OptimHire
OptimHire Job Auto-Applier
★★★★★
Install
We're excited to introduce the all new OptimHire Mobile App
Trusted by 50,000+ job seekers
Looking to Hire?
Urgently hiring

Delivery Lead

Refer and Earn up to $20000 USD

Refer & Earn

Get 10x More Interviews
Without Applying Manually

OptimHire’s Job Auto-Applier that will automatically apply to jobs at 300,000+ Companies.

★ ★ ★ ★ ★
Trusted by 50,000+ job seekers

Description

Delivery Lead – DevOps and SRE

 

Job description:


Experience: 12+ years (with 3+ years in a delivery leadership role)


Overview:


Experienced Delivery Lead to manage a team of Site Reliability Engineers (SREs) and DevOps Engineers responsible for automation, CI/CD, observability, incident response, and production operations in a complex, cloud-native environment.


This role leads delivery across infrastructure, platform, and operations functions, with deep involvement in GCP Cloud, Terraform-based IaC, Azure DevOps, and full-stack monitoring across network, infrastructure, and applications using tools such as AppDynamics, GCP Operations Suite, and Grafana.


Responsibilities:

• Lead the design, implementation, and continuous improvement of DevOps and IaC pipelines using Terraform on GCP, integrated with Azure DevOps.

• Own end-to-end delivery of infrastructure and application releases, change propagation, and production deployments.

• Oversee monitoring, alerting, and observability across:

• Network: VPC flow logs, firewall ingress/egress, latency, and DNS resolution

• Infrastructure: VM/compute instance health, memory/CPU thresholds, autoscaling, IAM, and storage

• Applications: Performance, error rates, response times, and transaction health using AppDynamics

• Drive the setup and maintenance of observability dashboards and alerts, using tools such as:

• GCP Operations Suite (Stackdriver) – Logs, Monitoring, Uptime checks

• Prometheus & Grafana – Infrastructure and custom application metrics

• AppDynamics – End-to-end application performance monitoring (APM), user journey insights, and deep-dive diagnostics

• Azure Monitor & Log Analytics – Visibility into hybrid cloud resources

• Implement and manage a 16x5 incident response model, including:

• On-call rotation and coverage planning

• Triage, escalation, and war room management

• Incident communication and stakeholder updates

• RCA documentation and preventive action tracking

• Ensure documentation of runbooks, SOPs, and proactive health checks.

• Track and report SLO/SLA compliance, error budgets, and availability KPIs.

• Coach and lead the DevOps/SRE team on modern delivery practices and reliability engineering.

• Collaborate with cloud architects, product teams, and security leads to ensure resilient and compliant delivery practices.


Required Skills & Experience

• 10+ years in IT, with 3+ years leading DevOps/SRE or cloud infrastructure delivery teams.

• Strong experience with Terraform for automating GCP infrastructure provisioning.

• Hands-on with Azure DevOps for CI/CD pipeline development and release automation.

• Deep understanding of GCP services (GKE, GCE, IAM, VPC, CloudSQL, etc.).

• Experience with AppDynamics for application performance monitoring (APM) and business transaction tracing.

• Familiarity with infrastructure and network monitoring tools such as GCP Monitoring, Prometheus, and Grafana.

• Proven ability to manage incident response processes in a 16x5 support environment.

• Excellent communication, team leadership, and cross-functional stakeholder management skills.


Nice to Have:

• Certifications: GCP Professional Cloud DevOps Engineer, Terraform Associate, or Azure DevOps Engineer Expert.

• Experience in regulated domains like healthcare

• Knowledge of security and compliance integrations in CI/CD (e.g., Veracode, Checkov).

• Exposure to Kubernetes, service mesh, or GitOps-based workflows.

 

Position

DevOps Engineers

Site Reliability Engineer

Deployment Manager

Delivery Manager

Must have skills

Google Cloud - 3 years

CI/CD - 4 years

Azure - 3 years

Site Reliability Engineering - 4 years

Nice to have skills

Terraform - 10 years

DevOps - 10 years

AppDynamics - 10 years

Grafana - 3 years


About the Company

Engineering, Technology, Training and Staffing solutions provider.

Access to 10M Jobs Get the OptimHire App

Your Dream Job is
Just a Tap Away

Search smarter, apply faster, and track application on the go with the OptimHire App.

AI Powered Job Search

AI Powered
Job Search

Access to 10M Jobs

Access to
10M Jobs

Auto-Apply to Job

Auto-Apply
to Job

AI Powered Custom Resumes

AI Powered
Custom Resumes

Trusted by 50,000+ job seekers
OptimHire App Preview