Senior Support Engineer
On site Β· I. Β· FULL TIME
Job Summary π»
We are seeking a Senior Support Engineer to serve as the technical backbone for our infrastructure and platform operations, ensuring that internal and external stakeholders receive expert-level support for cloud infrastructure, DevOps tooling, and platform services. You will partner closely with Site Reliability Engineering (SRE), DevOps, and product engineering teams to resolve complex escalations, drive root-cause analysis, and reduce time-to-resolution for critical platform incidents. As a Senior Support Engineer, you will establish support processes, own Tier 3 escalation paths, and mentor junior support engineers across the team. This role requires deep expertise in cloud platforms, Linux systems, observability tooling, and infrastructure-as-code (IaC), combined with strong communication skills to bridge technical and operational teams.
Key Responsibilities π§
Own and resolve Tier 3 infrastructure and platform escalations, including incidents affecting cloud services, Kubernetes clusters, CI/CD pipelines, and developer tooling.
Investigate and perform root-cause analysis on production incidents, documenting findings and driving structured post-incident reviews.
Partner with SRE and engineering teams to translate recurring support issues into permanent fixes, automation improvements, or platform enhancements.
Monitor platform health using observability tools and respond proactively to alerts before they escalate into customer-impacting incidents.
Develop and maintain runbooks, troubleshooting guides, and knowledge-base articles to accelerate resolution and enable team self-service.
Define and improve escalation workflows, on-call rotations, and incident management processes.
Mentor and coach junior and mid-level support engineers on debugging techniques, tooling, and professional communication standards.
Collaborate with product and infrastructure teams to provide structured feedback on platform stability, documentation gaps, and tooling usability.
Required Qualifications π¨βπ»π©βπΌ
5+ years of professional experience in technical support, site reliability engineering, or infrastructure engineering, with at least 3 years in a senior or escalation-owner capacity.
Deep proficiency with at least one major cloud platform (AWS, GCP, or Azure), including compute, networking, storage, and managed services.
Strong experience with Linux/Unix systems administration, including process management, file systems, networking, and performance tuning.
Hands-on experience with containerization and orchestration technologies, particularly Docker and Kubernetes.
Proficiency in scripting languages such as Python or Bash for automation, diagnostics, and tooling development.
Experience with observability and monitoring platforms (e.g., Datadog, Prometheus, Grafana, PagerDuty) for incident detection and root-cause analysis.
Demonstrated experience managing production incidents, including structured post-incident reviews and blameless retrospectives.
Preferred Qualifications
Experience with infrastructure-as-code (IaC) tools such as Terraform, Ansible, or Pulumi.
Familiarity with CI/CD platforms (e.g., GitHub Actions, Jenkins, ArgoCD) and supporting engineering teams that use them.
Background in SRE practices, including service-level objectives (SLOs), error budgets, and reliability engineering principles.
Relevant cloud or infrastructure certifications (e.g., AWS Solutions Architect, GCP Professional Cloud Engineer, Certified Kubernetes Administrator).
Experience with log aggregation and distributed tracing platforms (e.g., ELK Stack, Jaeger, OpenTelemetry).
Technical Skills Required
Category
Technologies
Cloud Platforms
AWS, GCP, or Azure (compute, networking, storage, IAM)
Operating Systems
Linux (RHEL, Ubuntu, Debian), Unix
Containers & Orchestration
Docker, Kubernetes, Helm
Scripting & Automation
Python, Bash
Observability & Monitoring
Datadog, Prometheus, Grafana, PagerDuty, CloudWatch
Infrastructure-as-Code
Terraform, Ansible, Pulumi
CI/CD
GitHub Actions, Jenkins, ArgoCD, CircleCI
Networking
TCP/IP, DNS, HTTP/S, VPN, load balancers, firewalls
Logging & Tracing
ELK Stack, OpenTelemetry, Jaeger, Splunk
Ticketing & ITSM
Jira, ServiceNow, Zendesk, or equivalent
Soft Skills Required
Strong written and verbal communication skills β able to convey complex infrastructure issues clearly to both technical and non-technical audiences.
High ownership under pressure β comfortable driving resolution during critical incidents with minimal guidance.
Systematic and methodical problem-solving approach, applying structured root-cause analysis to unfamiliar or ambiguous failures.
Collaborative and empathetic approach to working with engineering teams and end users, maintaining professionalism during high-stress escalations.
Detail-oriented documentation habits that improve institutional knowledge and reduce repeat escalations.
Ability to mentor and transfer knowledge effectively to junior engineers.
Continuous improvement mindset β proactively identifies patterns in support volume and drives systemic fixes rather than repeat resolutions.
Education Requirements
Bachelor's degree in Computer Science, Information Technology, Systems Engineering, or a related technical field β or equivalent professional experience. Relevant certifications (e.g., AWS Solutions Architect, Certified Kubernetes Administrator, ITIL Foundation) are a plus but not required.