[Remote] Site Reliability Engineering Technical Leader
Auto ImportNote: The job is a remote job and is open to candidates in USA. Cisco develops solutions that connect and protect organizations across physical and digital infrastructure. The Site Reliability Engineering Technical Leader will serve as the most senior technical individual contributor on the CloudOps team, owning critical incident strategy, complex customer environments, automation, infrastructure architecture, and operational improvements while providing technical leadership across the organization.
Responsibilities
- You will be the most senior technical individual contributor on the team — setting the bar for technical judgment, architectural thinking, and operational excellence across the organization
- You will own the technical strategy for how the team responds to, learns from, and prevents critical incidents, serving as the designated point of escalation during P1/P2 events and making high-stakes, risk-informed decisions that protect customer uptime and trust
- You will hold deep ownership of the most complex customer stacks and shape the team's automation direction and infrastructure architecture decisions with regional and business-wide impact
- Your insights will directly shape how Splunk Cloud's operational processes evolve, and you will be the trusted counterpart that Engineering, Customer Success, and Release Management turn to when navigating the most demanding customer situations
- You will raise the technical floor of everyone around you — senior engineers will calibrate their judgment against yours, and the frameworks you build will outlast any single incident
Skills
- Bachelors + 12 years of related experience, or Masters + 8 years of related experience, or PhD + 5 years of related experience
- 7+ years of experience in SRE, cloud operations, and systems engineering with Linux administration
- 6+ years of hands-on experience across AWS, GCP, or Azure
- 5+ years of experience in at least one scripting or automation language such as Python or Go/golang
- 5+ years of experience communicating across engineering, customer-facing, and senior leadership audiences
- 4+ years experience leading post-mortems and root cause analysis for high-severity events driving systemic improvements, closing process gaps
- Demonstrated track record of managing high-severity customer escalations and resolving critical production issues across major cloud platforms (AWS, GCP, Azure), with enterprise-scale experience spanning architecture, deployment, infrastructure changes, migrations, and operational management — including the ability to manage multiple concurrent high-priority escalations with composure and sound judgment
- Deep technical ownership of premier and strategic enterprise customer environments, with experience providing authoritative guidance on complex configurations, upgrades, and feature rollouts, while partnering with Account Managers and cross-functional stakeholders to translate customer insights into measurable improvements to the cloud service experience
- Extensive background in Linux systems administration and advanced expertise in large-scale distributed systems architecture — with Splunk-specific knowledge (SPL, indexer clustering, Search Head Clusters, KVStore, and migration patterns at scale) considered a strong asset
- Proven experience operating and improving monitoring, alerting, and observability platforms (such as Splunk Observability or equivalent technologies), with a demonstrated ability to close operational process and runbook gaps — driving recurrence prevention, not just incident resolution
- Exceptional communication and technical leadership skills, with the ability to engage credibly across customers, cross-functional partners, and senior leadership, combined with prior experience in customer-facing technical advisory roles (Professional Services, Global Services, or Technical Account Management) and a track record of growing senior technical talent without a formal management mandate
Benefits
- Medical, dental and vision insurance, subject to Cisco’s plan eligibility rules.
- A 401(k) plan with a Cisco matching contribution.
- Paid parental leave.
- Short- and long-term disability coverage.
- Basic life insurance.
- Employees may be eligible to receive grants of Cisco restricted stock units, which vest following continued employment with Cisco for defined periods of time.
- 10 paid holidays per full calendar year, plus 1 floating holiday for non-exempt employees.
- 1 paid day off for the employee’s birthday, paid year-end holiday shutdown, and 4 paid days off for personal wellness determined by Cisco.
- Non-exempt employees receive 16 days of paid vacation time per full calendar year, accrued at a rate of 4.92 hours per pay period for full-time employees.
- Exempt employees participate in Cisco’s flexible vacation time off program, which has no defined limit on how much vacation time eligible employees may use, subject to availability and some business limitations.
- 80 hours of sick time off provided on the hire date and each January 1st thereafter, and up to 80 hours of unused sick time carried forward from one calendar year to the next.
- Additional paid time away may be requested to deal with critical or emergency issues for family members.
- Optional 10 paid days per full calendar year to volunteer.
- For non-sales roles, employees are also eligible to earn annual bonuses subject to Cisco’s policies.
- Employees on sales plans earn performance-based incentive pay on top of their base salary, subject to the applicable Cisco plan.
Company Overview
Company H1B Sponsorship