Lead Site Reliability Engineer
Leidos — Remote
Posted: 2026-09-22
Job Description
• Lead reliability engineering for a mission-critical USAF multi-cloud environment, ensuring availability, performance, scalability, and resilience.
• Drive automation and Infrastructure as Code using AWS, Azure, Kubernetes, and CI/CD pipelines.
• Establish observability and performance management across monitoring, logging, tracing, SLOs, dashboards, alerting, and capacity planning.
• Lead incident response, complex troubleshooting, root-cause analysis, corrective actions, and continuous improvement.
• Requires a bachelor’s degree with 12+ years of experience or a master’s degree with 10+ years, U.S. citizenship, Secret clearance, and CompTIA Security+ or equivalent IAT Level 2 certification.