Site Reliability Engineer - Production Support (25491) Cupertino, California

Salary: USD60 - USD70 per hour

Site Reliability Engineer - Production Support

W2 Contract

Pay Rate: $60 - $70 per hour

Location: Cupertino, CA - Remote Role

Duties and Responsibilities:

  • Production support / application support / NOC / service reliability / incident management. Screen out: security analysts whose Splunk work is log-based threat detection.
  • Watch service health across production; catch, triage, and drive issues to resolution.
  • Run production support against SLAs and work the incident lifecycle directly with the incident management (IM) team. 
  • Monitor and investigate using Splunk (and Datadog, where in play).
  • Support the application stack in cloud and datacenter environments.
  • Non-prod and prod deployments using our in-house CD tool; monitor deployments and roll back / debug when they fail.
  • Keep Git repos current; debug deployment scripts.

Requirements and Qualifications:

  • Hands-on production support/application support in an SLA-bound environment. Must be able to describe what they personally did during a real incident.
  • Incident management experience: triage, severity assessment, escalation, bridge participation, driving to resolution, RCA follow-up. Direct interaction with an IM team.
  • Splunk is required. They will be asked what commands/searches they run.
  • SOC/production monitoring
  • Troubleshooting depth demonstrable at the command line: log tracing, connection failures, latency, service crashes. Conceptual answers fail here.
  • Linux/Unix fluency.
  • Containerized workloads: Docker + Kubernetes (managing, not just using)
  • Cloud: AWS primary; Azure/AKS acceptable and precedented. Datadog - Naveen called it out as a plus point. New to this team's stack vs. prior reqs in Lighthouse is likely a growing surface.
  • Payments/banking production environment (card networks, Visa/Mastercard, wallets, issuer/acquirer processing).
  • Java application troubleshooting.

Preferred Qualifications

  • CI/CD tooling - Jenkins, specifically, can debug the scripts behind a deployment, not just trigger one.
  • Python and/or shell scripting for automation and runbook work.
  • Load balancers, SSL/TLS and certificate management, DNS, and general network troubleshooting.
  • Config management (Ansible/Chef/Puppet), autoscaling design in Kubernetes.
  • Terraform / infrastructure-as-code exposure.
  • ServiceNow / Jira or equivalent ITSM incident tooling; ITIL familiarity.

 

Bayside Solutions, Inc. is not able to sponsor any candidates at this time. Additionally, candidates for this position must qualify as a W2 candidate

Bayside Solutions, Inc. may collect your personal information during the position application process. Please reference Bayside Solutions, Inc.'s CCPA Privacy Policy at www.baysidesolutions.com.