seniorremoteRemote

Site Reliability Engineer (SRE)

Hola!

Sobre el puesto

About the role

We're looking for a seasoned Site Reliability Engineer who treats the systems they own like their own. You'll be the person who spots the problem before it becomes an incident, keeps our CI/CD pipelines humming, and makes sure our cloud stacks stay healthy around the clock.

This is a hands-on, high-ownership role embedded in the engineering org. You'll work closely with development teams to bridge the gap between shipping fast and staying reliable — a true partner, not just a support function.

What you'll do

  • Monitor and maintain production infrastructure across AWS and GCP
  • Own the reliability and performance of CI/CD pipelines end-to-end
  • Manage and operate cloud stacks, ensuring uptime, scalability, and cost efficiency
  • Lead incident response: detect, triage, resolve, and write clear post-mortems
  • Proactively identify risks and toil, and drive initiatives to eliminate them
  • Collaborate daily with software engineers to embed reliability best practices early in the development cycle
  • Define and track SLIs, SLOs, and error budgets

The kind of person we're looking for

Technical chops matter, but attitude matters more. We want someone who takes ownership without being asked, raises their hand when something looks off, and communicates openly with the people around them. If you thrive in a collaborative, fast-moving environment and take pride in keeping things running smoothly, you'll fit right in.

Requisitos

Must have

  • 5+ years of experience in a SRE, DevOps, or cloud infrastructure role
  • Hands-on expertise with AWS and GCP — provisioning, networking, security, cost management
  • Proven experience managing and troubleshooting CI/CD pipelines in production
  • Strong scripting skills (Python, Bash, or similar) for automation and tooling
  • Experience with Infrastructure as Code — Terraform or equivalent
  • Solid understanding of observability: logging, metrics, alerting, and tracing
  • Experience running containerised workloads (Kubernetes, Docker)
  • Demonstrated ownership mindset: you drive problems to resolution, not just escalation

Nice to have

  • Experience with monitoring tools such as Datadog, Prometheus, or Grafana
  • Familiarity with on-call rotations and incident management frameworks (e.g. PagerDuty)
  • SRE or cloud certifications (AWS, GCP)
  • Background working in a product-led or startup environment

Beneficios Talentic Flex: queremos que encuentres el balance entre tu vida personal y profesional, por eso trabajamos en modalidad 100% remoto Cumple: te regalamos el día libre para que lo disfrutes como más te guste. Salud: la mejor medicina privada es para ti Talentic Kit: un regalo especial que incluye mochila, botella, taza, cuaderno, lapicera, tote bag, notebook con su funda y soporte, te espera en tu onboarding. TalentiGifts: celebramos tus momentos especiales haciéndolos aún más memorables con un hermoso regalo en cada uno de ellos.

Postular a este puesto

Inicia sesión para postular — solo toma un momento.

Enviar postulación