Accountabilities
- Define and implement the reliability strategy across the platform, including SLOs, SLIs, error budgets, incident practices, and reliability standards adopted by engineering teams.
- Drive major architectural decisions as infrastructure evolves, evaluating technologies and designing systems that remain scalable, resilient, observable, and maintainable.
- Design and own event-driven communication and messaging infrastructure, including the transition from synchronous patterns to durable asynchronous architectures.
- Manage and evolve cloud infrastructure on AWS, using Infrastructure as Code to automate provisioning, configuration, deployment, and operational processes.
- Ensure Kubernetes and containerized workloads scale reliably as transaction volumes and AI workloads increase.
- Build and maintain comprehensive observability through monitoring, dashboards, alerting, application performance monitoring, and distributed tracing.
- Serve as the senior escalation point for complex production incidents, leading incident response, root-cause investigations, and blameless postmortems.
- Turn incident findings into permanent improvements through architectural changes, automation, operational controls, and resilience patterns.
- Establish a continuous chaos engineering and resilience testing practice through fault injection, game days, and controlled failure experiments.
- Mentor senior and mid-level engineers while raising the technical bar for reliability engineering and influencing engineering practices across teams.
- Use AI-assisted tooling for automation, runbooks, incident analysis, and root-cause investigations, while helping establish effective AI-enabled engineering practices.
- Within the first 6–12 months, establish the platform reliability strategy, lead at least one major architectural evolution, and drive adoption of the SLO and error-budget framework across engineering teams.
Requirements
- Extensive experience in Site Reliability Engineering, Platform Engineering, DevOps, or a closely related discipline, with demonstrated ownership of production-scale systems.
- Deep expertise in event-driven architecture and messaging systems such as Kafka, NATS, or RabbitMQ, including at-least-once delivery, consumer groups, dead-letter queues, backpressure, and migrations from synchronous to asynchronous architectures.
- Strong AWS expertise across services such as EC2, VPC, IAM, S3, and RDS, combined with solid networking fundamentals.
- Hands-on Infrastructure as Code experience using Terraform, Pulumi, or similar tools, with infrastructure managed through version-controlled workflows and code reviews.
- Strong production experience with Kubernetes and Docker, including container lifecycle management, resource limits, health checks, and orchestration at scale.
- Proven observability expertise using Datadog or equivalent platforms, including dashboards, monitoring, APM, distributed tracing, and alerting.
- Demonstrated experience defining and operating SLOs, SLIs, and error budgets across multiple services.
- Hands-on experience with chaos engineering, fault injection, game days, or resilience experiments using tools such as Gremlin, Chaos Mesh, AWS FIS, or similar technologies.
- Strong distributed systems debugging skills, with experience diagnosing asynchronous workflows, cascading failures, and complex production incidents.
- Ability to code for automation and engineering tooling using Go, Python, or a similar programming language.
- Solid database knowledge across SQL and NoSQL technologies, particularly PostgreSQL, MongoDB, and Redis, including indexing, replication, and performance optimization.
- Proven technical leadership experience, including setting reliability standards, influencing architecture across teams, and mentoring engineers.
- Advanced written and spoken English communication skills.
- Experience with AI or MLOps infrastructure, including model serving, LLM inference, GPU/resource management, or AI agent observability, is highly advantageous.
- Familiarity with multi-tenant container platforms and customer workload infrastructure is a plus.
- Experience with data pipelines and orchestration tools such as Airflow or Prefect, and data platforms such as Databricks, Snowflake, or BigQuery, is beneficial.
- Familiarity with incident management platforms such as PagerDuty, Opsgenie, or incident.io is an advantage.
- Experience in the payments industry is preferred.
- Additional experience with ECS, s6-overlay, AI agent frameworks, or Spanish proficiency is a plus.
Benefits
- Competitive compensation.
- Fully remote working environment with the flexibility to work from different locations.
- One-time home office allowance to help create an effective workspace.
- Company-provided work equipment.
- Stock options.
- Health plan available wherever you are.
- Flexible days off.
- Access to language, professional, and personal development courses.
- Opportunity to work on globally scaled infrastructure supporting complex payment and AI workloads.
- Significant technical ownership and influence over reliability strategy, architecture, and engineering standards.
- Collaborative international environment with opportunities to mentor engineers and shape organization-wide engineering practices.
🇧🇷 Essa vaga exige inglês. Você está pronto?
A DevSpeak Academy prepara desenvolvedores brasileiros para conquistar vagas internacionais. Domine o inglês técnico com professores que entendem o mundo dev.
Conheça a DevSpeak Academy