ITDS Portugal

ITDS is a leader in outsourcing IT engineers and works with various web and mobile technologies for over 30 global clients. It has been recognized as one of the 1000 fastest-growing companies in Europe for three consecutive years, Great Place to Work, and the Forbes Diamond award in 2023. ITDS currently has more than 600 IT professionals working in Portugal, Poland, and the Netherlands.
About company

Site Reliability Engineer – Cloud Infrastructure

On-site

location

date August 21, 2026

Empower Cloud Resilience — Drive Innovation and Reliability at Scale!

Portugal-based opportunity with a fully remote work arrangement (up to 5 days per week).

As a Site Reliability Engineer (AWS), you will own the reliability, scalability, and operational health of our client's AWS-based platform — a leading player in the cloud and infrastructure domain. This is a hands-on engineering role with a systems mindset: you'll define and defend SLOs, drive down toil through automation, lead incident response, and build the tooling and infrastructure-as-code that lets development teams ship safely and fast.

Your main responsibilities:

  • Define, measure, and enforce SLOs, SLIs, and error budgets, and use them to drive engineering priorities and release decisions.
  • Own the incident lifecycle — detection, response, mitigation, and blameless post-mortems — reducing mean time to recovery through better tooling and runbooks.
  • Design and maintain highly available, fault-tolerant AWS architectures across compute, networking, storage, and data layers.
  • Build and maintain infrastructure-as-code and CI/CD pipelines that make deployments repeatable, auditable, and low-risk.
  • Systematically identify and eliminate toil through automation, replacing manual operational work with self-healing systems.
  • Own observability — metrics, logging, tracing, alerting — so problems surface before they reach customers.
  • Partner with development teams on capacity planning, performance tuning, cost optimization, and production-readiness reviews.
  • Improve the security and compliance posture of the cloud environment in collaboration with security teams.
  • Participate in and improve the on-call rotation, tuning alerts to reduce noise and fatigue.

You're ideal for this role if you have:

  • Strong production experience operating systems at scale on AWS, including core services (EC2, ECS/EKS, Lambda, VPC, IAM, S3, RDS, CloudWatch).
  • Solid grounding in SRE principles — SLOs, error budgets, toil reduction, blameless post-mortems.
  • Hands-on expertise with infrastructure-as-code (Terraform preferred; CloudFormation or CDK acceptable).
  • Proficiency with container orchestration (Kubernetes / EKS) and CI/CD tooling.
  • Strong scripting and automation skills in at least one language (Python, Go, or Bash).
  • Deep experience with observability stacks (Prometheus, Grafana, Datadog, ELK, or equivalents).
  • Demonstrated ownership of incident response and on-call in a production environment.
  • Sound understanding of networking, Linux internals, and distributed-systems failure modes.

Nice to have (not required):

  • AWS certifications (Solutions Architect, DevOps Engineer, or SysOps).
  • Experience with multi-account / multi-region AWS setups and cost governance.
  • Chaos engineering or resilience-testing experience.
  • Service mesh, GitOps (ArgoCD / Flux), or policy-as-code (OPA).
  • Regulated-industry background (financial services, healthcare).

Eligibility:

  • Only candidates with a legal right to work in Portugal or the wider European Union will be considered.
  • Only candidates based in Portugal will be considered.

Interested? Apply now and include your CV (preferably in English) along with a statement confirming your consent to the processing and storage of your personal data.

Apply here

Contacts and Address