Qutwo

Senior Site Reliability Engineer

Helsinki, FinlandHybridiKokoaikainenSenioriTekoälyrooli · hyödyntää tekoälyäTekoälyyritysMLOps, infra ja laskentaIlmoitus: englantitänään

Taidot

KubernetesTerraformHelmGitOpsPythonGoRayObservability

Kuvaus

Alkuperäinen ilmoitusteksti (englanti).

We are looking for a Senior Site Reliability Engineer to lay the foundations of a dependable, secure, and portable platform across multiple clouds.

Our product combines ML pipelines with small, security-critical SaaS services. We're building toward our first production deployments, so this is a chance to shape how reliability is done here from the start, rather than inherit someone else's choices. One week you might be designing how distributed ML workloads on Ray are scheduled and scaled on Kubernetes. The next you might be setting up observability or defining the first service-level objectives with the team.

We are not looking for someone who has used every tool in our stack. We want an engineer with the systems judgement to decide how our infrastructure should be built and why, and the depth to debug it and take responsibility for how it behaves.

Role Description

  • Building a multi-cloud platform. Design and operate Kubernetes-based infrastructure that runs consistently across multiple clouds, with sensible abstractions where portability matters.
  • Observability and reliability practices. Build the metrics, logging, tracing, and alerting we need, and help define pragmatic SLOs and incident practices that grow with the product.
  • Infrastructure as code and GitOps. Make every environment reproducible, reviewable, and automated end to end.
  • Running ML workloads well. Own scheduling, autoscaling, and cost efficiency for distributed training and inference workloads, including Ray clusters on Kubernetes.
  • Security by design. Own secrets management, network policy, workload identity, and supply-chain security.
  • Enabling the team. Build paved roads for CI/CD and deployment that make the reliable path the easy path for engineers and coding agents alike, and document decisions clearly enough that both can act on them.

Requirements

  • At least 7 years of experience building and operating production systems, including several years in an SRE, platform, or infrastructure role.
  • Deep, hands-on Kubernetes experience, including cluster operations, networking, storage, and troubleshooting.
  • Experience running systems across more than one cloud provider, and a clear view of the trade-offs involved.
  • Strong infrastructure-as-code and automation skills (e.g. Terraform, Helm, GitOps), and fluency in Python or Go for tooling.
  • Solid grasp of distributed systems, their failure modes, and the trade-offs behind reliability targets.
  • Security-first mindset: you think about trust boundaries, identity, secrets, and supply chain by default.
  • Comfort with AI-assisted development: you use coding agents well and review their output critically, including infrastructure changes.
  • Fluent in English, with a proven ability to thrive in agile, cross-functional teams.

Experience in one or more of the following areas is a plus

  • Operating Ray or other distributed compute frameworks in production.
  • Running GPU workloads, including scheduling and utilization optimization.
  • Working closely with ML researchers or running ML pipelines in production.
  • Meeting formal security or compliance requirements (e.g. ISO 27001, SOC 2) in a small company.
  • Cloud cost management and FinOps practices.
  • Familiarity with quantum computing concepts at a high level, or curiosity about them.

Lähde: Teamtailor (työnantajan rekrytointisivu)

Tilaa tämä haku

Saat uudet paikat haulla "MLOps, infra ja laskenta" sähköpostiisi. Vain kun uusia paikkoja on.

Lähetämme vahvistuslinkin. Tallennamme vain osoitteen ja hakuehdot, ja voit perua tilauksen jokaisen viestin lopusta.

Tietoa yrityksestä

Samankaltaiset paikat