Senior Platform Engineer - Data Infrastructure
Employer · es
Senior Platform Engineer - Data Infrastructure at Employer, based in es. This is a permanent role.
- Salary
- Competitive
- Location
- es
- Contract
- Permanent
- Posted
- 1 hour ago
- Closes
- 24 Sep 2026
- Sector
- DevOps Engineer
Reference j_9803aaa4
About the role
About the team The Data Infrastructure team looks after the data backbone: how those events are taken in, processed, stored, and served, and everything that keeps them reliable, fast, and affordable to run. When a piece of data infrastructure becomes important enough that several teams depend on it, looking after it properly becomes our job. Today that estate covers Kafka (Amazon MSK), ClickHouse (on Amazon EKS), Amazon Aurora, and MongoDB (Atlas) across several AWS regions. We are now adding search, vector, and in-memory stores (OpenSearch, Redis, ElastiCache) to support new product work. You would join a small group of experienced engineers who own their areas from the first design through to running them in production. We like DevOps and CI/CD, we automate the things worth automating, and above all we care about doing the work well. Why this role matters Almost everything our customers rely on sits on top of this layer. When a product team promises customers high availability for what they have built, that promise only holds if the Kafka, the databases, and the streaming underneath hold too. Your work allows Nexthink to move quickly and trust the ground they are standing on. What you will do Design and improve our data infrastructure alongside the Architecture, Product Engineering, and Security teams, following cloud-native good practice. Build the tooling and automation that provisions and scales it, with a real focus on resilience and elasticity, and make it self-service (Crossplane) so product teams can build on it with confidence. Bring in and run new kinds of data store (search, vector, in-memory and caching) to the same standard as everything else we look after. Spend time with the product and feature teams, understand what they actually need, and bring that back to shape the platform. Plan for the bad days: disaster recovery and cross-region replication, with clear RPO and RTO targets. Keep an eye on availability, performance, and observability (Datadog) so you spot trouble before it turns into an incident. Handle incidents from start to finish: spot them, work out what happened, fix them to SLA, and write the post-mortem. You will share the on-call rotation with the rest of the team. How we work A few things that are true about this team and, we think, make it a good place to build: It is a small team with real ownership. You look after your own areas and make the call on what is best for the platform, rather than working through a queue of tickets someone else has written. We do not work in a silo. We spend a lot of time with the teams who build on us, including a good deal of design work together, so the platform grows around what people genuinely need. The scale keeps it interesting. We run across several regions on modern AWS tooling, which is a good deal more involved, and more rewarding, than a single-region setup. We are improving something that already works. The platform is live and doing its job. Our task is to take it from good to really good, so the problems are about pushing things further rather than wiring up the basics. Like everyone here we use AI tools day to day, and we are slowly moving the more repetitive operational work onto automation. We do it carefully, because infrastructure is not somewhere to be careless. What you will bring The essentials A solid cloud infrastructure background and good software-engineering habits, usually around 5+ years as a Software, DevOps, Platform, or Site Reliability Engineer. A Computer Science degree or the equivalent in real experience. Hands-on AWS across several regions, with infrastructure as code (Terraform or OpenTofu) and pipelines (GitHub Actions, GitOps / Flux). Our platform runs entirely on AWS. Kubernetes and Amazon EKS on Linux with containers (Docker): running, upgrading, and patching workloads in production. Confidence in Go and/or Python. Clear thinking and good communication, and the organisation to keep several things moving at once, all in English. Nice to have Experience running distributed and streaming systems, especially Apache Kafka (AWS MSK). Running database clusters (Amazon Aurora, ClickHouse, MongoDB and the like): configuration, backup and restore, and day-to-day production operations, including operators for data systems on EKS. Building self-service internal developer platforms, ideally with Crossplane or something similar. Familiarity with search or vector databases (OpenSearch, Elasticsearch) and in-memory or caching stores (Redis, Amazon ElastiCache). Some exposure to disaster recovery and cross-region replication (RPO / RTO). Curious what we build on? Have a look at how Nexthink scales to trillions of events per day with Amazon MSK , powers real-time alerts with Amazon Managed Service for Apache Flink , and builds enterprise AI agents with fine-tuned LLMs on Amazon SageMaker . We are the pioneers and trailblazers of a global IT Market Category (DEX) that is shaping the future of how the world works, giving…
Reference: j_9803aaa4 · Posted 1 hour ago · Closes 24 Sep 2026 · Listed via Employer
Apply for this job
This role is listed via Employer. Applications are handled on the employer's site.
Apply on employer siteOpens the employer's website in a new tab.
Safe applying: a genuine employer will never ask you to pay for a DBS check, training or equipment, or move you onto WhatsApp before you are hired. If this listing does, report it and do not pay anything.