About the role
Your Impact
We are recruiting for a Senior or Principal AI Platform Engineer to help shape, build and operate the infrastructure and services that underpin our AI platform, providing engineers with secure, reliable access to AI models and inference services.
What you’ll do
Take ownership of the platform layer above our managed GPU infrastructure, including model serving, inference runtimes, AI gateways and supporting services. You will help make these capabilities easy for engineering teams to consume while maintaining appropriate controls around security, performance and availability.
Using technologies such as containers, Kubernetes, vLLM and AI gateway platforms, you will deploy and operate scalable inference services, improve observability and performance, and investigate complex technical issues across the platform. You will also help establish engineering standards for operating AI services within a secure enterprise environment, including cases where the target system has no route to the internet.
We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. You do not need to meet every requirement. What matters most is sound engineering judgement, adaptability and the ability to make a meaningful contribution.
What you’ll bring
Experience operating shared or business-critical platforms
Familiarity with large language models and the practical considerations of running them
Experience troubleshooting distributed systems and performance issues
Familiarity with infrastructure as code and CI/CD
Effective communication and stakeholder engagement
Curiosity and a drive for continuous improvement
Key areas we value experience in:
Linux, application and server administration
Containers and container orchestration, such as Docker and Kubernetes
AI inference and model serving platforms, such as vLLM, SGLang or similar
AI gateways and API management, such as LiteLLM, Bifrost or similar
GPU workloads, including performance, utilisation and resource management
Programming and scripting, such as Python, Bash or Go
Monitoring, observability and SRE practices
Secure, resilient and scalable service design
Authenticat…