This position is no longer open for applications

Principal Platform Engineer / Team Lead

Principal Platform Engineer / Team Lead (R0015924) London, England

  • Principal Platform Engineer / Team Lead - Azure Data & AI Platform

  • The Role

  • You'll lead a team senior platform engineers building and operating our Azure data and AI platform. This is a technical leadership role, not a people manager hiding from the work. You'll be 60% hands-on (architecture, complex problems, unblocking the team) and 40% leadership (mentoring, standards, strategy, team development).

    You have a mandate for continual improvement. Your job is to make the platform better, the team stronger, and the engineering practices more mature - month over month. You'll define what "good" looks like and hold the line on quality while keeping delivery moving.

    This role reports to the Head of Data Platform Engineering & Operations and has direct reports.

     

  • What You'll Actually Do

    Technical Leadership & Architecture (35%)

  • Run a full Agile workstream.
  • Own the technical vision and roadmap for the Azure platform
  • Work with the Head of Data Platform Engineering & Operations to make architectural decisions on network design, data architecture, MLOps patterns, and security models
  • Review and approve significant infrastructure changes and RFCs
  • Solve the hardest technical problems the team faces - the ones that require years of hard-won experience
  • Represent platform engineering in leadership discussions about technical direction and priorities
  • Stay hands-on: you're still writing Terraform, reviewing PRs, debugging production issues
  • Evaluate and pilot new technologies/services that could improve the platform
  • Make build-vs-buy decisions and challenge vendor claims with evidence
  • Standards, Best Practice & Quality (25%)

  • Define and document platform engineering standards - Terraform patterns, pipeline structures, security controls, documentation requirements
  • Establish and enforce code review quality bars without creating review bottlenecks
  • Implement and monitor SLIs/SLOs for platform services - what does "platform is working" mean?
  • Drive adoption of DevSecOps practices: security scanning, vulnerability management, secrets rotation, least privilege
  • Run retrospectives and blameless post-mortems - turn incidents into improvements, not finger-pointing
  • Champion technical debt management - maintain the backlog, prioritise paydown, prevent accumulation
  • Build consensus around best practices while remaining pragmatic about exceptions
  • Create and maintain architecture decision records (ADRs) and RFCs for significant choices
  • Team Development & Mentoring (20%)

  • Mentor the senior engineers - career development, technical growth, leadership skills for those who want it
  • Run 1-on-1s that matter - career conversations, skill gaps, blockers, workload balance
  • Create growth plans and ensure people have challenging work that develops them
  • Build a learning culture: lunch-and-learns, RFC reviews, pair programming, knowledge sharing
  • Identify skill gaps in the team and address through hiring, training, or reorganisation
  • Coach engineers on communication, stakeholder management, and influence without authority
  • Develop succession planning – ensuring there’s always someone to cover when needed
  • Handle performance issues directly and promptly - no festering problems
  • Process & Continual Improvement (15%)
  • Drive platform maturity - move from reactive to proactive, from manual to automated, from tribal knowledge to documented
  • Implement and iterate on team processes: sprint planning (if you sprint), incident response, on-call rotation, knowledge management
  • Track and improve key metrics: deployment frequency, lead time, MTTR, change failure rate
  • Run quarterly improvement initiatives based on team retros and pain points
  • Establish platform team rituals that add value: design reviews, demo days, incident reviews
  • Remove blockers and organisational friction that slow the team down
  • Build relationships with stakeholder teams (data engineering, ML, security, compliance) to smooth collaboration
  • Push back on unreasonable demands and protect the team from organisational chaos
  • Stakeholder Management & Communication (5%)

  • Translate technical work into business value for leadership
  • Communicate platform roadmap, incidents, and status to stakeholders
  • Manage expectations and negotiate priorities with product, data science, and other engineering teams
  • Escalate and resolve cross-team conflicts and dependencies
  • Advocate for platform investment (budget, headcount, tooling) with data-driven arguments

 

  • What We Need From You

    Required Experience

  • 10+ years in platform/infrastructure engineering with at least 3 years leading technical teams
  • Deep Azure expertise - you've architected multi-subscription environments with complex networking, security, and governance requirements
  • Databricks and data platform experience - you understand data architecture, not just infrastructure
  • MLOps/AI platform knowledge - you've built production ML systems and know the operational challenges
  • Proven track record establishing standards and best practices that teams followed
  • Mentoring and developing senior engineers - you've grown people into better engineers
  • Infrastructure as Code mastery - Terraform (or Bicep) at scale, with modules, state management, and testing
  • DevSecOps implementation - you've built secure pipelines and embedded security into engineering workflows
  • Incident response and production operations - you've been on-call, managed incidents, and improved systems based on failures
  • Critical Leadership Qualities

  • Technical credibility - the senior engineers respect your technical judgment because you've earned it
  • Clear decision-making - you gather input, make decisions, explain your reasoning, and commit
  • Comfortable with conflict - you have hard conversations about quality, performance, and standards
  • Servant leadership mindset - your job is to make the team successful, not to be the hero
  • Teaching ability - you can explain complex technical concepts and help people develop mastery
  • Intellectual humility - you admit mistakes, change your mind with new evidence, and don't have ego tied to being right

 

  • About The Role

    Howden Group Services is expanding its AI & Data Science capabilities and is looking for an AI Deployment Engineer to help accelerate our transformation and build enterprise-grade solutions with AI at their core that will be used by hundreds of colleagues across the Group.

     

    You will have a dual reporting line into the Group Head of Data Science and the Group Head of Data Operations and will bring deep technical expertise on cloud engineering and SRE with a focus on AI applications. You will be given freedom to experiment, test and bring new technologies that push the envelope on using AI to solve enterprise problems and apply them to the complex business domain of commercial insurance.

     

    Role Responsibilities

     

  • Design and develop scalable and secure infrastructure and CI/CD pipelines for AI solutions

  • Engineer and maintain production-ready RAG infrastructures and Vector Databases for our AI use cases and implement efficient retrieval strategies for data

  • Work with our AI Engineers and Data Scientists to develop and maintain a highly reliable Model Serving Layer and make our models available as scalable and reliable services, including for LLM access

  • Act as a technical and platform authority for AI Solutions and provide thought leadership to the rest of the Data Science and Data Platform team on the latest AI technologies and solutions

  • Engineer and maintain infrastructure and data pipelines for custom model fine-tuning

  • Implement and manage standardised Agent Frameworks, multi-agent systems and autonomous decision-making frameworks

  • Work with AI Engineers and Data Scientists to build appropriate observability, logging and monitoring solutions for our AI use cases, including model performance KPIs, token usage, drift and hallucination detection

  • Develop, build and maintain the Howden Enterprise AI Platform, including robust security, networking, and access strategies

  • Implement a robust FinOps framework for our AI use cases allowing for precise cost controls and chargebacks

  • Contribute to an excellent developer experience by building robust SDKs, APIs as well as drafting clear documentation and knowledge-sharing artefacts