Production Engineering Lead in Latchmere, Greater London

Posted by Ncounter Limited

Production Engineering Lead

£200,000-£220,000
London, New York, Montreal or Singapore | Permanent

Ncounter is looking for a Production Engineering Lead to take technical and leadership ownership of a global engineering team responsible for the reliability of large-scale, business-critical production infrastructure.

This is an opportunity for an experienced Production Engineer, SRE, Platform Engineer or Infrastructure Engineer who has remained technically hands-on while progressing into leadership. Financial services experience is not essential. We are particularly interested in engineers from highly scaled technology environments where availability, automation and operational excellence are fundamental.

You'll lead engineers across multiple regions, maintaining reliable production services while improving how the environment is operated. This includes leading major incident recovery, removing recurring operational problems, building automation and self-service capabilities, and reducing manual intervention across the platform.

The environment spans Linux infrastructure, Kubernetes, CI/CD, messaging, storage and distributed platform services. You'll need genuine engineering depth rather than simply managing people, with the credibility to troubleshoot complex production issues alongside your team when required.

We're looking for experience across:

  • Deep Linux/Unix systems engineering and troubleshooting
  • Python, Go and/or Bash for automation and tooling
  • Kubernetes and containerised production environments
  • CI/CD and configuration management
  • Infrastructure automation using Ansible, Chef, Puppet or similar
  • Observability using Prometheus, Grafana, ELK/OpenSearch or comparable tooling
  • Networking fundamentals including TCP/IP and production diagnostics
  • Distributed systems, messaging technologies such as Kafka, and large-scale storage
  • Leading engineers responsible for high-availability production systems
  • Major incident response, root-cause analysis and eliminating repeat failures

The successful person will lead a globally distributed function operating across critical production environments, setting engineering standards, improving processes and developing the team's capability while remaining close enough to the technology to lead from the front.

Apply Now →

Application opens at the source listing. Free for jobseekers.