Site Reliability Engineering (SRE) Lead - VP
Dublin, L, IE, D02 KF20
Who we are
Sumitomo Mitsui Finance Dublin Limited (SMFD) is a wholly owned subsidiary of SMBC and is growing rapidly as a Centre of Excellence for the bank’s universal banking business across EMEA. It provides a range of technology and operational support services, aligned to SMBC’s growth, innovation, and transformation strategies.
Purpose of the Role
We are establishing a Site Reliability Engineering (SRE) function within the Infrastructure group to enhance our delivery and support of production services to the business.
This role will establish our SRE capability then build and manage a team who will operate across all infrastructure domains, supporting systems in both on-premises datacentres and in Cloud, ensuring that our systems remain reliable, secure, scalable and performant.
Background and Organizational Context
Working in close alignment with Operations, Automation and Engineering teams as well as other key stakeholders, such as Security and Architecture groups, this is a key role within the IT function, providing a bridge between the operational and engineering aspects of Infrastructure delivery.
The successful candidate will bring significant previous experience and provide a hands-on leadership style, as well as a structured and strategic approach to their work. The ability to effectively context-switch between leading on incident response, establishing and maturing standards and fostering effective relationships across the group will be crucial.
The role is distinct from pure operations management or engineering roles, and owns technical reliability outcomes including release readiness, platform stability, and stakeholder confidence.
The EMEA business overall is approximately 5000 seats and spans it’s IT operations across two UK datacentres.
Reliability is the core outcome owned by this role.
Key Accountabilities and Responsibilities
- Establish and lead the SRE capability within Infrastructure, including building, mentoring, and managing a high-performing EMEA SRE team across on-premises datacentres and cloud environments.
- Own technical reliability outcomes, including release readiness, platform stability, reliability risk visibility, and stakeholder confidence, with reliability positioned as the core outcome of the role.
- Define and implement core SRE practices, including reliability reviews, toil reduction, operational readiness assessments, SLAs, SLOs, error budgets, and embedding reliability into service and platform design from the outset.
- Drive monitoring, observability, capacity management, and automation adoption, using metrics, logging, tracing, Infrastructure as Code, CI/CD, and operational tooling to improve system health and reduce manual effort.
- Act as a bridge between Operations, Engineering, Automation, Security, Architecture, and other stakeholders, embedding SRE practices into BAU operations and supporting effective service governance, release governance, and longer-term planning.
- Lead operational resilience and risk management activities, including participation in change management forums, promoting blameless post-mortems, ensuring compliance with regulatory and internal governance standards, mitigating operational risks, and acting as an escalation point via the on-call rota.
Knowledge, Skills and Experience
- Proven experience in SRE leadership in complex, enterprise scale environments
- Strong technical understanding of:
- Infrastructure platforms including compute, virtualisation, storage, and networking across on-premises and cloud environments
- Monitoring and observability platforms
- Infrastructure as Code – Terraform, Ansible
- Version control / CICD – Git etc
- Hands-on technical capability, with the ability to engage in tooling, processes, and remediation approaches where required
- Ability to operate in tightly regulated environments, ideally within financial services
- Experience building or maturing SRE functions at enterprise scale, including defining processes, governance, and reporting
- Excellent communication and stakeholder management skills, with the ability to engage across engineering, operations, and IT security teams
- Relevant industry certifications across infrastructure technology stacks (e.g VMware, Kubernetes, modern datacentre networks - on-premises and cloud) or the equivalent, demonstrable level of professional experience
- Experience with regulatory frameworks such as DORA, PRA, or ECB operational resilience requirements or similar