Join a team that builds secure, scalable data pipelines and platforms for collection, access and analytics.
As a Lead data Engineer, within the Corporate Sector you will lead architecture and engineering for scalable backup, recovery, immutable-data protection, and recovery-assurance services across data platforms and storage tiers and deliver data collection, storage, access, and analytics platform solutions in a secure, stable, scalable way, with clear SLOs/SLAs and operational readiness gates.
Job Responsibilities
Observability (Dashboards: Grafana, Dynatrace, Tableau):
- Design, build, and maintain role-based observability dashboards for data platforms and backup/recovery services using Grafana, Dynatrace, and Tableau, providing real-time and historical visibility into platform health, performance, and risk posture.
- Define and govern dashboard standards and operating model (golden dashboards, drill-down paths, naming/tagging conventions, ownership, and lifecycle management) to drive consistent adoption across Engineering, SRE, Operations, and Risk.
- Ensure end-to-end instrumentation and telemetry quality so dashboards surface actionable signals (availability, latency, throughput, error rates, saturation) and platform KPIs (job success/failure, lag/backlog, data freshness, restore success, retention coverage).
- Establish and publish SLIs/SLOs and operational health KPIs; translate these into Grafana/Dynatrace visualizations and Tableau reporting for executive and governance stakeholders.
Data Architecture, Modeling, and Governance:
- Generate and govern data models using firmwide tooling; apply linear algebra, statistical, and geometric algorithms where relevant for modeling and optimization.
- Own and continuously improve the strategy for database backup, recovery, archiving, retention, and restore testing across relational and NoSQL estates.
Platform Engineering & Automation:
- Drive an automation-first, API-led, self-service approach that reduces operational toil and improves resilience and customer experience.
- Build cloud-native capabilities using AWS services, infrastructure as code, CI/CD, and modern software engineering practices.
Reliability Engineering (SRE) & Operational Excellence:
- Embed SRE principles by defining and tracking reliability metrics (SLIs/SLOs), implementing observability standards, leading incident learning, and maintaining runbooks and recovery playbooks.
- Participate in an on-call and incident leadership rotation, acting as an escalation point for complex platform and data reliability issues.
Security, Controls, and Recovery Assurance:
- Assess and report on access control effectiveness and data asset security posture; partner with security/risk to remediate gaps.
- Design preventive/detective controls, policy guardrails, automated validation, and audit-ready evidence for platform controls and recovery readiness.
- Deliver telemetry, reporting, and operational intelligence for backup health, compliance, and recovery assurance (coverage, success rates, RPO/RTO attainment).
Technical Leadership & Delivery Management:
- Set engineering standards, perform high-quality code reviews, mentor engineers, and influence stakeholders across product, architecture, security, and SRE.
- Coordinate cross-team delivery with explicit dependencies, milestones, and measurable outcomes; manage technical debt and balance reliability/security work alongside feature delivery.
Core Engineering Scope (Hands-on):
- Data architecture & modeling (lakehouse/warehouse patterns, dimensional modeling, data contracts)
- Advanced SQL (performance tuning, warehousing design)
- Pipeline engineering & orchestration (reliable batch workflows, backfills, SLAs)
- Distributed processing (e.g., Spark; partitioning, joins/shuffles, file formats)
Data quality & testing (automated checks: schema/freshness/volume/business rules)
Required Qualifications, Capabilities, and Skills
- Typically 8+ years of applied engineering experience (software engineering or related discipline), including leading production-critical systems; experience with both relational and NoSQL databases; Strong capability in: observability & operations (monitoring, lineage, incident response/runbooks), security/privacy/governance (least privilege, encryption concepts, auditability), system design & tradeoff analysis (scalability, latency, reliability, cost), and technical leadership (standards, mentoring, stakeholder alignment).
- Proficient across the data lifecycle (ingestion, modeling, storage, serving, governance, operations) with demonstrated experience building distributed systems, platform services, APIs, and automation frameworks.
- Strong knowledge of enterprise data protection, backup/recovery, cyber resilience, immutable storage, retention, and recovery testing.
- Advanced AWS experience spanning backup/data protection services, IAM/security, networking fundamentals, observability, and infrastructure automation.
- Hands-on experience with observability/monitoring platforms such as Dynatrace and Grafana and proficiency in Python and at least one additional language (e.g., Java, Go, C#).
- Experience with CI/CD platforms (e.g., Jenkins, GitLab, GitHub Actions) and implementing database backup/recovery/archiving strategies with measurable RPO/RTO targets.
- Proficient knowledge of linear algebra, statistical, and geometric algorithms; ability to translate security, control, and regulatory requirements into engineered solutions and operational processes.
Preferred Qualifications, Capabilities, and Skills
- Experience with Kubernetes/OpenShift and containerized platform operations.
- Experience with database platforms and cyber-recovery / isolated recovery solutions.
- Experience building operational analytics (health/compliance dashboards, evidence automation) for regulated financial services environments.