Senior Data Engineer (London, UK, United Kingdom, Aberdeen City)
About this opportunity
THE AI ACCELERATOR
Most diseases are still poorly understood at a biological level. Despite decades of research, the causal mechanisms driving many conditions remain unclear, limiting our ability to identify the right targets, design the right interventions and bring the right medicines to patients.
The AI Accelerator exists to change that. Based in London and sitting within Computational Innovation (@computationalinnovation), a global organisation spanning computational biology, human genetics, data excellence and AI, the Accelerator’s mission is to build production-quality AI capabilities that deepen our understanding of disease biology and increase probability of success.
We do this by applying neural-based methods across the biomedical data landscape to integrate heterogeneous, multimodal data sources, infer biological relationships and embed causal thinking into what we build. The goal is not just to predict but to explain and understand why disease occurs.
It could be electronic health records and medical imaging to support patient segmentation. It could be ‘omics data to identify novel therapeutic targets. It could be predicting transcriptional change for a given disease-causing variant. It could be simulating the effect of modulating a target of interest.
A core component of the AI Accelerator is AI Enablement, a team that provides the support framework to make our ambitions a technical reality. It could be provisioning integrated, multimodal biomedical data for model training and inference. It could be managing the lifecycle of models provided by AI Systems. It could be working with IT to ensure the right infrastructure and tooling are in place. AI Enablement ensures that the model builders can focus on the technology and that Computational Innovation’s downstream users can leverage accelerator capabilities for real portfolio impact.
THE POSITION
We are seeking two Senior Data Engineers, one with a medical imaging focus and one with a multi-omics focus, to join the AI Enablement team and contribute to the design and delivery of robust data engineering pipelines that transform harmonised biomedical datasets into AI-ready, integrated assets across multi-omics, clinical and health records, and medical imaging data.
You will be an experienced, independent data engineer within AI Enablement, owning significant data engineering workstreams within the broader technical direction and architecture set by the Senior Staff Data Engineer. The pipelines and integrated datasets you build will enable model training, fine-tuning and inference in a production setting.
Key Responsibilities
Build and maintain entity linking pipelines that connect patients, samples and other biomedical entities across modalities such as imaging, clinical and multi-omics records
Build and maintain cross-modal integration pipelines that combine linked imaging, mulit-omics and clinical records, each with different formats, scales and structures, into unified assets ready for multimodal model training, fine-tuning and inference
Ensure pipelines and datasets are built and operated in accordance with data access permissions, consent conditions and usage restrictions, including within Trusted Research Environments or other controlled access settings
Build and maintain biomedical benchmark datasets with versioning and documentation
Write clean, well-tested, well-documented code that meets the required engineering standards and contribute to code reviews within the data engineering team
Stay current with advances in data engineering tooling and practices relevant to biomedical AI
Required Qualifications
PhD in Machine Learning, Computer Science, Bioinformatics, Computational Biology or a related quantitative field
Strong hands-on experience in data engineering for machine learning with proficiency in modern data engineering tech stacks
Experience working with medical imaging modalities (radiology and/or histopathology) or multi-omics modalities (transcriptomics, proteomics) and a working knowledge of other biomedical data modalities sufficient to support cross-modal integration
Practical experience with entity linking or record linkage, ideally in a biomedical or clinical context, and a strong understanding of biomedical data characteristics such as variant data formats, expression matrices and clinical coding standards such as SNOMED and ICD-10
Familiarity with data governance frameworks applicable to biomedical and clinical data and Trusted Research Environments or controlled access biomedical data environments
This is a hybrid role with approximately 3 days a week in the office
WHY THIS IS A GREAT PLACE TO WORK
Boehringer Ingelheim has been recognised as a Top Employer in the UK, demonstrating our commitment to building an exceptional workplace through strong people practices and supportive HR policies.
To learn more about why BI is a great place to work, visit:
https://www.boehringer-ingelheim.co.uk/careers/uk-careers/why-great-place-work
]]>
Job details
How this role compares
Computed from every other active Data & Digital role in our database, not just this employer's listings.
We currently track 92 comparable Senior Data & Digital roles across 23 biopharma companies.
Salary context
16 of 92 peers report a salary range (USD, annualized)
Peers share this role's job function and a matching or adjacent seniority level -- not necessarily the same therapeutic area or country.
Where these roles are based
Top locations among the 92 comparable roles
+ 7 more countries
Seniority mix
92 of 92 peers have a known seniority level
Therapeutic area mix
2 of 92 peers have a known therapeutic area; the rest are genuinely unlabeled, not hidden
Similar opportunities
The closest matches from our peer group, ranked by how similar they are, not how well you'd qualify for them -- treat this as market context, not a guaranteed shortlist; a weak match is labeled as one below.
How we calculate "similar"
No black box, no LLM guesswork: a deterministic score built from four normalized attributes. Here's this role's own peer group at different match levels, so you can see the mechanism, not just the result.
Every comparison starts from the same 100-point budget: 25 for working in the same function, 40 for the same therapeutic area, 20 for the same or adjacent seniority, 15 for the same country. A dimension we can't confirm on both sides contributes nothing, never a guess, never a free pass.
0 points, never a partial guess. A role we know almost nothing about beyond its function bottoms out at 25%; it never inflates to 100% just because there's little to compare against. Seniority uses a defined ladder (Associate → Manager → Associate Director → Senior → Principal → Director → Senior Director → Executive/VP) so "Director" and "Senior Director" count as adjacent, but "Director" and "Executive/VP" do not.