Lilly Posted August 8, 2026

Sr. Principal Data Engineer - Lakehouse Architecture

Indianapolis, Indiana, United States of America FULL_TIME
Data & Digital Principal

Lilly is the source of truth for this posting and owns the application process. We surface normalized context and market comparison you won't find on the original listing.

About this opportunity

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work, but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

Tech at Lilly is seeking a highly skilled Senior Data Engineer who can implement and optimize large-scale Lakehouse solutions and drive the evolution of our modern data platform while providing technical leadership to a growing team. The ideal candidate will have hands-on experience with modern data engineering technology stack and a proven track record of leading engineering talent in fast-paced environments .

What you will be doing

Design and implement comprehensive Lakehouse architecture solutions using technologies like Databricks and Snowflake platforms

Define and carry out medallion architecture standards (Bronze / Silver / Gold) across data domains, ensuring data quality, lineage, and discoverability

Lead Unity Catalog governance design: schemas, access control policies, and data contracts

Build and maintain real-time and batch data processing systems using Apache Spark (PySpark/Scala), Kafka, and Databricks Structured Streaming, and similar technologies

Architect scalable data pipelines that handle structured, semi-structured, and unstructured data to deliver AI ready data.

Develop data transformation workflows using tools like DBT, Airflow, or Databricks

Implement data governance frameworks, including data quality monitoring, lineage tracking, data time travel and security protocols.

Build data pipeline testing frameworks: unit tests, data quality assertions (Great Expectations / dbt tests), and schema validation

Define and publish data SLAs/SLOs in collaboration with data product owners; own incident response and root-cause analysis for pipeline failures

Drive adoption of modern data engineering standard processes including Infrastructure as Code, CI/CD, and automated testing

Collaborate with data scientists, analysts, and business collaborators to translate requirements into robust technical solutions

Mentor a team of 3-5 data engineers

Foster a collaborative team culture focused on continuous learning and innovation

How You Will Succeed

Proven ability to mentor junior engineers and facilitate knowledge sharing

Strong project management skills with experience leading multi-functional initiatives

Coordinate multi-functional projects and ensure effective communication between technical and business teams

Demonstrated ability to make architectural decisions and drive technical consensus

Embrace a growth mindset and actively seek opportunities to expand your leadership capabilities and technical mastery

What You Should Bring

Knowledge in the pharmaceutical o r life sciences domain

Experience with streaming data technologies ( Kafka ,Databricks Structured Streaming)

Familiarity with data cataloging tools

Familiarity with high performance data service framework ( Arrow Flight )

Expert-level proficiency in Python and SQL for data transformation and pipeline development

Strong experience with Apache Spark (PySpark or Scala) for big data processing and analytics

Hands-on experience with cloud platforms ( AWS or Azure ) and their data services, including S3, Glue, Redshift, IAM, and CloudWatch

Proficiency with Infrastructure as Code tools ( CloudFormation )

Experience with containerization ( Docker , Kubernetes ) and orchestration platforms

Knowledge of data modeling techniques for both analytical and operational workloads

Hands-on expertise with Databricks: clusters, Delta Lake, Unity Catalog, Workflows, and MLflow integration

Understanding of data governance, security, and compliance requirements including GxP, HIPAA, or GDPR data frameworks in regulated industries

Experience orchestrating pipelines with Apache Airflow

S trong command of dbt for modular, tested, and version-controlled transformations

Familiarity with data quality testing frameworks such as Great Expectations or dbt tests for schema validation

Experience with data observability: freshness, volume, schema drift, and anomaly detection (Databricks Lakehouse Monitoring or equivalent)

Familiarity with open table format interoperability: Delta Sharing or Apache Iceberg

Knowledge of DataOps practices and data mesh / data product principles

Exposure to ML platform integration: MLflow experiment tracking, feature stores, or model serving

Your Basic Qualifications

Master’s degree in computer science, Engineering, or related technical field

3+ years of hands-on experience with Lakehouse architectures (Databricks, Snowflake, or similar)

7+ years of overall data engineering experience with large-scale distributed systems, including at least 3 years in a senior or lead capacity

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form ( https://careers.lilly.com/us/en/workplace-accommodation ) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.

Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.

Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).

Actual compensation will depend on a candidate’s education, experience, skills, and geographic location.  The anticipated wage for this position is

$132,000 - $244,200

Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

#WeAreLilly

Job details

Seniority
Principal
Function
Data & Digital
Therapeutic area
Not listed
Location
Indianapolis, Indiana, United States of America
Employment type
FULL_TIME

How this role compares

Computed from every other active Data & Digital role in our database, not just this employer's listings.

We currently track 88 comparable Principal Data & Digital roles across 23 biopharma companies.

88Comparable roles tracked
82Currently active
23Companies hiring similar roles
11Countries represented

Salary context

17 of 88 peers report a salary range (USD, annualized)

Peers share this role's job function and a matching or adjacent seniority level -- not necessarily the same therapeutic area or country.

This roleSubject $132,000/yr – $244,200/yr
Lowest disclosed · Senior Associate, Metadata Engineering · Pfizer $79,400/yr – $132,400/yr
Highest disclosed · Director and Group Head, Applied AI · Novartis $194,600/yr – $361,400/yr
Peer group range $105,900 – $278,000 (median $222,900)

Where these roles are based

Top locations among the 88 comparable roles

India46
United States22
United Kingdom5
Ireland3
Switzerland3
China2

+ 5 more countries

Seniority mix

88 of 88 peers have a known seniority level

Senior57
Director16
Principal15

Therapeutic area mix

2 of 88 peers have a known therapeutic area; the rest are genuinely unlabeled, not hidden

Oncology2

Similar opportunities

The closest matches from our peer group, ranked by how similar they are, not how well you'd qualify for them -- treat this as market context, not a guaranteed shortlist; a weak match is labeled as one below.

60%similar
Regeneron Pharmaceuticals, Inc (USA) Warren, United States Principal
Same function Same seniority Same country
60%similar
GlaxoSmithKline LLC Upper Providence, United States Principal
Same function Same seniority Same country
60%similar
Sanofi US Services Inc. Framingham, United States Principal
Same function Same seniority Same country
60%similar
Sanofi US Services Inc. Cambridge, United States Principal
Same function Same seniority Same country
50%similar
Gilead Sciences, Inc. Foster City, United States Director
Same function Adjacent seniority Same country
50%similar
Novartis Remote Position (USA), United States Director
Same function Adjacent seniority Same country

How we calculate "similar"

No black box, no LLM guesswork: a deterministic score built from four normalized attributes. Here's this role's own peer group at different match levels, so you can see the mechanism, not just the result.

Every comparison starts from the same 100-point budget: 25 for working in the same function, 40 for the same therapeutic area, 20 for the same or adjacent seniority, 15 for the same country. A dimension we can't confirm on both sides contributes nothing, never a guess, never a free pass.

60%
Principal Statistical Programmer
Regeneron Pharmaceuticals, Inc (USA) · Warren, United States · Principal
Function Therapeutic area Seniority Country
60%
Principal Scientist - CMC Statistician
Sanofi US Services Inc. · Framingham, United States · Principal
Function Therapeutic area Seniority Country
50%
Director Analytics Infrastructure, Pipeline Operations
Novartis · Remote Position (USA), United States · Director
Function Therapeutic area Adjacent seniority Country
50%
Senior Data Engineer
E.R. Squibb & Sons,L.L.C. · Princeton, United States · Senior
Function Therapeutic area Adjacent seniority Country
Unmatched or unknown dimensions score exactly the same: 0 points, never a partial guess. A role we know almost nothing about beyond its function bottoms out at 25%; it never inflates to 100% just because there's little to compare against. Seniority uses a defined ladder (Associate → Manager → Associate Director → Senior → Principal → Director → Senior Director → Executive/VP) so "Director" and "Senior Director" count as adjacent, but "Director" and "Executive/VP" do not.