Open to work — Senior/Lead/Staff Data Engineer · immediate joiner

Lead Data Engineer
building lakehouse & semantic platforms at scale

Senior / Lead / Staff roles · 9+ years · Pharma & SaaS. Building enterprise-grade data lakes, lakehouse platforms and semantic analytics layers across pharma (Eli Lilly, 200K+ clinical assets) and SaaS (Kenko AI, 500+ multi-tenant clients). Consistent 40–60% improvements in pipeline efficiency, data quality and cost.

python · pyspark · sql · aws · apache-iceberg · cube-cloud-ldm · airflow

Data Lakes Lakehouse Architecture Real-time Streaming Multi-tenant Analytics AWS Full Stack Pharma & SaaS Mentoring
astra-pipeline ~ bash
0%
Query latency cut
Kenko AI
0
Clinical assets catalogued
Eli Lilly
0%
SLA compliance
achieved
0
Multi-tenant clients
Kenko AI

About

Data engineer, problem solver,
occasional chess player

I'm Amit Kumar Sahu — Senior Data Engineer based in Bengaluru. 9+ years building data infrastructure that delivers: 40–60% pipeline efficiency gains, dramatically reduced data incidents, and real cost savings on BI tooling.

My career spans regulated pharma (Eli Lilly — 200K+ clinical assets with Veeva Vault and Axon Data Catalog), high-growth SaaS (Kenko AI — full lakehouse from scratch, 500+ tenants), IoT (SpanIdea — real-time MQTT sensor pipelines), consulting (AspireNXT), and enterprise automotive (Infosys × Daimler AG).

Beyond engineering I mentor early-career data folks, play competitive chess, and read tech blogs voraciously. Open to Senior / Lead / Staff Data Engineer roles and freelance data engineering engagements.

↓ Download Resume LinkedIn ↗
Location
Bengaluru, Karnataka, India
Open to
Remote · Hybrid · Freelance
Degree
B.Tech CS&E — VSSUT Burla
Achievement
ACM-ICPC World Semifinalist 2013
Interests
Competitive programming · Chess · Tech blogs · Mentoring · Poetry

How I Lead

Leadership isn't a title change — it's already the job

Three examples of owning outcomes beyond my own code.

Team leadership
Led a 5-member team — AspireNXT
Owned delivery of AWS data lake pipelines across finance, industrial IoT, and healthcare clients — delivery velocity +30% — plus weekly knowledge-sharing sessions that lifted team productivity +20% and cut onboarding time −50%.
Cross-team standards
Standardised platforms for 5+ teams — Eli Lilly
Built 20+ reusable Data Products and DWS APIs for the Enterprise Data Backbone, adopted across 5+ teams — cut incident response time −40% via ServiceNow integration.
Platform architecture
Architected a multi-tenant lakehouse — Kenko AI
Designed the lakehouse and Cube Cloud semantic layer serving 500+ clients end-to-end — query latency −45%, SLA 70%→98%, BI licensing costs −35%.

Technical Skills

Core stack, built in production

Every tool here has been used in live, production-grade systems.

🐍 Languages & Databases
PythonPySparkSQLBashPostgreSQLMySQLOracleDynamoDB
☁️ Cloud
AWS S3GlueDMSLambdaRedshiftRDSAthenaKinesisStep FunctionsLake FormationBedrockAzure IoT Hub
🏗️ Data Engineering
Apache IcebergDelta LakeDatabricksApache Airflowdbt
📊 Semantic Layer & BI
Cube Cloud LDMCubesPre-aggregationsMulti-tenant RLSJWT AuthGoodDataQuickSight
🤖 AI & Dev Tools
ClaudeChatGPTGitHub CopilotCursorWindsurfDockerGitHubLinuxVS Code
⚙️ Workflow
LinearJiraConfluenceAgile ScrumGitDevSecOps
Python & PySpark90%
SQL & Data Modeling88%
AWS Data Stack85%
Data Engineering & Lakehouse85%
Semantic Layer — Cube Cloud LDM80%
CI/CD & DevOps75%

Experience

9+ years across pharma, SaaS, IoT & enterprise

Consistent delivery of production data platforms that cut costs, raise quality, and scale.

Feb 2026 – Apr 2026

Senior Data Engineer

Kenko AI · SaaS Fitness Tech · Bengaluru

  • Architected AWS Data Lakehouse for 500+ clients — query latency −45%, SLA 70% → 98%.
  • Designed Cube Cloud LDM — 18 cubes, 11 pre-aggregated views, 58 KPIs across Finance, Memberships, Classes, and Marketing; JWT-based multi-tenant RLS for 500+ operators.
  • Built 9-check DQ framework — data incidents −60%, confidence 82% → 97%.
  • Gold layer replacing GoodData — BI licensing costs −35%.
  • Role eliminated in a company restructuring.

Jun 2024 – Jan 2026

Planned career break

Bengaluru

  • Upskilling in lakehouse architecture — Apache Iceberg, Delta Lake.
  • GenAI on AWS Bedrock; building independent data projects.

Feb 2021 – May 2024

Engineer II — Software Configuration & Development

Eli Lilly and Company · Pharma · Bengaluru

  • Standardised 20+ reusable Data Products and DWS APIs for the Enterprise Data Backbone (PySpark, Airflow) across 5+ teams — incident response time −40% via ServiceNow integration.
  • Migrated 5+ TB CTMS clinical data IMPACT → Veeva Vault — onboarding +60%.
  • Axon Data Catalog for 200K+ clinical assets — discoverability +50%.
  • Won Lilly Global Ideas & Innovation Award for LillyTV.

Aug 2019 – Jun 2020

Data Engineer

AspireNXT Pvt. Ltd. · Consulting · Bengaluru

  • Led 5-member team — AWS Data Lakes across finance, IoT, healthcare; velocity +30%.
  • Migrated 10+ TB — latency −40%, infra costs −25%.
  • Established weekly knowledge-sharing sessions — team productivity +20%, onboarding time −50%.

Jan 2019 – Jul 2019

Senior Software Engineer

SpanIdea Systems · IoT · Bengaluru

  • Span Park smart parking — Python, C, Raspberry Pi, PostgreSQL, Azure IoT Hub; real-time availability tracking cut parking search time −50%.

Aug 2015 – Dec 2018

Senior Systems Engineer

Infosys Limited · Client: Daimler AG · Bengaluru

  • Automated critical ETL pipelines with RCA and on-call support across 12+ projects — production downtime −20%, automation efficiency +15%, manual effort −30%.

Education

Academic background

2011 – 2015

B.Tech — Computer Science & Engineering

Veer Surendra Sai University of Technology (VSSUT), Burla, Odisha

  • Founded ENIGMA coding club.
  • ACM-ICPC World Semifinalist 2013 — IIT Kharagpur Regionals, Asia.

2010

AISSCE

Jawahar Navodaya Vidyalaya, Sarang, Dhenkanal, Odisha

Case Studies

Problem → Architecture → Role → Outcome

Four projects, from a 500+ tenant SaaS lakehouse to a personal medallion-architecture build.

Kenko AI — Multi-tenant Lakehouse & Cube Cloud Semantic Layer

Problem. 500+ fitness-studio clients needed fast, reliable analytics; reporting on the legacy stack was slow (SLA at 70%) and BI licensing costs were high.

Architecture. AWS Data Lakehouse: DMS CDC → S3 Bronze → Apache Iceberg Silver (S3 Tables) → Athena Gold. Cube Cloud semantic layer — 18 cubes, 11 pre-aggregated views, 58 KPIs across Finance, Memberships, Classes, and Marketing — with JWT-based multi-tenant row-level security for 500+ operators. 20-table Gold analytics layer replacing GoodData.

My role. Senior Data Engineer — architected and built the lakehouse end-to-end and designed the Cube Cloud LDM.

Outcome. Query latency −45%, SLA 70%→98%, BI licensing costs −35%.

Eli Lilly — Clinical Data Migration & Axon Data Catalog

Problem. 5+ TB of CTMS clinical trial data sat in a legacy IMPACT system, slowing onboarding across 6+ clinical trial projects, and 200K+ clinical assets had no searchable catalog.

Architecture. PySpark/Airflow migration pipeline moving CTMS data from IMPACT to Veeva Vault; Axon Data Catalog indexing 200K+ clinical assets; 20+ standardised, reusable Data Products and DWS APIs shared across 5+ teams.

My role. Engineer II, Software Configuration and Development — led the migration engineering and built the catalog and reusable Data Products.

Outcome. Onboarding 60% faster across 6+ trials, discoverability +50%, validation efficiency +30%, incident response time −40%.

Kenko AI — 9-Check Data Quality Framework

Problem. Growing multi-tenant data volume across 500+ clients risked silent data quality issues undermining trust in the Gold analytics layer.

Architecture. A 9-check PySpark DQ framework with a quarantine path for failing records and SNS alerting, wired into the Bronze→Silver pipeline.

My role. Designed and built the framework as part of the Kenko AI lakehouse build.

Outcome. Data incidents −60%, data confidence 82%→97%.

dq/quality_engine.pyPython · PySpark (illustrative)
# One of 9 checks in the quarantine-and-alert DQ pipeline
def run(self, df, quarantine_path=None):
    total = df.count(); bad_mask = F.lit(False)
    for rule in self.rules:
        fail_cond = self._build_condition(rule)
        bad_mask = bad_mask | fail_cond
    clean_df = df.filter(~bad_mask)
    if quarantine_path:
        df.filter(bad_mask).write.format("delta") \
            .mode("append").save(quarantine_path)
    return clean_df
Astra Data Platform — Personal Lakehouse Project

Problem. Wanted a hands-on way to prove out lakehouse and data-quality patterns learned during a self-directed upskilling period, outside production constraints.

Architecture. Bronze/Silver/Gold medallion architecture with Kafka streaming ingestion, a PySpark DQ engine with quarantine, SCD Type 2 dimensions, and Delta Lake storage.

My role. Independent creator and maintainer — architecture, pipeline code, and CI/CD.

Outcome. CI green, v0.1.0 released. github.com/nikuamit/astra-data-platform ↗

Projects

Independent builds

Personal projects built during a self-directed upskilling period.

Python · PySpark · SQL
Astra Data Platform
Bronze/Silver/Gold medallion, Kafka streaming, PySpark DQ engine with quarantine, SCD2, Delta Lake, CI green, v0.1.0 release.
Vercel · Claude API
Amavya
Founder-built AI assistant for 22 Indian languages.

Freelance Services

What I can build for you

Time-boxed data engineering work. Remote-first, available now.

🏗️
Data Lakehouse Architecture
Bronze/Silver/Gold on AWS — Apache Iceberg or Delta Lake, DMS CDC ingestion, PySpark transforms, Athena Gold layer.
Apache IcebergAWSPySpark
⚡
Real-time Pipelines
Streaming via Kinesis or DMS CDC — end-to-end with SNS monitoring, schema evolution, and checkpointing.
KinesisDMS CDCAWS Glue
📊
Semantic Layer & Analytics
Cube Cloud LDM with cubes, pre-aggregations, KPIs, JWT multi-tenant RLS. Replace expensive BI tools.
Cube CloudJWT RLSSQL
🧪
Data Quality Framework
9-rule automated DQ with SNS alerting and quarantine. Proven: −60% incidents, confidence to 97%.
Great ExpectationsPython
🔄
ETL / ELT Development
PySpark, Python, SQL — batch or micro-batch. Migration, CDC, Airflow orchestration, performance tuning.
PythonPySparkSQL
🎓
Data Engineering Mentoring
1:1 — interview prep, code reviews, architecture walkthroughs, career guidance.
1:1 SessionsCareer Guidance
Ready to start a project?
Typical: 2–12 weeks · Remote-first · Available now
Get a free scoping call

Contact

Let's build something together

Full-time role or freelance — I reply within 24 hours.

Location
Bengaluru, Karnataka, India

Sent directly to my inbox. I reply within 24 hours on business days.