Senior / Lead / Staff roles · 9+ years · Pharma & SaaS. Building enterprise-grade data lakes, lakehouse platforms and semantic analytics layers across pharma (Eli Lilly, 200K+ clinical assets) and SaaS (Kenko AI, 500+ multi-tenant clients). Consistent 40–60% improvements in pipeline efficiency, data quality and cost.
python · pyspark · sql · aws · apache-iceberg · cube-cloud-ldm · airflow
About
I'm Amit Kumar Sahu — Senior Data Engineer based in Bengaluru. 9+ years building data infrastructure that delivers: 40–60% pipeline efficiency gains, dramatically reduced data incidents, and real cost savings on BI tooling.
My career spans regulated pharma (Eli Lilly — 200K+ clinical assets with Veeva Vault and Axon Data Catalog), high-growth SaaS (Kenko AI — full lakehouse from scratch, 500+ tenants), IoT (SpanIdea — real-time MQTT sensor pipelines), consulting (AspireNXT), and enterprise automotive (Infosys × Daimler AG).
Beyond engineering I mentor early-career data folks, play competitive chess, and read tech blogs voraciously. Open to Senior / Lead / Staff Data Engineer roles and freelance data engineering engagements.
How I Lead
Three examples of owning outcomes beyond my own code.
Technical Skills
Every tool here has been used in live, production-grade systems.
Experience
Consistent delivery of production data platforms that cut costs, raise quality, and scale.
Feb 2026 – Apr 2026
Senior Data Engineer
Kenko AI · SaaS Fitness Tech · Bengaluru
Jun 2024 – Jan 2026
Planned career break
Bengaluru
Feb 2021 – May 2024
Engineer II — Software Configuration & Development
Eli Lilly and Company · Pharma · Bengaluru
Aug 2019 – Jun 2020
Data Engineer
AspireNXT Pvt. Ltd. · Consulting · Bengaluru
Jan 2019 – Jul 2019
Senior Software Engineer
SpanIdea Systems · IoT · Bengaluru
Aug 2015 – Dec 2018
Senior Systems Engineer
Infosys Limited · Client: Daimler AG · Bengaluru
Education
2011 – 2015
B.Tech — Computer Science & Engineering
Veer Surendra Sai University of Technology (VSSUT), Burla, Odisha
2010
AISSCE
Jawahar Navodaya Vidyalaya, Sarang, Dhenkanal, Odisha
Case Studies
Four projects, from a 500+ tenant SaaS lakehouse to a personal medallion-architecture build.
Problem. 500+ fitness-studio clients needed fast, reliable analytics; reporting on the legacy stack was slow (SLA at 70%) and BI licensing costs were high.
Architecture. AWS Data Lakehouse: DMS CDC → S3 Bronze → Apache Iceberg Silver (S3 Tables) → Athena Gold. Cube Cloud semantic layer — 18 cubes, 11 pre-aggregated views, 58 KPIs across Finance, Memberships, Classes, and Marketing — with JWT-based multi-tenant row-level security for 500+ operators. 20-table Gold analytics layer replacing GoodData.
My role. Senior Data Engineer — architected and built the lakehouse end-to-end and designed the Cube Cloud LDM.
Outcome. Query latency −45%, SLA 70%→98%, BI licensing costs −35%.
Problem. 5+ TB of CTMS clinical trial data sat in a legacy IMPACT system, slowing onboarding across 6+ clinical trial projects, and 200K+ clinical assets had no searchable catalog.
Architecture. PySpark/Airflow migration pipeline moving CTMS data from IMPACT to Veeva Vault; Axon Data Catalog indexing 200K+ clinical assets; 20+ standardised, reusable Data Products and DWS APIs shared across 5+ teams.
My role. Engineer II, Software Configuration and Development — led the migration engineering and built the catalog and reusable Data Products.
Outcome. Onboarding 60% faster across 6+ trials, discoverability +50%, validation efficiency +30%, incident response time −40%.
Problem. Growing multi-tenant data volume across 500+ clients risked silent data quality issues undermining trust in the Gold analytics layer.
Architecture. A 9-check PySpark DQ framework with a quarantine path for failing records and SNS alerting, wired into the Bronze→Silver pipeline.
My role. Designed and built the framework as part of the Kenko AI lakehouse build.
Outcome. Data incidents −60%, data confidence 82%→97%.
# One of 9 checks in the quarantine-and-alert DQ pipeline def run(self, df, quarantine_path=None): total = df.count(); bad_mask = F.lit(False) for rule in self.rules: fail_cond = self._build_condition(rule) bad_mask = bad_mask | fail_cond clean_df = df.filter(~bad_mask) if quarantine_path: df.filter(bad_mask).write.format("delta") \ .mode("append").save(quarantine_path) return clean_df
Problem. Wanted a hands-on way to prove out lakehouse and data-quality patterns learned during a self-directed upskilling period, outside production constraints.
Architecture. Bronze/Silver/Gold medallion architecture with Kafka streaming ingestion, a PySpark DQ engine with quarantine, SCD Type 2 dimensions, and Delta Lake storage.
My role. Independent creator and maintainer — architecture, pipeline code, and CI/CD.
Outcome. CI green, v0.1.0 released. github.com/nikuamit/astra-data-platform ↗
Projects
Personal projects built during a self-directed upskilling period.
Freelance Services
Time-boxed data engineering work. Remote-first, available now.
Contact
Full-time role or freelance — I reply within 24 hours.