graphby datavruti

Satyaki Sengupta

Verified by 7 captured moments since Jul 2016
Senior Data Engineer·Architect

Kolkata, IndiaBuilding in public since Jul 2016 (~10.1 yrs)

Headline

Satyaki

Senior Data Engineer

Across five consulting-style roles, you've repeatedly led migrations from legacy ETL and Hadoop systems onto modern Lakehouse platforms - a specialization in transformation, not just maintenance.

1.5+ TBdaily data processed
40%SLA breaches reduced
45%pipeline execution time reduced
Growth trajectory

Satyaki's career, mapped.

A timeline of milestones captured directly from real, verified work.

  1. Jul 2016
    Big Data Foundations
  2. Apr 2019
    Real-Time Streaming
  3. Dec 2020
    Senior Big Data Engineer
  4. Feb 2022
    Senior Consultant at Deloitte
  5. 2021
    Cloudera Certified
  6. Apr 2024
    Specialist Data Engineer
Signals

Live in Production.

Each signal ties back to a specific, evidenced moment - not a self-rating.

Built4

Systems shipped end-to-end, from schema to deploy.

Evidence4 moments →
Migrated Marsh McLennan off Informatica onto Databricks in 4 monthsUnified 9+ enterprise systems into one Hadoop data lake for CITIBuilt real-time streaming pipelines for BHP mining analyticsDelivered fleet telemetry analytics platform for PepsiCo
Solved1

Gaps closed - false positives, edge cases, dead ends.

Evidence
Cut data quality incidents 40% with a silver-to-gold trust layer
Scaled1

Taken from prototype to real production volume.

Evidence
Piped 1.5+ TB of daily SAP HANA data into Snowflake for Siemens
Learned1

New tools and depth picked up on the job.

Evidence
Outside Work Exploration
Stack

Daily tools.

What Satyaki actually reaches for, pulled from the tags attached to captured moments - not a keyword list.

Python
PySpark
SQL
Databricks
Delta Lake
Delta Lake
S3
AWS S3
Glue
AWS Glue
Lambda
AWS Lambda
Snowflake
Hadoop
Kafka
Hive
Sqoop
Sqoop
Jenkins
Bitbucket
CI/CD
CloudFormation
CloudFormation
Work

Key Projects By Role

Case studies, not bullet points - impact numbers attached to each one.

LTIMindtree

Specialist Data Engineer · 2024 - Present
Apr 2024

Legacy ETL to Lakehouse Migration

Production-grade migration of 20+ AWS Glue ETL workflows from Informatica/Glue to Databricks Lakehouse for Marsh McLennan.

Reduced pipeline execution time by 45%Cut operational maintenance effort by 35%Migrated 20+ ETL workflows within 4 months
DatabricksAWS GlueDelta LakePySpark
Apr 2024

Gold-Layer Data Trust Platform

Scalable Databricks data validation frameworks reconciling silver and gold layers, with reusable PySpark modules for dynamic table comparison and data quality validation.

Reduced production data quality incidents by 40%Improved pipeline performance by 30-35% for multi-million record datasetsSaved approx 50 hrs per week
DatabricksPySparkData QualityAWS S3

Deloitte India

Senior Consultant (AWS Data Engineer) · 2022 - 2024
Feb 2022

Scalable Lakehouse Platform

AWS Glue ETL pipelines using PySpark ingesting SAP HANA data into S3 and Snowflake for Siemens finance and billing use cases, processing 1.5+ TB daily.

Processed 1.5+ TB of data dailyReduced batch execution time by 30%
AWS GluePySparkSnowflakeSAP HANA

LTI

Senior Big Data Engineer · 2020 - 2022
Dec 2020

Finance Data Platform

Kafka-based ingestion pipelines consolidating 9+ upstream enterprise systems into a centralized Hadoop data lake for CITI.

Reduced SLA breaches by 40%Reduced batch execution time by 20%
KafkaHadoopHDFSPySparkHive

Cognizant (CTS)

Associate · 2019 - 2020
Apr 2019

Operational Event Data Platform

Real-time data ingestion pipelines using Kafka, Azure Event Hub, SQL Server, and S3 for BHP mining and natural resources streaming analytics.

Enabled near real-time analytics for mining and natural resources workloads
KafkaAzure Event HubPySparkStreaming

HCL Technologies

Associate Software Engineer (Big Data Developer) · 2016 - 2019
Jul 2016

Industrial Event Analytics Platform

Data ingestion and transformation pipelines using Sqoop, NiFi, Kafka, Spark, Hive and Elasticsearch for PepsiCo vehicle telemetry analytics.

Improved reporting performance and data accessibility for business users
SqoopApache NiFiKafkaElasticsearchHive
Outside Work

Side projects, open source & publications.

Independent builds, open-source contributions, technical writing, and experiments beyond the day job.

Learning

Formal record.

  • B.Tech in Computer Science & Engineering
    West Bengal University of Technology (SMIT)
    2010 - 2014
  • CCA-175: Cloudera Certified Spark and Hadoop Developer, 2021
Activity heatmap

A record of momentum.

A combined view of your activities across the tools that you use daily (GitHub, Medium, etc. - coming soon). Add earlier projects as a graph entry if dates available.

LessMore