graph
graph.id/sampledataengineerCopy graph.id/sampledataengineer
Priya Patel's profile photo

Priya Patel

Identity verified · work history partially verified against employer & source records
Data Engineer·System Builder / Architect

Mumbai, Maharashtra  ·  Building in public since Jul 2021 (~5.0 yrs)

PortfolioGitHubMediumSubstackLinkedIn
200GB+Data ingested per day
15K/minPeak event throughput
80%Fewer pipeline failures
99.9%Data accuracy achieved
Growth trajectory

From scripts to streaming systems.

A clear shift from writing one-off ETL scripts to owning event-driven, production-grade data infrastructure.

  1. Jul 2021
    Wrote first production ETL scripts
  2. Feb 2022
    Shipped first AWS Glue pipeline
  3. Nov 2022
    Owned CDC ingestion end-to-end
  4. Jun 2023
    Built the streaming layer
  5. Mar 2024
    Led the lakehouse migration
  6. Jan 2025
    Owns pipeline reliability
Signals that matter

Live in Production.

Each signal ties back to a specific, evidenced moment - not a self-rating.

Built2

Data pipelines and ingestion systems shipped end-to-end.

EvidenceBuilt a CDC pipeline from Postgres to Redshift
Scaled2

Taken from a working prototype to real production volume.

EvidenceReal-time event pipeline handling 15,000 events/min
Led1

Orchestration and architecture decisions owned end-to-end.

EvidenceReplaced cron scripts with Step Functions + Airflow
Solved1

Data-accuracy and reliability issues closed under pressure.

EvidenceSolved a duplicate-record bug in a CDC pipeline
Stack

Daily tools.

What Priya actually reaches for, pulled from the tags attached to captured moments - not a keyword list.

AWS
PostgreSQL
MongoDB
Python
Apache Kafka
Apache Airflow
Apache Spark
Docker
Work

Key Projects By Role

Case studies, not bullet points - impact numbers attached to each one.

Nimbus Analytics

Senior Data Engineer · Jan 2023 - Present
4 months

Built a CDC pipeline from Postgres to Redshift

Designed a change-data-capture ingestion pipeline using AWS DMS to stream Postgres transactional data into Redshift, replacing a brittle nightly batch export analysts couldn't rely on.

12 tables migrated, zero downtimelag cut from 24h to under 5 min
AWS DMSPostgreSQLRedshiftCDC
Verified via GitHubView case study
3 months

Real-time event pipeline on Kinesis + Lambda

Built an event-driven ingestion layer using Kinesis Data Streams and Lambda to process transaction events in near real-time, feeding the warehouse and a downstream fraud-alerting service.

15,000 events/min at peaksub-second processing
KinesisLambdaPython
Verified via MediumView case study
2 months

Migrated MongoDB event logs into an S3 lakehouse

Used AWS Glue to build ingestion jobs pulling semi-structured MongoDB event data into a partitioned S3 lakehouse, with the Glue Data Catalog making it queryable via Athena.

200GB+ ingested daily60% faster queries
GlueS3MongoDB
Verified via Public ProfileView case study
3 months

Replaced cron scripts with Step Functions + Airflow

Redesigned a tangle of cron-scheduled scripts into a proper orchestration layer - Step Functions for AWS-native workflows, Airflow for cross-system dependencies - with retry logic and alerting built in.

80% fewer pipeline failures
Step FunctionsAirflow
Self-reportedView case study
6 weeks

Optimized EMR / PySpark batch jobs

Profiled and rewrote nightly PySpark batch jobs on EMR that were running over budget - repartitioning skewed joins and moving non-critical stages to spot instances.

40% lower EMR cost3x faster nightly batch
EMRPySpark
Self-reportedView case study

Verlio Data Labs

Data Engineer · Jul 2021 - Dec 2022Verified · employer record
3 weeks

Solved a duplicate-record bug in a CDC pipeline

Traced double-counted financial records downstream to a race condition in the CDC pipeline's retry logic, and fixed it with idempotent upserts keyed on source transaction ID.

99.9% data accuracy
CDCData Quality
Verified · employer recordView case study
2 months

Introduced Kafka as an event bus

Proposed and built a proof-of-concept using Kafka to decouple ingestion from downstream analytics consumers, replacing a point-to-point integration that broke every time a new consumer was added.

3 consumers decoupled from direct DB access
KafkaArchitecture
Verified · employer recordView case study
Outside work

What gets built off the clock.

Experiments, writing and talks - separate from paid work, tagged as such.

Open source

pg-cdc-toolkit

A small CLI wrapping AWS DMS task management - the tool Priya wished existed while debugging replication tasks by hand.

Verified · GitHubStar count and commit history pulled directly from GitHub.68
Article

CDC pipelines fail quietly - what nobody tells you about replication lag

On the gap between a "healthy" DMS task status and data actually being current downstream.

Published Feb 20255 min read
Speaker

Event-Driven Data Pipelines on AWS

Talk on trade-offs between CDC, batch and streaming ingestion for teams migrating off nightly ETL.

Data Engineers Mumbai Meetup · Apr 2025
Experiment

Local DuckDB-based data quality checker

A lightweight tool for spot-checking pipeline outputs against source before promoting to production.

Prototype · unshipped
Publication / book

Nothing published yet - this will fill in the moment it happens.

Learning

Formal record.

  • B.Tech, Computer Engineering
    Sardar Patel Institute of Technology, Mumbai
    2017 - 2021
  • AWS Certified Data Analytics - Specialty
    Verified · issuer
    2023
  • Data Engineering on AWS
    Coursera
    2022
Captured moments

Not a work log. A record of momentum.

Each shaded day is when Priya captured something built, solved, scaled or led - density here reflects real activity, not tenure.

LessMore
Open to

Open to the right opportunity.

Roles, speaking, mentoring, or a good conversation - Priya is listening.

Open to rolesOpen to speakingOpen to mentoringOpen to collaborating
Connect with Priya