Cohort 4Register

New book · 111 pages · PDF

Data Engineering System DesignZero to Hero

How to think, design, defend and scale data systems in a real interview. One framework, 26 worked designs with real follow-ups, and the reasoning that separates a hire from a no-hire.

By Durgesh Yadav · Data Engineer · Instructor · Founder, PrepNPlaced · 2026 edition

₹299Launch price

Until 11:59 pm IST on 14 October 2026. From 15 October it is ₹499.

One payment. PDF, yours to download again any time.

Read 20 pages free

Included free for learners in the AI-Powered Data Engineering cohort and in 1:1 mentorship: sign in with the email you enrolled with.

Data Engineering System Design: Zero to Hero, by Durgesh Yadav: the cover
reported interview questions, each tagged by evidence
100reported interview questions, each tagged by evidence
companies, from MAANG to Razorpay and Flipkart
30+companies, from MAANG to Razorpay and Flipkart
designs worked end to end with follow-ups
26designs worked end to end with follow-ups
cheat sheets plus one master checklist
19cheat sheets plus one master checklist

What is inside

Built from what interviewers actually asked

The book starts from 100 design prompts that candidates reported from real loops, sorts them by what they test, and teaches the method and the designs those prompts call for.

  • 100 reported interview questions

    Design and modeling prompts from 2020-2026 loops at 30 named companies, from MAANG to Razorpay and Flipkart. Each is tagged VERIFIED, REPORTED, PATTERN or INFERRED, so you know how much weight it can carry.

  • The PIPELINE method

    Eight steps in a 45-minute timebox, with something on the board after every step. When the interviewer pushes or you lose the thread, you know which step you are on and what is left.

  • 26 worked designs

    Six flagships in fifteen sections each, from clarifying questions to the follow-ups that sink most answers, and twenty compact designs on one page each.

  • Numbers you can say out loud

    One planning number sheet and seven worked estimates, each ending with the design decision the number forces.

  • Follow-ups and answer levels

    A follow-up simulator, bad, good and excellent answers side by side, and what senior and staff candidates are expected to raise on their own.

  • 19 cheat sheets and a master checklist

    Nineteen tools and ideas, seven lines each in the same order, then the whole method on one page for the night before.

Three of the pages

  • Data Engineering System Design: Zero to Hero, page 15
    Page 15 · The PIPELINE method: eight steps, a 45-minute timebox, and what goes on the board after each.
  • Data Engineering System Design: Zero to Hero, page 59
    Page 59 · Flagship 02, ride-hailing: the evidence on the header, ten clarifying questions, the numbers, the diagram.
  • Data Engineering System Design: Zero to Hero, page 94
    Page 94 · Requirement to architecture: eight questions, asked in this order, and what each answer rules out.

Contents

Eight parts and an appendix

Chapter 01 comes first, on what the 100 questions say. Then the book moves from what the round tests to the night before it.

  1. 01What 100 reported interview questions tell usfree to readp. 3
  1. Part 1 · from page 6

    Foundations

    What a data engineering design round is testing, the five stages every data system is built from, and how the round runs at different companies. It starts with one kirana store and grows it to 5,000.

    1. 02What system design means for a data engineerp. 7
    2. 03Anatomy of a data systemp. 10
    3. 04What the round actually looks likep. 12
  2. Part 2 · from page 14

    System design thinking

    The method used for every design in the book, the questions that turn a vague prompt into numbers, and back-of-envelope maths you can do out loud.

    1. 05The PIPELINE methodfree to readp. 15
    2. 06The art of the clarifying questionfree to readp. 16
    3. 07Scale estimation you can do out loudfree to readp. 19
  3. Part 3 · from page 22

    Distributed systems fundamentals

    How data is split, how it is copied, what consistent means when machines disagree, and how bytes are laid out on disk, each taught through something you will draw in the round.

    1. 08Partitioning and shardingp. 23
    2. 09Replication, consensus and coordinationp. 25
    3. 10CAP, PACELC and consistencyp. 27
    4. 11The scaling toolkitp. 28
    5. 12Row, columnar and table formatsp. 29
  4. Part 4 · from page 32

    Data engineering architecture

    The patterns you will be asked to compare, where each kind of data should live, how to model it, and what you say when the interviewer asks what breaks, who can see it, whether it is right and what it costs.

    1. 13Architecture patterns, and when not to use themp. 33
    2. 14Storage decision guidep. 36
    3. 15Data modeling for design roundsp. 37
    4. 16Reliability engineeringp. 39
    5. 17Data quality in productionp. 41
    6. 18Security and governancep. 42
    7. 19Costp. 42
  5. Part 5 · from page 44

    Core technologies

    What each building block is for and where it breaks, then deep chapters on the two that interviewers probe hardest: Spark and Kafka.

    1. 20Building blocksp. 45
    2. 21Spark system designp. 47
    3. 22Kafka system designp. 52
  6. Part 6 · from page 55

    Real interview designs

    Twenty-six designs, most traced to a prompt candidates reported, with the evidence printed on each header. Six flagships in fifteen sections, twenty compact designs on one page each.

    1. 23The 26 designs at a glancep. 56
    2. 24Twenty compact designsp. 69
  7. Part 7 · from page 91

    Interview mastery

    The 45 minutes themselves: drawing a design that reads at a glance, turning a vague prompt into choices, holding your ground through follow-ups, and what moves an answer from good to a strong hire.

    1. 25How to draw the architecturefree to readp. 92
    2. 26From question to architecturefree to readp. 93
    3. 27The follow-up simulatorp. 95
    4. 28Bad, good and excellent answersp. 96
    5. 29Senior and staff thinkingp. 98
  8. Part 8 · from page 99

    Cheat sheets and the master checklist

    The night before the round: nineteen cards, seven lines each in the same order, then the whole method on one page.

    1. 30Cheat sheetsp. 100
    2. 31The system design master checklistp. 105
  9. Appendix A · from page 106

    The question database

    Every interview question behind the book, with an honest label for how much weight each one can carry.

    1. 32All 100 questions with sourcesp. 107

Part 6

The 26 designs

6 flagships are worked in fifteen sections each, from the clarifying questions to the follow-ups. The other 20 fit on one page each. Every header says where the prompt was reported and how strong that evidence is.

The 26 designs in the book
#DesignReported atEvidencePage
01Real-time clickstream and ad-click analytics pipelineD. StreamingFlagshipFree to readGoogle 2026, Apple 2024, Uber blog, Zepto blogREPORTED57
02Ride-hailing data model and trip analyticsC. Warehouse, K. BusinessFlagshipFree to readMeta 2026, Lyft 2026, Uber prep sitesREPORTEDPATTERN59
03Self-healing ingestion from 1,000+ databases into the lakehouseA. IngestionFlagshipRazorpay c. 2024, Salesforce 2026, Microsoft 2024REPORTED61
04Food-delivery order analytics with batch and real-time pathsK. Business, D. StreamingFlagshipDoorDash 2026, Swiggy blog, Airbnb blogREPORTED63
05Viewing analytics: top-10 titles and hourly trending by locationJ. Analytics, K. BusinessFlagshipUber Hyderabad 2024, Netflix 2025REPORTED65
06Payment ledger and transaction analytics with reconciliationG. Reliability, K. PaymentsFlagshipStripe 2025-26, Flipkart 2021, Razorpay blogPATTERNREPORTED67
07Product analytics platform like MixpanelJ. AnalyticsWalmart GT Bengaluru 2023REPORTED69
08Count-metrics monitoring with duplicate and late eventsD. Streaming, G. ReliabilityStripe 2026REPORTED70
09User activity counts for the last minute, hour and dayD. Streaming, J. AnalyticsLinkedIn 2026REPORTED71
10Real-time driver heatmap and surge signalD. StreamingUber 2025REPORTEDPATTERN72
11A compaction service that does what Delta OPTIMIZE doesB. LakehouseRazorpay c. 2024REPORTED73
12Booking metrics: daily partitions, cumulative tables, first and last bookingE. Batch, C. WarehouseAirbnb 2024 / 2026, Snowflake 2025REPORTED74
13Retail supply chain model: vendor to warehouse to customerC. Warehouse, K. BusinessAmazon 2025REPORTED75
14Subscription data model for a membership like Amazon PrimeC. WarehouseAmazon 2026REPORTED76
15Data reconciliation and validation systemH. Data qualityMicrosoft 2024REPORTED77
16Logging and a data mart for a new marketplace featureC. Warehouse, J. AnalyticsMeta 2026VERIFIEDREPORTED78
17Short-video (Reels) engagement analytics and a metric that movedJ. Analytics, K. BusinessMeta 2024-26PATTERN80
18Real-time feature pipeline for recommendationsD. Streaming, ML dataApple 2025, Meesho blog, Swiggy blogREPORTED81
19Real-time fraud detection for UPI and card paymentsD. StreamingSwiggy blog, Razorpay blogINFERRED82
20Log ingestion platform at 50 TB a dayA. IngestionUber 2026, Zomato blogREPORTED83
21Lakehouse with bronze, silver and gold for an e-commerce companyB. LakehousePaytm 2025-26, Walmart 2026, Morgan Stanley 2021REPORTED84
22Internal data platform with a job scheduler: retries, SLAs, backfillsF. PlatformLyft 2026, Robinhood 2026REPORTED85
23Metadata catalog with lineage and data discoveryF. Platform, H. GovernanceLinkedIn blogINFERRED86
24User-profile store that merges attributes from many teamsF. Platform, B. LakehouseJioHotstar 2024REPORTED87
25Ad spend tracking and budget pacingK. AdsNetflix 2026, Zepto blogREPORTED88
26Real-time inventory across stores and warehousesK. BusinessGoogle AE 2024, Walmart 2025REPORTED89

Who it is for

Data engineers with a design round coming up

Freshers who want to see the bar, and engineers with two to eight years of experience who are expected to drive the round. The book is clear about what each level is expected to do on their own.

  • SDE-2 / L4

    2-4 years

    Expected to drive: A correct, buildable design for a scoped prompt. A clear model and grain. Duplicates and late data handled when the interviewer asks.

    Common miss: Vague boxes, no grain, no numbers.

  • Senior / L5

    4-8 years

    Expected to drive: Runs the scoping unprompted, does the numbers, raises failure modes before being asked, names costs, defends the design at 10x.

    Common miss: Waits for the interviewer to bring up failure; cannot say what breaks at 10x.

  • Staff / L6

    8+ years

    Expected to drive: Questions the requirement itself. Designs for many teams: contracts, self-serve, quotas. Plans migrations and org-wide cost.

    Common miss: A good pipeline with no platform view and no migration path.

How much time do you have?

  • Two weeks

    One part every two days. In Part 6, answer section 01 of each design on paper before reading the rest.

  • Three days

    Part 2 (the PIPELINE method), the Spark and Kafka chapters, the six flagship designs, then Part 7.

  • One night

    The Part 7 decision tree, the cheat sheets and the master checklist. Then talk through one flagship design out loud.

Chapter 05

The PIPELINE method, in one table

Eight steps in a 45-minute timebox, and something on the board after each one. The full chapter, with what to draw at every step, is one of the free pages.

  1. Pmin 0-3

    Purpose

    Who consumes the data, and which decisions or queries it serves.

  2. Imin 3-5

    Inputs

    Sources, event shape, append-only or mutable, ordering, and how late data arrives.

  3. Pmin 5-8

    Peak and volume

    Average and peak events a second, event size, daily volume, retention, growth.

  4. Emin 8-10

    Expectations

    Freshness, correctness, availability, consistency and compliance, each with a number.

  5. Lmin 10-17

    Layout

    Data model and grain, keys, partitioning, file and table format.

  6. Imin 17-29

    Implementation

    Ingest, process, store and serve, batch or streaming, drawn left to right.

  7. Nmin 29-37

    Negative paths

    Failures, retries, late, duplicate and bad data, replay and backfill, data quality.

  8. Emin 37-42

    Economics and evolution

    Cost levers, security, the 10x plan, and the trade-offs you made.

Read the PIPELINE chapter free

Free sample

Read 20 of the 111 pages free

Pages 1-5, 13, 15-21, 57-60 and 92-94: the contents, chapter 01 on the 100 questions, the level table, the PIPELINE method, clarifying questions and scale estimation, the first two flagship designs, and the chapters on drawing the architecture and turning a question into one.

Checking your access to the book…

Durgesh Yadav

About the author

Durgesh Yadav

Data Engineer · Instructor · Founder, PrepNPlaced

Durgesh teaches the AI-Powered Data Engineering cohort at PrepNPlaced and mentors data engineers one to one. This book is the design round as he teaches it: one method, worked designs, and the follow-ups to expect.

More about Durgesh

Questions

Before you buy

What format is the book in?

A 111-page PDF in A4 portrait. It reads on a laptop, a tablet or a phone, and prints on A4 paper.

What does it cost?

₹299 until 11:59 pm IST on 14 October 2026, then ₹499. One payment, no subscription.

How do I get the PDF after paying?

Sign in with Google, pay through Cashfree, and a Download button appears on this page. The download is for life: the book stays on your account, and you can download it again whenever you sign in.

Can I read some of it first?

Yes. After a Google sign-in you can read 20 of the 111 pages here: the contents, how to use the book, chapter 01 on the 100 questions, the level table, the PIPELINE method, clarifying questions, scale estimation, the first two flagship designs, and the chapters on drawing the architecture and turning a question into one.

Is it included in my cohort or mentorship?

Yes, if you are a learner in the AI-Powered Data Engineering cohort or in 1:1 mentorship with Durgesh. Sign in with the email you enrolled with and the Download button appears in place of the price.

Is it part of the Learn Pass, the notes vault or a plan?

No. The book is sold on its own. The Learn Pass, the notes vault and the paid plans do not include it.

Are the 100 questions real?

Each one is tagged. VERIFIED means an official company document says it, REPORTED means a candidate wrote it up after their own interview, PATTERN means the prompt repeats across several posts with no single report behind it, and INFERRED is the author's own reasoning. Only 2 of the 100 rows are VERIFIED, and the book says so.

Who is it for?

Data engineers preparing for a design round, from freshers who want to see the bar to engineers with 2 to 8 years of experience who are expected to drive the round. Staff-level expectations are covered too.

Can I get a refund?

Refunds are handled under our Refund Policy. If something is wrong with your copy, email support@prepnplaced.com from the address you paid with and include your order id. Read the Refund Policy