New book · 111 pages · PDF
Data Engineering System DesignZero to Hero
How to think, design, defend and scale data systems in a real interview. One framework, 26 worked designs with real follow-ups, and the reasoning that separates a hire from a no-hire.
By Durgesh Yadav · Data Engineer · Instructor · Founder, PrepNPlaced · 2026 edition
₹299Launch price
Until 11:59 pm IST on 14 October 2026. From 15 October it is ₹499.
One payment. PDF, yours to download again any time.
Included free for learners in the AI-Powered Data Engineering cohort and in 1:1 mentorship: sign in with the email you enrolled with.

- reported interview questions, each tagged by evidence
- 100reported interview questions, each tagged by evidence
- companies, from MAANG to Razorpay and Flipkart
- 30+companies, from MAANG to Razorpay and Flipkart
- designs worked end to end with follow-ups
- 26designs worked end to end with follow-ups
- cheat sheets plus one master checklist
- 19cheat sheets plus one master checklist
What is inside
Built from what interviewers actually asked
The book starts from 100 design prompts that candidates reported from real loops, sorts them by what they test, and teaches the method and the designs those prompts call for.
100 reported interview questions
Design and modeling prompts from 2020-2026 loops at 30 named companies, from MAANG to Razorpay and Flipkart. Each is tagged VERIFIED, REPORTED, PATTERN or INFERRED, so you know how much weight it can carry.
The PIPELINE method
Eight steps in a 45-minute timebox, with something on the board after every step. When the interviewer pushes or you lose the thread, you know which step you are on and what is left.
26 worked designs
Six flagships in fifteen sections each, from clarifying questions to the follow-ups that sink most answers, and twenty compact designs on one page each.
Numbers you can say out loud
One planning number sheet and seven worked estimates, each ending with the design decision the number forces.
Follow-ups and answer levels
A follow-up simulator, bad, good and excellent answers side by side, and what senior and staff candidates are expected to raise on their own.
19 cheat sheets and a master checklist
Nineteen tools and ideas, seven lines each in the same order, then the whole method on one page for the night before.
Three of the pages

Page 15 · The PIPELINE method: eight steps, a 45-minute timebox, and what goes on the board after each. 
Page 59 · Flagship 02, ride-hailing: the evidence on the header, ten clarifying questions, the numbers, the diagram. 
Page 94 · Requirement to architecture: eight questions, asked in this order, and what each answer rules out.
Contents
Eight parts and an appendix
Chapter 01 comes first, on what the 100 questions say. Then the book moves from what the round tests to the night before it.
- 01What 100 reported interview questions tell usfree to readp. 3
Part 1 · from page 6
Foundations
What a data engineering design round is testing, the five stages every data system is built from, and how the round runs at different companies. It starts with one kirana store and grows it to 5,000.
- 02What system design means for a data engineerp. 7
- 03Anatomy of a data systemp. 10
- 04What the round actually looks likep. 12
Part 2 · from page 14
System design thinking
The method used for every design in the book, the questions that turn a vague prompt into numbers, and back-of-envelope maths you can do out loud.
- 05The PIPELINE methodfree to readp. 15
- 06The art of the clarifying questionfree to readp. 16
- 07Scale estimation you can do out loudfree to readp. 19
Part 3 · from page 22
Distributed systems fundamentals
How data is split, how it is copied, what consistent means when machines disagree, and how bytes are laid out on disk, each taught through something you will draw in the round.
- 08Partitioning and shardingp. 23
- 09Replication, consensus and coordinationp. 25
- 10CAP, PACELC and consistencyp. 27
- 11The scaling toolkitp. 28
- 12Row, columnar and table formatsp. 29
Part 4 · from page 32
Data engineering architecture
The patterns you will be asked to compare, where each kind of data should live, how to model it, and what you say when the interviewer asks what breaks, who can see it, whether it is right and what it costs.
- 13Architecture patterns, and when not to use themp. 33
- 14Storage decision guidep. 36
- 15Data modeling for design roundsp. 37
- 16Reliability engineeringp. 39
- 17Data quality in productionp. 41
- 18Security and governancep. 42
- 19Costp. 42
Part 5 · from page 44
Core technologies
What each building block is for and where it breaks, then deep chapters on the two that interviewers probe hardest: Spark and Kafka.
- 20Building blocksp. 45
- 21Spark system designp. 47
- 22Kafka system designp. 52
Part 6 · from page 55
Real interview designs
Twenty-six designs, most traced to a prompt candidates reported, with the evidence printed on each header. Six flagships in fifteen sections, twenty compact designs on one page each.
- 23The 26 designs at a glancep. 56
- 24Twenty compact designsp. 69
Part 7 · from page 91
Interview mastery
The 45 minutes themselves: drawing a design that reads at a glance, turning a vague prompt into choices, holding your ground through follow-ups, and what moves an answer from good to a strong hire.
- 25How to draw the architecturefree to readp. 92
- 26From question to architecturefree to readp. 93
- 27The follow-up simulatorp. 95
- 28Bad, good and excellent answersp. 96
- 29Senior and staff thinkingp. 98
Part 8 · from page 99
Cheat sheets and the master checklist
The night before the round: nineteen cards, seven lines each in the same order, then the whole method on one page.
- 30Cheat sheetsp. 100
- 31The system design master checklistp. 105
Appendix A · from page 106
The question database
Every interview question behind the book, with an honest label for how much weight each one can carry.
- 32All 100 questions with sourcesp. 107
Part 6
The 26 designs
6 flagships are worked in fifteen sections each, from the clarifying questions to the follow-ups. The other 20 fit on one page each. Every header says where the prompt was reported and how strong that evidence is.
| # | Design | Reported at | Evidence | Page |
|---|---|---|---|---|
| 01 | Real-time clickstream and ad-click analytics pipelineD. StreamingFlagshipFree to read | Google 2026, Apple 2024, Uber blog, Zepto blog | REPORTED | 57 |
| 02 | Ride-hailing data model and trip analyticsC. Warehouse, K. BusinessFlagshipFree to read | Meta 2026, Lyft 2026, Uber prep sites | REPORTEDPATTERN | 59 |
| 03 | Self-healing ingestion from 1,000+ databases into the lakehouseA. IngestionFlagship | Razorpay c. 2024, Salesforce 2026, Microsoft 2024 | REPORTED | 61 |
| 04 | Food-delivery order analytics with batch and real-time pathsK. Business, D. StreamingFlagship | DoorDash 2026, Swiggy blog, Airbnb blog | REPORTED | 63 |
| 05 | Viewing analytics: top-10 titles and hourly trending by locationJ. Analytics, K. BusinessFlagship | Uber Hyderabad 2024, Netflix 2025 | REPORTED | 65 |
| 06 | Payment ledger and transaction analytics with reconciliationG. Reliability, K. PaymentsFlagship | Stripe 2025-26, Flipkart 2021, Razorpay blog | PATTERNREPORTED | 67 |
| 07 | Product analytics platform like MixpanelJ. Analytics | Walmart GT Bengaluru 2023 | REPORTED | 69 |
| 08 | Count-metrics monitoring with duplicate and late eventsD. Streaming, G. Reliability | Stripe 2026 | REPORTED | 70 |
| 09 | User activity counts for the last minute, hour and dayD. Streaming, J. Analytics | LinkedIn 2026 | REPORTED | 71 |
| 10 | Real-time driver heatmap and surge signalD. Streaming | Uber 2025 | REPORTEDPATTERN | 72 |
| 11 | A compaction service that does what Delta OPTIMIZE doesB. Lakehouse | Razorpay c. 2024 | REPORTED | 73 |
| 12 | Booking metrics: daily partitions, cumulative tables, first and last bookingE. Batch, C. Warehouse | Airbnb 2024 / 2026, Snowflake 2025 | REPORTED | 74 |
| 13 | Retail supply chain model: vendor to warehouse to customerC. Warehouse, K. Business | Amazon 2025 | REPORTED | 75 |
| 14 | Subscription data model for a membership like Amazon PrimeC. Warehouse | Amazon 2026 | REPORTED | 76 |
| 15 | Data reconciliation and validation systemH. Data quality | Microsoft 2024 | REPORTED | 77 |
| 16 | Logging and a data mart for a new marketplace featureC. Warehouse, J. Analytics | Meta 2026 | VERIFIEDREPORTED | 78 |
| 17 | Short-video (Reels) engagement analytics and a metric that movedJ. Analytics, K. Business | Meta 2024-26 | PATTERN | 80 |
| 18 | Real-time feature pipeline for recommendationsD. Streaming, ML data | Apple 2025, Meesho blog, Swiggy blog | REPORTED | 81 |
| 19 | Real-time fraud detection for UPI and card paymentsD. Streaming | Swiggy blog, Razorpay blog | INFERRED | 82 |
| 20 | Log ingestion platform at 50 TB a dayA. Ingestion | Uber 2026, Zomato blog | REPORTED | 83 |
| 21 | Lakehouse with bronze, silver and gold for an e-commerce companyB. Lakehouse | Paytm 2025-26, Walmart 2026, Morgan Stanley 2021 | REPORTED | 84 |
| 22 | Internal data platform with a job scheduler: retries, SLAs, backfillsF. Platform | Lyft 2026, Robinhood 2026 | REPORTED | 85 |
| 23 | Metadata catalog with lineage and data discoveryF. Platform, H. Governance | LinkedIn blog | INFERRED | 86 |
| 24 | User-profile store that merges attributes from many teamsF. Platform, B. Lakehouse | JioHotstar 2024 | REPORTED | 87 |
| 25 | Ad spend tracking and budget pacingK. Ads | Netflix 2026, Zepto blog | REPORTED | 88 |
| 26 | Real-time inventory across stores and warehousesK. Business | Google AE 2024, Walmart 2025 | REPORTED | 89 |
Who it is for
Data engineers with a design round coming up
Freshers who want to see the bar, and engineers with two to eight years of experience who are expected to drive the round. The book is clear about what each level is expected to do on their own.
SDE-2 / L4
2-4 years
Expected to drive: A correct, buildable design for a scoped prompt. A clear model and grain. Duplicates and late data handled when the interviewer asks.
Common miss: Vague boxes, no grain, no numbers.
Senior / L5
4-8 years
Expected to drive: Runs the scoping unprompted, does the numbers, raises failure modes before being asked, names costs, defends the design at 10x.
Common miss: Waits for the interviewer to bring up failure; cannot say what breaks at 10x.
Staff / L6
8+ years
Expected to drive: Questions the requirement itself. Designs for many teams: contracts, self-serve, quotas. Plans migrations and org-wide cost.
Common miss: A good pipeline with no platform view and no migration path.
How much time do you have?
Two weeks
One part every two days. In Part 6, answer section 01 of each design on paper before reading the rest.
Three days
Part 2 (the PIPELINE method), the Spark and Kafka chapters, the six flagship designs, then Part 7.
One night
The Part 7 decision tree, the cheat sheets and the master checklist. Then talk through one flagship design out loud.
Chapter 05
The PIPELINE method, in one table
Eight steps in a 45-minute timebox, and something on the board after each one. The full chapter, with what to draw at every step, is one of the free pages.
Pmin 0-3
Purpose
Who consumes the data, and which decisions or queries it serves.
Imin 3-5
Inputs
Sources, event shape, append-only or mutable, ordering, and how late data arrives.
Pmin 5-8
Peak and volume
Average and peak events a second, event size, daily volume, retention, growth.
Emin 8-10
Expectations
Freshness, correctness, availability, consistency and compliance, each with a number.
Lmin 10-17
Layout
Data model and grain, keys, partitioning, file and table format.
Imin 17-29
Implementation
Ingest, process, store and serve, batch or streaming, drawn left to right.
Nmin 29-37
Negative paths
Failures, retries, late, duplicate and bad data, replay and backfill, data quality.
Emin 37-42
Economics and evolution
Cost levers, security, the 10x plan, and the trade-offs you made.
Free sample
Read 20 of the 111 pages free
Pages 1-5, 13, 15-21, 57-60 and 92-94: the contents, chapter 01 on the 100 questions, the level table, the PIPELINE method, clarifying questions and scale estimation, the first two flagship designs, and the chapters on drawing the architecture and turning a question into one.
Checking your access to the book…
Questions
Before you buy
What format is the book in?
A 111-page PDF in A4 portrait. It reads on a laptop, a tablet or a phone, and prints on A4 paper.
What does it cost?
₹299 until 11:59 pm IST on 14 October 2026, then ₹499. One payment, no subscription.
How do I get the PDF after paying?
Sign in with Google, pay through Cashfree, and a Download button appears on this page. The download is for life: the book stays on your account, and you can download it again whenever you sign in.
Can I read some of it first?
Yes. After a Google sign-in you can read 20 of the 111 pages here: the contents, how to use the book, chapter 01 on the 100 questions, the level table, the PIPELINE method, clarifying questions, scale estimation, the first two flagship designs, and the chapters on drawing the architecture and turning a question into one.
Is it included in my cohort or mentorship?
Yes, if you are a learner in the AI-Powered Data Engineering cohort or in 1:1 mentorship with Durgesh. Sign in with the email you enrolled with and the Download button appears in place of the price.
Is it part of the Learn Pass, the notes vault or a plan?
No. The book is sold on its own. The Learn Pass, the notes vault and the paid plans do not include it.
Are the 100 questions real?
Each one is tagged. VERIFIED means an official company document says it, REPORTED means a candidate wrote it up after their own interview, PATTERN means the prompt repeats across several posts with no single report behind it, and INFERRED is the author's own reasoning. Only 2 of the 100 rows are VERIFIED, and the book says so.
Who is it for?
Data engineers preparing for a design round, from freshers who want to see the bar to engineers with 2 to 8 years of experience who are expected to drive the round. Staff-level expectations are covered too.
Can I get a refund?
Refunds are handled under our Refund Policy. If something is wrong with your copy, email support@prepnplaced.com from the address you paid with and include your order id. Read the Refund Policy
