Cohort 4Register

Data Engineering System Design Interview Questions

These are 100 design and data modeling questions that candidates reported from data engineering loops at 30+ companies between 2020 and 2026. Each row gives the company, the role, the year and a tag for how strong the evidence is. The questions are paraphrased and cut to the core ask.

The list is the appendix of Data Engineering System Design: Zero to Hero, where 26 of these prompts are worked into full designs.

Published 7 October 2026 · By Durgesh Yadav

Read this first

How to read the evidence tags

A question that one candidate wrote up is not the same as a question a company publishes. Every row carries a tag that says which it is, so you can decide how much weight to give it.

VERIFIED
An official company document says it. Only 2 of the 100 rows qualify, and neither is a design prompt.
REPORTED
A candidate wrote it up after their own interview: one candidate's account of one interviewer on one day.
PATTERN
The same prompt repeats across several community posts, but no single report can be pinned down.
INFERRED
The author's own reasoning. Useful, but treat it as an opinion.

Some rows also carry a note beside the tag: weak means a second-hand snippet or a topic list rather than a prompt; aggregator means a prep site says a candidate reported it but does not link the report; pre-2022 means the interview is older than 2022, so the loop may have changed.

Questions 1 to 40

MAANG + Microsoft (40)

Meta 11, Amazon 8, Microsoft 6, Netflix 6, Apple 5, Google 4.

Evidence in this group: 1 VERIFIED, 32 REPORTED, 7 PATTERN.

MAANG + Microsoft: 40 reported data engineering design questions
#CompanyRoleYearQuestionEvidence
1MetaDEundatedRound format, not a prompt: a product's data needs, then its logging, a data mart and reporting SQL.VERIFIED3rd-party copy
2MetaDE2026Model a new housing-marketplace feature, define KPIs, then SQL for view-to-lead conversion by location.REPORTED
3MetaDE2025, 2021Data model for a gaming company like Epic Games, with the schema trade-offs.REPORTED
4MetaDE2026Model an Uber-style business, then KPI SQL, Python data processing and chart choices on the same model.REPORTED
5MetaDE (Product Analytics), E42024-26Define success metrics for a product: ride matching, Twitter, a video platform, Reels, Marketplace.PATTERN
6MetaProduct DE, L52025Design facts and dimensions, state each relationship's cardinality, then SQL on a given schema.REPORTED
7MetaDE, E4n/sDefend a star schema and argue the pros and cons of the alternatives.REPORTED
8MetaDE, Analytics2025How would you redesign your data model to make it more granular?REPORTEDweak
9MetaDEn/sModel events that never happened, such as products never ordered, using factless fact tables.PATTERN
10MetaDE2024-26Simulate a stream with lists and dicts: a running average, or ride requests per 15-minute window.PATTERN
11MetaDE2025-26A core metric moves unexpectedly (a DAU spike, an Instagram drop): find the cause and the data needed.PATTERNprep sites
12AmazonDE2026Design a data model for Amazon Prime from scratch; expect many follow-ups.REPORTED
13AmazonDE2025Model retail from vendor to Amazon warehouse to customer; a separate round covered an AWS pipeline.REPORTED
14AmazonDE2024Design a streaming data model and a batch one. Have you built SCD Type 2?REPORTED
15AmazonDE2022ETL pipeline built for scale; model and query from scratch; ETL vs ELT; how you test integrity.REPORTED
16AmazonDE II, L52021Design a data warehouse for a clothing store and write the DDL.REPORTEDpre-2022
17AmazonDE II, L52021Model your latest project, then slowly changing dimensions and surrogate keys.REPORTEDpre-2022
18AmazonDE, L52020-25Pipeline for a large dataset: justify the technology, explain how it scales and how you monitor it.PATTERN
19AmazonDE, L52020Fact-dimension model: design the primary key and handle partial data.REPORTEDpre-2022
20GoogleDE2026Real-time pipeline for clickstream from millions of users, such as ad-click analytics.REPORTED
21GoogleDE2026Design and build an end-to-end batch data pipeline.REPORTED
22GoogleAE (Data), L42024Central ERP integration that tracks warehouse inventory across a global network.REPORTED
23GoogleAE (Data), L42024Car-rental reservations with vehicle tiers, discounts and multi-tenancy, down to the data model.REPORTED
24MicrosoftDE2024Schema for employee check-ins and check-outs that yields hours worked, then query it.REPORTED
25MicrosoftDE2024E-commerce schema (products, orders, customers), SQL on it, then Spark performance tuning.REPORTED
26MicrosoftDE2024Data reconciliation and validation system; size Spark resources for different volumes.REPORTED
27MicrosoftSenior SWE (Data)2024Daily incremental S3-to-Oracle ingestion; estimate executor memory for 1 TB.REPORTED
28MicrosoftDE2024Design a real-time pipeline that processes streaming data.REPORTED
29MicrosoftDE2024Incremental loads in Azure Data Factory, normalization trade-offs, Z-ordering; redo a SQL answer in PySpark.REPORTED
30NetflixSenior DE, L52025Rewrite a single-threaded Scala solution in Spark, then as a streaming job, weighing trade-offs.REPORTED
31NetflixSenior DE, L52025Design a pipeline architecture, then work through debugging scenarios.REPORTED
32NetflixSenior DE2026Design a data pipeline, then defend whether it scales to Netflix's volume.REPORTED
33NetflixDE, L52025A real scenario from the team's own work; Spark depth and modeling of unstructured data.REPORTED
34NetflixDEn/sSketch the data or object model of a system you built; the round is fitted to the team.PATTERN
35NetflixSWE, Senior+2026Data model to track a direct-sold demand order on an advertising DSP.REPORTEDweak
36AppleDE2024Data model for a given scenario, then medium-to-hard SQL on it.REPORTED
37AppleDE2024Design a clickstream processing system for log data.REPORTED
38AppleAnalytics DE2025Design a pipeline, then a model that copes with schema evolution and slowly changing dimensions.REPORTED
39AppleDE, Software DEto 2025Spark depth: partitioning, skew, executor count and memory, dynamic allocation, MapReduce vs Spark.PATTERN
40AppleSoftware DE, Media Products2025Scalable, reliable pipelines, and how the data feeds ML for personalized recommendations.REPORTED

Questions 41 to 70

US tech and finance (30)

Airbnb, DoorDash, Uber, Lyft, Stripe, LinkedIn, Pinterest, Snowflake, Databricks, Salesforce, Adobe, JPMorgan, Morgan Stanley, Robinhood.

Evidence in this group: 27 REPORTED, 3 PATTERN.

US tech and finance: 30 reported data engineering design questions
#CompanyRoleYearQuestionEvidence
41UberSWE2025Live driver heatmap: GPS pings into grid cells, density pushed to clients, top-K hottest cells per region.REPORTEDaggregator
42UberDE2025-26Rides or Eats data model serving city revenue, driver and merchant payouts, funnel and retention reports.PATTERNprep sites
43UberDE2025-26Real-time pipeline for driver location, matching or surge pricing with sub-second freshness.PATTERNprep sites
44AirbnbSenior DE2024Design a bookings table and a properties table, then query them.REPORTED
45AirbnbSenior DE2024A SQL dump lands in Hive daily partitions; from two given tables, write the ETL query for the target output.REPORTED
46AirbnbDE2025Design several tables, choose their partitioning, explain scaling, then SQL on the model.REPORTED
47AirbnbDE2021SQL that loads a new date into a ds-partitioned table from incremental data in earlier partitions.REPORTEDpre-2022
48AirbnbDE2026Model cumulative booking metrics plus each account's first and latest booking time.REPORTED
49LinkedInSenior DE, IC32025End-to-end data system for several use cases: ETL, storage optimization, query performance.REPORTED
50LinkedInSWE2026Collect user activity events and answer counts for the last minute, hour and day.REPORTEDaggregator
51LyftDE2026Design the driver and rider data models.REPORTED
52LyftDE2026Backend API that is the entry point for real-time driver request processing.REPORTED
53LyftDE2026Model a use case around a given query pattern; system design leaned toward a data platform.REPORTED
54DoorDashDE2024, 2026Dimensional model, facts and dimensions, for a fitness app.REPORTEDtwo mentions
55DoorDashDE2026End-to-end pipeline with both a batch and a real-time path.REPORTEDaggregator
56DoorDashDE2026How do you handle late-arriving data?REPORTED
57DoorDashDEundatedSchema for a customer's address history over time, including past occupants and moves.REPORTEDaggregator
58PinterestStaff SWE (Data)2026Analytics platform that produces derived data; hot partitions and skew from popular boards.REPORTED
59PinterestStaff SWE (Data)2026Top boards by engagement over a 5-minute sliding window; Kafka rebalancing and Spark shuffle internals.REPORTED
60StripeSWE, Senior+2026Count-metrics monitoring with rates, labels and alerts that copes with duplicate and late events.REPORTEDaggregator
61StripeSWE2025Near-real-time activity counter per user, device and region; offline clients sync later.REPORTEDaggregator
62StripeSWE, mixed2025-26Idempotent double-entry ledger, or an auditable ledger with reconciliation.PATTERN
63SnowflakeDE2025Keep a cumulative historical sales table up to date from current data, in SQL.REPORTED
64SnowflakeDE (per IQ)2026Metadata service for columnar storage: sharding, concurrent updates, min/max pruning, compaction.REPORTEDaggregator
65DatabricksDE (per IQ)2026Design a distributed file system without S3 or any commercial blob store.REPORTEDaggregator
66SalesforceDE2026Your warehousing approach for bringing in data from many sources.REPORTED
67AdobeDE2026End-to-end solution for a loosely defined problem; state assumptions, plan for scale.REPORTED
68JPMorganDEundatedArchitecture spec under hardware limits; requirements changed mid-presentation, so offer an alternative.REPORTED
69Morgan StanleyDE2021Modeling and normalization, Databricks lakehouse, PySpark ETL design, batch vs streaming.REPORTEDpre-2022
70RobinhoodSWE2026Distributed job scheduler with retries and SLA guarantees.REPORTED

Questions 71 to 100

India product and retail (30)

Flipkart, Razorpay, Walmart GT, Meesho, Paytm, Target, Nike, JioHotstar, Uber Hyderabad, VISA, Swiggy, unnamed startups.

Evidence in this group: 1 VERIFIED, 25 REPORTED, 4 PATTERN.

India product and retail: 30 reported data engineering design questions
#CompanyRoleYearQuestionEvidence
71RazorpayDE, 2.3 yrsc. 2024A service that does Delta OPTIMIZE's job on S3 for a real-time workload, with monitoring and failover.REPORTED
72RazorpayDE, 2.3 yrsc. 2024Self-healing system that loads hundreds to thousands of SQL and NoSQL databases into a warehouse.REPORTED
73RazorpayDE, 2.3 yrsc. 2024Resume projects end to end: challenges, data quality metrics tracked, Spark internals.REPORTED
74RazorpayDE, ~3 yrsc. 2022SQL, data modeling and Spark internals in one round (topics, no prompt).REPORTEDweak
75FlipkartDE II2024Model a cricket tournament where players turn out for several teams; SQL for cumulative player scores.REPORTED
76FlipkartDE II2024Pipeline failure scenarios and consuming from Kafka in depth; Spark dynamic allocation, caching, joins.REPORTED
77FlipkartDE-I to DE II2021-24Machine coding: join several (nested) JSON files in Spark and answer SQL-style questions.PATTERN3 reports
78FlipkartDE-I2021Financial database: models, schema, fields, partition keys and query optimization.REPORTEDpre-2022
79FlipkartDE2021E-commerce data model with each table and relationship explained, then three SQL questions.REPORTEDpre-2022
80FlipkartDE2021Explain an ETL scenario end to end, from ingestion to the warehouse.REPORTEDpre-2022
81FlipkartDE2026Web crawling, API calls and anomaly detection, plus system design questions.REPORTEDweak
82MeeshoSDE-2 DE, ~3 yrsc. 2022SQL, modeling and Spark internals, then Spark coding with Meesho's panel (topics, no prompt).REPORTEDweak
83MeeshoSDE III Data, 5-8 yrs2026Job description: fault-tolerant batch and streaming pipelines on Spark, Flink and Kafka; deep Spark internals.VERIFIEDjob post only
84PaytmDE, 1 yr2023Design an online ticketing system: data flow and architecture.REPORTED
85PaytmTech Lead2025-26Design a data warehouse or a data lake.REPORTEDtwo reviews
86Walmart GTDE-32023Design Mixpanel: capture events from Android, iOS and web apps.REPORTED
87WalmartSenior DE2026End-to-end data platform: batch vs streaming, schema, storage layer; ingestion and recovery earlier.REPORTEDtopics only
88Walmart GTDE to Senior DE2022-26Own design round; recurring topics: Scala Spark, Kafka, Airflow, Spark tuning, schema evolution.PATTERN
89Walmart GTSDE2025Design an order schema, then an inventory management system.REPORTEDSDE loop
90Target IndiaDE2022What is a data pipeline and how would you build one? Warehouse vs lake.REPORTED
91Target IndiaDE2022Walk through an end-to-end project pipeline, then design and service comparisons.REPORTEDweak
92NikeDE2021Bucketing external Hive tables, Parquet file sizing, bucket joins and map joins.REPORTEDpre-2022
93NikeDE2025Normalization, SQL optimization, partitioning and indexing.REPORTEDweak
94JioHotstarSDE-2 DE2024Profile store merging daily attributes from many teams, each with a min, max, first or last rule.REPORTED
95JioHotstarSDE-2 DE2024Design a threaded comment system like YouTube's.REPORTED
96Uber HyderabadSDE-2 DE, L42024Top-10 movies by category and window; a view counts at 80% watched, so merge overlapping intervals.REPORTED
97Uber HyderabadSDE-2 DE, L42024Clickstream pipeline for an hourly trending-by-location dashboard, scored on completed views.REPORTED
98VISA BangaloreSenior DE2024Every step from a CSV file to a partitioned table; update 200 of 10,000 Hive rows with no timestamp.REPORTED
99Startups (unnamed)DE, platform, SDE-Datan/sDistributed-systems prompts: a Kafka-like log, an event-driven system, a distributed cache.PATTERNone thread
100SwiggyDE2026Deep dive on pipelines you ran at scale, framed on Swiggy order and delivery data.PATTERNweak

Chapter 01

What real loops asked, in 11 categories

The book sorts the 100 questions into these categories. The grouping is the author's own reading, and the last column lists the designs in the book that answer each one.

Data engineering design questions by category
CategoryWhat was askedEvidenceDesigns in the book
A. Data ingestionRazorpay c. 2024: self-healing loads from thousands of SQL and NoSQL databases. Microsoft 2024: an incremental S3-to-Oracle load sized for 1 TB. VISA 2024: a CSV to a partitioned table.Good: several direct reports03, 20
B. Lake / lakehouseRazorpay c. 2024: a service that does the job of Delta OPTIMIZE on S3. Paytm 2025-26: design a lake or warehouse. Morgan Stanley 2021: a Databricks lakehouse.Moderate: few full prompts11, 21, 24
C. Warehouse / modelingMeta 2026: model a housing-marketplace feature, then conversion SQL. Amazon 2026: Prime from scratch. Flipkart 2024: a cricket tournament, players in several teams.Strongest: a verified guide plus most loops02, 12, 13, 14, 16
D. StreamingGoogle 2026: a real-time clickstream or ad-click pipeline. Uber Hyderabad 2024: hourly trending titles by location. Stripe 2026: counts that survive duplicate and late events.Strong: many reports, some via aggregators01, 04, 08, 09, 10, 18, 19
E. BatchGoogle 2026: an end-to-end batch pipeline. Airbnb 2024: a SQL dump lands in Hive daily partitions, write the ETL. Amazon 2022: ETL vs ELT, and testing integrity.Good12; batch paths in 04, 21
F. PlatformLyft 2026: a design round that leaned toward a data platform. Walmart 2026: an end-to-end platform and the storage choice. Robinhood 2026 (SWE): a job scheduler with SLAs.Moderate: senior or SWE loops22, 23, 24
G. ReliabilityNetflix 2025: design a pipeline, then debug failure scenarios. Flipkart 2024: failures while reading Kafka. Stripe 2025-26: an idempotent double-entry ledger (a pattern).Good, and the most common follow-up06, 08
H. Data quality / governanceMicrosoft 2024: a reconciliation and validation system. Razorpay c. 2024: the data quality metrics you tracked. Amazon 2022: testing integrity. No catalog or lineage prompt turned up.Moderate15, 23
I. Distributed systemsDatabricks 2026: a distributed file system without S3. Snowflake 2026: a metadata service for columnar storage. Indian startups: build a Kafka-like log (a pattern).Thin: aggregators and one forum threadNone of its own: Part 3, and rows 64, 65 and 99
J. AnalyticsWalmart GT 2023: design Mixpanel. LinkedIn 2026: activity counts for the last minute, hour and day. Meta 2024-26: success metrics for a product (a pattern).Good05, 07, 09, 16, 17
K. Business systemsAmazon 2025: vendor to warehouse to customer. Google AE 2024: inventory across a global warehouse network. Paytm 2023: an online ticketing system.Good, but the domain changes each time02, 04, 05, 06, 13, 17, 25, 26

What the list says

Six things to take into the room

  1. 01

    Modeling shows up in nearly every loop

    Meta's guide gives modeling a round of its own, and candidates report modeling prompts at Meta, Amazon, Airbnb, Lyft, DoorDash, Flipkart, Microsoft and Apple in 2024-2026. Most then ask SQL on the model you drew, so a wrong grain costs you twice.

  2. 02

    Late data and duplicates are the near-certain follow-up

    DoorDash asked it outright in 2026. Stripe, LinkedIn and Airbnb list retries, double counts or late bookings, and Google's 2026 clickstream round covered both. Have the answer ready before you draw.

  3. 03

    Requirements can change mid-round

    A JPMorgan candidate had the requirements changed during the architecture presentation and had to offer an alternative on the spot. Writing your requirements on the board is what lets you see what changed.

  4. 04

    Spark internals get tested next to design

    Apple, Flipkart, Microsoft and Razorpay all went into executors, memory, skew, dynamic allocation or broadcast joins after the diagram. Expect "how many executors?" right after you draw.

  5. 05

    Your past project is a design prompt

    Netflix asks you to sketch the model of a system you built, and Razorpay, Target India and Amazon had candidates walk through a past pipeline. Bring one project already drawn and sized.

  6. 06

    Some popular targets have almost no evidence

    No data engineering design prompt turned up for Swiggy, Zomato, PhonePe or CRED. Google is thin for pure data engineering roles, and finance reports mostly list topics rather than prompts.

What this list cannot tell you

  • Rows are not all independent: rows 75, 76 and 94-98 come from one candidate's 2024 write-up, rows 74 and 82 from another, and rows 30-32 may be one candidate. There are fewer independent reports than rows.
  • Of the 84 REPORTED rows, 16 are weak snippets or older than 2022, and 8 US rows reach us through prep aggregators.
  • No data engineering design prompt was found for Swiggy, Zomato, PhonePe, CRED, Zepto, Myntra or Ola, and no first-person Uber data engineering prompt. Finance loops gave topic lists or SWE-style low-level design.
  • Some sources could not be read when the research was done in October 2026 (blocked, timed out or behind a login wall), so coverage has holes.

The worked answers are in the book

Data Engineering System Design: Zero to Hero

  • 100 reported questions from 30+ companies, each tagged by evidence
  • The PIPELINE method: eight steps for the 45 minutes
  • 26 designs worked end to end, 19 cheat sheets and a master checklist

₹299 until 14 October, then ₹499. PDF, 111 pages.

More on the round: system design interview questions, system design prep by company and the system design mock interview.