New · Cohort 3Engineering Analytics Cohort 3 goes live 26 September — only 30 seatsRegister Now
← All webinars
Live Webinar Sunday, 16 August 12:00 PM IST 2 hours · Free on Zoom

Data Engineering · Big Data & Spark

What is Spark? Introduction to PySpark

Understand big data processing from zero — then write your first PySpark job, live.

Your pandas script dies on a 20 GB file. Spark is how real companies process terabytes — and PySpark is how you write it in Python. Two hours. Zero big data background needed. You'll leave having written your first Spark job — with a working PySpark setup and a clear next step.

Agenda

What we'll cover

  • Why pandas breaks — and how Spark fixes it
  • How Spark really works: driver, executors, partitions, lazy evaluation
  • RDDs vs DataFrames — and which one you'll actually use
  • Transformations vs actions — the concept every beginner gets wrong
  • Live coding: set up PySpark and build a real job — load, filter, group, aggregate

Perfect for

Who it's for

Python users, analysts, and anyone heading into data engineeringAnyone whose pandas scripts are hitting real-data limitsBeginners — zero big data background needed; basic Python is enough

Every attendee

What you get

  • Live, hands-on session on Zoom
  • Live Q&A with the instructor
  • Session notes & key resources — emailed after
  • Optional: full session recording — add for ₹99
Durgesh Yadav

Your host

Durgesh Yadav

Sr. Data Engineer @ 7-Eleven · Instructor @ Bosscoder, Scaler & GeeksforGeeks

Free seat No spam

Reserve your free seat

One quick step: sign in with Google, then reserve your free seat — your Zoom link lands in your email instantly.

Free forever · Sunday 16 August · 12:00 PM IST

Free live seat

Sunday · 12:00 PM IST