Skip to main content

Command Palette

Search for a command to run...

Day 24 : Apache Spark for Beginners: RDDs, DataFrames, Transformations and Actions Explained

Updated
•8 min read•View as Markdown
Day 24 : Apache Spark for Beginners: RDDs, DataFrames, Transformations and Actions Explained
S

Developer based in India. Passionate learner and blogger. All blogs are basically Notes of Tech Learning Journey.

Meta Description: Learn Apache Spark fundamentals with a simple real-life analogy. Understand RDDs, DataFrames, transformations, and actions through hands-on PySpark code in Jupyter Notebook.

Introduction

Imagine you worked at that same warehouse from Day 22 and Day 23, and your manager handed you a huge pile of invoices to process by hand. It would take forever using one person.

Now imagine instead you hire a team of 10 workers. You don't ask them to process invoices immediately though. First, you hand out instructions: "sort these by date," "filter out the cancelled ones," "add up the totals." The workers don't actually start working the moment you give instructions. They wait until you say "go," and only then do they execute everything in one efficient pass.

This is almost exactly how Apache Spark works. Spark is a fast, distributed data processing engine built on top of ideas similar to Hadoop's MapReduce, but much faster because it processes data in memory instead of constantly reading and writing to disk.

The "instructions you hand out" (sort, filter, add) are called transformations, and the "go" command that actually triggers the work is called an action. The data structures that hold your invoices while workers process them are called RDDs and DataFrames.

Let's unpack all of this properly.

Theory

Why Spark exists: Hadoop's MapReduce (Day 22) writes intermediate results to disk after every step, which is reliable but slow. Spark keeps data in memory (RAM) across operations wherever possible, making it up to 100 times faster for many workloads.

A Brief History

Spark was created in 2009 at UC Berkeley's AMPLab by Matei Zaharia, and became an Apache top level project in 2014. It was designed specifically to address MapReduce's speed limitations while keeping the same core idea of distributed, parallel data processing.

RDD (Resilient Distributed Dataset)

An RDD is Spark's original core data structure. Breaking down the name tells you everything:

  • Resilient: can recover automatically from node failures

  • Distributed: split across multiple machines in a cluster

  • Dataset: simply, a collection of data

RDDs are low level and give you fine grained control, but they require more code to work with compared to newer structures.

DataFrame

A DataFrame is a higher level, more structured way to work with data in Spark, similar to a table in a database or a Pandas DataFrame, but distributed across a cluster. Most modern Spark work uses DataFrames instead of raw RDDs because they're easier to use and are automatically optimized by Spark's query engine (called Catalyst).

💡
You can think of RDDs as the "engine room" and DataFrames as the "dashboard." DataFrames are built on top of RDDs internally, but you rarely need to touch RDDs directly in modern Spark work.

Transformations vs Actions

This is the single most important concept to internalize in Spark.

Transformations are lazy. They don't run immediately, they just build a plan. Examples: map(), filter(), select(), groupBy(). Actions trigger the actual execution of that plan. Examples: collect(), count(), show(), take().

Why laziness matters: Spark waits until an action is called so it can look at the entire chain of transformations and optimize the execution plan as a whole, rather than running each step wastefully one at a time. This is exactly like the warehouse manager waiting to say "go" until all instructions are handed out, so the workers can plan the most efficient order to do things.

Practical

Good news: unlike Day 23, Spark can run locally on your own machine without needing a full cluster or Hadoop setup. We'll use PySpark, the Python API for Spark, directly in Jupyter Notebook.

Cell 1: Install PySpark (run once)

python !pip install pyspark

Cell 2: Create a Spark Session

from pyspark.sql import SparkSession

spark = SparkSession.builder.appName("Day24_SparkIntro").getOrCreate()
print("Spark session created successfully")
spark

A SparkSession is your entry point into all Spark functionality, similar to opening a connection before you can start working.

Cell 3: Create your first RDD

python data = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10] 
rdd = spark.sparkContext.parallelize(data) print(rdd)

parallelize() takes a normal Python list and distributes it into an RDD across partitions.

Cell 4: Apply a transformation (lazy, nothing runs yet)

python squared_rdd = rdd.map(lambda x: x * x) print("Transformation defined, but not executed yet")

Cell 5: Trigger an action (this actually runs the computation)

python result = squared_rdd.collect() print(result)

Expected output: [1, 4, 9, 16, 25, 36, 49, 64, 81, 100]

Cell 6: Filter transformation example

python even_rdd = rdd.filter(lambda x: x % 2 == 0) print(even_rdd.collect())

Expected output: [2, 4, 6, 8, 10]

Cell 7: Create a DataFrame from sample data

python data = [("Alice", 25, "HR"), ("Bob", 30, "IT"), ("Charlie", 28, "IT"), ("Diana", 35, "Finance")] columns = ["Name", "Age", "Department"]

df = spark.createDataFrame(data, columns) df.show()

Cell 8: Select and filter using DataFrame syntax

python df.select("Name", "Department").filter(df.Age > 27).show()

Notice this reads almost like SQL. This readability is a big reason DataFrames are preferred over raw RDDs today.

Cell 9: GroupBy and aggregation

python df.groupBy("Department").count().show()

Cell 10: Running actual SQL queries on a DataFrame

python df.createOrReplaceTempView("employees") spark.sql("SELECT Department, AVG(Age) as avg_age FROM employees GROUP BY Department").show()

Spark lets you write plain SQL directly against DataFrames once you register them as a temporary view, extremely useful if you're already comfortable with SQL.

Cell 11: Stop the Spark session when done

python spark.stop()

Key Takeaways

Spark processes data in memory, making it significantly faster than Hadoop's disk based MapReduce for most workloads RDD is the low level, fault tolerant, distributed data structure at Spark's core DataFrame is the higher level, structured, optimized way to work with data, similar to a SQL table or Pandas DataFrame Transformations (map, filter, select, groupBy) are lazy and only build an execution plan Actions (collect, count, show, take) trigger the actual computation Laziness allows Spark to optimize the entire chain of operations before running anything You can query DataFrames directly using SQL after registering them as a temporary view Interview Questions to Look Out For

  1. What is Apache Spark and why is it faster than Hadoop MapReduce? Apache Spark is a distributed data processing engine that performs computations in memory rather than writing intermediate results to disk at every step, the way MapReduce does. This in memory processing is the main reason Spark can be significantly faster for iterative and multi step workloads.

  2. What does RDD stand for and what do each of its parts mean? RDD stands for Resilient Distributed Dataset. Resilient means it can recover from node failures automatically, Distributed means the data is spread across multiple machines, and Dataset simply refers to the collection of data being processed.

  3. What is the difference between a transformation and an action in Spark? A transformation defines an operation on data but does not execute immediately, it only builds up a logical execution plan. An action triggers the actual execution of all the transformations that were chained before it and returns a result.

  4. Why does Spark use lazy evaluation? Lazy evaluation allows Spark to analyze the full sequence of transformations before executing anything, so it can optimize the overall plan, combine steps where possible, and avoid unnecessary computation, rather than executing each transformation immediately and wastefully.

  5. What is the difference between an RDD and a DataFrame? An RDD is a low level, unstructured, distributed collection of objects that gives fine grained control but requires more manual code. A DataFrame is a higher level, structured collection organized into named columns, similar to a database table, and benefits from automatic query optimization through Spark's Catalyst engine.

  6. Give examples of transformations and actions in Spark. Common transformations include map, filter, select, groupBy, and join. Common actions include collect, count, show, take, and reduce.

  7. What is a SparkSession and why is it needed? A SparkSession is the unified entry point for interacting with Spark functionality, including creating DataFrames, running SQL queries, and configuring settings. It replaced the older SparkContext and SQLContext as the single starting point in modern Spark applications.

  8. Can you run SQL queries directly on a Spark DataFrame? How? Yes, by registering the DataFrame as a temporary view using createOrReplaceTempView, and then using spark.sql to run standard SQL queries against that view, with results returned as a new DataFrame.

  9. What happens if you call a transformation like map or filter without ever calling an action afterward? Nothing actually executes. Since transformations are lazy, Spark only builds up the logical plan internally and nothing is computed or returned until an action like collect or show is called.

  10. Why might you choose DataFrames over RDDs in most modern Spark projects? DataFrames offer a higher level, more readable API similar to SQL or Pandas, automatic performance optimization through Spark's Catalyst query optimizer, and better memory management through Tungsten's binary format, whereas raw RDDs require manual optimization and more verbose code for the same tasks.

Conclusion

Today we moved from Hadoop's disk based processing model into Spark's faster, in memory approach, understanding RDDs as Spark's foundational data structure and DataFrames as the modern, structured way most real world Spark work actually happens. We also saw the critical distinction between lazy transformations and the actions that trigger them, and got hands on with PySpark directly in Jupyter Notebook, running everything from basic RDD operations to SQL queries on DataFrames.

Next up: Day 25 – Spark MLlib for Machine Learning