Data Science

Data Engineering with Spark

Master data engineering with Apache Spark through Vidya LLC's hands-on training course. Learn to build scalable, fault-tolerant data pipelines, process big data, and leverage Spark's powerful features for real-world data engineering challenges.

Description

Apache Spark gets your analytics developed fast and running fast, but large-scale distributed computing is hard. You cannot set a breakpoint on code spread across a cluster, so monitoring and optimization matter, and you have to rethink architecture, security, and engineering practices like testing and DevSecOps.

Data Engineering with Spark teaches you just enough Scala to build powerful pipelines into your existing architecture, and to give those pipelines the same testing, continuous integration, containerization, and security as the rest of your enterprise.

What makes this course different

Spark is a huge topic, and the typical course crams in too much too fast. You leave overwhelmed, and you never touch the architectural patterns and software engineering that separate a garage experiment from production architecture.

We take a practical approach. Rich code examples give you deep insight into the API, and the exercises use real datasets from Data.gov and encourage you to collaborate with your peers, the Spark Scaladoc, generative AI, and other sources just as you would at work. We also help you decide where to configure Spark yourself and where to hand configuration to a provider so you can focus on what matters.

Agile development and DevSecOps build quality in through automation, testing, and continuous delivery, and a distributed environment makes them even more valuable. Data Engineering with Spark shows you how to apply these techniques to improve the quality and reliability of your analytics.

Collaborating software engineers
What You'll Learn

Course Syllabus

About 8 hours total across two lessons of roughly 4 hours each. Run them on separate days or pair them into one intensive day.

1

Lesson 1: Mastering the Spark API

4 hours

Learn just enough Scala to write your own Spark jobs and navigate the ecosystem with confidence.

  • MapReduce: The Phantom Menace
  • Advantages of Spark
  • Just Enough Scala
  • Using the Spark Shell
  • Writing You Own Spark Jobs
  • The Spark Ecosystem
2

Lesson 2: Professional Spark

4 hours

Test, optimize, secure, and deploy Spark on Docker and Kubernetes like a professional.

  • Just Enough Hadoop
  • Testing Your Spark Jobs
  • Optimizing Spark and When to Stop Trying
  • Spark on Docker
  • Deploying Spark to Kubernetes
  • Spark Security
  • Visualizing Your Spark Jobs
Meet Your Guide

Your Instructor

Course instructor Neil Chaudhuri

"I have built several Scala and Spark applications currently in production and I worked with the original Spark team, AMPLab at UC Berkeley, on a research project for DARPA known as XDATA. Somehow I have helped enough developers around the world to earn Spark and Scala badges on Stack Overflow. I am passionate about Spark and look forward to helping you harness its power."

Neil Chaudhuri
Course Instructor

Want to transform your business?Get in touch today!

Partner with Vidya to modernize your technology, empower your teams, and accelerate your mission.