Doesn't suit? No problem! You can return items for up to 30 days
You won't go wrong with a gift voucher. The gift recipient can choose anything from our offer.
Up to 30 days for returns
Build distributed data systems for real-time analytics, large-scale processing, and production machine learning
Modern data systems operate at enormous scale.
Organizations process terabytes of logs, events, transactions, sensor streams, and machine learning workloads that must remain fast, fault tolerant, and continuously available.
Apache Spark has become one of the most important technologies for handling these large-scale distributed workloads.
"Spark in the Wild" is a practical, engineering-focused guide to building scalable data processing systems, streaming pipelines, and machine learning infrastructure using Spark and modern cloud-native tooling.
This book teaches engineers how to design reliable distributed systems that transform massive volumes of data into actionable intelligence.
Modern organizations face challenges such as:
Distributed data systems must balance scalability, reliability, and operational simplicity.
Throughout the book, you will learn how to:
Each chapter focuses on practical engineering workflows used in real-world data infrastructure teams.
These examples reflect real-world distributed data engineering challenges.
If you want to build scalable, fault-tolerant, and production-ready big data systems using Spark, this book provides the roadmap.
Process at scale.
Stream intelligently.
Engineer distributed data systems that last.