Skip to product information
1 of 1

Akash Tandon,Sandy Ryza,Uri Laserson,Sean Owen,Josh Wills

Advanced Analytics with PySpark: Patterns for Learning from Data at Scale Using Python and Spark

Advanced Analytics with PySpark: Patterns for Learning from Data at Scale Using Python and Spark

đź’Ž Earn 187 Points (ÂŁ1.87) on this item.

Low Stock: Only 2 copies remaining
Regular price ÂŁ37.51 GBP
Regular price ÂŁ52.99 GBP Sale price ÂŁ37.51 GBP
Sale Sold out
Taxes included. Shipping calculated at checkout.

YOU SAVE ÂŁ15.48

  • Condition: Brand new
  • UK Delivery times: Usually arrives within 2 - 3 working days
  • UK Shipping: Fee starts at ÂŁ3.89. Subject to product weight & dimension

Bulk ordering. Want 15 or more copies? Get a personalised quote and bigger discounts. Learn more about bulk orders.

  • More about Advanced Analytics with PySpark: Patterns for Learning from Data at Scale Using Python and Spark


Apache Spark is the de facto tool for analyzing big data, and this updated guide teaches you how to approach analytics problems using PySpark and other best practices. It covers common techniques such as classification, clustering, collaborative filtering, and anomaly detection in fields such as genomics, security, and finance.

Format: Paperback / softback
Length: 275 pages
Publication date: 24 June 2022
Publisher: O'Reilly Media


The sheer volume of data being generated today is truly astounding, and it continues to expand at an unprecedented rate. In this landscape, Apache Spark has emerged as the go-to tool for analyzing large datasets, playing a pivotal role in the field of data science. This comprehensive guide, updated for Spark 3.0, is designed to empower data scientists with the knowledge and skills necessary to tackle complex analytics problems using PySpark, Spark's Python API, and other best practices in Spark programming.

Led by a team of experienced data scientists, including Akash Tandon, Sandy Ryza, Uri Laserson, Sean Owen, and Josh Wills, this guide provides a comprehensive introduction to the Spark ecosystem. It delves into various patterns and techniques that apply to diverse fields such as genomics, security, and finance, utilizing common machine learning and statistical approaches.

Furthermore, this updated edition expands its coverage to include natural language processing (NLP) and image processing, making it an invaluable resource for data scientists working with text and visual data. Whether you have a solid foundation in machine learning and statistics or are new to the field, this book will guide you through the process of conducting large-scale data analysis.

By familiarizing yourself with Spark's programming model and ecosystem, you will gain a deep understanding of how to leverage its powerful capabilities. You will explore general approaches in data science, examine complete implementations that analyze large public datasets, and discover which machine learning tools are most suitable for specific problems. Additionally, you will explore code that can be adapted to a wide range of applications, enabling you to apply your knowledge to real-world scenarios.

In conclusion, if you are looking to unlock the full potential of big data and advance your data science skills, this updated guide to Apache Spark is an essential resource. With its comprehensive coverage, practical examples, and expert guidance, it will empower you to tackle complex analytics problems and make informed decisions based on data. So, whether you are a seasoned data scientist or just starting your journey, this book is your key to success in the world of big data analytics.

Weight: 410g
Dimension: 176 x 234 x 15 (mm)
ISBN-13: 9781098103651

UK and International shipping information

We deliver throughout the United Kingdom and to 128 countries and territories worldwide, including the United States, Australia, Canada, Germany, Spain and France.

View full UK and international delivery information.

View full details