$39.99
Serverless ETL and Analytics with AWS Glue
Design scalable data lakes, optimize ETL pipelines, and accelerate analytics on AWS
Use AWS Glue to integrate growing data sources with serverless ETL, building secure, observable pipelines that support reliable analytics while managing performance and cost across a governed AWS data platform as workloads grow
Key Features
Use runnable code, console walkthroughs, and downloadable examples for core AWS Glue workflows
Apply DataOps practices with AWS CDK, Docker, and CI/CD in real-world scenarios
Learn from six data specialists with AWS, Spark, Apache Iceberg, and data lake expertise
Book Description
Whether you build data pipelines, design cloud architectures, or deliver analytics on AWS, bringing data together is only part of the challenge. You must also keep this data clean, trustworthy, and available while controlling costs. AWS Glue offers serverless data integration, but using it effectively requires decisions about storage, metadata, security, orchestration, monitoring, and performance.
This book guides you from modern data management and core AWS Glue features through ingestion from files, streams, SaaS applications, and JDBC sources, preparation, storage layout, metadata, security, sharing, and pipeline operations. Console walkthroughs and runnable examples show how to manage schemas and lineage in AWS Glue Data Catalog, apply AWS Lake Formation access controls, monitor workloads, tune Spark jobs, troubleshoot failures, and manage development with AWS CDK, Docker, and CI/CD. You will also examine analytics, machine learning and generative AI integrations, real-world data lake scenarios, and cost optimization. Learn how Apache Iceberg, Apache Hudi, and Delta Lake add transactions, schema evolution, and efficient data management to data lakes.
By the end, you will be able to design, build, operate, and continuously improve a serverless data platform that fits your organization's scale, structure, and priorities.
What you will learn
Design scalable serverless ETL pipelines with AWS Glue
Ingest data from files, streams, SaaS, and JDBC sources
Optimize file formats, partitions, compression, and layouts
Manage schemas, partitions, and lineage in AWS Glue Data Catalog
Secure data with access control, encryption, and auditing
Automate testing and multi account CI/CD using AWS CDK and Docker
Monitor, tune, and troubleshoot AWS Glue and Spark workloads
Apply Apache Iceberg, Hudi, and Delta Lake to data lakes with AWS Glue
Who this book is for
This book is for data engineers, ETL developers, cloud architects, and analytics professionals who build or operate data platforms on AWS. It suits readers working on serverless data lakes, Spark ETL, governance, data sharing, reliability, or cost control. It is especially useful if you aim to improve pipeline reliability, governance, or cost visibility as workloads grow. Basic familiarity with the AWS Management Console, Amazon S3, and IAM is recommended. Experience with Python, SQL, or Apache Spark will help with the code examples, and an AWS account is useful for following the walkthroughs.