Overview

In this intermediate course, you will learn to design, build, and optimize robust batch data pipelines on Google Cloud. Moving beyond fundamental data handling, you will explore large-scale data transformations and efficient workflow orchestration, essential for timely business intelligence and critical reporting.

Get hands-on practice using Dataflow for Apache Beam and Serverless for Apache Spark (Dataproc Serverless) for implementation, and tackle crucial considerations for data quality, monitoring, and alerting to ensure pipeline

reliability and operational excellence. A basic knowledge of data warehousing, ETL/ELT, SQL, Python, and Google Cloud concepts is recommended.

Read more +

Prerequisites

Participants should have:

  • Basic proficiency with Data Warehousing and ETL/ELT concepts
  • Basic proficiency in SQL
  • Basic programming knowledge (Python recommended)
  • Familiarity with gcloud CLI and the Google Cloud console
  • Familiarity with core Google Cloud concepts and services

Target audience

This course is designed for:

  • Data Engineers
  • Data Analysts
Read more +

Delegates will learn how to

By the end of this course, learners will be able to:

  • Determine whether batch data pipelines are the correct choice for your business use case.
  • Design and build scalable batch data pipelines for high-volume ingestion and transformation.
  • Implement data quality controls within batch pipelines to ensure data integrity.
  • Orchestrate, manage, and monitor batch data pipeline workflows, implementing error handling and observability using logging and monitoring tools.
Read more +

Outline

Module 1 When to choose batch data pipelines

Topics

  • You will learn the critical role of a data engineer in developing and maintaining batch data pipelines, understand their core components and lifecycle, and analyze common challenges in batch data processing. You'll also identify key Google Cloud services that address these challenges.

Objectives

  • Explain the critical role of a data engineer in developing and maintaining batch data pipelines.
  • Describe the core components and typical lifecycle of batch data pipelines from ingestion to downstream consumption.
  • Analyze common challenges in batch data processing, such as data volume, quality, complexity, and reliability, and identify key Google Cloud services that can address them.

Module 2 Design and build batch data pipelines

Topics

  • You will design scalable batch data pipelines for high-volume data ingestion and transformation. You'll also optimize batch jobs for high throughput and cost-efficiency using various resource management and performance tuning techniques.

Objectives

  • Design scalable batch data pipelines for high-volume data ingestion and transformation.
  • Optimize batch jobs for high throughput and cost-efficiency using various resource management and performance tuning techniques.

Module 3 Control data quality in batch data pipelines

Topics

  • You will develop data validation rules and cleansing logic to ensure data quality within batch pipelines. You'll also implement strategies for managing schema evolution and performing data deduplication in large datasets.

Objectives

  • Develop data validation rules and cleansing logic to ensure data quality within batch pipelines.
  • Implement strategies for managing schema evolution and performing data deduplication in large datasets.

Module 4 Orchestrate and monitor batch data pipelines

Topics

  • You will orchestrate complex batch data pipeline workflows for efficient scheduling and lineage tracking. You'll also implement robust error handling, monitoring, and observability for batch data pipelines.

Objectives

  • Orchestrate complex batch data pipeline workflows for efficient scheduling and lineage tracking
  • Implement robust error handling, monitoring, and observability for batch data pipelines

Exams and assessments

There is no specific certification related to this course.

Hands-on learning

There are four practical labs in this course.

Read more +

Why choose QA

Need to know

Frequently asked questions

How can I create an account on myQA.com?

There are a number of ways to create an account. If you are a self-funder, simply select the "Create account" option on the login page.

If you have been booked onto a course by your company, you will receive a confirmation email. From this email, select "Sign into myQA" and you will be taken to the "Create account" page. Complete all of the details and select "Create account".

If you have the booking number you can also go here and select the "I have a booking number" option. Enter the booking reference and your surname. If the details match, you will be taken to the "Create account" page from where you can enter your details and confirm your account.

Find more answers to frequently asked questions in our FAQs: Bookings & Cancellations page.

How do QA’s virtual classroom courses work?

Our virtual classroom courses allow you to access award-winning classroom training, without leaving your home or office. Our learning professionals are specially trained on how to interact with remote attendees and our remote labs ensure all participants can take part in hands-on exercises wherever they are.

We use the WebEx video conferencing platform by Cisco. Before you book, check that you meet the WebEx system requirements and run a test meeting to ensure the software is compatible with your firewall settings. If it doesn’t work, try adjusting your settings or contact your IT department about permitting the website.

How do QA’s online courses work?

QA online courses, also commonly known as distance learning courses or elearning courses, take the form of interactive software designed for individual learning, but you will also have access to full support from our subject-matter experts for the duration of your course. When you book a QA online learning course you will receive immediate access to it through our e-learning platform and you can start to learn straight away, from any compatible device. Access to the online learning platform is valid for one year from the booking date.

All courses are built around case studies and presented in an engaging format, which includes storytelling elements, video, audio and humour. Every case study is supported by sample documents and a collection of Knowledge Nuggets that provide more in-depth detail on the wider processes.

When will I receive my joining instructions?

Joining instructions for QA courses are sent two weeks prior to the course start date, or immediately if the booking is confirmed within this timeframe. For course bookings made via QA but delivered by a third-party supplier, joining instructions are sent to attendees prior to the training course, but timescales vary depending on each supplier’s terms. Read more FAQs.

When will I receive my certificate?

Certificates of Achievement are issued at the end the course, either as a hard copy or via email. Read more here.

Let's talk

A member of the team will contact you within 4 working hours after submitting the form.

By submitting this form, you agree to QA processing your data in accordance with our Privacy Policy and Terms & Conditions. You can unsubscribe at any time by clicking the link in our emails or contacting us directly.