Google Cloud Professional Data Engineer (PDE) Study Guide: Pipelines and Analytics

Prepare for GCP-PDE with data system design, ingestion and processing, storage, analytics, and workload operations.

The Professional Data Engineer (PDE) exam tests how to design, build, and operate data platforms that securely and reliably support analytics and business decisions. Questions ask you to choose among batch and streaming designs, storage services, governance, performance, and cost.

GCP-PDE is the shorthand used on this site; the official name is Professional Data Engineer. This guide follows the official certification page and current exam guide. Confirm the version that applies to your exam date.

Exam domains and weights

Official domain Main topics Weight
Designing data processing systems Security, reliability, migration, architecture ~22%
Ingesting and processing the data Batch/stream pipelines, transformation, orchestration ~25%
Storing the data Storage selection, warehouses/lakes, data platforms ~20%
Preparing and using data for analysis BI, AI/ML preparation, data sharing ~15%
Maintaining and automating data workloads Optimization, capacity, automation, recovery ~18%

How to study the domains

Design data processing systems

Derive an architecture from security, compliance, residency, governance, reliability, and migration goals. Translate business needs into sources, consumers, freshness, and recovery objectives, then plan migration validation.

Ingest and process data

Compare Pub/Sub, Dataflow/Apache Beam, Dataproc/Spark, BigQuery, and Cloud Data Fusion. Distinguish batch from streaming and understand windows, late data, transformations, cleansing, and orchestration.

Store and analyze data

Choose among BigQuery, BigLake, Cloud Storage, Spanner, Bigtable, Cloud SQL, and Firestore based on access patterns. Review warehouse modeling, lake governance, partitioning/clustering, query performance, BI Engine, and preparing data for feature engineering or RAG.

Operate and automate

Know Cloud Composer, Workflows, CI/CD, BigQuery capacity/reservations, monitoring, quotas, recovery, and cost controls. Build repeatable, observable workflows that handle failures.

Common confusion: Dataflow provides managed batch/stream processing, while Dataproc suits Spark/Hadoop ecosystems. Pub/Sub transports messages; it does not perform the full transformation pipeline. Identify the processing model and operational needs first.

Study sequence and readiness check

Study in this order: architecture, pipelines, storage, analytics, operations. For each scenario, note volume, latency, format, access pattern, governance, recovery objectives, and budget. Build at least one batch and one streaming workflow.

  • Choose batch or streaming based on freshness and data characteristics.
  • Explain the roles of Dataflow, Dataproc, Pub/Sub, and BigQuery.
  • Select storage based on access patterns and explain cost/performance trade-offs.
  • Design governance, monitoring, scheduling, retries, and recovery.

Keep practicing

Use data scenarios to derive pipelines, storage, and analytics from sources and freshness requirements. Start GCP-PDE practice.

WeChat mini program

IT知习 mini program QR code

Search WeChat for: IT知习