Google Cloud Professional Data Engineer (PDE) Study Guide: Pipelines and Analytics
Prepare for GCP-PDE with data system design, ingestion and processing, storage, analytics, and workload operations.
The Professional Data Engineer (PDE) exam tests how to design, build, and operate data platforms that securely and reliably support analytics and business decisions. Questions ask you to choose among batch and streaming designs, storage services, governance, performance, and cost.
GCP-PDE is the shorthand used on this site; the official name is Professional Data Engineer. This guide follows the official certification page and current exam guide. Confirm the version that applies to your exam date.
Exam domains and weights
| Official domain | Main topics | Weight |
|---|---|---|
| Designing data processing systems | Security, reliability, migration, architecture | ~22% |
| Ingesting and processing the data | Batch/stream pipelines, transformation, orchestration | ~25% |
| Storing the data | Storage selection, warehouses/lakes, data platforms | ~20% |
| Preparing and using data for analysis | BI, AI/ML preparation, data sharing | ~15% |
| Maintaining and automating data workloads | Optimization, capacity, automation, recovery | ~18% |
How to study the domains
Design data processing systems
Derive an architecture from security, compliance, residency, governance, reliability, and migration goals. Translate business needs into sources, consumers, freshness, and recovery objectives, then plan migration validation.
Ingest and process data
Compare Pub/Sub, Dataflow/Apache Beam, Dataproc/Spark, BigQuery, and Cloud Data Fusion. Distinguish batch from streaming and understand windows, late data, transformations, cleansing, and orchestration.
Store and analyze data
Choose among BigQuery, BigLake, Cloud Storage, Spanner, Bigtable, Cloud SQL, and Firestore based on access patterns. Review warehouse modeling, lake governance, partitioning/clustering, query performance, BI Engine, and preparing data for feature engineering or RAG.
Operate and automate
Know Cloud Composer, Workflows, CI/CD, BigQuery capacity/reservations, monitoring, quotas, recovery, and cost controls. Build repeatable, observable workflows that handle failures.
Common confusion: Dataflow provides managed batch/stream processing, while Dataproc suits Spark/Hadoop ecosystems. Pub/Sub transports messages; it does not perform the full transformation pipeline. Identify the processing model and operational needs first.
Study sequence and readiness check
Study in this order: architecture, pipelines, storage, analytics, operations. For each scenario, note volume, latency, format, access pattern, governance, recovery objectives, and budget. Build at least one batch and one streaming workflow.
- Choose batch or streaming based on freshness and data characteristics.
- Explain the roles of Dataflow, Dataproc, Pub/Sub, and BigQuery.
- Select storage based on access patterns and explain cost/performance trade-offs.
- Design governance, monitoring, scheduling, retries, and recovery.
Keep practicing
Use data scenarios to derive pipelines, storage, and analytics from sources and freshness requirements. Start GCP-PDE practice.