AI & Data - Introduction to Data Engineering on Google Cloud
A practical course focused on Introduction to Data Engineering on Google Cloud. Participants build skills across three priorities: Tasks and components of data engineering, Replication and migration of data, The Extract-Load pipeline model.
Detailed programme
Tasks and components of data engineering
- Understand the role of a data engineer.
- Understand the differences between a data source and a data sink.
- Discover the different types of data formats.
- Explain the storage solution options on Google Cloud.
- Discover metadata management options on Google Cloud.
- Learn how to easily share datasets using Analytics Hub.
- Find out how to load data in BigQuery using the Google Cloud console or the CLI gcloud.
Hands-on work
- Lab: data loading in BigQuery.
Replication and migration of data
- Discover Google Cloud's basic data replication and migration architecture.
- Understand options and cases of using the gcloud command line tool.
- Discover the features and usage cases of Storage Transfer Service.
- Discover the features and usage cases of Transfer Appliance.
- Discover the features and deployment of Datastream.
Hands-on work
- Lab: PostgreSQL replication to BigQuery.
The Extract-Load pipeline model
- Extract and load architecture.
- Understand the options of the command line tool bq.
- Discover the features and cases of use of BigQuery data transfer service.
- Discover the features and cases of use of BigLake as a load model without extraction.
Hands-on work
- Lab: Introduction to BigLake.
The Extract, Load, Transform pipeline model
- Explain the basic diagram of extraction, loading and processing architecture.
- Understand an ELT pipeline running on Google Cloud.
- Discover BigQuery's SQL programming and script features.
- Explain the functionality and usage cases of Dataform.
Hands-on work
- Lab: Create and run SQL workflow in Dataform.
The Extract, Transform, Load pipeline model
- Discover the basic extraction, transformation and loading architecture diagram.
- Discover the graphical user interface tools on Google Cloud used for ETL data pipelines.
- Explain batch data processing using Dataproc.
- Use Dataproc Serverless for Spark for ETL..
- Explore streaming data processing options.
- Understand Bigtable's role in data pipelines.
Hands-on work
- Lab: Use Dataproc Serverless for Spark to load BigQuery. Lab:
- Create a continuous data pipeline for a real-time dashboard with Dataflow.
Automation techniques
- Discover the automation models and options available for pipelines.
- Discover Cloud Scheduler and Workflows.
- Discover Cloud Composer.
- Discover the functions of Cloud Run.
- Discover the use of functionality and automation cases for Eventarc.
Hands-on work
- Lab: Use Cloud Run functions to load BigQuery.
Tell us about your project
Our offices
- Exceev Consulting
61 Rue de Lyon
75012, Paris, France - Exceev Technology
332 Bd Brahim Roudani
20330, Casablanca, Morocco