Data Engineering Zoomcamp
Free course by @DataTalksClub: https://github.com/DataTalksClub/data-engineering-zoomcamp/
🐳 Module 1 done! (2026-05-02)
Docker containers
Postgres & SQL
Terraform & GCP
NYC taxi data pipeline
🚀 Module 2 done! (2026-05-21)
@kestra_io workflow orchestration
ETL pipelines for taxi data
Backfill & scheduling
Variables & dynamic flows
📊 Module 3 done! (2026-06-02)
BigQuery & GCS
External vs materialized tables
Partitioning & clustering
Query optimization
📈 Module 4 done! (2026-06-12)
Analytics Engineering with dbt
Transformation models & tests
Data lineage & dependencies
NYC taxi revenue analysis
📊 Module 5 done! (2026-07-08)
Data Platforms with Bruin
End-to-end ELT pipelines
Data quality & lineage
Deployment to BigQuery
⚡ Module 6 done! (2026-07-17)
Batch processing with Spark 🔥
PySpark & DataFrames
Parquet file optimization
Spark UI on port 4040
Spark Cluster không có mạng nên không cài được Docker và bị lặp => Mất 1 tiếng để tìm lỗi và cài Cloud NAT
Default Setting cấu hình thấp nên tràn RAM sập mà không báo lỗi => Mất 1 tiếng rưỡi mới tìm ra lỗi và tăng lên 16GB (yeah)
Module 7 of Data Engineering Zoomcamp done! (2026-07-27)
Kafka producers and consumers
PyFlink tumbling and session windows
Real-time taxi data analysis
Redpanda as Kafka replacement