This comprehensive course teaches you how to design and implement a complete data engineering workflow using real-world COVID-19 data. The project is divided into three parts, guiding learners from raw data extraction to advanced data analysis.
In Part 1, you’ll learn how to collect and clean COVID-19 datasets, explore data quality issues, and structure the data for storage. Part 2 focuses on building scalable pipelines using Python and AWS services, including setting up databases, automating data ingestion, and preparing data for analytics. Part 3 teaches advanced analysis and visualization techniques in Jupyter Notebooks, helping you derive meaningful insights from the data.
Throughout the course, learners gain practical experience in building robust, end-to-end data engineering projects. By completing this project, you’ll be prepared to handle large-scale datasets, optimize workflows, and showcase a fully functional data engineering project to potential employers.