1. Data Engineering Fundamentals
Data engineering focuses on designing and maintaining systems that collect, process, transform, store and deliver data for applications, analytics, reporting and artificial intelligence. Data engineers create the infrastructure that allows organisations to turn raw information into reliable and usable data.
Key topics
- What Is Data Engineering?
- Data Engineering vs. Data Science vs. Data Analytics
- The Role of a Data Engineer
- Understanding the Data Engineering Lifecycle
- Structured, Semi-Structured and Unstructured Data
- Data Engineering Architecture
- Common Data Engineering Challenges
- Essential Data Engineering Skills
- Data Engineering Tools and Technologies
- How Data Engineers Support Business Intelligence
2. Databases & Data Storage
Databases and storage systems are fundamental components of data engineering. They provide the infrastructure needed to store, organise, retrieve and manage large volumes of information efficiently.
Key topics
- Introduction to Databases for Data Engineers
- SQL Fundamentals Every Data Engineer Should Know
- Relational vs. NoSQL Databases
- OLTP vs. OLAP
- Database Normalization and Denormalization
- Data Warehouses vs. Data Lakes
- What Is a Data Lakehouse?
- Columnar Databases
- Database Indexing
- Query Optimization
- Designing Scalable Data Storage Systems
3. Data Pipelines & ETL
Data pipelines move information from source systems into storage, processing and analytics environments. ETL and ELT processes allow organisations to extract raw data, transform it into useful formats and load it into appropriate destinations.
Key topics
- What Is a Data Pipeline?
- ETL vs. ELT
- Building Your First ETL Pipeline
- Extract, Transform and Load
- Batch Processing vs. Real-Time Processing
- Data Pipeline Architecture
- Data Pipeline Automation
- Handling Failures in Data Pipelines
- Data Pipeline Monitoring
- Data Pipeline Logging
- Best Practices for Reliable Data Pipelines
4. Big Data Engineering
Big data engineering focuses on designing systems capable of storing and processing datasets that are too large, fast or complex for traditional data-processing systems. Distributed computing allows workloads to be divided across multiple machines.
Key topics
- What Is Big Data?
- The Three Vs of Big Data
- Introduction to Apache Hadoop
- Apache Spark Explained
- Hadoop vs. Apache Spark
- Distributed Data Processing
- Handling Massive Datasets
- Data Partitioning
- Distributed Computing
- Fault Tolerance
- Scaling Data Processing Systems
5. Cloud Data Engineering
Cloud data engineering involves building and operating data platforms using cloud infrastructure and managed services. Cloud platforms allow organisations to scale storage and processing resources according to their workloads.
Key topics
- Introduction to Cloud Data Engineering
- AWS Data Engineering Services
- Microsoft Azure Data Engineering Services
- Google Cloud Data Engineering Services
- Cloud Data Warehouses
- Cloud Data Lakes
- Cloud Data Storage
- Building Data Pipelines in the Cloud
- Serverless Data Processing
- Cloud Data Engineering Architecture
- Scalable Cloud Data Platforms
6. Data Processing & Streaming
Data processing and streaming technologies enable organisations to work with information as it is generated. Real-time systems are particularly useful for applications that require immediate insights, monitoring, alerts and automated responses.
Key topics
- Batch Processing vs. Stream Processing
- Introduction to Real-Time Data Processing
- Apache Kafka Explained
- Event-Driven Data Architecture
- Real-Time Data Pipelines
- Data Streaming vs. Traditional Data Pipelines
- Building Real-Time Data Processing Systems
- Stream Processing with Apache Spark
- Handling High-Velocity Data
- Real-Time Analytics and Data Engineering