DATA ENGINEER · HIGH CODE PROGRAMME · BUILD ROBUST DATA PIPELINES →

High Code technical programme · 11 modules
Data Engineer
Build and maintain robust data pipelines
An intensive and highly code-focused programme, designed for current and future data engineers who want to master Python, SQL, data modelling, and CI/CD. Across eleven modules you will learn to develop, automate, and maintain reliable and scalable data pipelines, combining self-guided study with live sessions and hands-on practice on the EDUKATE.AI platform.
The programme
Practical and end-to-end data engineering
The itinerary covers the entire work of a data engineer: processing and modelling data, building databases and pipelines, automating with CI/CD, and deploying scalable data products in the cloud, whilst keeping quality, governance, and ethics in mind. Each module combines self-guided study with live sessions and assessed practical work.
11
Modules, including a hands-on hackathon.
~20 h
Commitment per module.
30
Maximum participants per cohort.
High Code
Python, SQL, Docker, Kubernetes, Terraform and Spark.
Competencies
What you will master
Python and data processing
Python, data manipulation and cleaning with Pandas, version control with Git and unit testing.
SQL and databases
SQL from basics to advanced, window functions, and the differences between SQL and NoSQL structures.
Pipelines and automation
ETL pipelines, batch and streaming processing, and automation with Apache Spark, Kafka and Luigi.
DevOps and containerisation
Software life cycle, CI/CD with GitHub Actions, Docker containers, Kubernetes, and Terraform.
Cloud and scalable architectures
Cloud data services (AWS, GCP and Azure), distributed file systems and continuity plans.
Quality, governance and ethics
The six dimensions of quality (DAMA), encryption, data ethics principles, and regulatory compliance.
Syllabus
The syllabus, module by module
Eleven modules that take you from code and databases to pipelines, the cloud, and a final hackathon. Each with its key technologies and its level of code.
1
Python and Pandas for data processing
Python
Pandas
Git and testing
Code level
2
Data engineering concepts
Life cycle
Schemes
Data warehousing
Code level
3
Databases, SQL and NoSQL
SQL
Window functions
NoSQL
Code level
4
Data product management
Roadmapping
Stakeholders
Product life cycle
Code level
5
DevOps and containerisation
CI/CD
Docker
Kubernetes and Terraform
Code level
6
Data quality, governance and ethics
Lady
Encryption
Compliance
Code level
7
Data pipelines and automation
ETL
Spark
Kafka and Luigi
Code level
8
Data product implementation
Requirements
Documentation
Architectures
Code level
9
Advanced data engineering techniques
AWS · GCP · Azure
Distributed systems
Business Continuity (BCP)
Code level
10
Emerging trends and technologies
Cloud-native
MLOps
DataOps
Code level
11
Hands-on hackathon
Team project
Real dataset
Presentations
Code level

Do you want to train your team as Data Engineers?
We adapt the itinerary to your organisation's starting level and objectives. Tell us about your context and we will prepare a proposal with the modules, format and schedule that best suit your team.