DATA ENGINEER · HIGH CODE PROGRAMME · BUILD ROBUST DATA PIPELINES →

High Code technical programme · 11 modules

Data Engineer

Build and maintain robust data pipelines

An intensive and highly code-focused programme, designed for current and future data engineers who want to master Python, SQL, data modelling, and CI/CD. Across eleven modules you will learn to develop, automate, and maintain reliable and scalable data pipelines, combining self-guided study with live sessions and hands-on practice on the EDUKATE.AI platform.

The programme

Practical and end-to-end data engineering

The itinerary covers the entire work of a data engineer: processing and modelling data, building databases and pipelines, automating with CI/CD, and deploying scalable data products in the cloud, whilst keeping quality, governance, and ethics in mind. Each module combines self-guided study with live sessions and assessed practical work.

11

Modules, including a hands-on hackathon.

~20 h

Commitment per module.

30

Maximum participants per cohort.

High Code

Python, SQL, Docker, Kubernetes, Terraform and Spark.

Competencies

What you will master

Python and data processing

Python, data manipulation and cleaning with Pandas, version control with Git and unit testing.

SQL and databases

SQL from basics to advanced, window functions, and the differences between SQL and NoSQL structures.

Pipelines and automation

ETL pipelines, batch and streaming processing, and automation with Apache Spark, Kafka and Luigi.

DevOps and containerisation

Software life cycle, CI/CD with GitHub Actions, Docker containers, Kubernetes, and Terraform.

Cloud and scalable architectures

Cloud data services (AWS, GCP and Azure), distributed file systems and continuity plans.

Quality, governance and ethics

The six dimensions of quality (DAMA), encryption, data ethics principles, and regulatory compliance.

Syllabus

The syllabus, module by module

Eleven modules that take you from code and databases to pipelines, the cloud, and a final hackathon. Each with its key technologies and its level of code.

1
Python and Pandas for data processing

Python

Pandas

Git and testing

Code level

2
Data engineering concepts

Life cycle

Schemes

Data warehousing

Code level

3
Databases, SQL and NoSQL

SQL

Window functions

NoSQL

Code level

4
Data product management

Roadmapping

Stakeholders

Product life cycle

Code level

5
DevOps and containerisation

CI/CD

Docker

Kubernetes and Terraform

Code level

6
Data quality, governance and ethics

Lady

Encryption

Compliance

Code level

7
Data pipelines and automation

ETL

Spark

Kafka and Luigi

Code level

8
Data product implementation

Requirements

Documentation

Architectures

Code level

9
Advanced data engineering techniques

AWS · GCP · Azure

Distributed systems

Business Continuity (BCP)

Code level

10
Emerging trends and technologies

Cloud-native

MLOps

DataOps

Code level

11
Hands-on hackathon

Team project

Real dataset

Presentations

Code level

Do you want to train your team as Data Engineers?

We adapt the itinerary to your organisation's starting level and objectives. Tell us about your context and we will prepare a proposal with the modules, format and schedule that best suit your team.