Hands-On Data Engineering

Modern analytics and artificial intelligence rely on robust data engineering foundations. From data collection to analytics-ready outputs, well-designed data platforms are essential for transforming raw data into actionable insights in real-world systems.

Course description

This course offers a practical introduction to modern data engineering by guiding you through the end-to-end development of a realistic data platform. Participants will build a complete system in Docker, starting with batch ETL pipelines using PostgreSQL, dbt, and Airflow, and extending it with real-time streaming using Kafka and Apache Flink.

During the first week, participants focus on designing and implementing batch data pipelines and a simple analytics data warehouse. In the second week, the platform is enhanced with real-time capabilities, introducing streaming data ingestion and processing.

The course is entirely hands-on. Each day combines short theoretical sessions with guided labs and a mini project. By the end of the course, participants will have implemented both batch and streaming pipelines, understood the trade-offs between different architectures, and gained the ability to discuss how modern data platforms support analytics and machine-learning use cases.

General course information

Course dates27 July - 7 August, 2026 (A two-week course, 10 study days)
Course fee800 EUR
Course formatSummer course
Study fieldComputer Science, Data Science, Information Technology
LanguageEnglish
Study groupmaster's and PhD students
Assessment / ECTSPass/Fail (3 ECTS)
Location

Tartu

University of Tartu Delta Centre,
Narva mnt 18

Course lecturers

Course lecturerDescription
Kristo Raun

Lecturer of Data Engineering at the University of Tartu. His research focuses on streaming conformance checking and data engineering, with interests in data management systems, business processes, and provenance. He has more than 10 years of industry experience in data and analytics, primarily in consulting, and regularly works with modern data platforms and engineering practices. He teaches data engineering courses and supervises BSc and MSc students.

Riccardo Tommasini

Visiting Professor at Tartu University and Associate Professor at INSA Lyon (France) . His research interests cover data management and engineering with a focus on stream processing and graph analysis. At UT, he is part of the Data System group within the Data Science chair, and contributes to several courses in the area. He has more than 10 years experience in research and education, has published and presented in several top-tier data management conferences including: VLDB, SIGMOD, ICDE, ISWC, and EDBT.

Study information about the course

Participants are expected to have:

  • basic SQL knowledge (select, join, group by);
  • basic Python programming skills (functions, simple scripts);
  • general understanding of databases or data analysis
  • familiarity with Linux command line, Git and Docker is beneficial but not strictly required; key commands will be introduced.

NB! This is a preliminary programme. The final schedule will be sent to the participants two weeks before the course starts.

Week 1: Foundations and Batch Data Pipelines

Day 1: Monday, 27 July

Introduction, course setup, Docker basics, running services in containers

Day 2: Tuesday, 28 July

Postgres as analytical storage, basic data modelling, loading data

Day 3: Wednesday, 29 July

dbt fundamentals, modeling layers (staging / intermediate / marts)

Day 4: Thursday, 30 July

Airflow basics, DAGs, scheduling and monitoring batch pipelines

Day 5: Friday, 31 July

Mini-project I – design and implement a small batch data pipeline (teamwork)

Saturday, 1 August: free day
Sunday, 2 August: free day

Week 2 – Streaming and Real-Time Data

Day 6: Monday, 3 August

Introduction to event-driven architectures, Kafka concepts and setup

Day 7: Tuesday, 4 August

Producing and consuming streams, integrating Kafka with the existing stack

Day 8: Wednesday, 5 August

Flink basics, simple stream processing jobs and windowed aggregations

Day 9: Thursday, 6 August

Mini-project II – build a small end-to-end streaming pipeline (teamwork)

Day 10: Friday, 7 August

Project polishing, presentations, reflection and discussion of real-world use cases

1. Daily hands-on exercises

  • Short lab tasks for Docker, Postgres, dbt, Airflow, Kafka and Flink.
  • Checked mainly for completion and basic correctness.

2. Mini-project I (Batch pipeline)

  • Team assignment: design a small warehouse schema and implement a batch pipeline from raw data to analytics-ready tables using dbt and Airflow.
  • Short demo of architecture and design decisions.

3. Mini project II (Streaming pipeline)

  • Team assignment: design and implement a simple streaming use case with Kafka and Flink, optionally integrating it with the batch warehouse.
  • Short demo and 5–10-minute presentation on the last day

4. Individual reflection (short report, 1–2 pages)
What was implemented, main lessons learned, and how these tools differ from what participants have used before (if anything).

After successful completion of the course, participants will be able to:

  • explain the role of data engineering in analytics and AI projects;
  • run and manage a small data platform using Docker;
  • design and implement simple batch ETL pipelines using Postgres, dbt and Airflow;
  • describe the main principles of event-driven and streaming architectures;
  • build basic Kafka and Flink pipelines for real-time data ingestion and processing;
  • compare batch and streaming approaches and choose appropriate patterns for typical use cases;
  • work effectively in small teams on an end-to-end data engineering mini-project and present the results.

Course registration info

  • Application period: 20 March – 20 April
  • Notification of acceptance: accepted participants will be informed after the application period, by 30 April at the latest
  • Deadline for paying the course fee: 31 May
  • Confirmation of courses taking place: 5 June
  • UniTartu Summer School in Tartu: 27 July – 7 August

Only fully completed applications, including all required annexes, received by the deadline (20 April) will be considered for selection.

Applicants must submit the following:

  • Online application form
    (application period: 20 March–20 April 2026)
  • Motivation letter (maximum 1 page), explaining:
    • your motivation to participate,
    • your expectations for the programme,
    • how the summer course relates to your studies and interests,
    • how you plan to use the knowledge and experience gained in the future
  • Transcript of academic records
  • Copy of your passport
  • Proof of the application fee payments (25 EUR)

The participants of the UniTartu Summer School courses are required to pay:

  • The application fee of 25 EUR must be paid by the application deadline (20 April) at the latest.
    The application fee is non-refundable.
  • The course fee is 800 EUR.
    Includes: Study materials, academic work with lecturers, Certificate of completion, cultural events in the evenings
    Not included: meals, transportation and accommodation

Please note that the course fee is payable only after you have been accepted into the course. Once accepted, you will receive a confirmation of acceptance together with an invoice. The course fee can only be paid based on the invoice issued to you.

By paying the application fee, course fee and cultural events fee, you accept the terms and conditions information document. You are required to tick the box in the credit card payment form to confirm you have read and agree to terms and conditions. If you choose to pay by bank transfer, you will be informed of the same conditions.

Please note that by paying the fees, you are considered to have accepted the Terms and Conditions.

Cultural programme

Arrival and housing

Visit us virtually

Future study options