The course discusses various concepts of data engineering, such as big data, Apache Spark, data lakes, and technologies for storing and processing large datasets. Students gain hands-on experience in developing data processing pipelines in a cloud-based environment.
Course contents
- Principles of big data and big data processing platforms
- The Apache Spark programming model
- Databricks data platform
- Data formats, including Apache Parquet and Delta Lake
- Data storage architectures, including data lakes and data warehouses
- Strengths, limitations, and appropriate use cases of different data engineering solutions
Learning outcomes
After completing the course, the student
- knows about programming tools for managing and analysing big data
- understands the modern solutions for data-intensive programming, their use cases, and the purpose and limits of the solutions
- can use Apache Spark and Databricks to implement data-processing solutions
- can choose and apply suitable technologies introduced in the course to solve data-processing problems
Course material
Course material will be distributed through Moodle.
Teaching schedule
Lectures on Mondays 14:15-16:00 (streamed online with recordings available afterwards). In addition, possibly additional guest lectures. Exercise deadlines on every Monday.
Completion methods
Compulsory requirements:
- A programming assignment submitted at the end of the course. The assignment can be done either solo or in groups of 2-3.
- Electronic exam done at campus (at any available EXAM room) within the exam window (5.12.-20.12.). (Using Finnish to answer exam questions is allowed.)
Other graded work:
- Weekly exercises with smaller programming tasks.
More information in the Tampere University study guide.
You can get a digital badge after completing this course.