Postgraduate study

Postgraduate taught 

Advanced Statistics MSc

Multivariate Statistics and Machine Learning 1 (Level M) STATS5021

  • Academic Session: 2026-27
  • School: School of Mathematics and Statistics
  • Credits: 10
  • Level: Level 5 (SCQF level 11)
  • Typically Offered: Semester 1
  • Available to Visiting Students: Yes
  • Collaborative Online International Learning: No
  • Curriculum For Life: No

Short Description

The course provides an introduction to multivariate statistics and classical machine learning with an overview of standard clustering methods, basic linear and non-linear classification methods, multivariate regression, dimension reduction and data visualisation. Students will learn to critically assess the strengths of weaknesses of various modelling paradigms, including algorithmic versus probabilistic methods and linear versus nonlinear models.

Timetable

- 20 x 1 hour lectures

- 5 x 1 hour tutorials

- 5 x 1 hour labs

Excluded Courses

STATS4046 Multivariate Statistics and Machine Learning 1

Assessment

90-minute, end-of-course examination (85%)

Coursework (15%)

Main Assessment In: December

Course Aims

To provide an appreciation of the types of problems and questions which arise with multivariate data;
to provide a good understanding of the application of classical multivariate techniques for:

■ the graphical exploration of multivariate data

■ the reduction of dimensionality of multivariate data

■ analysis in unsupervised and supervised settings

■ classical clustering methods

■ linear and nonlinear classification methods

■ multivariate regression

■ bias-variance trade-off of the generalisation error

■ cross validation

■ regularisation

■ inference of latent variables with the Expectation Maximization (EM) algorithm

Intended Learning Outcomes of Course

By the end of this course students will be able to:

- display multivariate data in a variety of graphical ways and interpret such displays;

- apply and interpret methods of dimension reduction including principal component analysis, multidimensional scaling, the biplot, factor analysis, canonical variates;

- apply and interpret classical methods for cluster analysis and discrimination;

- use formal criteria for model selection in prediction and model fitting;

- interpret the output of statistical programming tools for multivariate statistics;

- understand the conceptual difference between algorithmic and probabilistic methods;

- understand the pros and cons of linear versus non-linear methods;

- investigate further into one topic related to the course and use these concepts to solve a real-world problem;

■ - understand the concept of and infer latent variables.

Minimum Requirement for Award of Credits

No exceptions