Statistical Learning and Inference


Overview

This advanced course develops an integrated understanding of statistical learning and statistical inference, with emphasis on the distinction between predictive performance, explanatory analysis, and causal or population-level claims. Students examine regression, classification, model selection, regularization, resampling, uncertainty quantification, tree-based methods, ensemble learning, support vector methods, dimensionality reduction, clustering, and model interpretability.

The course addresses likelihood-based and Bayesian inference, causal estimands, confounding, study design, propensity methods, missing data, multiple comparisons, statistical power, calibration, fairness, and reproducibility. Students use R or Python to implement, tune, validate, and interpret models; diagnose assumptions; quantify uncertainty; and evaluate the effects of leakage, overfitting, sampling bias, measurement error, and distribution shift.

Through problem sets, coding laboratories, model critiques, simulations, and an applied analysis project, students develop the capacity to select methods appropriate to data structure and research aims, assess robustness and limitations, and communicate evidence responsibly in reproducible technical reports and presentations.

Learning Outcomes

  • Formulate statistical, predictive, explanatory, and causal questions appropriate to a specified research or decision-making context.
  • Evaluate data structures, sampling processes, measurement characteristics, and research designs to identify relevant sources of bias and uncertainty.
  • Select and justify statistical learning and inferential methods in relation to research aims, assumptions, and data limitations.
  • Fit, tune, validate, and compare regression, classification, ensemble, kernel-based, and unsupervised learning models using R or Python.
  • Diagnose overfitting, leakage, misspecification, distribution shift, calibration problems, and violations of model assumptions.
  • Quantify uncertainty using resampling, likelihood-based, Bayesian, and simulation-based methods where appropriate.
  • Interpret model parameters, predictions, variable importance measures, and inferential results without conflating association, prediction, and causation.
  • Evaluate confounding, missingness, multiple comparisons, statistical power, and propensity-based methods in observational and experimental analyses.
  • Assess model robustness, fairness, reproducibility, and ethical implications across relevant populations and deployment contexts.
  • Synthesize a reproducible applied analysis that distinguishes empirical evidence from speculation and communicates conclusions to technical and non-technical audiences.

Timetable

TypeLengthFrequencyPeriod
Lecture2 hoursWeeklyAll semester
Lab2 hoursWeeklyAll semester
Tutorial1 hourFortnightlyAll semester

Assessment Schedule

TypeDescriptionWeighting
AssignmentProblem Sets (5 × 6%)30.00%
DeliverableCoding Laboratories (4 × 5%)20.00%
AssignmentModel Critique10.00%
TestMid-Semester Test15.00%
CapstoneApplied Analysis Project25.00%

Teaching Staff & Programs

This course is delivered jointly by faculty from the participating programs listed below. In line with the Douchewater Way, the University of Sexology tailors core instruction directly to each cohort's specific discipline — adapting curriculum to program needs rather than forcing students into a one-size-fits-all model. Learn more about our approach at The Douchewater Way.