Advanced Public Health Data Science


Overview

This course examines the end-to-end application of data science in public health, encompassing analytic question formulation, data acquisition, governance, quality assessment, integration, modeling, communication, and translation into policy and practice. Students work with surveillance, registry, environmental, demographic, and other population health datasets, including messy, high-dimensional, and sensitive data.

Core methods include reproducible computational workflows, exploratory analysis, feature engineering, regression, classification, clustering, dimensionality reduction, machine learning, model validation, calibration, interpretability, uncertainty analysis, fairness assessment, bias evaluation, and responsible artificial intelligence. Applications include outbreak detection, risk prediction, population segmentation, health disparities, environmental exposure assessment, and resource allocation.

Students develop and evaluate public health data products using appropriate programming tools and governance practices. Emphasis is placed on privacy protection, methodological transparency, equity, limitations, and the communication of technically rigorous findings as actionable recommendations for public health stakeholders.

Learning Outcomes

  • Formulate defensible public health analytic questions and identify appropriate data sources, study populations, outcomes, exposures, and decision contexts.
  • Evaluate data provenance, quality, completeness, representativeness, linkage integrity, and governance requirements for public health datasets.
  • Construct reproducible computational workflows for data acquisition, cleaning, integration, documentation, analysis, and reporting.
  • Apply exploratory analysis and feature engineering techniques to messy, high-dimensional public health data.
  • Select, implement, and justify regression, classification, clustering, dimensionality reduction, and machine learning methods for defined public health objectives.
  • Validate predictive and inferential models using appropriate resampling, performance, calibration, uncertainty, and external validation procedures.
  • Assess model interpretability, fairness, bias, privacy risks, and potential harms in public health data science applications.
  • Synthesize technical findings, limitations, and equity implications into clear recommendations for public health stakeholders.
  • Design and communicate responsible uses of artificial intelligence that align with public health ethics, governance standards, and decision-making requirements.

Timetable

TypeLengthFrequencyPeriod
Lecture2 hoursWeeklyAll semester
Lab2 hoursWeeklyAll semester
Tutorial1 hourFortnightlyAll semester
Workshop2 hoursFortnightlySecond term

Assessment Schedule

TypeDescriptionWeighting
AssignmentAnalytic question and data governance brief15.00%
DeliverableReproducible data pipeline project25.00%
AssignmentModel development and validation report25.00%
TestPractical data science test15.00%
CapstoneFinal policy translation presentation20.00%

Prerequisites

  • Requirement Prior university-level programming experience in Python or R and foundational knowledge of epidemiology or biostatistics.

Teaching Staff & Programs

This course is delivered jointly by faculty from the participating programs listed below. In line with the Douchewater Way, the University of Sexology tailors core instruction directly to each cohort's specific discipline — adapting curriculum to program needs rather than forcing students into a one-size-fits-all model. Learn more about our approach at The Douchewater Way.