Data Science Systems and Deployment


Overview

This course examines the engineering practices required to transform data science notebooks and prototypes into reliable, maintainable production systems. Topics include system architecture, data ingestion, batch and streaming pipelines, storage formats, databases, data warehouses, feature stores, workflow orchestration, service interfaces, containerization, cloud infrastructure, and resource management.

Students apply software engineering principles to data science systems, including modular design, testing, type checking, dependency management, documentation, configuration, logging, security, and version control. Machine-learning operations topics include experiment tracking, data and model versioning, continuous integration and delivery, automated training and validation, model registries, deployment patterns, rollback strategies, and infrastructure as code.

The course addresses model serving through APIs and batch jobs, latency, throughput, scalability, reliability, cost, privacy, access control, and secure handling of sensitive data. Production monitoring covers data quality, drift, performance degradation, fairness, service health, alerting, retraining, and incident response. A substantial project requires deployment of a working system supported by architecture diagrams, operational documentation, and post-deployment evaluation.

Learning Outcomes

  • Design an end-to-end data science system that integrates ingestion, storage, processing, model execution, serving, monitoring, and operational controls.
  • Construct reproducible batch or streaming data pipelines using appropriate storage formats, databases, orchestration tools, and configuration practices.
  • Apply software engineering methods to data science systems, including modular design, automated testing, type checking, dependency management, documentation, logging, and version control.
  • Package and deploy analytical or machine-learning services through APIs, batch jobs, containers, and cloud infrastructure.
  • Automate experiment tracking, data and model versioning, validation, continuous integration, continuous delivery, model registration, and controlled release processes.
  • Evaluate deployment patterns using latency, throughput, scalability, reliability, cost, privacy, security, and access-control requirements.
  • Implement monitoring for data quality, distributional drift, model performance, fairness, service health, and operational alerting.
  • Diagnose pipeline, service, infrastructure, and model failures, and formulate appropriate rollback, retraining, and incident-response procedures.
  • Synthesize architecture diagrams, operational documentation, and post-deployment evaluations that communicate technical trade-offs to specialist and non-specialist stakeholders.

Timetable

TypeLengthFrequencyPeriod
Lecture2 hoursWeeklyAll semester
Lab2 hoursWeeklyAll semester
Tutorial1 hourFortnightlyAll semester
Workshop2 hoursFortnightlyAll semester

Assessment Schedule

TypeDescriptionWeighting
AssignmentSystems architecture and trade-off analysis15.00%
DeliverableAutomated testing, packaging, and deployment exercise15.00%
TestPractical systems implementation test15.00%
CapstoneProduction data science system, documentation, and evaluation35.00%
ExamFinal examination20.00%

Prerequisites

Teaching Staff & Programs

This course is delivered jointly by faculty from the participating programs listed below. In line with the Douchewater Way, the University of Sexology tailors core instruction directly to each cohort's specific discipline — adapting curriculum to program needs rather than forcing students into a one-size-fits-all model. Learn more about our approach at The Douchewater Way.