Overview
This course examines the engineering practices required to transform data science notebooks and prototypes into reliable, maintainable production systems. Topics include system architecture, data ingestion, batch and streaming pipelines, storage formats, databases, data warehouses, feature stores, workflow orchestration, service interfaces, containerization, cloud infrastructure, and resource management.
Students apply software engineering principles to data science systems, including modular design, testing, type checking, dependency management, documentation, configuration, logging, security, and version control. Machine-learning operations topics include experiment tracking, data and model versioning, continuous integration and delivery, automated training and validation, model registries, deployment patterns, rollback strategies, and infrastructure as code.
The course addresses model serving through APIs and batch jobs, latency, throughput, scalability, reliability, cost, privacy, access control, and secure handling of sensitive data. Production monitoring covers data quality, drift, performance degradation, fairness, service health, alerting, retraining, and incident response. A substantial project requires deployment of a working system supported by architecture diagrams, operational documentation, and post-deployment evaluation.
Learning Outcomes
- Design an end-to-end data science system that integrates ingestion, storage, processing, model execution, serving, monitoring, and operational controls.
- Construct reproducible batch or streaming data pipelines using appropriate storage formats, databases, orchestration tools, and configuration practices.
- Apply software engineering methods to data science systems, including modular design, automated testing, type checking, dependency management, documentation, logging, and version control.
- Package and deploy analytical or machine-learning services through APIs, batch jobs, containers, and cloud infrastructure.
- Automate experiment tracking, data and model versioning, validation, continuous integration, continuous delivery, model registration, and controlled release processes.
- Evaluate deployment patterns using latency, throughput, scalability, reliability, cost, privacy, security, and access-control requirements.
- Implement monitoring for data quality, distributional drift, model performance, fairness, service health, and operational alerting.
- Diagnose pipeline, service, infrastructure, and model failures, and formulate appropriate rollback, retraining, and incident-response procedures.
- Synthesize architecture diagrams, operational documentation, and post-deployment evaluations that communicate technical trade-offs to specialist and non-specialist stakeholders.
Timetable
| Type | Length | Frequency | Period |
|---|---|---|---|
| Lecture | 2 hours | Weekly | All semester |
| Lab | 2 hours | Weekly | All semester |
| Tutorial | 1 hour | Fortnightly | All semester |
| Workshop | 2 hours | Fortnightly | All semester |
Assessment Schedule
| Type | Description | Weighting |
|---|---|---|
| Assignment | Systems architecture and trade-off analysis | 15.00% |
| Deliverable | Automated testing, packaging, and deployment exercise | 15.00% |
| Test | Practical systems implementation test | 15.00% |
| Capstone | Production data science system, documentation, and evaluation | 35.00% |
| Exam | Final examination | 20.00% |
Prerequisites
Teaching Staff & Programs
This course is delivered jointly by faculty from the participating programs listed below. In line with the Douchewater Way, the University of Sexology tailors core instruction directly to each cohort's specific discipline — adapting curriculum to program needs rather than forcing students into a one-size-fits-all model. Learn more about our approach at The Douchewater Way.
