Skip to content
STAT411Undergraduate

Statistical Data Mining

Printed in the catalogue as STATISTICAL DATA MINING

Course content

Descriptive and predictive mining. Data preprocessing: cleaning transformation. outlier detection, missing data imputation. Dimension reduction, Principal Component Analysis (PCA). Sampling, oversampling. Exploratory data analysis (EDA). Clustering methods: partitioning, hierarchical, density-based, model-based. Predictive modeling. Regression. Variable selection. Robust and nonlinear regression. Nonparametric regression. Classifiers. Logistic regression. Decision trees. Random Forest. Model evaluation and validation. Real-life applications using recent available software.

Where it sits in a curriculum

Programs whose published curriculum lists this course, and the term it falls in. Your own curriculum is the one that counts.

All STAT courses