MAS3928: Statistical Modelling
Module information
Lecturer information
Module schedule
Course materials
Assessment
Relevant texts
1
Introduction
1.1
Multiple linear regression
1.2
Matrix form of the model
Example - Matrix form for pre-diabetes data
1.3
Parameter estimation
1.3.1
Estimation of
\(\underline{\beta}\)
1.3.2
Estimation of
\(\sigma_{\epsilon}^2\)
1.3.3
Residuals, fitted values and the ‘hat matrix’
1.3.4
Properties of the hat matrix
Example: Multiple linear regresion analysis of bodyweight data
1.4
Expectations, variances and inference
1.4.1
Expectation of
\(\underline{\hat{\beta}}\)
1.4.2
Variance of
\(\underline{\hat{\beta}}\)
1.4.3
Inference for
\(\underline{\hat{\beta}}\)
1.4.4
Expectation and variance of the fitted values
1.4.5
Expectation and variance of the residuals
1.5
Multiple linear regression in
R
1.5.1
Using data in
R
Example: Analysis of bodyweight data using
R
1.6
The role of the intercept
Example: Analysis of bodyweight data without an intercept term
Example: Analysis of men’s Premier League football data - the role of the intercept
1.6.1
Interpretability of the intercept and extrapolation
1.6.2
Mean-centering of covariates
Example: Mean-centering (men’s Premier League football data)
1.7
Properties of
\(\left(\mathrm{X}^T\mathrm{X}\right)^{-1}\)
: multicollinearity
Example: Multicollinearity in men’s Premier League football data
2
Regression diagnostics
2.1
Standardised residuals
2.1.1
Residual plots
Correlation between residuals and fitted values
2.1.2
Outliers
Example: Residual analysis for pre-diabetes data
2.1.3
Normality of the residuals
2.1.4
Anderson-Darling test
A cautionary note
2.2
Regression diagnostics
2.2.1
Leverage values
Example: Leverage values for the pre-diabetes data
2.2.2
Influential observations
2.2.3
Dealing with unusual observations
Example: Model checking for the Premier League data
3
Inference for the multiple linear regression model
3.1
Assessing the fit
3.2
The basic anova table
Cochran’s Theorem
The coefficient of determination
Example: Cheddar cheese study
Solution
Anova in
R
3.3
The extra sum of squares method
3.3.1
The extended anova table
Example: Extra sum of squares for cheese data
Example: Warfarin study
3.4
The general extra sum of squares method
Example: Extended extra sum of squares for cheese data
Summary: extra sum of squares method
Example: Anova for crime data based on summary information
Solution
3.5
Confidence and prediction intervals for the fitted values
Example: Confidence and prediction intervals for the cheese data
3.6
Polynomial models
Example: Polynomial model
3.6.1
Choosing the order of a polynomial model
4
Categorical variables and interactions
4.1
Indicator and dummy variables
Example: Gasoline data
Model interpretation
4.2
Model selection criteria
4.2.1
Model selection criteria: adjusted
\(R^2\)
4.2.2
Model selection criteria: Akaike’s Information Criterion
4.3
Reducing the number of variables
4.3.1
Backward elimination
4.3.2
Forward selection
4.3.3
Stepwise selection
4.4
Multicollinearity and the variance inflation factor
4.4.1
Variance inflation factors
Example: Variance inflation factors
4.5
Transformations
4.5.1
The Box-Cox transformation
Example: Box-Cox transformation
5
Analysis of designed experiments
5.1
Completely randomised design
5.1.1
Handling factors with
\(k\)
levels
Example: One-way anova
5.1.2
Analysis of completely randomised design data in
R
5.1.3
Interpretation of results: multiple comparisons
Example: Multiple comparisons
5.1.4
Model checking
5.1.5
Completely randomised design: dealing with quantitative variables
Example: Randomised design with a quantitative variable
5.2
Randomised block design
5.2.1
Two-way analysis of variance
5.2.2
Orthogonality and testing of blocks in a two-way analysis of variance model
Example: Two-way anova on nitrate data
Example: Chicken egg production
5.3
Factorial experiments
Example: Two-way anova with interactions for the yeast data
5.3.1
Exploratory plots for interactions
Example: Transforming the response
Published with bookdown
MAS3928: Statistical Modelling