Describe the effects of multicollinearity on the estimated coefficients, theassociated standard errors and the significance of the coefficients using theordinary maximum likelihood method.ii.
Describe the effects of multicollinearity on the estimated coefficients, theassociated standard errors and the significance of the coefficients using theordinary maximum likelihood method.ii.
July 22, 2020 Comments Off on Describe the effects of multicollinearity on the estimated coefficients, theassociated standard errors and the significance of the coefficients using theordinary maximum likelihood method.ii. Uncategorized Assignment-helpThe following is an example of the coursework that will be expected to be delivered within 12 hours, This coursework contains four questions. Answer ALL FOUR. All questions will be given equal weight (25%).Time allowed – Expected Writing Time: 2 hours (you would have 12 hours to answer)In this exam is (a) Suppose that yi ∼ N(µ, 1) for i = 1, . . . , n and that the yi’s are independent.i. Show that the sample mean estimator ˆµ1 =1/n ∑yi is obtained fromminimising the least squares criterion [7 marks]µˆsub(1) = argmin.∑(yi-µ)^2, and that ^µsub(1) an unbiased estimator of µ. Also find the variance of ^µsub(1)ii. Consider adding a penalty term to the least squares criterion, and therefore using the estimator that minimises µˆ2 = argmin∑(yi-µ)^2+ λ(µ)^2 for the mean, where λ is a non-negative tuning parameter. Derive ˆµ2, find it bias and show that its variance is lower than that of ˆµ1Consider the multiple linear regression model yi = β0 + ∑βsub(j)x(sub)ij + e(sub)i, i = 1, . . . , n, j = 1, dots, p, where β = (β1, …, βp)^T and error-term= (e(sub)1….e(sub)n)^T∼ N(0, σ^2 I(sub)n).i. When p is comparable to n, the multicollinearity becomes an issue. Describe the effects of multicollinearity on the estimated coefficients, theassociated standard errors and the significance of the coefficients using theordinary maximum likelihood method.ii. The ridge regression estimate of β can be obtained by minimising a particular expression with respect to β. Write down this expression as well asan alternative formulation of it.iii. Explain why ridge regression can potentially correct the problems ofmulticollinearity. [2 marks]iv. Provide an advantage and a disadvantage of ridge regression over the standard linear regression.2. Let x = (x1, . . . , x100), with ∑xi = 20, be a random sample from the Exponential(λ)distribution with probability density function given byf(x(sub)i|λ) = 1/λ exp(−x(sub)i/λ), x(sub)i > 0, λ > 0. Note that E(xi) = λ.(a) Assign the IGamma(0.1, 0.1) prior to λ and find the corresponding posterior distribution. (b) Find the Jeffreys’ prior for λ. Which is the corresponding posterior distribution.(c) Find a Bayes estimator for λ based on the priors of parts (a) and (b)(d) Let y represent a future observation from the same model. Find the predictivedistribution of y based either on the prior of part (a) or (b).(e) Describe how you can calculate the mean the of the predictive distribution insoftware such as R.3. (a) i. Suppose a non-linear model that can be written as Y = f(X) + e,where e has zero mean and variance σ^2, and is independent of X. Showthat the expected test error, conditional on X can be decomposed into thefollowing three parts:E[(Y − ˆf(X))^2] = σ^2 + Bias [f(x)]^2 + Var [f(x)] , where f(·) is estimated from the training data.ii. To estimate the test error rate, one can use the 10-fold Cross Validation(CV) approach or the information criterion approach, e.g. AIC, BIC. Whatare the main advantage and disadvantage of using the 5-fold CV approachin comparison with AIC or BIC?iii. State which one of AIC and BIC tends to select smaller size model andexplain the reason(b) i. The tree in Figure 1 provides a regression tree based on a dataset of patient visits for upper respiratory infection. The aim is to identify factorsassociated with a physicians rate of prescribing, which is a continuous variable. The variables appearing in the regression tree are private: percentof privately insured patients a physician has, black: the percent of blackpatients a physician has, and fam whether or not the physician specialisesin family medicine. Provide an interpretation of this tree.ii. Consider the regression tree of Figure 2 where the response variable is thelog salary of a baseball player, based on the number of years that he hasplayed in the major leagues (Years) and the number of hits that he madein the previous year (Hits). Create a diagram that represent the partitionof the predictors spaces according to this tree4 (a) i. Consider the following data: 10 20 40 80 85 121 160 168 195.Use the k-means algorithm with k = 3 to cluster the data set. Use theEuclidean distance to measure the distance between the data points. Suppose that the points 160, 168, and 195 were selected as the initial clustermeans. Work from these initial values to determine the final clustering forthe data. Provide results from each iteration.ii. What are the main disadvantages of k-means clustering? Why one maywant to consider hierarchical clustering as an alternative? (b) i. Data are available for students taking BSc degree in Data Science andin particular the variables X1: average mark on project coursework, X2:average hours studied per course, and Y : get a degree with distinction. Theestimated coefficients of a logistic regression model were β0 =?5, β1 = 0.02,β2 = 0.1. Estimate the probability that a student who takes on average50% on project coursework and studies 30 hours on average for each coursegets a degree with distinction? How many hours would the student in part(a) need to study on average to have a 50 % chance of getting a degreewith distinction ? ii. Suppose that we wish to predict whether a high quality chip produced ina factory will pass the quality control (‘Pass’ or ‘Fail’) based on x, themeasurement of its diameter. Diameter measurements are available for alarge number of chips. After examining them it turns out that the meanvalue of x for chips that passed the quality control was 5mm, while themean for those that didn’t was 7mm. Moreover, the variance of x forthese two sets of companies was σ^2 = 1. Finally, 70% of the producedchips passed the quality control. Assuming that x follows the normaldistribution, predict the probability that a chip with x = 5.8 will pass thequality control.If any writter is interested I have the solution to this coursework and two other that can be used as preparation. Let me know to send the answer.


