Engineering Mathematics · Lesson 16 of 28
Statistics: Regression, Correlation, and Sampling
Fit a least-squares line to paired data, measure the fit with the correlation coefficient, then move to sampling distributions, the standard error, confidence intervals, and a first hypothesis test, every step carried to a finished number.
15 min read · Super EaFree lesson
Statistics questions on the CELE reward a steady hand with formulas more than deep theory. Given a small table of paired data you fit a straight line, measure how tightly the points cluster around it, and then step up to the language of sampling: how a sample mean behaves, how wide to draw a confidence interval, and how to decide whether a claim survives the evidence. This lesson carries each of those tools all the way to a finished number so you can check your own arithmetic against a worked answer.
Least-squares linear regression
Given n paired observations (x, y), linear regression fits the straight line y = a + b x that comes closest to the points. "Closest" has a precise meaning: the least-squares line is the one line that minimizes the sum of the squared vertical distances from the points to the line. Those vertical gaps are the residuals, and squaring them keeps positive and negative gaps from cancelling and penalizes the far-off points hardest. The slope and intercept fall out of two formulas:
b = [n sum(xy) - sum(x) sum(y)] / [n sum(x^2) - (sum x)^2]
a = mean(y) - b mean(x)
Compute the slope first, then feed it into the intercept. The intercept formula guarantees the line passes through the centroid (mean x, mean y), a useful check.
Worked example: Fit a line to the five points (1, 2), (2, 4), (3, 5), (4, 4), (5, 5). Build the running sums in a table before touching the formulas.
| x | y | x*y | x^2 | y^2 |
|---|---|---|---|---|
| 1 | 2 | 2 | 1 | 4 |
| 2 | 4 | 8 | 4 | 16 |
| 3 | 5 | 15 | 9 | 25 |
| 4 | 4 | 16 | 16 | 16 |
| 5 | 5 | 25 | 25 | 25 |
| sum = 15 | sum = 20 | sum = 66 | sum = 55 | sum = 86 |
Here n = 5, sum x = 15, sum y = 20, sum xy = 66, and sum x^2 = 55. The slope is b = [5(66) - 15(20)] / [5(55) - 15^2] = [330 - 300] / [275 - 225] = 30 / 50 = 0.6. The means are mean x = 15/5 = 3 and mean y = 20/5 = 4, so the intercept is a = 4 - 0.6(3) = 4 - 1.8 = 2.2. The fitted line is y = 2.2 + 0.6 x. To predict at a new x, substitute: at x = 8 the model gives y = 2.2 + 0.6(8) = 2.2 + 4.8 = 7.0.
The correlation coefficient
The slope tells you the direction and steepness of the trend, but not how well the points actually line up. That job belongs to the linear correlation coefficient r, which measures the strength and direction of the straight-line association:
r = [n sum(xy) - sum(x) sum(y)] / sqrt{ [n sum(x^2) - (sum x)^2] [n sum(y^2) - (sum y)^2] }
The value of r always lies between -1 and 1. A value near +1 means the points sit almost perfectly on an upward line, near -1 means almost perfectly on a downward line, and near 0 means no linear pattern. Notice the numerator of r is the same bracket that sits on top of the slope formula, so the two quantities always share a sign. Squaring gives the coefficient of determination r^2, which reports the fraction of the variation in y that the line explains.
Worked example: Use the same five points, where the sums add sum y^2 = 86 to the earlier totals. The numerator is n sum(xy) - sum(x) sum(y) = 5(66) - 15(20) = 330 - 300 = 30. The first bracket in the denominator is n sum(x^2) - (sum x)^2 = 5(55) - 225 = 50, and the second is n sum(y^2) - (sum y)^2 = 5(86) - 20^2 = 430 - 400 = 30. So r = 30 / sqrt(50 times 30) = 30 / sqrt(1500) = 30 / 38.73 = 0.775. Squaring, r^2 = 0.60, so about 60 percent of the variation in y is explained by the linear relationship with x. A strong correlation like this signals a tight linear association, but it never by itself proves that x causes y.
Sampling distributions and the standard error
You almost never measure a whole population, so you work from a sample and its mean, xbar. If you drew sample after sample of the same size n, their means would themselves scatter around the true population mean mu. That spread of sample means is the sampling distribution of the mean, and its standard deviation has a special name, the standard error:
SE = sigma / sqrt(n)
The larger the sample, the smaller the standard error, so bigger samples give more reliable means. Because the square root sits in the denominator, the payoff shrinks: to halve the standard error you must quadruple n. The Central Limit Theorem is what makes this usable on the board: for a large sample size the sampling distribution of the mean is approximately normal, no matter what shape the underlying population has, as long as sigma is finite.
Worked example: Concrete cylinder strengths in a population have mean mu = 28 MPa and standard deviation sigma = 3 MPa. For samples of n = 36 cylinders, the standard error of the mean is SE = 3 / sqrt(36) = 3 / 6 = 0.5 MPa. If you instead tested n = 144 cylinders, the standard error would fall to 3 / sqrt(144) = 3 / 12 = 0.25 MPa. Quadrupling the sample from 36 to 144 exactly halved the standard error, the square-root payoff in action.
Confidence intervals
A single sample mean is a point estimate; a confidence interval turns it into a range that plausibly contains the true mean. When the population standard deviation sigma is known and the sample is large, the interval for the mean is:
xbar +/- z (sigma / sqrt(n))
The z multiplier is the critical value for the confidence level you want, read from the standard normal distribution. The product E = z (sigma / sqrt(n)) is the margin of error, half the total width of the interval. Raising the confidence level pushes z up, so a more confident statement is a wider, less precise interval. The most-used critical values are worth memorizing:
| Confidence level | Two-sided z critical value |
|---|---|
| 90 percent | 1.645 |
| 95 percent | 1.960 |
| 99 percent | 2.576 |
When sigma is unknown and must be estimated from a small sample, you swap the z value for a Student t value with n - 1 degrees of freedom, which is slightly larger and so widens the interval to pay for the extra uncertainty. For the large samples the board usually poses, z is the safe default.
Worked example: A sample of n = 36 cylinders gives a mean strength of xbar = 28.0 MPa, and the population standard deviation is known to be sigma = 3 MPa. For a 95 percent confidence interval the critical value is z = 1.96, so the margin of error is E = 1.96(3 / sqrt(36)) = 1.96(3 / 6) = 1.96(0.5) = 0.98 MPa. The interval is 28.0 +/- 0.98, that is from 27.02 MPa to 28.98 MPa. You can be 95 percent confident the true mean strength lies in that band.
A first look at hypothesis testing
Hypothesis testing flips the confidence interval around: instead of estimating the mean, you test a specific claim about it. State two competing hypotheses. The null hypothesis H0 asserts no effect or no departure from a claimed value, and the alternative hypothesis Ha states the departure you suspect. You pick a significance level alpha, usually 0.05, which is the probability you will wrongly reject a true H0. When sigma is known you form the test statistic:
z = (xbar - mu0) / (sigma / sqrt(n))
This counts how many standard errors the observed mean sits from the claimed mean mu0. Compare it to the critical value for alpha, or compute the p-value, the probability, assuming H0 is true, of seeing a result at least as extreme as the one observed. If the p-value is below alpha, or the test statistic falls past the critical value, you reject H0; otherwise you fail to reject it. Failing to reject is not the same as proving H0 true, it only means the evidence was not strong enough to overturn it.
Worked example: A supplier claims the mean strength of its cylinders is mu0 = 30 MPa. You test n = 36 cylinders, find xbar = 28.5 MPa, and know sigma = 3 MPa. You suspect the true mean is lower, so Ha: mu < 30, a one-sided test at alpha = 0.05, whose critical value is z = -1.645. The standard error is 3 / sqrt(36) = 0.5 MPa, so the test statistic is z = (28.5 - 30) / 0.5 = -1.5 / 0.5 = -3.0. Since -3.0 < -1.645, the statistic falls in the rejection region and you reject H0. The p-value, the area below z = -3.0, is about 0.0013, far under 0.05, so the sample gives strong evidence the true mean is below the claimed 30 MPa.
Exam-day strategy
- Build the sums table first. Fill columns for x, y, x*y, x^2, and y^2 and total each before you touch the slope, intercept, or correlation formula, since all three reuse the same running sums.
- Compute the slope before the intercept, then use a = mean(y) - b mean(x); confirm the line passes through the centroid (mean x, mean y) as a quick self-check.
- Keep r and r^2 straight: r runs from -1 to 1 and carries the sign of the slope, while r^2 runs from 0 to 1 and reports the fraction of variation explained. If a choice for r is above 1 or below -1, it is a trap.
- Remember that the standard error divides sigma by sqrt(n), not by n; to cut the standard error in half you must quadruple the sample size, not double it.
- For a confidence interval, lock the critical values 1.645, 1.96, and 2.576 for 90, 95, and 99 percent; a higher confidence level always gives a wider interval, so if raising the confidence made your interval narrower you slipped.
- On a hypothesis test, write H0 and Ha first, form z = (xbar - mu0) / (sigma / sqrt n), and reject H0 only when the statistic passes the critical value or the p-value falls below alpha; failing to reject never proves H0 true.
Marking it done updates your Exam-Ready progress.
Lesson quiz
Check you actually have it
20 items on this lesson alone, randomized each try, with the reasoning on every answer.
Statistics: quick check
Item 01 / 20 · Score 0
For a one-sided test with Ha: mu < 30 at alpha = 0.05 the critical value is -1.645. If the computed test statistic is z = -3.0, the correct decision is to:
This whole first section is free
Read every lesson in Engineering Mathematics and take its quizzes free. The full CELE reviewer unlocks the other 5 subjects, all section tests, and the timed mock exams — one payment, lifetime access, ₱399.
Unlock the full reviewer