Statistics Notes

Math 121 - Fall 2026

Jump to: Math 121 homepage, Week 1, Week 2, Week 3, Week 4, Week 5, Week 6, Week 7, Week 8, Week 9, Week 10, Week 11, Week 12, Week 13, Week 14, Week 15

Week 1 Notes

Day Section Topic
Mon, Aug 24 1.2 Data tables, variables, and individuals
Wed, Aug 26 2.1.3 Histograms & skew
Fri, Aug 28 2.1.5 Boxplots

Mon, Aug 24

Today we covered data tables, individuals, and variables. We also talked about the difference between categorical and quantitative variables.

  1. We looked at a case of a nurse who was accused of killing patients at the hospital where she worked for 18 months. One piece of evidence against her was that 40 patients died during the shifts when she worked, but only 34 died during shifts when she wasn’t working. If this evidence came from a date table, what would be the most natural individuals (rows) & variables (columns) for that table?
  1. In the data table in the example above, who or what are the individuals? What are the variables and which are quantitative and which are categorical?

  2. If we want to compare states to see which are safer, why is it better to compare the rates instead of the total fatalities?

  3. What is wrong with this student’s answer to the previous question?

Rates are better because they are more precise and easier to understand.

I like this incorrect answer because it is a perfect example of bullshit. This student doesn’t know the answer so they are trying to write something that sounds good and earns partial credit. Try to avoid writing bullshit. If you catch yourself writing B.S. on one of my quizzes or tests, then you can be sure that you a missing a really simple idea and you should see if you can figure out what it is.

Wed, Aug 26

We talked briefly about making bar charts for categorical data.

  1. Exercise 2.21

Then we introduced stem & leaf plots (stemplots) and histograms for quantitative data. We started by making a stemplot and a histogram for the weights of the students in the class. We also talked about how to tell if data is skewed left or skewed right.

  1. Can you think of a distribution that is skewed left?

  2. Why isn’t this bar graph from the book a histogram?

Then we did this workshop:

We finished by reviewing the mean and the median.

Fri, Aug 28

We introduced the five number summary and box-and-whisker plots (boxplots). We started by reviewing the median.

Median versus Average

The median of NN numbers is located at position N+12\dfrac{N+1}{2}.

The median is not affected by skew, but the average is pulled in the direction of the skew. So the average will be bigger than the median when the data is skewed right, and smaller when the data is skewed left.

We also talked about the interquartile range (IQR) and how to use the 1.5×IQR1.5 \times \text{IQR} rule to determine if data is an outlier. We also mentioned that a statistic is robust if it is not affected by skew and outliers. Both the median and the IQR are robust statistics.

We started with this simple example:

  1. An 8 man crew team actually includes 9 men, the 8 rowers and one coxswain. Suppose the weights (in pounds) of the 9 men on a team are as follows:

     120  180  185  200  210  210  215  215  215

    Find the 5-number summary and draw a box-and-whisker plot for this data. Is the coxswain who weighs 120 lbs. an outlier?


Week 2 Notes

Day Section Topic
Mon, Aug 31 2.1.4 Standard deviation
Wed, Sep 2 4.1 Normal distribution
Fri, Sep 4 4.1.4 Normal distribution computations

Mon, Aug 31

We also introduced the standard deviation. We did this one example of a standard deviation calculation by hand, but you won’t ever have to do that again in this class.

  1. 11 students just completed a nursing program. Here is the number of years it took each student to complete the program. Find the standard deviation of these numbers.

     3  3  3  3  4  4  4  4  5  5  6

From now on we will just use software to find standard deviation. In a spreadsheet (Excel or Google Sheets) you can use the =STDEV() function.

  1. Which of the following data sets has the largest standard deviation?

    1. 1000, 998, 1005
    2. 8, 10, 15, 20, 22, 27
    3. 30, 60, 90

We finished by looking at some examples of histograms that have a shape that looks roughly like a bell. This is a very common pattern in nature that is called the normal distribution.

The normal distribution is a mathematical model for data with a histogram that is shaped like a bell. The model has the following features:

  1. It is symmetric (left & right tails are same size)
  2. The mean (μ\mu) is the same as the median.
  3. It has two inflection points (the two steepest points on the curve)
  4. The distance from the mean to either inflection point is the standard deviation (σ\sigma).
  5. The two numbers μ\mu and σ\sigma completely describe the model.

The normal distribution is a theoretical model that doesn’t have to perfectly match the data to be useful. We use Greek letters μ\mu and σ\sigma for the theoretical mean and standard deviation of the normal distribution to distinguish them from the sample mean x‾\bar{x} and standard deviation ss of our data which probably won’t follow the theoretical model perfectly.

Wed, Sep 2

We talked about z-values and the 68-95-99.7 rule.

We also did these exercises before the workshop.

  1. In 2020, Farmville got 61 inches of rain total (making 2020 the second wettest year on record). How many standard deviations is this above average?

  2. The average high temperature in Anchorage, AK in January is 21 degrees Fahrenheit, with standard deviation 10. The average high temperature in Honolulu, HI in January is 80°F with σ = 8°F. In which city would it be more unusual to have a high temperature of 57°F in January?

Fri, Sep 4

Today we used the Probability Distributions app (android version, iOS version) to calculate normal distribution probabilities.

  1. (Percent below) SAT verbal scores are roughly normally distributed with mean μ = 500, and σ = 100. Estimate the percentile of a student with a 560 verbal score.

  2. (Percent above) What percent of students get above a 560 verbal score on the SATs?

  3. (Percent to locations) What SAT score is in the 90th percentile?

  4. (Percent between) What percent of years do we get between 40 and 50 inches of rain in Farmville?

We also talked about the probability shorthand notation P(X<x)P(X < x) which literally means “the probability that the outcome X is less than x”. Then we did this workshop.


Week 3 Notes

Day Section Topic
Mon, Sep 7 No class (Labor day)
Wed, Sep 9 2.1, 8.1 Scatterplots and correlation
Fri, Sep 11 8.2 Least squares regression introduction

Wed, Sep 9

We introduced scatterplots and correlation coefficients with these examples:

  1. What would the correlation between husband and wife ages be in a country where every man married a woman exactly twice his age?

Important concept: correlation does not change if you change the units or apply a simple linear transformation to the axes. Correlation just measures the strength of the linear trend in the scatterplot.

Another thing to know about the correlation coefficient is that it only measures the strength of a linear trend. The correlation coefficient is not as useful when a scatterplot has a clearly visible nonlinear trend.

We finished by introducing the least squares regression line which has these features:

  1. Slope m=Rsysxm = R \frac{s_y}{s_x}
  2. Point (x‾,y‾)(\bar{x}, \bar{y})

The two main applications of a least squares regression line are:

It is important to be able to describe the units of the slope.

  1. What are the units of the slope of the regression line for predicting BAC from the number of beers someone drinks?

Fri, Sep 11

We started by deriving the slope-intercept formula for a least squares regression line.

Least Squares Regression Line

y=mx+by = m x + b

where m=Rsysxm = R \frac{s_y}{s_x} is the slope and b=y‾−mx‾b = \bar{y} - m \bar{x} is the y-intercept.

  1. What are the slope and y-intercept to predict someone’s weight based on their height?

  2. What are the units of the slope for predicting someone’s weight from their height?

We also introduced the following concepts.

The coefficient of determination. R2R^2 represents the proportion of the variability of the yy-values that follows the trend line. The remaining 1−R21-R^2 represents the proportion of the variability that is above and below the trend line.

Regression to the mean. Extreme xx-values tend to have less extreme predicted yy-values in a least squares regression model.

  1. The regression line for predicting midterm 2 grades based on midterm 1 is: y=0.58x+29.7.y = 0.58 x + 29.7. with a coefficient of determination R2=0.371R^2 = 0.371.

    1. What is the predicted average midterm 2 grade for students who got a 100 on midterm 1?
    2. What is the predicted average midterm 2 grade for students who got a 50 on midterm 1?
    3. What percent of the variability in midterm 2 grades is not predicted by midterm 1 grades?

Week 4 Notes

Day Section Topic
Mon, Sep 14 8.2 Least squares regression practice
Wed, Sep 16 1.3 Sampling: populations and samples
Fri, Sep 18 1.3 Bias versus random error

Mon, Sep 14

Before the workshop, we started with this warm-up exercise.

  1. Suppose that the correlation between the heights of fathers and adult sons is R=0.5R = 0.5. Given that both fathers and sons have normally distributed heights with mean 7070 inches and standard deviation 3 inches, find an equation for the least squares regression line.

Wed, Sep 16

We talked about the difference between samples and populations. The central problem of statistics is to use sample statistics to answer questions about population parameters.

We looked at an example of sampling from the Gettysburg address, and we talked about the central problem of statistics: How can you answer questions about the population using samples?

The reason this is hard is because sample statistics usually don’t match the true population parameter. There are two reasons why:

We looked at this case study:

Important Concepts

  1. Bigger samples have less random error.

  2. Bigger samples don’t reduce bias.

  3. The only sure way to avoid bias is a simple random sample.

We finished with this practice problem:

  1. Suppose you want to determine the percent of grocery stores that carry a specific brand of pasta. To find out, you visit a random selection of nearby grocery stores. In this situation, identify the following:

    1. The population
    2. The sample
    3. The variable(s) of interest
    4. The statistic
    5. The parameter

Fri, Sep 18

We did this workshop.


Week 5 Notes

Day Section Topic
Mon, Sep 21 1.4 Randomized controlled experiments
Wed, Sep 23 Review
Fri, Sep 25 Midterm 1

Mon, Sep 21

One of the hardest problems in statistics is to prove causation. Here is a diagram that illustrates the problem.

The explanatory variable might be the cause of a change in the response variable. But we have to watch out for other variables that aren’t part of the study called lurking variables. When researchers take a variable into account in a study, we say it is controlled.

A lurking variable that might be associated with both the explanatory and response variable is called a confounding variable.

We say that correlation is not causation because you can’t assume that there is a cause and effect relationship between two variables just because they are strongly associated. The association might be caused by lurking variables or the causal relationship might go in the opposite direction of what you expect.

Experiments versus Observational Studies

An experiment is a study where individuals are put into different treatment groups. An experiment is randomized if the individuals are randomly assigned to the treatment groups. An observational study is one where the researchers do not place the individuals into different treatment groups.

Proving Cause and Effect

We looked at these examples.

  1. A study tried determine whether cellphones cause brain cancer. The researchers interviewed 469 brain cancer patients about their cellphone use between 1994 and 1998. They also interviewed 469 other hospital patients (without brain cancer) who had the same ages, genders, and races as the brain cancer patients.

    1. What was the explanatory variable?
    2. What was the response variable?
    3. Which variables were controlled?
    4. Was this an experiment or an observational study?
    5. Are there any possible lurking variables?
  2. In 1954, the polio vaccine trials were one of the largest randomized controlled experiments ever conducted. Here were the results.

    1. What was the explanatory variable?
    2. What was the response variable?
    3. This was an experiment because it had a treatment variable. What was that?
    4. Which variables were controlled?
    5. Why don’t we have to worry about lurking variables?

We talked about why the polio vaccine trials were double blind and what that means.

  1. Do magnetic bracelets work to help with arthritis pain?

    1. What is the explanatory variable?
    2. What is the response variable?
    3. How hard would it be to design a randomized controlled experiment to answer the question above?

We finished by talking about anecdotal evidence.

Wed, Sep 23

We talked about the midterm 1 review problems.


Week 6 Notes

Day Section Topic
Mon, Sep 28 3.1 Defining probability
Wed, Sep 30 3.1 Multiplication and addition rules
Fri, Oct 2 3.4 Weighted averages & expected value

Mon, Sep 28

Today we introduced probability models which always have two parts:

  1. A list of possible outcomes called a sample space.
  2. A probability function P(E)P(E) that gives the probability for any subset EE of the sample space.

A subset of the sample space is called an event. We already intuitively know lots of probability models, for example we described the following probability models:

  1. Flip a coin.

  2. Roll a six-sided die.

  3. If you roll a six-sided die, what is P(result at least 5)?P(\text{result at least 5})?

  4. The proportion of people in the US with each of the four blood types is shown in the table below.

    Type O A B AB
    Proportion 0.45 0.40 0.11 ?

    What is P(Type AB)?P(\text{Type AB})?

Wed, Sep 30

Today we talked about the multiplication and addition rules for probability. We also talked about independent events and conditional probability. We started with these examples.

  1. If you roll two six-sided dice, the results are independent. What is the probability that both dice land on a six?

  2. Suppose you shuffle a deck of 52 playing cards and then draw two cards from the top. Find

    1. P(First card is an ace)P(\text{First card is an ace})
    2. P(Second card is an ace | first is an ace)P(\text{Second card is an ace } | \text{ first is an ace})
    3. P(Second card is an ace | first is not an ace)P(\text{Second card is an ace } | \text{ first is not an ace})

Then we did this workshop:


Week 7 Notes

Day Section Topic
Mon, Oct 5 3.4 Random variables
Wed, Oct 7 7.1 Sampling distributions
Fri, Oct 9 5.1 Sampling distributions for proportions

Week 8 Notes

Day Section Topic
Mon, Oct 12 No class (Fall break)
Wed, Oct 14 5.2 Confidence intervals for a proportion
Fri, Oct 16 5.2 Confidence intervals for a proportion - con’d

Week 9 Notes

Day Section Topic
Mon, Oct 19 5.3 Hypothesis testing for a proportion
Wed, Oct 21 Review
Fri, Oct 23 Midterm 2

Week 10 Notes

Day Section Topic
Mon, Oct 26 6.1 Inference for a single proportion
Wed, Oct 28 5.3.3 Decision errors
Fri, Oct 30 6.2 Difference of two proportions (hypothesis tests)

Week 11 Notes

Day Section Topic
Mon, Nov 2 6.2.3 Difference of two proportions (confidence intervals)
Wed, Nov 4 7.1 Introducing the t-distribution
Fri, Nov 6 7.1.4 One sample t-confidence intervals

Week 12 Notes

Day Section Topic
Mon, Nov 9 7.2 Paired data
Wed, Nov 11 7.3 Difference of two means
Fri, Nov 13 7.3 Difference of two means - con’d

Week 13 Notes

Day Section Topic
Mon, Nov 16 Choosing the right technique
Wed, Nov 18 Review
Fri, Nov 20 Midterm 3

Week 14 Notes

Day Section Topic
Mon, Nov 23 7.4 Statistical power
Wed, Nov 25 No class (Thanksgiving break)
Fri, Nov 27 No class (Thanksgiving break)

Week 15 Notes

Day Section Topic
Mon, Nov 30 6.3 Chi-squared statistic
Wed, Dec 2 6.4 Testing association with chi-squared
Fri, Dec 4 Chi-squared caveats
Mon, Dec 7 Last day, recap & review