Algebra 1 (HS pathway)
Descriptive Statistics: Data Distributions and Two-Way Tables
Algebra 1
- ✓By the end of this lesson students will be able to describe the shape, center, and spread of a data distribution.
- ✓By the end of this lesson students will be able to construct and interpret two-way frequency tables.
- ✓By the end of this lesson students will be able to calculate and interpret joint, marginal, and conditional relative frequencies from two-way tables.
- ✓By the end of this lesson students will be able to identify possible associations between two categorical variables using two-way tables.
Key concepts
A data distribution describes how data values are spread out and clustered. We analyze distributions by their shape, center, and spread. Common visual representations include dot plots, histograms, and box plots.
The shape describes the overall pattern of the data. Key shapes include:\n- **Symmetric**: The left and right sides of the distribution are approximate mirror images.\n- **Skewed Right (Positively Skewed)**: The tail of the distribution extends to the right, meaning there are more data values on the lower end and fewer, larger values pulling the mean to the right.\n- **Skewed Left (Negatively Skewed)**: The tail of the distribution extends to the left, meaning there are more data values on the higher end and fewer, smaller values pulling the mean to the left.\n- **Uniform**: All data values or intervals have roughly the same frequency, resulting in a flat shape.
The center describes a typical or central value of the data.\n- **Mean**: The arithmetic average of all data values. It is sensitive to outliers and skewness.\n Formula: Sum of all values / Number of values.\n- **Median**: The middle value when the data is ordered from least to greatest. If there's an even number of data points, it's the average of the two middle values. It is resistant to outliers and skewness, making it a better measure of center for skewed distributions.
The spread (or variability) describes how much the data values vary from each other.\n- **Range**: The difference between the maximum and minimum values in the dataset. It is simple but highly affected by outliers.\n Formula: Maximum value - Minimum value.\n- **Interquartile Range (IQR)**: The range of the middle 50% of the data. It is the difference between the third quartile (Q3) and the first quartile (Q1). It is resistant to outliers and is a good measure of spread for skewed distributions.\n Formula: IQR = Q3 - Q1.
A two-way frequency table (or contingency table) displays the frequencies of two categorical variables. It shows how many observations fall into each combination of categories for the two variables.
The entries in the body of a two-way table that represent the count of observations sharing two specific characteristics (one from each variable).
The totals in the margins (rows and columns) of a two-way table. They represent the total count for each category of a single variable, ignoring the other variable.
Frequencies expressed as proportions or percentages of a total. They help compare distributions across different sample sizes.\n- **Joint Relative Frequency**: The ratio of a joint frequency to the grand total of all observations. It tells you the proportion of the total sample that has both characteristics.\n Formula: Joint Frequency / Grand Total.\n- **Marginal Relative Frequency**: The ratio of a marginal frequency to the grand total of all observations. It tells you the proportion of the total sample that has a specific characteristic of one variable.\n Formula: Marginal Frequency / Grand Total.\n- **Conditional Relative Frequency**: The ratio of a joint frequency to a marginal frequency. It tells you the proportion of observations with a specific characteristic GIVEN that they also have another specific characteristic. The 'condition' defines the denominator.
Key facts to remember
- 1Data distributions are described by their shape (symmetric, skewed, uniform), center (mean, median), and spread (range, IQR).
- 2The mean is sensitive to outliers and skewness, while the median is resistant.
- 3For symmetric distributions, the mean and median are close; for skewed distributions, the mean is pulled towards the tail.
- 4A two-way frequency table displays the counts of two categorical variables.
- 5Joint frequencies are counts within the body of the table, representing the intersection of two categories.
- 6Marginal frequencies are the row and column totals, representing the total counts for each category of a single variable.
- 7Relative frequencies (joint, marginal, conditional) express frequencies as proportions or percentages of a total.
- 8Conditional relative frequencies are used to examine associations between variables by comparing proportions within specific subgroups.
Worked examples
Example 1
A math teacher recorded the scores on a 10-point quiz for 20 students: 7, 8, 5, 9, 7, 10, 6, 8, 7, 9, 8, 6, 7, 8, 5, 9, 7, 8, 6, 7. Describe the shape, estimate the center, and discuss the spread of this data distribution.
Answer
The data distribution of quiz scores is roughly symmetric. The center (median) is 7, and the mean is 7.3. The spread is relatively small, with a range of 5 points and an interquartile range (IQR) of 1.5 points, indicating that the middle 50% of scores are tightly clustered.
For symmetric distributions, the mean and median are usually very close. For skewed distributions, the median is generally a better measure of center.
Example 2
A survey asked 150 high school students about their preferred method of transportation to school: Car, Bus, or Walk. The results are summarized below:\n- 90 students are Freshmen or Sophomores (underclassmen).\n- 60 students are Juniors or Seniors (upperclassmen).\n- 40 underclassmen prefer Car.\n- 30 underclassmen prefer Bus.\n- 20 underclassmen prefer Walk.\n- 25 upperclassmen prefer Car.\n- 15 upperclassmen prefer Bus.\n- 20 upperclassmen prefer Walk.\n\nConstruct a two-way frequency table and then calculate the joint and marginal relative frequencies.
Answer
The two-way frequency table and relative frequency table are shown above. For example, the joint relative frequency of being an underclassman and preferring a car is approximately 0.267 (26.7%), and the marginal relative frequency of preferring a bus is 0.300 (30%).
Always ensure your relative frequencies sum to 1 (or 100%) for the grand total, and for row/column totals if calculated correctly, allowing for minor rounding differences.
Example 3
Using the two-way frequency table from the previous example (Problem 2):\n\n| Transportation | Car | Bus | Walk | Total |\n|----------------|-----|-----|------|-------|\n| Underclassmen | 40 | 30 | 20 | 90 |\n| Upperclassmen | 25 | 15 | 20 | 60 |\n| Total | 65 | 45 | 40 | 150 |\n\na) What is the conditional relative frequency that a student prefers to walk, GIVEN that they are an underclassman?\nb) What is the conditional relative frequency that a student is an upperclassman, GIVEN that they prefer a car?
Answer
a) The conditional relative frequency that a student prefers to walk, given they are an underclassman, is approximately 0.222 (or 22.2%).\nb) The conditional relative frequency that a student is an upperclassman, given they prefer a car, is approximately 0.385 (or 38.5%).
Pay close attention to the 'GIVEN that' part of the question. This specifies which marginal total becomes the denominator for your calculation.
Common mistakes
- ✗Confusing 'skewed left' with data clustered on the left; skewed left means the tail is on the left.
- ✗Incorrectly calculating the median for an even number of data points (forgetting to average the two middle values).
- ✗Using the mean as the measure of center for highly skewed distributions when the median would be more appropriate.
- ✗Confusing joint, marginal, and conditional relative frequencies, especially misidentifying the correct denominator for conditional frequencies.
- ✗Failing to interpret the meaning of calculated frequencies in the context of the problem.
Exam tips
- ★Always order data from least to greatest before finding the median or quartiles.
- ★Clearly label all rows, columns, and totals when constructing two-way tables.
- ★When calculating relative frequencies, double-check that your denominator is correct based on whether you need a joint, marginal, or conditional frequency.
- ★Practice interpreting the meaning of each type of frequency in plain language, as this is often required in exam questions.
Ready to practise?
Try a problem on this topic
Snap a photo or type a question — get step-by-step working instantly.
