Geometry & Statistics

Analyzing Bivariate Data: Scatter Plots and Lines of Best Fit

Grade 8

  • ✓By the end of this lesson students will be able to construct and interpret scatter plots for bivariate measurement data.
  • ✓By the end of this lesson students will be able to describe patterns of association in scatter plots, including identifying positive, negative, linear, non-linear associations, and recognizing outliers and clusters.
  • ✓By the end of this lesson students will be able to informally fit a straight line to a scatter plot that suggests a linear association.
  • ✓By the end of this lesson students will be able to use the equation of a linear model to solve problems in the context of bivariate data, interpreting the slope and intercept.

Key concepts

Bivariate Data

Bivariate data is data that involves two different variables. We collect measurements for two characteristics for each item or individual in a sample. The purpose of analyzing bivariate data is to see if there is a relationship or association between the two variables.

Scatter Plot

A scatter plot is a graph that displays bivariate data as a set of points. Each point on the scatter plot represents a pair of data values for the two variables. The independent variable (or explanatory variable) is typically plotted on the horizontal (x) axis, and the dependent variable (or response variable) is plotted on the vertical (y) axis. Scatter plots help us visualize the relationship or association between the two variables.

Association (Correlation)

Association describes the relationship between two variables shown in a scatter plot. We look for patterns in the way the points are arranged. \n\n* **Positive Association:** As one variable increases, the other variable also tends to increase. The points generally rise from left to right.\n* **Negative Association:** As one variable increases, the other variable tends to decrease. The points generally fall from left to right.\n* **No Association:** There is no clear pattern between the two variables. The points appear randomly scattered.\n\nAssociations can also be described as **linear** (the points tend to follow a straight line) or **non-linear** (the points follow a curve or another pattern that is not a straight line).

Outliers

An outlier is a data point that lies far away from the general pattern of the other data points in a scatter plot. Outliers can sometimes indicate an error in data collection or an unusual circumstance.

Clusters

Clusters are groups of data points in a scatter plot that are tightly packed together, indicating a concentration of data in a particular region.

Line of Best Fit (Trend Line)

When a scatter plot shows a linear association, we can draw a straight line that best represents the trend of the data. This line is called the line of best fit or trend line. For Grade 8, we informally fit this line by visually estimating a line that passes through the 'middle' of the data points, with roughly an equal number of points above and below it. The line of best fit helps us summarize the relationship and make predictions.

y = mx + b (where m is the slope and b is the y-intercept)
Using the Line of Best Fit for Prediction

Once a line of best fit is drawn, its equation (y = mx + b) can be determined. The slope (m) tells us the rate of change of the dependent variable with respect to the independent variable. The y-intercept (b) tells us the predicted value of the dependent variable when the independent variable is zero. We can use this equation to make predictions about one variable given a value for the other, especially for values within the range of the observed data (interpolation).

y = mx + b

Key facts to remember

  • 1Bivariate data involves two variables, and scatter plots are used to visualize their relationship.
  • 2A scatter plot shows each pair of data values as a single point on a coordinate plane.
  • 3Associations can be positive (both variables increase), negative (one increases, other decreases), or no association (no clear pattern).
  • 4Associations can be linear (points form a straight line) or non-linear (points form a curve).
  • 5Outliers are data points that lie far away from the general trend of the data.
  • 6A line of best fit (trend line) is a straight line drawn to represent the general trend of data in a scatter plot with a linear association.
  • 7The equation of the line of best fit (y = mx + b) can be used to make predictions within the range of the observed data.
  • 8The slope of the line of best fit describes the rate of change, and the y-intercept describes the predicted value when the independent variable is zero.

Worked examples

Example 1

The table below shows the number of hours students spent studying for a math test and their scores on the test. Create a scatter plot for this data, describe the association, and identify any outliers or clusters.\n\n| Hours Studied (x) | Test Score (y) |\n|--------------------|----------------|\n| 2 | 65 |\n| 4 | 78 |\n| 3 | 70 |\n| 5 | 85 |\n| 1 | 60 |\n| 6 | 92 |\n| 3 | 75 |\n| 7 | 95 |\n| 2 | 50 |

I**Step 1: Set up the axes.** Draw a horizontal axis for 'Hours Studied' (x) and a vertical axis for 'Test Score' (y). Label the axes clearly and choose an appropriate scale for each, starting from 0 or a relevant minimum value.
II**Step 2: Plot the data points.** For each pair of (Hours Studied, Test Score), plot a point on the graph. For example, for (2, 65), go 2 units right on the x-axis and 65 units up on the y-axis.
III**Step 3: Describe the association.** Observe the general trend of the points. Do they tend to go up, down, or show no clear direction? Do they form a straight line or a curve?
IV**Step 4: Identify outliers or clusters.** Look for any points that are far away from the general pattern (outliers) or groups of points that are close together (clusters).

Answer

The scatter plot shows the data points:\n(2, 65), (4, 78), (3, 70), (5, 85), (1, 60), (6, 92), (3, 75), (7, 95), (2, 50).\n\n**Description of Association:** There is a strong, positive, linear association between the number of hours studied and the test score. As the hours studied increase, the test scores generally tend to increase.\n\n**Outliers/Clusters:** The point (2, 50) appears to be an outlier, as it is significantly lower than other scores for a similar number of hours studied. There are no obvious distinct clusters, but the data points generally cluster around the increasing trend.

When describing association, always mention direction (positive/negative), strength (strong/weak), and form (linear/non-linear).

Example 2

Using the scatter plot from the previous example (Hours Studied vs. Test Score):\n1. Draw an informal line of best fit.\n2. Write the equation of your line of best fit.\n3. Use your equation to predict the test score for a student who studies for 4.5 hours.

I**Step 1: Draw an informal line of best fit.** On the scatter plot, draw a straight line that visually represents the general trend of the data. Try to have roughly an equal number of points above and below the line, and make sure it follows the direction of the data.
II**Step 2: Find two points on your line of best fit.** Choose two points that lie directly on the line you drew (they don't have to be original data points). For example, let's say our line passes through (2, 65) and (6, 90).
III**Step 3: Calculate the slope (m) of the line.** Use the formula m = (y2 - y1) / (x2 - x1).\n m = (90 - 65) / (6 - 2) = 25 / 4 = 6.25
IV**Step 4: Find the y-intercept (b) of the line.** Use the slope (m) and one of the points (x, y) from your line in the equation y = mx + b.\n Using (2, 65):\n 65 = 6.25(2) + b\n 65 = 12.5 + b\n b = 65 - 12.5\n b = 52.5
V**Step 5: Write the equation of the line of best fit.** Substitute the calculated m and b into y = mx + b.\n y = 6.25x + 52.5
VI**Step 6: Predict the test score for 4.5 hours of study.** Substitute x = 4.5 into your equation.\n y = 6.25(4.5) + 52.5\n y = 28.125 + 52.5\n y = 80.625

Answer

1. (A line of best fit would be drawn on the scatter plot, visually representing the trend.)\n2. The equation of the line of best fit is approximately **y = 6.25x + 52.5**.\n3. The predicted test score for a student who studies for 4.5 hours is approximately **80.6** (or 81, rounded to the nearest whole number for a test score).

Your line of best fit and its equation might vary slightly from this example, as fitting is informal. However, it should be close and follow the general trend of the data.

Common mistakes

  • ✗Incorrectly plotting data points on the scatter plot, leading to misinterpretations of the association.
  • ✗Confusing positive association with negative association, or incorrectly identifying 'no association'.
  • ✗Drawing a line of best fit that does not accurately represent the general trend of the data, or drawing it too steeply/flatly.
  • ✗Assuming that a strong association shown in a scatter plot automatically implies causation (that one variable causes the other). Correlation does not imply causation.
  • ✗Extrapolating (making predictions far outside the range of the observed x-values) using the line of best fit, which can lead to unreliable predictions.

Exam tips

  • ★Always label your axes clearly and choose appropriate scales when creating a scatter plot.
  • ★Use a ruler to draw your line of best fit; it should be a straight line that visually balances the points above and below it.
  • ★When asked to describe an association, be specific: mention if it's positive or negative, linear or non-linear, and if it appears strong or weak.
  • ★If asked to find the equation of a line of best fit, choose two points *on your drawn line* (not necessarily original data points) to calculate the slope and y-intercept.

Ready to practise?

Try a problem on this topic

Snap a photo or type a question — get step-by-step working instantly.