Scatter plots, explanatory and response variables and the statistical investigation process
What this note covers
- Where bivariate statistics sits in the Stage 2 course
- The statistical investigation process
- Explanatory and response variables — and why you must decide first
- Constructing a scatter plot that earns full marks
- Describing an association: direction, form and strength
- Reading a scatter plot for outliers, clusters and gaps
- A worked example from data table to description
- How this is examined in the SACE paper — and what separates a top answer
8 sections · 12 key terms & formulas · 6 common mistakes
1. Where bivariate statistics sits in the Stage 2 course
Your external examination is built from exactly three topics: Topic 3 Statistical models, Topic 4 Financial models and Topic 5 Discrete models. Topics 1, 2 and 6 are school-assessed only and never appear on the paper. Within Topic 3, Subtopic 3.1 Bivariate statistics is the larger of the two subtopics, and it supplies roughly half of the statistical marks in a typical paper.
Bivariate means two variables measured on the same individuals. Every question in this subtopic starts from a set of paired data — a table of years against populations, hours studied against test scores, engine size against fuel consumption, rainfall against crop yield. Your job across the subtopic is a single connected chain: identify which variable explains which, plot the pairs, describe what you see, quantify the association with r and r², fit a model with technology, check the model with residuals, then use the model to predict and comment honestly on how much you trust the prediction.
This note is the first link in that chain. It looks unglamorous — axes, labels, direction, form, strength — but examiners award real marks for it every year, and those marks are the easiest in the whole paper to bank. A scatter plot part is usually 2 to 3 marks; a description of association is usually 2 to 3 marks; both are routinely lost by students who can operate a graphics calculator perfectly but cannot write a precise sentence about what the plot shows.
Two habits matter from the very first line of a statistics question. First, read the context sentence and name the two variables with their units before you touch the calculator. Second, keep the variable names in every later answer. The single biggest separator between a mid-range and a top response in Topic 3 is whether the student writes about 'the number of hours of training' or just about 'x'.
2. The statistical investigation process
The subject outline frames all of Subtopic 3.1 inside the statistical investigation process. It is the cycle a statistician actually works through, and the outline expects you to recognise which stage a question is asking about. The stages run:
- Pose the question. A statistical question is one that data can answer and that names the variables — not 'is exercise good for you?' but 'is there an association between the number of hours of weekly exercise and resting heart rate in adults aged 18 to 25?'
- Collect or obtain the data. Where did these numbers come from? A survey, a census, an experiment, a published data set? How the data were collected controls what conclusions are legitimate later.
- Display and analyse the data. The scatter plot comes first, then the numerical summaries: r, r², the least squares line, the residual plot.
- Interpret the results and communicate them. Translate the numbers back into the context, state the limitations, and answer the original question.
Why does this matter for an examination that mostly asks you to calculate? Because the interpretation and communication marks — which is where Reasoning and Communication is assessed — are marks for closing the loop. When a question ends with 'Comment on the reliability of this prediction' or 'Discuss whether the model is appropriate', it is asking you to step back into the final stage of the cycle, not to do more arithmetic.
A practical consequence: never end a Topic 3 question with a bare number. If a question gives you a context, the last line of your answer should mention that context. A prediction of 43.7 is not an answer; 'the model predicts approximately 43.7 litres of fuel used for a 520 km trip' is.
The cycle also explains why the outline stresses the source of the data. If a data set covers only Adelaide households in one year, any conclusion is about Adelaide households in that year. Saying so in one clause is often worth a mark in a 'Discuss' or 'Explain' part.
3. Explanatory and response variables — and why you must decide first
The outline uses the pair independent (explanatory) variable and dependent (response) variable. The explanatory variable is the one you think does the explaining, or the one that is set or observed first; the response variable is the one that responds to it, or that you would want to predict.
The convention that governs everything downstream is:
- The explanatory variable goes on the horizontal axis and is called x.
- The response variable goes on the vertical axis and is called y.
Getting this backwards does not change r or r² — they are symmetric in the two variables — but it does produce a completely different least squares line, a different slope, a different intercept and wrong predictions. Because the regression parts of a question follow on from the plotting part, one reversal at the start can cost marks in three or four later parts.
How to decide when the question does not tell you outright:
- Time is almost always explanatory. Year, age, number of months since launch, hours elapsed — these go on the horizontal axis. Nothing responds by changing the year.
- Follow the prediction. If a later part asks you to predict the mass of a seedling from the amount of fertiliser, then fertiliser is explanatory and mass is the response.
- Follow the plausible mechanism. Engine capacity can plausibly affect fuel consumption; fuel consumption does not resize the engine.
- Follow the wording. Phrases such as 'the effect of A on B', 'B depends on A', 'use A to estimate B' all make A explanatory.
Where the paper hands you the roles directly — 'the number of visits (x) and the total spend (y)' — accept them without argument, even if you would have chosen differently. Marks follow the question's assignment, not yours.
Write the choice down explicitly if the question asks. A one-line answer such as 'Explanatory variable: number of hours of training per week; response variable: 5 km run time in minutes' earns the mark cleanly.
4. Constructing a scatter plot that earns full marks
When the paper asks you to draw or complete a scatter plot, it supplies grid space and often partly drawn axes. A full-mark plot has five features, and it is worth checking them like a list.
- Correct axis assignment. Explanatory horizontal, response vertical.
- Labelled axes including units. 'Distance travelled (km)', not 'distance'. Marks are lost for missing units more often than for misplotted points.
- An appropriate, uniform scale. Each axis must have equal intervals for equal steps. Choose a scale that spreads the data across most of the grid — a plot squeezed into the bottom-left corner makes the form impossible to describe and can cost the description mark that follows.
- Every data point plotted accurately with a small cross or dot. Count your points against the table before moving on; a missing point is a common and avoidable loss.
- No joining of the points and no freehand curve. A scatter plot is a cloud of points, not a line graph. Drawing a connecting zig-zag is treated as a misunderstanding of the display.
On axis scaling, one subtlety appears regularly. If the data start a long way from zero — years 2015 to 2024, say, or heights from 155 cm to 190 cm — you do not have to begin the axis at zero, and starting at a convenient value such as 150 cm makes a better plot. If you do this, mark the break or simply start the axis label at that value; do not leave equal gaps that imply the axis runs from zero.
Where the plot is already drawn and you are asked to complete it, read the existing scale carefully before adding points. The examiner has usually chosen a scale where each small square is 2 or 5 or 0.5 units, and misreading it produces points that sit visibly off the pattern.
Finally, keep the plot. Later parts of the same question often ask you to identify an outlier visually, sketch the least squares line, or explain why a linear model is or is not appropriate — all of which are answered from the plot you have just drawn.
5. Describing an association: direction, form and strength
The outline is explicit that an association is described by direction, form and strength. When a question says 'Describe the association between …', it is asking for all three, and the mark allocation usually tells you so — a 3-mark description part is three components, not one long sentence about one of them.
Direction is positive or negative. Positive: as the explanatory variable increases, the response variable tends to increase (the cloud rises left to right). Negative: as the explanatory variable increases, the response tends to decrease. The outline allows direction to be judged from the scatter plot and/or from the sign of r.
Form is linear or non-linear. Linear means the points cluster about a straight line. Non-linear means they cluster about a curve — for this course, most often a curve that rises ever more steeply or falls away towards a floor, which points to an exponential model. Say which one; 'the form is linear' is a complete answer to that component.
Strength is strong, moderate or weak, and describes how tightly the points hug the underlying pattern. The outline allows strength to be judged with r and/or r². A widely used classroom guide in Australian senior general mathematics treats |r| of about 0.75 or more as strong, roughly 0.5 to 0.75 as moderate and roughly 0.25 to 0.5 as weak, but the safe examination habit is to state the descriptor and quote the statistic that justifies it.
Put together, a full-mark description reads like this: There is a strong, negative, linear association between the age of the car in years and its resale value in dollars (r = −0.912). That single sentence contains strength, direction, form, both variables in context and the supporting statistic.
Two refinements that mark out a top answer. First, name the variables rather than x and y. Second, use tentative language — 'tends to decrease', 'is associated with' — rather than causal language such as 'causes' or 'makes'. The distinction is worth marks in its own right, and it is developed further when correlation and causation are examined directly.
6. Reading a scatter plot for outliers, clusters and gaps
Beyond the three-part description, a scatter plot carries information that the summary statistics hide. Examiners ask about these features because they are exactly what a statistician would notice.
Outliers. The outline is specific that outliers in bivariate data are identified visually from the scatter plot. There is no formula to apply here. An outlier is a point that sits well away from the pattern formed by the rest of the data. Note that this is not the same as being extreme in one variable: a point can have an ordinary x value and an ordinary y value and still be a clear outlier because that combination is far off the trend. When you identify one, describe it by its coordinates and in context: 'the point (12, 38) is an outlier — the 12-year-old car sold for far more than the trend predicts, possibly because it is a restored or collectable model.'
Clusters. Sometimes the cloud separates into two or more groups. This usually signals a hidden third variable dividing the data — petrol versus diesel vehicles, metropolitan versus regional stores, two different production lines. Pointing this out is a strong observation in a 'Comment' part, because it warns that fitting a single line to the whole set may be misleading.
Gaps and limited range. Look at where the data actually live along the horizontal axis. If every observation has x between 5 and 40, then the model you fit is only supported over that interval. That observation is the seed of every later reliability comment, so make a mental note of the minimum and maximum explanatory values as soon as you have plotted them.
A short discipline that pays: after plotting, spend ten seconds saying to yourself, in words, what the plot shows — direction, form, strength, anything unusual. If a later part asks 'Explain why the linear model may not be suitable', you will already have the answer from the picture rather than having to hunt for it under time pressure.
7. A worked example from data table to description
The following is an original example written to the style of the paper; it is not a past examination question.
A gardening supplier records, for ten plots, the amount of fertiliser applied and the mass of tomatoes harvested.
| Fertiliser (kg) | 1.0 | 1.5 | 2.0 | 2.5 | 3.0 | 3.5 | 4.0 | 4.5 | 5.0 | 5.5 |
| Yield (kg) | 8.2 | 10.6 | 12.1 | 14.8 | 15.9 | 18.4 | 19.6 | 22.1 | 23.4 | 25.8 |
Step 1 — assign the variables. The supplier controls how much fertiliser is applied and wants to predict yield from it, so fertiliser applied (kg) is the explanatory variable on the horizontal axis and yield (kg) is the response variable on the vertical axis.
Step 2 — choose scales. Fertiliser runs from 1.0 to 5.5, so a horizontal axis from 0 to 6 in steps of 0.5 fits comfortably. Yield runs from 8.2 to 25.8, so a vertical axis from 0 to 30 in steps of 5, with each small square worth 1, uses the grid well.
Step 3 — plot and label. Ten crosses, axes labelled 'Fertiliser applied (kg)' and 'Yield (kg)', no joining lines.
Step 4 — describe. The points rise steadily from lower left to upper right and lie very close to a straight line, with no point standing away from the others. A full description: There is a strong, positive, linear association between the mass of fertiliser applied and the tomato yield. If the question allows technology, add the supporting statistic: r = 0.998, confirming a very strong positive linear association.
Step 5 — note the range. The data cover fertiliser amounts from 1.0 kg to 5.5 kg only. Any later prediction outside that interval is extrapolation and must be flagged as such — a point worth remembering when the question moves on to the regression line.
8. How this is examined in the SACE paper — and what separates a top answer
The Stage 2 General Mathematics examination is a single question booklet of eight or nine compulsory multi-part questions totalling 90 marks in 130 minutes, with no sections and no multiple choice. Statistical, financial and discrete questions are interleaved through the paper, so a Topic 3 question can appear anywhere. There is no formula sheet; you may take one unfolded A4 sheet of handwritten notes (both sides) and your approved graphics calculator.
Bivariate statistics almost always opens with a context paragraph and a data table, then works through parts in the order of the investigation process. The plotting and describing parts are typically the first 4 to 6 marks of that question, and they are the marks most often left on the table by capable students.
What examiners reward here:
- Precision in the description. Three components — strength, direction, form — plus the variables in context. A one-word answer such as 'positive' cannot score a 3-mark description part.
- Correct axis assignment, carried consistently into the regression parts that follow.
- Labels with units on both axes.
- Statistical rather than causal language. 'As hours of training increase, run time tends to decrease' scores; 'training makes you faster' invites a lost mark.
- Answers pitched to the mark allocation. Two marks means two distinct statements, not one statement written twice.
What separates a top answer: it treats the plot as evidence rather than decoration. A top student names the outlier by coordinates and offers a plausible contextual reason for it; notes that the data span only a limited range of the explanatory variable and flags what that will mean for prediction later; and links each descriptive claim to the feature of the plot or the statistic that supports it. Command words matter — 'State' wants a short direct answer, 'Describe' wants the three components, 'Explain' and 'Discuss' want a reason attached to every claim. Reading the command word and the mark allocation together, before writing, is the cheapest technique available in this topic.
20 full-length practice exams with worked solutions, 20 revision notes, 64 practice questions and 200 flashcards.
Unlock General Mathematics — $20
Preview a sample note and question free on the SACE General Mathematics hub →