← Mathematics ApplicationsMathematics ApplicationsLog in

WACE Mathematics Applications ATAR exam: Fri 6 Nov, 9:20am — 27 days away

ATARMAxxing · WACE Mathematics Applications revision notes

The statistical investigation process and describing association in percentaged two-way tables

Statistical investigation process and two-way tables
3 · Topic 3.1: Bivariate data analysis

What this note covers

  1. Where this topic sits, and why the syllabus starts here
  2. Categorical variables, explanatory and response
  3. Constructing a two-way frequency table with row and column sums
  4. Percentaging the table: the decision that determines your whole answer
  5. Reading the percentaged table for evidence of association
  6. Writing the description that earns the marks
  7. A full investigation from question to conclusion
  8. How this is examined in the WACE paper

8 sections · 12 key terms & formulas · 6 common mistakes

Free sample

1. Where this topic sits, and why the syllabus starts here

Topic 3.1 is the largest single block of Unit 3 at 20 hours, and the syllabus is explicit that all of it is to be taught within the framework of the statistical investigation process. That framing matters for the examination: markers are not only checking whether you can compute a percentage or a correlation coefficient, they are checking whether you can move from a real-world problem to a defensible statement about data and back again.

The syllabus glossary describes the statistical investigation process as a cyclical four-step process. Step 1: clarify the problem and formulate one or more questions that can be answered with data. Step 2: design and implement a plan to collect or obtain appropriate data. Step 3: select and apply appropriate graphical or numerical techniques to analyse the data. Step 4: interpret the results of this analysis, relate the interpretation to the original question, and communicate findings in a systematic and concise manner. Dot point 3.1.1 asks you to review that process; dot point 3.1.19 asks you to implement it for questions about two categorical variables or two numerical variables.

Notice the word cyclical. The process does not end at step 4. A finding usually raises a new question, which sends you back to step 1 with a sharper problem. In an extended examination question you may be asked what further data would be needed, or what the next stage of the investigation should be, and the expected answer is a new question that the existing data cannot answer.

The phrase in a systematic and concise manner appears three separate times in Topic 3.1 (dot points 3.1.4, 3.1.16 and 3.1.18). Treat it as a marking instruction rather than as decoration. Systematic means your description follows a fixed structure with the relevant figures quoted. Concise means no padding: a two-mark description does not want a paragraph of introduction. Students who write three sentences of context before the actual comparison routinely run out of time in Section Two.

This note deals with the categorical half of the topic: two-way frequency tables, percentaging them, and describing the association they reveal. Numerical variables, scatterplots and least-squares lines are handled separately, but the reporting discipline you build here is exactly the discipline you will need there.

2. Categorical variables, explanatory and response

A categorical variable places each individual into a category rather than giving it a number: hair type (straight, curly), attitude to a proposal (agree, no opinion, disagree), suburb, brand of phone, whether a student is enrolled in an ATAR course. Some categorical variables have exactly two categories and are called binary or dichotomous; nothing in the analysis changes if there are three or four categories, the table simply gets bigger.

Beware of variables that look numerical but are being used as categories. Year level recorded as 10, 11 or 12 is being used categorically. A survey response recorded on a scale of 1 to 5 is ordinal categorical. If the numbers are labels for groups rather than measurements you can add, the correct display is a two-way table, not a scatterplot. Conversely, a variable such as time in minutes or mass in kilograms is numerical and belongs in a scatterplot, so a question that gives you a numerical variable and asks for a two-way table has usually already banded it for you into categories such as under 30 minutes, 30 to 60 minutes, over 60 minutes.

Dot point 3.1.8 asks you to identify the response variable and the explanatory variable. The syllabus glossary defines the explanatory variable as the variable used to explain or predict a difference in the response variable. In its bread example, time in the oven is the explanatory variable and the temperature of the loaf is the response variable, because it is sensible to say that time spent in the oven explains the temperature reached, and absurd to say it the other way round.

Three practical tests decide the roles. First, ask which variable could plausibly come first in time; a variable measured earlier is almost never the response. Second, read the question stem: wording such as does the mode of travel depend on distance from school names distance as the explanatory variable. Third, ask which direction the sentence makes sense in. You can meaningfully compare the percentage of long-distance students who catch a bus; comparing the percentage of bus users who live a long way away answers a different, usually less useful, question.

Getting these roles the right way round is worth real marks, because they determine which way you percentage the table in the next step. In the examination the roles are frequently implied rather than stated, so decide them deliberately before you touch a number.

3. Constructing a two-way frequency table with row and column sums

Dot point 3.1.2 asks you to construct two-way frequency tables and determine the associated row and column sums and percentages. A two-way frequency table displays the frequency distribution that arises when a group of individuals or objects is categorised according to two criteria at once. The categories of one variable form the rows, the categories of the other form the columns, and each interior cell holds the count of individuals with that combination of categories.

Build the table in a fixed order and it will not go wrong. Draw the grid with one extra row and one extra column for the totals. Label the rows with one variable and the columns with the other, and write the variable name outside the labels so the marker can see which is which. Tally the raw data into the interior cells. Add across each row to get the row sums and down each column to get the column sums. Finally, add the row sums and separately add the column sums; both must give the same grand total, which is your check that nothing has been miscounted or double-counted.

Take a survey of 200 Year 12 students who were classified by how far they live from school and by whether they usually travel to school by public transport.

Distance from schoolUses public transportDoes not use public transportTotal
Within 5 km3090120
More than 5 km522880
Total82118200

The row sums 120 and 80 tell you how many students fall in each distance category; they add to 200. The column sums 82 and 118 tell you how many students use each mode; they also add to 200. Every interior cell counts each student exactly once, which is the defining feature of the display and the reason percentages out of a row or a column are meaningful.

Read carefully whether a question gives you counts or already-percentaged values. If a table is described as showing percentages, the interior entries do not add to the number of people, and multiplying a percentage by the relevant total is how you recover a count. Examination items frequently give a percentaged table plus one total and ask you to reconstruct a missing frequency, which is exactly this step run backwards.

4. Percentaging the table: the decision that determines your whole answer

The syllabus glossary is precise about the vocabulary. If the table is percentaged using row sums, the results are row percentages; if it is percentaged using column sums, they are column percentages. Dot point 3.1.3 then requires an appropriately percentaged table, and the word appropriately is doing all the work.

The rule is short: percentage in the direction of the explanatory variable, then compare in the direction of the response variable. Practically, that means each category of the explanatory variable should have its own set of percentages adding to 100%. If the explanatory variable labels the rows, calculate row percentages so that each row totals 100%. If the explanatory variable labels the columns, calculate column percentages so that each column totals 100%.

In the transport table, distance from school is the explanatory variable and it labels the rows, so row percentages are appropriate. Within 5 km: 30 divided by 120 is 25%, and 90 divided by 120 is 75%. More than 5 km: 52 divided by 80 is 65%, and 28 divided by 80 is 35%.

Distance from schoolUses public transportDoes not use public transportTotal
Within 5 km25%75%100%
More than 5 km65%35%100%

Percentaging the same table by column would produce a valid table that answers a different question. Of the 82 public transport users, 30 divided by 82 is about 37% living within 5 km and 52 divided by 82 is about 63% living further away. That is a description of who the bus users are, not of whether distance changes travel behaviour, and it does not directly address a question phrased as whether mode of travel depends on distance.

Two habits keep this clean under examination pressure. Always write the 100% total in the margin of each percentaged row or column, because a row that does not total 100% (allowing for rounding) is a signal you divided by the wrong denominator. And state the base explicitly when you quote a figure: 65% of students living more than 5 km from school, not simply 65%. A percentage without its base is uninterpretable and marking keys for description questions expect the base to be identifiable.

5. Reading the percentaged table for evidence of association

Association is defined in the syllabus glossary as a general term for the relationship between two or more variables. For categorical variables the question is whether knowing an individual's category on one variable changes what you would expect of the other.

That gives a clean decision rule. Look along one category of the response variable and compare its percentage across the categories of the explanatory variable. If those percentages are similar, there is no evidence of an association: knowing the explanatory category tells you nothing extra. If those percentages are noticeably different, there is evidence of an association, and the size of the difference is your evidence for how strong it is.

In the transport table, 25% of students living within 5 km use public transport compared with 65% of students living more than 5 km away. That is a difference of 40 percentage points, which is large, so there is clear evidence of an association between distance from school and use of public transport.

Consider what the absence of an association would look like. If 44% of the nearby students and 43% of the distant students used public transport, the row percentages would be almost identical, both close to the overall figure of 41% (82 out of 200). You would then report that the percentages are similar across the two distance categories, so the data suggest no association between distance and mode of travel.

Two refinements matter. First, work in percentage points when quoting a difference: the gap between 25% and 65% is 40 percentage points, not 40%, and describing it as a percentage increase would need the calculation 65 divided by 25, giving a 160% increase, which is a different claim. Second, be cautious with very small category totals. A category containing 6 people produces percentages that jump by 17 points if a single person changes category, so a difference based on tiny row sums is weak evidence. If the question gives you the frequencies, glance at the row sums before you commit to the word strong.

The syllabus stops here for categorical variables. There is no chi-squared test, no formal significance test and no numerical measure of the strength of a categorical association in this course. Your evidence is the difference in percentages, described in words.

6. Writing the description that earns the marks

Dot point 3.1.4 requires you to describe an association in terms of differences observed in percentages across categories, in a systematic and concise manner, and to interpret it in the context of the data. That single sentence contains the whole marking key, so build your answer from its parts.

A reliable four-part template covers it. One: name the two variables. Two: quote at least two comparable percentages, each with its base. Three: state the direction of the difference in plain language. Four: conclude that this does or does not suggest an association, naming both variables again in context.

Applied to the transport data: The percentage of students who use public transport to get to school was much higher for those living more than 5 km from school (65%) than for those living within 5 km (25%). This suggests there is an association between distance from school and use of public transport. Two sentences, both figures quoted, both bases identified, a conclusion in context. That is a complete answer, and it takes under a minute to write.

Common ways to lose marks here are all avoidable. Quoting only one percentage gives the marker nothing to compare, so no comparison mark can be awarded. Quoting raw frequencies instead of percentages (52 students versus 30 students) is not a valid comparison, because the two distance groups contain different numbers of students; that is precisely why the table is percentaged in the first place. Describing the pattern without naming the variables (the second group is higher) fails the requirement to interpret in context.

Watch the strength of the verb you choose. If the difference is 40 percentage points, words such as much higher or considerably higher are justified. If the difference is 4 percentage points, write that the percentages are similar and that the data provide little evidence of an association. Do not use the language of causation: an association in a two-way table does not license the claim that distance causes students to catch a bus, and marking keys penalise causal wording where only association has been shown.

Finally, if a question asks you to interpret rather than describe, it wants what the pattern means for the people or context involved, not a restatement of the arithmetic. Add a clause such as students living further away are far more likely to depend on public transport, which is relevant when planning bus services.

7. A full investigation from question to conclusion

Run the whole process once, because Section Two questions increasingly wrap several dot points into a single scenario.

Step 1, the problem and the question. A school wants to know whether it should extend its bus service. The statistical question must be answerable with data: Is there an association between how far a Year 12 student lives from school and whether that student travels to school by public transport? Note what makes this a statistical question rather than a general one; it names two variables and a population, so data can settle it.

Step 2, obtaining the data. The school surveys all 200 Year 12 students, recording distance from school in two bands and mode of travel. Because the school collected these data itself for this purpose, they are primary data; had the school taken them from a Department of Transport report they would be secondary data. The syllabus mentions both in dot point 3.1.8, and you may be asked which you have been given.

Step 3, analysing. Construct the two-way frequency table with row and column sums, identify distance as the explanatory variable, and percentage by row so each distance category totals 100%. This gives 25% and 75% for students within 5 km, and 65% and 35% for students more than 5 km away.

Step 4, interpreting and communicating. Report that 65% of students living more than 5 km from school use public transport compared with 25% of those living within 5 km, a difference of 40 percentage points, which suggests an association between distance and mode of travel. Then relate that back to the original problem: demand for public transport is concentrated among students living further out, so an extended service would be most valuable on longer routes.

Then close the cycle honestly. The survey covers only Year 12 at one school, so the finding cannot be generalised to other year levels or other schools. Distance was banded coarsely into two categories, so a finer banding might reveal where the change occurs. And the association observed does not establish that distance is the cause; some other variable, such as whether a family has a second car, may be involved. Sentences of that kind are what separate a top answer on the final part of a long question.

8. How this is examined in the WACE paper

Two-way table work sits comfortably in Section One (Calculator-free), because the syllabus design brief restricts that section to content and procedures that can reasonably be completed without a calculator and without time-consuming calculations. Expect friendly denominators: totals of 20, 25, 40, 50, 80, 100 or 200, so that a division such as 30 out of 120 resolves to a clean 25%. Practise these divisions mentally, since no calculator and no notes are permitted in Section One. In Section Two (Calculator-assumed) the same skills appear inside a longer, multi-part scenario with less friendly numbers, often as the opening parts of a question that later moves on to a scatterplot or a least-squares line.

Command words tell you what to produce. Construct means draw the table with labelled rows, columns and totals. Determine or calculate means produce a number, with the division shown. Describe means apply the comparison template with figures quoted. Explain or justify means give the reason behind a claim, typically why one percentaging direction is appropriate or why an association does not establish causation. Interpret means say what a value or pattern means for the context.

The working rule printed on every paper applies here: for any question or part question worth more than two marks, valid working or justification is required to receive full marks, and incorrect answers given without supporting reasoning cannot be allocated any marks. So write the division 52/80 before writing 65%. If your arithmetic slips, the visible method still attracts method marks; a bare wrong percentage attracts none.

What separates a top answer is discipline rather than difficulty. Top scripts state which variable is explanatory before percentaging, percentage in that direction, quote both comparable percentages with their bases, use percentage points for the difference, choose a strength word that matches the size of that difference, name both variables in the conclusion, and stop. They avoid causal language unless the question specifically asks about causation, and they never substitute raw frequencies for percentages when the group sizes differ.

Two habits are worth building now. First, before writing anything, underline the two variables in the stem and mark which one is explanatory; a mis-percentaged table is a whole-question error, not a one-mark slip. Second, check the mark allocation and match your length to it: one mark buys one figure, two marks buy a comparison with figures, three marks buy a comparison plus an interpretation in context.

Included in the WACE Mathematics Applications Mastery Pack

20 full-length practice exams with worked solutions, 20 revision notes, 64 practice questions and 200 flashcards.

Unlock Mathematics Applications — $20

Preview a sample note and question free on the WACE Mathematics Applications hub →

WACE Mathematics Applications · revision note 1 of 20

Keep going