Sitemap

Teaching Scores, Storied Averages, and Small Classes

6 min readJul 1, 2025
Press enter or click to view image in full size
Image of an intricate bar chart showing student evaluation scores across five years.
From a blog post by my colleague Catherine Rawn, where she analyzes her student evaluation scores.

University administrations are very keen on student evaluations. Behind the scenes, they are a managerial lever in employment decisions and if an instructors' evaluation scores are low their job can be on the line. More front-facing, the managerial use of these scores is presented as an incentive: listen to students when they score your teaching, know that the results are relevant in your hiring and promotion, and — voilà — your teaching will improve. Even better, through systematic use of these scores instruction will become more excellent across the institution. It appears, however, that these processes might not actually work like this. It's possible, even, that most of the time they do not work at all.

In 2018, in a conflict between Toronto Metropolitan University's administration and faculty association, an arbitrator wrote a wave-making decision regarding the use of student evaluations in decisions of hiring and promotion. The arbitrator noted that neither side in the dispute challenged the expert views which asserted that student evaluations are not a measure of teaching effectiveness. It was a big deal to hear this point made so clearly. While the faculty association argued that making personnel decisions reliant on averaged student evaluations was actively harmful and possibly a human rights violation, the university defended these averages as useful information that "raised flags" which then require "further investigation." The arbitrator was unequivocal — these scores are biased in ways that are nearly impossible to adjust for; wide variation in course delivery creates score differences that have nothing to do with teaching quality; the use of score averages so as to compare people, departments, faculties is untenable and must stop.

This demand to stop relying on average scores when making hiring and promotion decisions tries to counter an old problem: when students are asked to provide numerical answers to qualitative questions, those numbers can then be used for mathematical maneuvers which are not necessarily warranted by the questions asked. At the time, the arbitrator's report recommended that instead of student evaluation scores the university rely on teaching dossiers and peer reviews of teaching in assessing instructors' teaching quality. Student evaluations "have a role to play" in such dossiers, says the report. They are valuable as a "source of information that comes directly from students." They "tell a story" but this story "needs to be carefully contextualized." I'd like to tell you a story here about pitfalls of such contextualization.

Where I work, we are encouraged to tell stories about our own and about others' student evaluations. We can use our teaching dossier — submitted in renewal and promotion processes — for such story-telling. Low teaching evaluation scores in one course? Come up with a story that explains them. High averages in one year but low ones in another? Recount some context about what happened that year. Generally lower scores than other instructors in your department? Maybe your courses are different or incur more bias — narrate away. You'll notice it's only low scores — noticeable through simple numerical comparisons — that require such stories. High scores, presumably, speak for themselves. These low-score stories are made up under pressure, selectively, and after the fact. The conscientious researcher's soul cringes at this approach. If it's not you inventing these stories on your own behalf, then those who review your file might be asked to create some narratives for you. Speculation ensues.

Because the issue is about numbers, I am prepared to give you some numbers. Imagine you are teaching first-year writing courses, like I do, with lower course caps than the many lectures which first-year students otherwise attend. Perhaps there are 70 students in a cohort program in which you teach, and they are divided across 3 sections of the same course. These students have joined the program through a selective entry process, they come with similar levels of skill, they take nearly the same roster of courses and get to know each other well, they are in it together for the full year. You are teaching these students in your three writing courses in the same way, it makes no sense to do it otherwise — same syllabus, same assessment approach, same assignment instructions, same lesson plans, same you. You'll adjust your teaching to the different dynamics in each group, of course— working with their strengths and supporting them in their difficulties — but you also ensure they get a similar learning experience across these sections. I phrase this scenario hypothetically, but it's my real case.

At the end of each term, teaching evaluations come in, averaged for each course section. And they look a little wild! The average score in response to the one, single question that your administration pays all attention to reads, let's say, 3.2 (out of 5) for one section, 4.0 for another, 4.8 for the third. The next term — again, same cohort of students, course taught the same way in each section— the numbers are 3.4 and 4.3 and 4.6 between sections. Your scores are automatically supplied to the committees that evaluate your file, and the committees will surely notice that these numbers appear to be all over the place. What kind of inconsistent teacher are you? The committee will need an explanation. What's your story? How can this veering and teetering be explained? And don't say you were drunk.

Remember, these are smallish classes (even if not nearly small enough for a writing course). Response rates for these evaluations are a general concern and they are especially low and unreliable across sections of smaller courses. But these response rates would not be as uneven if you took the numbers together for the same sections— like would be the case for your colleagues who teach these students in first-year lectures with all 70 or so cohort members in one room. In that case, you'd have a much calmer looking 4.0 for one term and 4.1 the next. Such numbers would have been quite all right. No extra explanations would have been needed. But given that you teach smallish sections you fret each year about these fluctuating numbers and how they look in your file; you worry about how to craft your teaching dossier to provide a story that helps smooth the numerical waves. Others are recruited into this process, too. Can your department committee sufficiently explain these numbers in their report? And if the next higher committee isn't satisfied and asks the lower committee for more narrative, what will they be able to say?

In the effort of contextualizing these numbers, goal-oriented analysis and speculative story-telling is asked of you and others in our department. It is extra work which, without the flagging concern of your fluctuating scores, would not need to be expended. The university's overall goal with this exercise is to produce excellent teaching. The collective agreement lists these components of teaching excellence: "command over subject matter, familiarity with recent developments in the field, preparedness, presentation, accessibility to students and influence on the intellectual and scholarly development of students." Clearly, whatever the student evaluation scores may indicate in this case, they do not provide reliable data on these criteria. Remember, the course is the exact same across each term's sections and designed for a known cohort of students: same command of material, same familiarity with recent research, same preparedness, same presentation, same accessibility to students. Yet students choose differently in attaching numbers to this sameness. You are the same teacher showcasing the same level of excellence (whatever it may be) for each group of students. They are their own people with their own thoughts and feelings about numbers.

In this case, the need to contextualize differences in scores seems to produce story-making work that bears no relation to pedagogical improvement. Rather, it causes ongoing friction which, in this case, falls repeatedly on the shoulders of writing instructors who teach many sections of first-year writing courses. Courses which, because of the intense labour they require, are capped at smaller numbers than other courses, but which in Canada are still not nearly small enough.