Course Evaluation Questions and Survey Template
Sixteen questions, weighted toward how the course was built rather than who stood at the front. Take it here, copy it, or open it in the editor. Then the harder half: reading the result honestly.
- 16questions
- About 4 minto complete
- 5 namedanswer categories
- 3 figuresthe honest reading
What a course evaluation score is, and why it is not an average
Three figures, not one: the spread, the count, and what share of the class answered. Leave the average out of it.
The answers are categories, not quantities. A rating scale orders answers from worst to best, but the numbers on it are labels. Averaging them treats a label as a measurement, so two answers of "neither agree nor disagree" come out level with one "strongly agree" and one "strongly disagree".[2] That is why every agreement item here offers strongly agree, agree, neither agree nor disagree, disagree and strongly disagree instead of a row of numbers: there is nothing to average by accident.
Report the distribution, the count and the response rate. That is the recommendation this page is built on: do not average scores or compare averages, report the percentage of answers in each category, how many people responded, and what share of the class that was.[2]
The one summary added here. Where a single figure is unavoidable, use the share of responders in the top two categories and print the count beside it. That top-two share is our summary of a distribution, not something the source recommends, and it is honest only while the count travels with it.
Do not set it beside a department average. An instructor at 4.2 against a departmental 4.5 is below average, and how bad depends on the spread behind the 4.5: if colleagues score sixes half the time and threes the rest, 4.2 sits well inside it.[2]
Check the unit first. A course has a syllabus, a sequence and an end date. If what finished was one session or workshop, you want the post training survey questions. If the thing being judged is several courses run toward an outcome, that is a program evaluation questionnaire.
The 16 course evaluation questions
Every item, the categories it offers, and what a low answer points at.
The course as it was built
Six items, and where most of the reading comes from. In one large Swedish study of more than 6,000 evaluations, course design predicted perceived quality more strongly than the teachers did.[7]
- The course objectives were clear from the start.A low answer makes every later answer harder to read.
- The topics were sequenced in an order that helped me learn.Cheap to change between cohorts, expensive to leave wrong.
- The course materials were worth the time they took.Separates a reading list that is too long from one that is the wrong list.
- The assessments tested what the course actually taught.Read it before any complaint about marks. A gap here is a design fault, not a marking one.
- In an average week, how many hours did this course take outside class?Asked as hours, not as an agreement statement, so it can be set against the workload the course was designed to carry.
- The course description matched the course I actually took.Low answers belong to whoever wrote the handbook entry, not to whoever taught the class.
How the course ran
Delivery, asked as attributes of the course rather than a verdict on a person. A survey whose subject is a named individual is a different instrument with a different confidentiality problem.
- Explanations in class were clear enough to follow.Read against question 11: clear but too fast needs a different fix from unclear.
- There was enough time in class to ask about things I did not understand.Cheap to fix, and it points at the shape of the session rather than the material.
- Feedback on my work came back in time to be useful.Sounds factual and is not. In an identity-swap experiment the same grading turnaround was rated 4.35 out of 5 under a male name and 3.55 under a female one.[4]
- I could find the course materials whenever I needed them.If this is your lowest item, the complaint is the platform rather than the course.
- The pace of the course was one I could keep up with.Read the spread. Too fast for a third of the room and too slow for another third leaves no useful centre.
What people leave with
Two self-reports, treated as such, then the item that decides how much weight the rest of somebody's answers carry.
- I can do something now that I could not do at the start of this course.A self-rating. It is not a test result and this page never treats it as one.
- What this course taught connects to work I expect to do later.Read it beside the open boxes, where the reason for a low answer is often written out.
- How much of the course did you attend or work through?Asked near the end so it does not colour the answers above it. Somebody who did half the course is answering about half a course.
In their own words
The shortest route on the form to something you can act on. Students are well placed to judge some parts of teaching, and comments carry what five categories throw away.[2]
- What was the single most useful part of this course?Asking for one thing rather than a list is what makes the answers countable into themes.
- What one change would improve this course for the next group?The only item that produces something to do without any arithmetic at all.
| Item | The answer categories |
|---|---|
| The twelve agreement items | Strongly agree; agree; neither agree nor disagree; disagree; strongly disagree |
| Hours the course took outside class | Under 2; 2 to 4; 5 to 8; 9 to 12; more than 12 |
| How much of the course you did | Almost all of it; most of it; about half; less than half; very little |
All sixteen are built into the survey above, so you can create a questionnaire online from them and change the wording first. Keep the categories if you plan to run the form again: a changed label ends one series and starts another.
Two things are deliberately missing. No overall-effectiveness item, because the source this page leans on hardest recommends dropping omnibus questions about overall teaching effectiveness and the value of a course, on the grounds that they mislead.[2] And no demographic block: age, gender, subject and course code together identify one student in a class of twenty.
When the survey really is about one named instructor rather than the course, the facilitator feedback survey template is built for it.
Turn your own answers into a reading
Enter the counts from one item and the size of the class.
What a missing answer costs you. Suppose half a class answers and all of them put an item in the bottom category. The figure for the whole class sits between that and the top of the scale, because the half who said nothing could have gone either way.[2] The first box prints that range.
How many answers you need. On deliberately generous assumptions, a 10 per cent sampling error, an even yes or no split and an 80 per cent confidence level rather than the usual 95, the requirement looks like this.[6] Conventionally it is far higher: the best rate reported for paper surveys, 65 per cent, is only adequate once a class passes roughly 500 students.[6]
| Students on the roll | Responses needed | Which is a response rate of |
|---|---|---|
| 20 | 12 | 58 per cent |
| 30 | 14 | 48 per cent |
| 50 | 17 | 35 per cent |
| 100 | 21 | 21 per cent |
Nulty's own caution travels with it: the formula assumes random sampling, which course evaluations do not meet.[6] It is a floor to clear, not a licence to publish. A small class also swings further at full response, because averages of small samples move with the luck of the draw.[2] Under about twenty on the roll, print the counts and skip the summary.
What the rating can and cannot decide
Where its own number stops, and who says so.
It is not a measure of learning. Across the studies comparing sections of one course sitting a common exam, ratings account for up to one per cent of the variance in what students learned, and once prior ability is controlled for the relationship is not distinguishable from zero.[1] The authors put it plainly: students do not learn more from professors with higher ratings.
It is not a forecast. At one United States institution where students were randomly assigned and follow-on courses compulsory, ratings predicted achievement in the course being rated and predicted it poorly in later ones. In mathematics, the instructors whose students did best at the time were the ones whose students did worst afterwards.[5]
A full response rate does not make it neutral. At a French university where students are randomly assigned to seminars, cannot get a transcript without completing the evaluation, and sit a common anonymously-marked exam, ratings moved with the seminar grade and were largely uncorrelated with the exam grade.[3] Male students there rated male instructors higher on preparation and organisation, materials and the usefulness of feedback.[3] A response rate fixes who is represented, not what is measured.
An item that looks factual can move on the name at the top. In an online course where two assistant instructors each taught under their own identity and the other's, the perceived male identity scored higher on all twelve measures, six significantly.[4] That is the study behind the figure at question 9. One small course is not the world, but it is enough to stop anyone reading a decimal place as a fact.
What it is good for. Mostly the course. The answers point hardest at objectives, sequence, materials and assessment, which is the part you can change before the next cohort.[7] If they say the wrong course was scheduled at all, that is a training needs analysis. To measure the change rather than describe it, take a baseline first with the training presurvey.
Where this page stops. None of this settles what an institution should do about promotion or tenure. What it supports is narrower: read the answers as evidence about the course, publish the spread and the response rate beside any figure, and never let one number stand in for a judgement about a person.[1][5]
Six decisions that belong to a course evaluation
What every survey needs is in the survey research guide. These six belong to this one.
Freeze the wording, then run it every term
Identical items are the only thing that makes one term comparable with the next. Whatever is specific to a course goes in the open boxes.
Publish the window with the survey
Say when it opens, when it closes, and that it closes before marks are released. A window nobody knew about is the cheapest missing answer there is.
Record the response rate every time
Write it beside the result, in the same sentence. A figure quoted without it cannot be checked by whoever reads it.
Publish counts per category, not the mean
A bar for each of the five carries everything a mean does and none of the false precision.
Refuse the departmental comparison
The comparison people reach for first is the one the evidence supports least. Compare a course against its own last run.
Say who reads the answers, before anyone answers
Whoever owns this survey sees each individual response, so tell people whose job that is when you ask. In a small group, add that a comment can identify whoever wrote it.
Course evaluation questions FAQ
How many questions should a course evaluation have?
Sixteen. Published lists run from eight items to more than eighty, and the long ones are menus rather than instruments: pick ten from a bank of sixty and you have written a survey nobody can compare with anything, including your own from last year.
What is the difference between a course evaluation and a training feedback form?
The unit being judged. A training feedback form is filled in after one session or workshop and asks whether that event worked. A course evaluation runs at the end of a course with a syllabus, and produces a rating somebody other than the teacher will read and act on.
Should a course evaluation be anonymous?
Say who reads the answers before anyone answers, and make sure that person is not marking the work. Class size does more than any promise: in a small group a comment can identify its author however the form is labelled, and students in small classes may already feel their anonymity is tenuous.[2]
Are student evaluations of teaching reliable?
They reliably measure something. What they do not measure is learning: ratings account for up to one per cent of the variance in what students learned.[1] Read them as evidence about the course, never about how much anyone learned.
What should an open ended course evaluation question ask for?
One thing, not everything. "What would you change?" invites a paragraph that says nothing. "What one change would improve this course for the next group?" invites a sentence you can count.
A course has a syllabus and runs over weeks. One talk on one afternoon has neither, so the seminar survey is a shorter instrument with a different last question.
Methods and sources
What this template is based on
The method for reading the result comes from an evaluation of course evaluations by two University of California statisticians: report the distribution, the number of responders and the response rate, and do not average or compare averages.[2] The response-rate table is Nulty's, under his liberal assumptions with his own caution attached.[6] The weighting toward course design follows a 2024 study of more than 6,000 evaluations.[7]
Question licensing
Every question here is original and free to use with this editor or without it. No licensed instrument is reproduced, adapted or imitated, and no item is copied from any published form.
Limits and disclosure
The evidence is contested and this page takes a narrow position rather than staging the debate. The findings cited come from multi-section validity studies[1], one French university[3], one United States academy[5] and one small online course[4], each described with its setting attached, and the 2024 design finding is stated at the level of its abstract.[7] Nothing here is advice about promotion or tenure.
SuperSurvey makes the editor this template opens in. The questions are free to use with it or without it.
References
- Uttl, B., White, C. A., and Wong Gonzalez, D. "Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related." Studies in Educational Evaluation 54 (2017): 22-42. home.miracosta.edu (accepted manuscript, PDF)
- Stark, P. B., and Freishtat, R. "An Evaluation of Course Evaluations." ScienceOpen Research, 26 September 2014. stat.berkeley.edu (PDF)
- Boring, A. "Gender biases in student evaluations of teaching." Journal of Public Economics 145 (2017): 27-41. unlv.edu (PDF)
- MacNell, L., Driscoll, A., and Hunt, A. N. "What's in a Name: Exposing Gender Bias in Student Ratings of Teaching." Innovative Higher Education 40, no. 4 (2015): 291-303. Figures quoted here are from the authors' own open-access presentation of the same experiment in the Journal of Collective Bargaining in the Academy. thekeep.eiu.edu (PDF)
- Carrell, S. E., and West, J. E. "Does Professor Quality Matter? Evidence from Random Assignment of Students to Professors." Journal of Political Economy 118, no. 3 (2010): 409-432. Quoted from the NBER working paper version. nber.org (PDF)
- Nulty, D. D. "The adequacy of response rates to online and paper surveys: what can be done?" Assessment and Evaluation in Higher Education 33, no. 3 (2008): 301-314. uaf.edu (PDF)
- Levinsson, H., Nilsson, A., Martensson, K., and Persson, S. D. "Course design as a stronger predictor of student evaluation of quality and student engagement than teacher ratings." Higher Education 88 (2024): 1997-2013. link.springer.com
Send it at the end of the course, and record what came back
Sixteen questions, about four minutes, and no number line to average.
Use this template