How We Design ESAT-Style Questions
Past papers are useful, but they are not a complete map of the ESAT. They show what has been asked before—not everything the specification allows an examiner to ask.
That distinction shapes how we build our question bank. If we only imitate the most common past-paper formats, the bank gets larger without becoming more useful. Students get quicker at recognising familiar questions, but they are still exposed when a known topic appears in an unfamiliar form. We want to provide both: enough repetition to build speed and enough variety to require real decisions.
Past papers are evidence, not a blueprint
We begin with the official specification because it defines what is examinable. We then compare it with older ESAT, ENGAA and NSAA questions, notes from people who sat the exams, and the questions already in our bank. Each source answers a different question: what can be tested, what has been tested, what felt difficult under time pressure, and where our own coverage is thin or repetitive.
This gives us a coverage map. Some specification points appear regularly, some only once or twice, and some have not appeared in the available papers at all. Historical frequency helps us decide how much practice to write, but it does not decide what belongs in the bank. A topic does not stop being examinable simply because it has not appeared recently.
The common material gets enough questions for students to become fluent. Less familiar material gets enough variation to show how the same idea can be used in different settings. Unseen but examinable material gets careful treatment so that practice does not turn into a prediction game.
Start with what the question should reveal
Before writing a stem, we answer a simple question: what should a student have to notice or decide here? Sometimes the answer is a piece of algebra. More often it is choosing the right model, identifying a constraint, connecting two basic ideas or resisting a tempting shortcut.
The brief records that reasoning step, the relevant specification point, a realistic completion time, the mistakes we expect, and what the solution should teach. “Make a hard physics question” is not a useful brief. “Make the student choose a model before calculating” gives the writer something concrete to build and the reviewer something concrete to test.
Turn the brief into a fair question
A strong draft usually has one main insight and a solution that can be explained without hand-waving. We try not to create difficulty by piling on arithmetic or hiding the task behind awkward prose. If a question is slow only because it is tedious, we shorten it or rethink it.
The wrong answers matter too. A distractor should come from a mistake a capable student might genuinely make: using the wrong relationship, overlooking a condition, confusing two quantities or following an attractive shortcut too far. Random wrong numbers tell us very little. A good distractor helps us understand the student’s reasoning.
We also remove details that do not earn their place. Extra information can be useful when the point is to decide what matters, but it should not be there merely to make the question look sophisticated.
What we automate—and what we do not
We use AI to compare wording, format approved material and flag inconsistencies between a stem, its options and its solution. Scripts handle jobs such as taxonomy, identifiers, file conversion, image processing, imports and rendering previews. This saves time on work that is repetitive and easy to check.
The important decisions remain human. A person chooses the purpose of the question. Reviewers solve it independently, look for alternative interpretations, test whether the distractors are credible and decide whether the difficulty feels appropriate. A generated draft can be a useful starting point, but it is never evidence that a question is ready.
Review the question as a student will see it
Review includes the answer key, possible alternative readings, calculator-free solvability, notation, diagrams, topic labels and the worked solution. We also test the question in the actual student interface. A diagram that is clear on a large monitor may be unreadable on a smaller screen; a line break or cramped option can make a straightforward question unnecessarily confusing.
The question is released only when there is one defensible answer, a clean route to it and a clear reason for every important design choice. If a reviewer cannot explain what an element is doing, we revise or remove it.
Publication is not the end of the process
Once students begin using a question, we can compare our intentions with what actually happens. Timing data and reports return to the editorial queue. A slow question might be doing exactly what we intended, or it might contain unclear wording, a poor diagram or too much mechanical work. The number alone cannot tell us which, so we read the question again and investigate.
When we change a question, we update its solution at the same time and record the reason. The solution should do more than justify the correct option. It should show a method or decision that the student can carry into a different problem.
That is the whole cycle: use the specification to find the gap, decide what reasoning to test, write a fair question, review it carefully, and learn from how students respond. The number of questions in the bank matters less than the range of thinking those questions actually develop.



