It's About The Journey Not The Destination: AI in Education Edition

28   |   By Leah Suchet

The practice of classifying students according to performance in graded examinations dates back to the early 1700s, spreading relatively quickly and becoming mainstream in schools and universities by the turn of the century1. However, criticism of this system has been around from the beginning, with one of the earliest critiques originating from the former president of Yale University in 1786, who felt that the fear induced by the pressure of examinations and grading resulted in poor scholarship ((Stiles 1786), quoted in1. Since then, mounting evidence has emerged showing the negative effects this has on students both in terms of performance and long-term emotional health. Interestingly, it has been said that very little academic literature exists in favour of grading. The system was developed with practical implementation at the forefront; the industrial revolution drove the need for scalable education, and the increase in class sizes made it unsustainable to rely on small-group models of teaching and assessment2. Using the resources available, these systems had to be clearly defined and, to try to achieve consistency across institutions, standardised. This is a very logical progression and one in which many intelligent, motivated individuals have committed significant efforts to making as good as possible. In this piece I explore how the unregulated use of generative AI challenges the current output-based assessment model and then imagine how agentic AI could truly shake up the education system: personalising some aspects of teaching, freeing up teachers to make small-group teaching achievable at scale and shifting from primarily grade-based assessments to personalised feedback. In this, I will unpack some of the practical considerations and risks relating to such a revolution in the education system.

In a snapshot of literature taken from3, the following core arguments against grading were extracted, relating to motivation456, emotional and social well-being4, collaboration8, and long-term effects and shaping ultimate social value910. On the latter point, it is interesting to clearly see the link between the learned meritocratic hierarchy of grading in school and an acceptance of a perceived meritocratic hierarchy within society, whereby performance in school leads ultimately to high or low social value. However, with society being far from a pure meritocracy, this can cause a blinding to and acceptance of systemic inequalities, hidden beneath the perception of meritocracy.

As someone with limited knowledge of the types of research practices needed to perform studies like this, I would caveat this with the inherent variability in such research and the danger of biasing the selection of literature to make an intended point.

It seems to me that for a long time, the education system has been perceived as flawed but the best that we have. And I think I would accept that statement, given that the resource constraints that lead to the system's creation are largely still in place -- as the movement of students between institutions is only increasing with globalisation, standardisation is important to manage entry and try to ensure fairness in that process. Furthermore, having sufficient teachers for widespread small-group teaching has not been a feasible pursuit, especially in developing countries.

However, the emergence of generative AI represents a uniquely significant challenge to the education system. While I don't think it challenges the concept of grading directly, it does ask whether examining output is fit for purpose to measure and assess the learning process. Output quality is now less an indication of an individual's capacity to understand content, and more an assessment of that individual's ability to interact with an AI model, as well as that model's ability to generate output. Of course, this disproportionately impacts assignment and essay-based assessments over traditional closed-book examinations but considering it will likely affect the entire learning process prior to examination, the effects may still carry through even to closed-book assessment. Furthermore, as with the use of calculators that was originally banned but has now become widespread, there is logic to the argument of learning to use the tools that will realistically be available in the 'real world', post-education. Reliance on calculators meant we lost the ability to do mental maths. Reliance on cell phones has meant a diminished ability to recall simple facts; directions; phone number. How much will we lose from a reliance on generative AI? Taking a pessimistic view, could this result in a new type of standardisation, with all the outputs will begin to look and sounds the same as individual styles, tones and creativity is lost? Or, considering another of the other key roles that grading plays in society (providing entry criteria and deciding who attends which institution), will this make it harder for programmes to select candidates based on target skills/abilities? This feels like it may blur the natural 'sorting' of people into the various stratifications of society, which, on some level, may be a good thing considering the imperfections and inequalities in the current system, but on the other hand, which may cause chaos and harm (at least in the short term) should this current 'order' begin to fall apart. Whether this intensifies or diffuses inequality of opportunity is an interesting question... I suppose the digital divide: accessibility to devices and this technology, education around the use of these tools etc., will perpetuate existing inequalities.

Thinking from a skills perspective, those that students will likely develop less are critical thinking, writing, editing, summarising key information and extracting key points. These are the equivalents of the 'mental maths' loss from the invention of the calculator. Of course, we have made huge leaps in the mathematical sciences, not least leading to the computational algorithms underlying computers and the algorithms behind AI, so there are of course benefits and perhaps they will outweigh the cost of what is lost. But I am straying into an assessment that is beyond the scope of this piece, and far beyond my knowledge/expertise.

What I do think, is that as generative AI challenges educators to adjust their output-based assessments, simultaneously, agentic AI is emerging and offering, for the first time, a feasible alternative to the large-group, standardised teaching model that has had to persist for so long. Finally, there is a scalable solution that can offer personalised teaching (and assessment), and that, in theory, does not have to be reserved for the wealthiest echelons of society who can afford schools able to provide such one-on-one tutelage. It seems to me that the weaving of agentic AI into the teaching process can help with both the the risks of generative AI and the traditional criticisms of the traditional grading system.

What I am imagining is a personalised agentic interface which can respond to an individual student's interests and progress, tailoring content as needed to maximise engagement and enjoyment (linked clearly to learning retention), showing links to the real world based on existing educational content videos/resources available online. Like any formal system, it will need to have clear learning outcomes, linked to clear metrics to measure progress toward these outcomes. While this sounds similar to the output-based assessments of the past, there will be the unique ability to track the quality of student's input prompts, and a tailoring of the returned content, which can monitor improvements in the quality of questions asked, and can prevent a reliance on AI to simply 'generate an answer'. This can be imagined as an assessment of the quality of the student's response to an AI generated prompt, flipping the current usage model on its head. By learning to recognise the individual's natural language style, these personalised agents could also likely detect inputs that may have been generated by external generative AI platforms in attempts to 'cheat' the system.

Imagining how this could link to/merge into/replace the existing grade-based assessment is an interesting thought... The commonly cited three core functions of grading are (1) an orientation and report function; (2) a pedagogical function; and (3) a selection, ranking, and entitlement function (cited in Nöel Rohde's book1 from Jörg Ziegenspeck's 'Grades and Reports in School' (1973)). It seems that the personalised AI tutor could perform functions 1 and 2 relatively easily in terms of providing feedback to the student about their performance during the learning process, and also to the learning coordinator and parents through trends in the student's performance over time. I would ideally imagine this system operating in conjunction with a human teacher, who could provide support on the interaction with the agent, tracking the emotional well-being of the student and facilitating practical activities, which should make up a core part of the learning process.

However, for the third 'selection, ranking, and entitlement function', this may require a shift in the current linear system from school -> university -> job, with ages tied to each educational stage. A danger of a fully personalised educational journey is the de-synchronising of age and the loss of peer-groups. This could also be a cause of competition and bullying between children if there is a perception of being 'behind' or 'ahead'. If these were to be no standardised examinations marking the end of distinct stages in the schooling process, one can imagine young adults entering the workplace at different times, depending on how quickly they learned which materials. Another question would be how much the students get to decide the direction of the studies, versus sometimes having to endure things they do not enjoy in that moment. If this is never enforced, they may end up pigeonholing into a single career path/direction from too young an age, which is already in issue in the current system, for example, the choice of subjects at AS/A-levels. As an employer, I think your biggest question is who will fill your role, demonstrating the required skills (technical and soft) and who matches your company culture. I suppose the personalised AI tutor could provide information on the student to the employer, but this does raise concerns over personal data protection and may make the process more stressful for students if they feel they must always present as 'perfect' a side as possible to their tutor knowing that all their behaviour is being recorded and may come back to haunt them...

Other concerns raised by this hypothetical situation is the lack of unity and shared experiences that the students would go through together. As I work through this scenario, I am inclined to think the structure and milestones of grades and the age-based cohort still has value and should be preserved, if possible. This means that within the personalised landscape of the AI tutor, there should be clear learning outcomes per term/year, and, by facilitating curiosity students who progress faster through materials should be kept stimulated without leaping so far ahead, perhaps, that they 'leave' the year group completely. And likewise, where possible those who may learn concepts slower should be supported to hit the learning outcomes as defined and to cover the required material with their peers. I think this is also important for groupwork, which, as much as it is challenging, simply is more representative of the real world and of most working environments. If collaboration were to be lost, that would be a major flaw in a new system. I believe that the models of hackathons can be quite successful here, bringing in a competitive element, but linking problem solving to real-world challenges, fostering interdisciplinary teams, requiring creativity and exposing students to a wide variety of skillset and industries.

Ultimately, what this all points to is a hybrid model, combining the structure of the existing system with the personalised teaching element of agentic AI for much of the content, and freeing up teachers to engage in distinct, small-group teaching sessions to foster collaborative learning, group innovation and giving that sense of shared learning experience. These sessions harken back to the approach followed before numerical grading took over; whereby small group teaching allowed for assessments based on group participation to gauge understanding and proficiency as opposed to standardised written examinations. It is also the model followed by undergraduate tutelage in institutions like Oxford and Cambridge, lauded for their small-group supervision system that prevents students from 'falling through the cracks' and is geared toward mandating engagement, rewarding curiosity and maximising each individual's performance, enjoyment and ultimately, mastery of the subject matter.

In closing, I believe that the considered integration of agentic AI can enable the goals of education to be met without the limitations of the traditional standardised system. Moreover, generative AI is challenging the validity of output-based numerical grading, and reform is required to ensure we can adequately assess the learning process itself, over the quality of output alone. I would like to convey an optimistic take on the future of education; one where the system can finally get a long-awaited revamp hopefully leading to better outcomes for students and society. A system where tailored experiences bring out the best in each student and help them find learning fun. Where collaboration and teamwork are encouraged, with links to real-world applications driving engagement and translation outside of the school environment. Where the journey is truly valued as much as, or even more, than the destination.

Note: these are the ramblings of someone who has always been interested in the education system but with no more qualification to write about it than having been through it. I imagine nothing in this paper is particularly novel to the field itself, so this is really more of a meandering for me to learn more about the history & limitations of the current system and, as I'm sure so many have done before, try to imagine the future with radical reforms. Furthermore, I acknowledge that it is not nearly as well-researched as it should be, so please, take what I say which a pinch of salt and if you have any relevant resources that you think I should read, please do pass them on!

Recommended reading: Nöel Rohde's "Marked: School Grades and the Quantified Life" is a fascinating comprehensive analysis of this topic, and well worth a read.

Limitations: I intended to consider PhD studies within this scope but ended up being taken down a route more focused on primary/secondary/undergraduate education. Due to the relatively unstructured nature of a PhD, it does not feel appropriate to discuss in the current framework and will require further thought.

References

  1. Noëlle Rohde. 1. The global triumph of grades. In: Marked: School Grades and the Quantified Life [Internet]. 2026 [cited 2026 Jul 12]. p. 9--20. Available from: https://www.jstor.org/content/oa_chapter_monograph/jj.43516545.10?seq=3
  2. Sabuj Ahmed. The History of Grading Systems: A ResearchGate Article. ResearchGate [Internet]. 2024 Dec 16 [cited 2026 Jul 11]. Available from: https://www.researchgate.net/publication/387090374_The_History_of_Grading_Systems_A_ResearchGate_Article
  3. Chris McNutt. Human Restoration Project [Internet]. 2022 [cited 2026 Jul 12]. A Brief History of Grades and Gradeless Learning. Available from: https://www.humanrestorationproject.org/writing/a-brief-history-of-grades-and-gradeless-learning/
  4. Butler R. Task-involving and ego-involving properties of evaluation: Effects of different feedback conditions on motivational perceptions, interest, and performance. Journal of Educational Psychology. 1987;79(4):474--82. doi:10.1037/0022-0663.79.4.474
  5. Carrell PL, Monroe LB. Learning Styles and Composition. The Modern Language Journal. 1993;77(2):148--62. doi:10.1111/j.1540-4781.1993.tb01958.x
  6. Lipnevich AA, Smith JK. Response to Assessment Feedback: The Effects of Grades, Praise, and Source of Information. ETS Research Report Series. 2008;2008(1):i--57. doi:10.1002/j.2333-8504.2008.tb02116.x
  7. Beck HP, Rorrer-Woody S, Pierce LG. The Relations of Learning and Grade Orientations to Academic Performance. Teaching of Psychology. 1991 Feb 1;18(1):35--7. doi:10.1207/s15328023top1801_10
  8. Hayek AS, Toma C, Oberlé D, Butera F. Grading Hampers Cooperative Information Sharing in Group Problem Solving. Social Psychology. 2015 May 27;46(3):121--31. doi:10.1027/1864-9335/a000232
  9. Brookhart SM. Determinants of Student Effort on Schoolwork and School-Based Achievement. The Journal of Educational Research. 1998 Mar 1;91(4):201--8. doi:10.1080/00220679809597544
  10. Poorthuis AMG, Juvonen J, Thomaes S, Denissen JJA, Orobio de Castro B, van Aken MAG. Do grades shape students' school engagement? The psychological consequences of report card grades at the beginning of secondary school. Journal of Educational Psychology. 2015;107(3):842--54. doi:10.1037/edu0000002

← Are We the Martians Under the Bed?Xenagogy →