Intelligenza Artificiale per la valutazione linguistica nell’ambito degli esami di certificazione Cambridge

Caso studio del B1 Preliminary

Autori

  • Luca Di Stefano Università degli studi di Bari Aldo Moro

DOI:

https://doi.org/10.15162/1970-1861/2661

Parole chiave:

Intelligenza Artificiale; Decodifica e Valutazione Linguistica; Cambridge Assessment English; Productive Skills; Human-in-the-loop

Abstract

Cambridge Assessment English sets global standards in language tests, strictly tying exams to
the CEFR framework, yet the rapid rise of AI poses new paths for test grading. While existing
literature covers proprietary automated assessment systems, there is a gap regarding the
interactional evaluation limits of off-the-shelf generative AIs. This study aims to fill this gap by exploring how commercial AIs, specifically NotebookLM and Gemini Pro, assess productive skills, comparing their outputs to human examiners’ comments in a blind study of ten B1-level Italian learners. Data suggest AI limitations in the tested models when grading candidates: in the Writing test, they demand higher-level structures, failing to stick to strict CEFR frameworks. Challenges emerge in the Speaking test, where the selected AIs relied on text transcriptions, struggling to process prosodic features like pitch contour or speech rate, missing key facts of
human interaction, and prioritizing pure grammar over it. All factors considered, AI is a highly efficient tool to track class progress regularly, but high-stakes formal exams must adhere to a human-in-the-loop methodology, for only real experts can currently grasp the deepest pragmatic
traits of speech and fix AI’s “blindness”, ensuring a fair and valid test.

Riferimenti bibliografici

Alderson J. C., Clapham C., Wall D., 1995, Language test construction and evaluation, Cambridge, Cambridge University Press.

Bachman L. F., 1990, Fundamental Considerations in Language Testing, Oxford, Oxford University Press.

Bachman L. F., Palmer A. S., 1996, Language Testing in Practice: Designing and Developing Useful Language Tests, Oxford, Oxford University Press.

Barni M., 2005, “La valutazione delle competenze linguistico-comunicative in L2”, in Vedovelli M. (a cura di), Manuale della certificazione dell’italiano L2, Roma, Carocci Editore, pp. 29-46.

Brown T. B., Mann B., Ryder N. et al., 2020, “Language Models are Few-Shot Learners”, in Larochelle H., Ranzato M., Hadsell R. (a cura di), NIPS ‘20: Proceedings of the 34th International Conference on Neural Information Processing Systems, New York, Curran Associates Inc., pp. 1877 - 1901, https://dl.acm.org/doi/abs/10.5555/3495724.3495883, consultato il: 10/05/2026.

Cambridge University Press & Assessment, 2024, B1 Preliminary Handbook for teachers for exams, Cambridge, CUP&A, https://res.cloudinary.com/swiss-exams/image/upload/v1711528439/168150_b1_preliminary_teachers_handbook_6d2aceb282.pdf, consultato il: 09/05/2026.

Cambridge University Press & Assessment, 2025, VOCABULARY LIST, B1 Preliminary, B1 Preliminary for Schools, Cambridge, CUP&A, https://www.cambridgeenglish.org/Images/506887-b1-preliminary-vocabulary-list.pdf, consultato il: 10/05/2026.

Chen J., Guo Z., Chun J. et al., 2026, “Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance”, in Demberg V., Inui K., Marquez L. (a cura di), Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics, Rabat, Association for Computational Linguistics, pp. 5848-5877, https://aclanthology.org/2026.eacl-long.274/, consultato il: 08/05/2026.

Cinganotto L., Montanucci G., 2025, Intelligenza artificiale per l’educazione linguistica, Torino, UTET Università.

Council of Europe, 2001, Common European Framework of Reference for Languages: Learning, Teaching, Assessment, Strasburgo, Council of Europe, https://rm.coe.int/common-european-framework-of-reference-for-languages-learning-teaching/16802fc1bf, consultato il: 09/05/2026.

Eckes T., 2012, “Operational Rater Types in Writing Assessment: Linking Rater Cognition to Rater Behavior”, in Iwashita N., Yu G., Language Assessment Quarterly, 9(3), Abingdon, Routledge, pp. 270-292.

Fulcher G., 2010, Practical Language Testing, London, Routledge.

Khalifa H., Weir C. J., 2009, Volume 29 - Examining Reading: Research and practice in assessing second language reading. Studies in Language Testing (SiLT), Cambridge, Cambridge University Press & Assessment, https://www.cambridgeenglish.org/Images/735100-studies-in-language-testing-volume-29.pdf, consultato il: 07/05/2026.

Kim E., Li S., Khalil S. et al., 2025, “STAIR-AIG: Optimizing the Automated Item Generation Process through Human-AI Collaboration for Critical Thinking Assessment”, in Kochmar E., Alhafni B., Bexte M. et al., Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications, Vienna, Association for Computational Linguistics, pp. 920-930, https://aclanthology.org/2025.bea-1/, consultato il: 08/05/2026.

Li H., He Z. X., Tian S. et al., 2026, “Martingale Foresight Sampling: A Principled Approach to Inference-Time LLM Decoding”, in Demberg V., Inui K., Marquez L., Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics. Volume 1: Long Papers, Rabat, Association for Computational Linguistics, pp. 3522-3533, https://aclanthology.org/2026.eacl-long.162/, consultato il: 10/05/2026.

McKnight S. W., Civeleko˘glu A., Gales M. J. F. et al., 2023, “Automatic Assessment of Conversational Speaking Tests”, in ISCA (a cura di), 9th Workshop on Speech and Language Technology in Education (SLaTE), Dublin, International Speech Communication Association, pp. 99-103, https://doi.org/10.21437/slate.2023-19, consultato il: 08/05/2026.

McNamara T. F., 2000, Language Testing, Oxford, Oxford University Press.

Nitti P. (a cura di), 2026, “AI-Generated texts: issues, challenges and operational proposals”, in Quaderni della Rassegna, 272, Firenze, Franco Cesati Editore.

North B., Goodier T., Piccardo E., 2021, Quadro comune europeo di riferimento per le lingue: apprendimento, insegnamento, valutazione. Volume complementare. Language Policy Programme, Education Policy Division, Education Department, Council of Europe. Italiano LinguaDue 12(2), Milano, Council of Europe - Milano University Press, https://doi.org/10.13130/2037-3597/15120, consultato il: 07/05/2026.

Taylor L., 2009, “Developing Assessment Literacy”, in Spolsky B. (a cura di), Annual Review of Applied Linguistics, 29, Cambridge, Cambridge University Press, pp. 21-36.

Tsvilodub P., Klumpp J.-F., Mohammadpour A., 2026, On Emergent Social World Models - Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language Models, (preprint), arXiv - Cornell University, https://doi.org/10.48550/arXiv.2602.10298, consultato il: 09/05/2026.

Weigle S. C., 2013, “English language learners and automated scoring of essays: Critical considerations” in Elliot N., Williamson D. M. (a cura di), Automated Assessment of Writing, 18(1), Amsterdam, Elsevier, pp. 85-99.

Weir C. J., 2005, Language Testing and Validation. An Evidence-Based Approach, Basingstoke, Palgrave Macmillan.

White J., Fu Q., Hays S. et al., 2023, “A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT”, in Vranić V., Brown K., Yoder J., et al. (a cura di), PLoP ‘23: Proceedings of the 30th Conference on Pattern Languages of Programs, Monticello, IL, USA, The Hillside Group, Articolo 5, pp. 1 - 31, https://dl.acm.org/doi/10.5555/3721041.3721046, consultato il: 10/05/2026.

Yu G., Xu J. (a cura di), 2024, Volume 52 - Language Test Validation in a Digital Age. Studies in Language Testing (SiLT), Cambridge, Cambridge University Press & Assessment, https://www.cambridgeenglish.org/Images/735163-studies-in-language-testing-volume-52.pdf, consultato il: 07/05/2026.

Zechner K., Higgins D., Xi X., et al., 2009, “Automatic scoring of non-native spontaneous speech in tests of spoken English”, in Eskenazi M., Alwan A., Strik H. (a cura di), Spoken Language Technology for Education: Spoken Language, 51(10), Amsterdam, Elsevier, pp. 883-895.

Pubblicato

2026-09-07 — Aggiornato il 2026-09-08

Fascicolo

Sezione

Articoli