Governing Generative AI in Educational Assessment: Stakeholder Perspectives on Validity, Fairness, and Integrity in Turkey’s High-Stakes Testing
Keywords:
Artificial intelligence, generative AI, educational measurement, high-stakes assessment, responsible AI governance, validity and fairness, reflexive thematic analysis, TurkeyAbstract
The fast proliferation of artificial intelligence (AI) specifically generative AI is transforming the measurement of education in various aspects, such as the development of items, the administration of the test, scoring, security, and reporting. Meanwhile, AI poses novel threats to validity, fairness, integrity, privacy, and societal trust- in high stakes settings, where the outcomes of the assessment are associated with consequential decisions. This qualitative research observes the ways the major educational measurement players in Turkey perceive the opportunities and threats of AI and what they regard as the governance conditions under which they feel responsibly adopt AI in assessment. The study employs a multi-stage qualitative design combining the analysis of documents and semi-structured interviews with experts using their reflexive thematic analysis, which produces explanatory thematic structure. The results have yielded five connected themes, namely purpose-first governance of AI in assessment, (2) validity provided by AI based on construct integrity, evidentiary expectation, and documentation, (3) fairness by continuous monitoring to address drift and subgroup effects, (4) integrity and security issues in the generative AI age necessitating redefined inference rules and assessment jobs, and (5) legitimacy, compliance with privacy, and institutional preparedness as pre-requisites to sustainable implementation. The paper concludes that the Turkish assessment systems that can be responsibly AI-powered involve risk tiering within contexts of use, having clear accountability, documentation that is auditory, ongoing monitoring of fairness, integrity-by-design solutions, and privacy-sensitive governance. The results are used to develop Turkey-specific standards of Responsible AI and monitoring indicators of defensible, equitable, and trustworthy AI-enabled assessment.
References
American Educational Research Association. (2014). Standards for educational and psychological testing. American Educational Research Association.
Aruğaslan, E. (2025). Artificial intelligence-based proctored online exams: A study on the experiences of distance education students. Journal of Buca Faculty of Education, 65, 2728–2748. https://doi.org/10.53444/deubefd.1543471
Aydemir, M., Erdogdu, E., & Ucar, H. (2024). Exploring online learners’ perspectives in relation to proctored exams. Turkish Online Journal of Distance Education, 25(4), Article e334. https://doi.org/10.17718/tojde.1459600
Bennett, R. E., & Zhang, M. (2016). Validity and automated scoring. In F. Drasgow (Ed.), Technology and testing: Improving educational and psychological measurement (pp. 142–173). Routledge. https://doi.org/10.4324/9781315871493-8
Bowen, S. A. (2015). Exploring the role of the dominant coalition in creating an ethical culture for internal stakeholders. Public Relations Journal, 9(1), 1–23.
Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa
Braun, V., & Clarke, V. (2019). Reflecting on reflexive thematic analysis. Qualitative Research in Sport, Exercise and Health, 11(4), 589–597. https://doi.org/10.1080/2159676X.2019.1628806
Braun, V., & Clarke, V. (2024). Supporting best practice in reflexive thematic analysis reporting in Palliative Medicine: A review and introduction to RTARG. Palliative Medicine, 38(6), 608–616. https://doi.org/10.1177/02692163241234800
Braun, V., & Clarke, V. (2025). Reporting guidelines for qualitative research: A values-based approach. Qualitative Research in Psychology, 22, 399–438. https://doi.org/10.1080/14780887.2024.2382244
British Educational Research Association. (2024). Ethical guidelines for educational research (5th ed.). British Educational Research Association.
Bulut, O., Beiting-Parrish, M., Casabianca, J. M., Slater, S. C., Jiao, H., Song, D., Ormerod, C. M., Fabiyi, D. G., Ivan, R., Walsh, C., Rios, O., Wilson, J., Yildirim-Erbasli, S. N., Wongvorachan, T., Liu, J. X., Tan, B., & Morilova, P. (2024). The rise of artificial intelligence in educational measurement: Opportunities and ethical challenges. arXiv. https://doi.org/10.59863/MIQL7785
Burstein, J. (2025). The Duolingo English Test responsible AI standards. Duolingo.
Burstein, J., & LaFlair, G. T. (2024). Where assessment validation and responsible AI meet. arXiv. https://doi.org/10.32038/ltrq.2025.50.09
Burstein, J., LaFlair, G. T., Yancey, K., von Davier, A. A., & Dotan, R. (2024). Responsible AI for test equity and quality: The Duolingo English Test as a case study. ArXiv.
Creswell, J. W., & Poth, C. N. (2018). A Book Review: Qualitative Inquiry & Research Design: Choosing Among Five Approaches. SAGE Publications. https://doi.org/10.13187/rjs.2017.1.30
Denzin, N. K. (1978). The research act (2nd ed.). McGraw-Hill.
European Parliament & Council of the European Union. (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act). Official Journal of the European Union. https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
Fereday, J., & Muir-Cochrane, E. (2006). Demonstrating rigor using thematic analysis. International Journal of Qualitative Methods, 5(1), 80–92. https://doi.org/10.1177/160940690600500107
Grant, S., & Khodyakov, D. (2025). Proposal for a critical appraisal tool for Delphi studies. BMJ, 391, Article e084509. https://doi.org/10.1136/bmj-2025-084509
Hasson, F., Keeney, S., & McKenna, H. (2025). Revisiting the Delphi technique. International Journal of Nursing Studies, 168, Article e105119. https://doi.org/10.1016/j.ijnurstu.2025.105119
Ho, A. D. (2022). Specifying the three Ws in educational measurement. Journal of Educational Measurement, 59(4), 418–422. https://doi.org/10.1111/jedm.12355
Ho, A. D. (2024). Artificial intelligence and educational measurement. Journal of Educational and Behavioral Statistics, 49(5), 715–722. https://doi.org/10.3102/10769986241248771
International Test Commission, & Association of Test Publishers. (2025). Guidelines for technology-based assessment (Version 1.1).
Johnson, M. S., & McCaffrey, D. F. (2023). Evaluating fairness of automated scoring. In V. Yaneva & A. von Davier (Eds.), Advancing natural language processing in educational assessment. Routledge.
Kallio, H., Pietilä, A. M., Johnson, M., & Kangasniemi, M. (2016). Systematic methodological review: developing a framework for a qualitative semi‐structured interview guide. Journal of Advanced Nursing, 72(12), 2954–2965. https://doi.org/10.1111/jan.13031
Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
Kitchen, H., Bethell, G., Fordham, E., Henderson, K., & Li, R. R. (2019). Student assessment in Turkey. OECD Publishing. https://doi.org/10.1787/5edc0abe-en
Lane, S. (2023). Validity, fairness, and technology-based assessment. In V. Yaneva & M. von Davier (Eds.), Advancing natural language processing in educational assessment (pp. 127–141). Routledge. https://doi.org/10.4324/9781003278658-11
Lincoln, Y. S., & Guba, E. G. (1985). Naturalistic inquiry. SAGE.
Litman, D., Zhang, H., Correnti, R., Matsumura, L. C., & Wang, E. (2021). Fairness evaluation of automated scoring. In I. Roll et al. (Eds.), AIED 2021 (pp. 255–267). Springer. https://doi.org/10.1007/978-3-030-78292-4_21
Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13–103). Macmillan.
Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO.
Nowell, L. S., Norris, J. M., White, D. E., & Moules, N. J. (2017). Thematic analysis. International Journal of Qualitative Methods, 16, 1–13. https://doi.org/10.1177/1609406917733847
OECD. (2019). Student assessment in Türkiye. OECD Publishing.
Okoli, C., & Pawlowski, S. D. (2004). The Delphi method as a research tool: An example, design considerations and applications. Information & management, 42(1), 15–29. https://doi.org/10.1016/j.im.2003.11.002
Patton, M. Q. (2015). Qualitative research & evaluation methods: Integrating theory and practice (4th ed.). Sage.
Salamanca, S. L. C., Oliveri, M. E., & Zenisky, A. L. (2025). Advancing good practices in a global digital future: ITC/ATP guidelines. International Journal of Testing, 25(2), 194–211. https://doi.org/10.1080/15305058.2025.2490234
Salamanca, Y. C. (2025). AI in educational measurement: A practitioner’s perspective. Journal of Educational Measurement, 62(2), 363–369.
Saldana, J. (2021). The coding manual for qualitative researchers (4th ed.). SAGE.
Saunders, C. (2018). How to teach, lead, and live well: A qualitative in-depth interview study with eight North Carolina teacher-leaders who flourish [Doctoral disseation Columbia University]. ProQuest and Dissertations.
Shaw, S. (2011). Tracing the evolution of validity in educational measurement. Cambridge Assessment.
Shenton, A. K. (2004). Strategies for ensuring trustworthiness in qualitative research projects. Education for Information, 22(2), 63–75. https://doi.org/10.3233/EFI-2004-22201
Squires, A. (2009). Methodological challenges in cross-language qualitative research: A research review. International Journal of Nursing Studies, 46(2), 277–287. https://doi.org/10.1016/j.ijnurstu.2008.08.006
Tabassi, E. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology. https://doi.org/10.6028/nist.ai.100-1
UNESCO. (2023). Guidance for generative AI in education and research. UNESCO.
Williamson, D. M., Xi, X., & Breyer, F. J. (2012). Automated scoring framework. Educational Measurement: Issues and Practice, 31(1), 2–13. https://doi.org/10.1111/j.1745-3992.2011.00223.x
Wutich, A., Beresford, M., & Bernard, H. R. (2024). Sample sizes in qualitative research. International Journal of Qualitative Methods, 23, 1–15 https://doi.org/10.1177/16094069241296206
Young, J. W., So, Y., & Ockey, G. J. (2013). Guidelines for test development. ETS.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Umer Farooq, Ali Asghar, Arslan Javed, Javaid Nasir

This work is licensed under a Creative Commons Attribution 4.0 International License.

