Objectives: Large language models (LLMs) have emerged as promising tools for generating educational assessment items; however, their ability to produce multiple-choice questions (MCQs) that accurately reflect intended cognitive levels remains uncertain. This study evaluated the educational quality and cognitive alignment of head and neck anatomy MCQs generated by GPT-5 (Generative Pre-trained Transformer 5) using Bloom’s Revised Taxonomy.Methods: Fifty head and neck anatomy MCQs were generated by GPT-5 using structured prompts specifying the anatomical topic, the intended Bloom level, and undergraduate learning objectives.
Two experienced anatomists independently evaluated each question for clarity, scientific accuracy, curriculum relevance, and suitability for its intended Bloom level using a five-point Likert scale. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs), and content validity was evaluated using item- and scale-level content validity indices (I-CVI and S-CVI).
To examine cognitive alignment independently, all MCQs subsequently underwent a blinded consensus Bloom-level classification, in which two anatomists independently assigned Bloom levels before resolving disagreements through discussion. Agreement between the intended and consensus-assigned Bloom levels was assessed using Cohen’s κ and weighted κ statistics.Results: GPT-5-generated MCQs received high pooled ratings for clarity (4.42±1.02), scientific accuracy (4.55±0.97), curriculum relevance (4.51±1.11), and suitability for the intended Bloom level (4.29±1.11).
Inter-rater agreement ranged from good to excellent across all domains (ICC=0.717–0.945). Content validity was high, with S-CVI/Ave values ranging from 0.82 to 0.89. Blinded consensus Bloom classification showed 92.0% exact agreement between the intended and expert-assigned Bloom levels (Cohen’s κ=0.900; linear weighted κ=0.913; quadratic weighted κ=0.926).
All four discrepancies represented reassignment to a lower cognitive level; two involved adjacent Bloom categories and two spanned more than one category.Conclusion: GPT-5 generated head and neck anatomy MCQs that experts considered clear, scientifically accurate, educationally relevant, and generally appropriate for their intended cognitive level.
The high level of agreement observed during blinded consensus Bloom classification is consistent with structured prompting largely preserving intended cognitive objectives, although this preservation was not uniform at higher cognitive levels. Expert review therefore remains essential, particularly for optimising higher-order cognitive complexity and distractor quality before incorporation into undergraduate assessments.