- Home
- About AALHE
- Board of Directors
- Committees
- Guiding Documents
- Legal Information
- Organizational Chart
- Our Institutional Partners
- Membership Benefits
- Member Spotlight
- Contact Us
- Member Home
- Symposium
- Annual Conference
- Resources
- Publications
- Donate
EMERGING DIALOGUES IN ASSESSMENTAligning AI and Human Judgment: How Prompt Design Shapes AI-Assisted Scoring of Ethical Reasoning
August 25, 2026
Abstract: This article presents a case study examining how prompt design influences the ability of large language models to score student ethical reasoning essays in alignment with trained human raters. Findings suggest that structured prompting—particularly strategies that include self-verification—improves score consistency. Practical guidance is provided for higher education assessment practitioners who are exploring AI-assisted scoring for complex constructs.
Aligning AI and Human Judgment: How Prompt Design Shapes AI-Assisted Scoring of Ethical ReasoningAssessment practitioners in higher education face a persistent tension: the constructs we value most are often the most labor-intensive to assess. Writing-based assessments, widely regarded as the gold standard for capturing complex reasoning, require significant investments in rater training, calibration, and ongoing quality monitoring. As institutional demands for accountability and scalable evidence of student learning intensify, many assessment offices are asking a practical question: Can artificial intelligence help reduce the burden of scoring while maintaining the quality and interpretive value of assessment results? The emergence of large language models (LLMs) has renewed interest in automated scoring. Yet for complex, process-oriented constructs like ethical reasoning, the key question is not simply whether AI can generate scores, but whether those scores are meaningful, consistent, and aligned with expert human judgment. This issue becomes even more pressing when the assessment task requires students to demonstrate nuanced reasoning in writing. To explore this challenge, this article presents a case study of how one institution investigated the use of AI to support scoring of student essays on ethical reasoning. Rather than treating AI scoring as a plug-and-play solution, we examined how prompt design shapes scoring quality and concluded by offering practical lessons for assessment practitioners. The Context: Assessing Ethical Reasoning at ScaleAt a public university in the Mid-Atlantic region, ethical reasoning has been assessed through the Ethical Reasoning in Action (ERiA) initiative, which centers on the Eight Key Questions (8KQ) to prompt moral considerations, such as fairness, outcomes, rights, responsibilities, and other ethical dimensions (Sanchez et al., 2017), to help with decision-making. Students write essays about ethical dilemmas they have encountered, and these essays are scored by trained faculty members using a structured rubric that captures multiple aspects of reasoning. The framework emphasizes the quality of reasoning rather than the correctness of conclusions. The ethical reasoning rubric measures five dimensions: (A) identification of the ethical issue, (B) engagement with relevant 8KQs, (C) explanation of their applicability, (D) analytical depth, and (E) justification and weighing of competing considerations. Each dimension is assessed on a five-level developmental scale, ranging from 0 to 4 (Sanchez et al., 2017; Linder et al., 2019). To explore whether AI could serve as a supplementary scoring tool, the Office of ERiA partnered with assessment specialists to investigate LLM-assisted scoring under controlled conditions. What We Tested: AI Prompting ApproachesThis study employed a comparative experimental design to examine how different prompting strategies influenced the AI scoring performance of student essays on ethical reasoning. Microsoft Copilot, based on GPT-5, was used as the scoring platform for all conditions and the same dataset was used throughout the study. Each essay was scored independently under each prompting condition. To evaluate AI scoring performance, we tested four prompting strategies (Sahoo et al., 2023) using a set of 100 student essays that had previously been scored by expert human raters:
This design allowed us to compare how different levels of reasoning structure shaped scoring outcomes while controlling for dataset and model variability. What We Found: Prompting Approaches Shaped Scoring QualityTo examine how closely AI-generated scores aligned with human ratings, five complementary indicators were used: systematic bias, exact agreement, mean absolute error (MAE), tolerance-based agreement (within ±1 point) and quadratic weighted kappa. Although the details of these analyses are not reported here, we found that prompt design had a substantial impact on scoring quality. Consistent differences emerged across prompting strategies. Zero-shot prompting produced the least consistent scoring, with greater variability between how the AI and humans approached the rubric. Few-shot prompting improved the human-AI alignment with evidence of reduced error, but scoring patterns depended on the specific calibration examples provided. ToT prompting further enhanced score stability by requiring AI to consider multiple evaluative perspectives before assigning a score, yet some bias persisted across rubric dimensions. CoVe prompting showed the strongest overall alignment with human raters. Across dimensions, CoVe produced bias values closest to zero, lower average deviation, and higher rates of both exact and near agreement than other strategies. For some rubric dimensions, agreement levels approached those observed between human raters. What This Means: AI Scoring Requires Calibration and ValidationJust as human raters must be calibrated to a rubric during rater training to produce consistent results, LLMs must also be carefully calibrated for scoring purposes. This case study suggests that AI can be used to score ethical reasoning writing under an established rubric. However, scoring outcomes vary substantially depending on how the model is prompted. When tested on the same platform and dataset, different prompting strategies produced different levels of alignment with human raters. Among the conditions examined, CoVe showed the closest and most consistent agreement with trained human scoring. These findings underscore a central insight: prompt design functions much like rater training. When prompts incorporated structured self-verification, alignment with human judgment improves substantially. AI scoring is therefore not a plug-and-play solution, but a design-sensitive process requiring careful development, testing, and validation. Implications for Assessment PractitionersThis study provides practical guidance for educators and assessment practitioners who are considering the use of AI to support scoring complex learning outcomes: Recognize that AI Scoring is not Plug-and-PlayDifferent prompts produced substantially different results, even when using the same LLM model and dataset. This means AI scoring outcomes are not a fixed property, but are instead shaped by how the task is framed and communicated to the system. Add Structure to Improve Scoring ConsistencyPrompts that require explanation, verification, and revision lead to more consistent scoring. Adding structure helps to guide the model’s reasoning process, which appears to reduce reliance on immediate surface-level judgments. Design Prompts as Carefully as Scoring ProtocolsPrompt design should be documented, tested, and refined analogous to how we document rater training strategies. Even small changes in wording or structure can meaningfully impact AI behavior, which makes iteration and standardization essential. Validate AI Scores against Human ScoringAlthough AI scoring technology will likely continue to improve, it is not yet at a point where it can replace human judgment for complex constructs. AI-generated scores should therefore always be compared to human raters before being used in practice to ensure alignment with the intended construct. Use AI as a Supplement—not a ReplacementAI may be most useful for preliminary scoring, formative feedback, or reducing workloads within a human-in-the-loop system. Human oversight remains essential for interpreting complex constructs. Exercise Caution in High-stakes ContextsAI scoring requires extensive validation and ongoing monitoring, particularly if one desires to use such scores for consequential decisions. The consequences of misalignment are amplified in high-stakes settings, and we should elevate our standards of evidence appropriately to account for consequential decisions. Looking AheadAs higher education continues to grapple with the demand for scalable, meaningful assessment, AI-assisted scoring represents a promising frontier—but one that requires the same rigor we apply to any measurement tool. The lesson from this study is clear: when it comes to AI scoring of complex constructs, how we ask matters as much as what we ask. For assessment practitioners, this means investing in prompt design, validation protocols, and careful implementation, while approaching AI as a carefully calibrated partner rather than a replacement for professional judgment.
REFERENCESLinder, G. F., Ames, A., Hawk, W. J., Smith, K. L., Fulcher, K. H., & Sanchez, E. R. H. (2019). Teaching ethical reasoning: Program design and initial outcomes of Ethical Reasoning in Action, a university-wide ethical reasoning program. Teaching Ethics, 19(2), 147–169. http://doi.org/10.5840/tej202081174 Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S. S., & Chadha, A. (2024). A systematic survey of prompt engineering in large language models: Techniques and applications. ArXiv. https://doi.org/10.48550/arXiv.2402.07927 Sanchez, E. R. H., Fulcher, K. H., Smith, K. L., Ames, A., & Hawk, W. J. (2017). Defining, teaching, and assessing ethical reasoning in action. Change: The Magazine of Higher Learning, 49(2), 30–36. https://doi.org/10.1080/00091383.2017.1286215 Xiao, C., Ma, W., Song, Q., Xu, S., Zhang, K., Wang, Y., & Fu, Q. (2024). Human-AI collaborative essay scoring: A dual-process framework with LLMs. In Proceedings of the 15th International Learning Analytics and Knowledge Conference (pp. 293-305). https://doi.org/10.1145/3706468.3706507 |