EMERGING DIALOGUES IN ASSESSMENT

Aligning AI and Human Judgment: How Prompt Design Shapes AI-Assisted Scoring of Ethical Reasoning
August 25, 2026
  • Suyu Wang, MA, PhD Student, James Madison University
  • John D. Hathcoat, Ph.D., Professor, James Madison University
  • Yu Bao, Ph.D., Associate Professor, James Madison University
  • Christian Early, Ph.D. , Director and Professor of Philosophy, James Madison University
AbstractThis article presents a case study examining how prompt design influences the ability of large language models to score student ethical reasoning essays in alignment with trained human raters. Findings suggest that structured prompting—particularly strategies that include self-verification—improves score consistency. Practical guidance is provided for higher education assessment practitioners who are exploring AI-assisted scoring for complex constructs.

Aligning AI and Human Judgment: How Prompt Design Shapes AI-Assisted Scoring of Ethical Reasoning

Assessment practitioners in higher education face a persistent tension: the constructs we value most are often the most labor-intensive to assess. Writing-based assessments, widely regarded as the gold standard for capturing complex reasoning, require significant investments in rater training, calibration, and ongoing quality monitoring. As institutional demands for accountability and scalable evidence of student learning intensify, many assessment offices are asking a practical question: Can artificial intelligence help reduce the burden of scoring while maintaining the quality and interpretive value of assessment results?

The emergence of large language models (LLMs) has renewed interest in automated scoring. Yet for complex, process-oriented constructs like ethical reasoning, the key question is not simply whether AI can generate scores, but whether those scores are meaningful, consistent, and aligned with expert human judgment. This issue becomes even more pressing when the assessment task requires students to demonstrate nuanced reasoning in writing. To explore this challenge, this article presents a case study of how one institution investigated the use of AI to support scoring of student essays on ethical reasoning. Rather than treating AI scoring as a plug-and-play solution, we examined how prompt design shapes scoring quality and concluded by offering practical lessons for assessment practitioners. 

The Context: Assessing Ethical Reasoning at Scale

At a public university in the Mid-Atlantic region, ethical reasoning has been assessed through the Ethical Reasoning in Action (ERiA) initiative, which centers on the Eight Key Questions (8KQ) to prompt moral considerations, such as fairness, outcomes, rights, responsibilities, and other ethical dimensions (Sanchez et al., 2017), to help with decision-making. Students write essays about ethical dilemmas they have encountered, and these essays are scored by trained faculty members using a structured rubric that captures multiple aspects of reasoning. The framework emphasizes the quality of reasoning rather than the correctness of conclusions.

The ethical reasoning rubric measures five dimensions: (A) identification of the ethical issue, (B) engagement with relevant 8KQs, (C) explanation of their applicability, (D) analytical depth, and (E) justification and weighing of competing considerations. Each dimension is assessed on a five-level developmental scale, ranging from 0 to 4 (Sanchez et al., 2017; Linder et al., 2019). To explore whether AI could serve as a supplementary scoring tool, the Office of ERiA partnered with assessment specialists to investigate LLM-assisted scoring under controlled conditions.     

What We Tested: AI Prompting Approaches

This study employed a comparative experimental design to examine how different prompting strategies influenced the AI scoring performance of student essays on ethical reasoning. Microsoft Copilot, based on GPT-5, was used as the scoring platform for all conditions and the same dataset was used throughout the study. Each essay was scored independently under each prompting condition.

To evaluate AI scoring performance, we tested four prompting strategies (Sahoo et al., 2023) using a set of 100 student essays that had previously been scored by expert human raters:

  1. Zero-shot prompting, which provided only the scoring rubric and rating instructions.
  2. Few-shot prompting, which included example essays representing different performance levels.
  3. Tree-of-Thoughts (ToT) prompting, which asked the model to generate and compare multiple evaluations of an essay before selecting a score.
  4. Chain-of-Verification (CoVe) prompting, which required the model to generate an initial score, check its reasoning against rubric criteria, and revise the score if inconsistencies were detected.

This design allowed us to compare how different levels of reasoning structure shaped scoring outcomes while controlling for dataset and model variability.

What We Found: Prompting Approaches Shaped Scoring Quality

To examine how closely AI-generated scores aligned with human ratings, five complementary indicators were used: systematic bias, exact agreement, mean absolute error (MAE), tolerance-based agreement (within ±1 point) and quadratic weighted kappa. Although the details of these analyses are not reported here, we found that prompt design had a substantial impact on scoring quality. 

Consistent differences emerged across prompting strategies. Zero-shot prompting produced the least consistent scoring, with greater variability between how the AI and humans approached the rubric. Few-shot prompting improved the human-AI alignment with evidence of reduced error, but scoring patterns depended on the specific calibration examples provided. ToT prompting further enhanced score stability by requiring AI to consider multiple evaluative perspectives before assigning a score, yet some bias persisted across rubric dimensions. CoVe prompting showed the strongest overall alignment with human raters. Across dimensions, CoVe produced bias values closest to zero, lower average deviation, and higher rates of both exact and near agreement than other strategies. For some rubric dimensions, agreement levels approached those observed between human raters.

What This Means: AI Scoring Requires Calibration and Validation

Just as human raters must be calibrated to a rubric during rater training to produce consistent results, LLMs must also be carefully calibrated for scoring purposes. This case study suggests that AI can be used to score ethical reasoning writing under an established rubric. However, scoring outcomes vary substantially depending on how the model is prompted. When tested on the same platform and dataset, different prompting strategies produced different levels of alignment with human raters. Among the conditions examined, CoVe showed the closest and most consistent agreement with trained human scoring. These findings underscore a central insight: prompt design functions much like rater training. When prompts incorporated structured self-verification, alignment with human judgment improves substantially. AI scoring is therefore not a plug-and-play solution, but a design-sensitive process requiring careful development, testing, and validation.

Implications for Assessment Practitioners

This study provides practical guidance for educators and assessment practitioners who are considering the use of AI to support scoring complex learning outcomes: 

Recognize that AI Scoring is not Plug-and-Play

Different prompts produced substantially different results, even when using the same LLM model and dataset. This means AI scoring outcomes are not a fixed property, but are instead shaped by how the task is framed and communicated to the system.

Add Structure to Improve Scoring Consistency

Prompts that require explanation, verification, and revision lead to more consistent scoring. Adding structure helps to guide the model’s reasoning process, which appears to reduce reliance on immediate surface-level judgments.

Design Prompts as Carefully as Scoring Protocols

Prompt design should be documented, tested, and refined analogous to how we document rater training strategies. Even small changes in wording or structure can meaningfully impact AI behavior, which makes iteration and standardization essential.

Validate AI Scores against Human Scoring

Although AI scoring technology will likely continue to improve, it is not yet at a point where it can replace human judgment for complex constructs. AI-generated scores should therefore always be compared to human raters before being used in practice to ensure alignment with the intended construct.

Use AI as a Supplement—not a Replacement

AI may be most useful for preliminary scoring, formative feedback, or reducing workloads within a human-in-the-loop system. Human oversight remains essential for interpreting complex constructs.

Exercise Caution in High-stakes Contexts  

AI scoring requires extensive validation and ongoing monitoring, particularly if one desires to use such scores for consequential decisions. The consequences of misalignment are amplified in high-stakes settings, and we should elevate our standards of evidence appropriately to account for consequential decisions. 

Looking Ahead

As higher education continues to grapple with the demand for scalable, meaningful assessment, AI-assisted scoring represents a promising frontier—but one that requires the same rigor we apply to any measurement tool. The lesson from this study is clear: when it comes to AI scoring of complex constructs, how we ask matters as much as what we ask. For assessment practitioners, this means investing in prompt design, validation protocols, and careful implementation, while approaching AI as a carefully calibrated partner rather than a replacement for professional judgment.

 

REFERENCES

Linder, G. F., Ames, A., Hawk, W. J., Smith, K. L., Fulcher, K. H., & Sanchez, E. R. H. (2019). Teaching ethical reasoning: Program design and initial outcomes of Ethical Reasoning in Action, a university-wide ethical reasoning program. Teaching Ethics, 19(2), 147–169. http://doi.org/10.5840/tej202081174

Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S. S., & Chadha, A. (2024). A systematic survey of prompt engineering in large language models: Techniques and applications. ArXiv. https://doi.org/10.48550/arXiv.2402.07927

Sanchez, E. R. H., Fulcher, K. H., Smith, K. L., Ames, A., & Hawk, W. J. (2017). Defining, teaching, and assessing ethical reasoning in action. Change: The Magazine of Higher Learning, 49(2), 30–36. https://doi.org/10.1080/00091383.2017.1286215

Xiao, C., Ma, W., Song, Q., Xu, S., Zhang, K., Wang, Y., & Fu, Q. (2024). Human-AI collaborative essay scoring: A dual-process framework with LLMs. In Proceedings of the 15th International Learning Analytics and Knowledge Conference (pp. 293-305). https://doi.org/10.1145/3706468.3706507