EMERGING DIALOGUES IN ASSESSMENT

Enhancing Assessment Feedback: Using AI as a Third Reviewer
August 25, 2026
AbstractThe University of North Dakota recently implemented a new strategy for providing feedback and recommendations to academic programs. Microsoft Copilot was used to generate a list of recommendations for each credential submitted based on assessment committee feedback, and submitted reports. The results were beneficial but included some pitfalls. This treatise describes the process of using Copilot to generate recommendations for academic programs and offers insights for improvement of the process and future directions.

 

Background

Providing feedback on annual assessment reports to academic programs is a vital part of the assessment process (Banta & Palomba, 2014). Continuous improvement of academic programs depends on a detailed and specific feedback loop. At UND the feedback loop was seldom closed in part because feedback from the campus office responsible for providing it did not exist.

The Office of Institutional Effectiveness & Accreditation (OIEA) is charged with overseeing the assessment of academic and cocurricular programs, as well as the general education program, known at UND as the essential studies program. The annual assessment feedback process is conducted each spring for selected academic credentials. All bachelor’s degrees follow an annual reporting cycle as approximately three-fourths of UND’s student body is comprised of undergraduates. Graduate degrees report on an alternating year cycle, along with both undergraduate and graduate certificates.

All submitted academic assessment reports receive feedback in the form of a document. Programs then schedule a meeting with the OIEA to discuss the contents of the document. The OIEA provides feedback to programs who submit their assessment reports during each annual cycle, as recent practice at UND did not include any feedback from the OIEA. 

To address the lack of assessment feedback for academic departments, the OIEA worked with the University Assessment Committee (UAC) to implement a review and feedback process. Each credential is now reviewed by two individuals: one member of the UAC and one member of the OIEA team. The review process is guided by a rubric that considers six elements of the assessment report (outcomes, measures, targets, results, analysis/findings, actions) and provides narrative feedback to the academic program. Meetings are scheduled during the spring semester with each academic department to review each of their credentials.

Introducing Copilot Recommendations

Beginning in the spring of 2026, the OIEA added an additional feedback element to the process. The OIEA began using artificial intelligence (AI) to generate recommendations for each academic credential reviewed as part of the annual assessment feedback process. The AI used for this was Microsoft’s Copilot platform, as it was already integrated into the larger Microsoft technology environment on campus.

Heretofore, any recommendations from the OIEA were provided as narrative feedback written by human reviewers. The feedback did not always contain constructive or detailed suggestions to make meaningful changes. The value in using Copilot to provide this feedback is considerable. For example, Copilot-generated recommendations can quickly and efficiently identify where students may be at risk so that academic programs are better positioned to engage in targeted intervention strategies (Sun et al., 2025). As displayed in Figure 1, the process involves providing Copilot with the three most recent years of assessment reports and feedback reports from the UAC and OIEA.

Figure 1: Feedback recommendation schematic.

Figure 1: Feedback recommendation schematic. 

Prompting

The assessment data and feedback documentation described above was attached to a prompt asking Copilot for recommendations. The initial prompt produced highly detailed results. The extensive details allowed OIEA staff to determine whether it was appropriate to retain these recommendations. The various follow-up prompts provided shorter descriptions of the patterns and trends identified from the reports and feedback documents, but specific examples of how to address the recommendations were not always included. This made it difficult to identify which recommendations were within or beyond OIEA’s scope. Creating detailed and specific prompts is a vital part of the process for creating effective feedback and recommendations (Tharp, 2026). Table 1 provides a summary of Copilot prompts and general output.

Table 1. Copilot prompts and results.

 Table 1. Copilot prompts and results.

Results

Efficiencies
Copilot was useful for identifying patterns within and across reports. While OIEA staff could have identified some patterns, it would have been a time-consuming process. Copilot identified patterns in a few minutes, saving OIEA staff considerable time. If assessment reports noted that changes were planned but no documentation was provided in the subsequent report, Copilot quickly identified these discrepancies. For example, when programs noted they planned curriculum changes in one report, it was not always noted in subsequent reports even if changes were made.

The output provided links to the reports and feedback documents Copilot based its recommendations on. Recommendations were drawn from various combinations of early reports, recent reports, and UAC feedback. This was helpful for the recommendation development process because it allowed the OIEA to eliminate or revise recommendations based on program status. For example, if an area of concern was identified in the 2022-2023 report but it was not evident in the 2023-2024 or 2024-2025 report, it was eliminated. As stated above, due to the difficulty of tracking changes, being able to identify which documents Copilot based its recommendations upon allowed OIEA to verify the validity of recommendations.

Inefficiencies
While there are many benefits of using Copilot to support the development of quality improvement recommendations, there are also several inefficiencies. The first inefficiency relates to context. Many of the recommendations provided by Copilot needed extensive revision or were eliminated due to missing context from the input. This result reinforced the notion that humans should always be the interpreters of any AI output and should maintain the responsibility for all that AI produces (Tharp, 2026).

Another context that Copilot excluded from its analysis related to how academic programs can drastically differ. For example, there are over 200 academic programs at UND. Some programs have external accreditors, while others do not. Some recommendations critiqued outcome language for accredited programs. This became problematic since the outcome language cannot be changed by the program due to accreditation requirements. This is an example of where Copilot lacked contextual knowledge.

Additionally, the chronological order of reports was not recognized well by Copilot. Some recommendations were drawn only from the oldest report and feedback documents. When the OIEA reviewed the recommendations, many were often eliminated because the recommendations provided were no longer relevant.

Finally, Copilot did not always link reports to recommendations, as in it did not reference where in the reports provided the recommendation originated. When this occurred, OIEA staff manually exported the reports to verify the relevance of recommendations. Similarly, the formatting of recommendations was not consistent across departments. As a result of these issues, the OIEA spent considerable amounts of time ensuring consistency of formatting.

Recommendations for Improvement         

After completing the process for all programs, several avenues or recommendations for improvement were evident. One area of improvement included the output. For best results, users should include program-specific context within the prompts. During spring meetings with department chairs and assessment leads, faculty turnover, curriculum changes, and student factors were frequently identified as having affected assessment data. Oftentimes the recommendations Copilot produced were very obvious, such as recommending that we collect more data for years where the data were missing, which is usually not possible. As a result, the prompting cycle was repeated several times to improve the results.

Another challenge encountered during the process included the time commitment required to accomplish this task. Given the size of our institution, with over 300 programs, it was very time intensive to manage the process. It is therefore recommended that users should ensure both adequate planning and time allocation for the process to be successful.  

Conclusions and Future Directions

The process of integrating recommendations from Copilot into our assessment feedback process based on both historical reports and UAC feedback provided many insightful recommendations for improvement. Although a sizable amount of these recommendations amounted to obvious drivel, the process did provide valuable insights. These insights often resulted from data patterns in previous reports that would have taken human reviewers hours of review time to glean from past assessment reports and UAC feedback.  

In subsequent years, fine-tuning the process for both time savings and higher-quality feedback will be necessary, given staffing limitations. Improving and focusing on prompt engineering will undoubtedly improve and streamline the process. The creation of AI agents to assist with the process, which could be specific to degree level or department, will likely enhance this process and provide further time savings and clarity in the results.

Given the busy schedules and lack of enthusiasm for assessment that describes many faculty at the present time, feedback and recommendations for assessment improvement can be provided efficiently and effectively using AI tools. As the OIEA at UND has emphasized in the past few years, the continual work of building and maintaining a culture of assessment on campus often comes about through relationship-building, offering assistance, and providing resources.

 

REFERENCES

Banta, T. W., & Palomba, C. A. (2014). Assessment essentials: Planning, implementing, and improving assessment in higher education (2nd ed.). Jossey-Bass.

Sun, X., Yang, Y., & Du, Q. (2025). Advancing total quality management in higher education through AI and big data analytics. Total Quality Management & Business Excellence, 36(9-10), 896–918. https://doi.org/10.1080/14783363.2025.2495103

Tharp, D. S. (2026). Musings on artificial intelligence’s (AI’s) place at the assessment table. Assessment Update, 38(1), 10–12. https://doi.org/10.1002/au.70011