7 Evaluating AI Outputs and Interactions

The Critical Eye

Consider an illustrative case. A colleague asks a chatbot to summarize recent research on student retention. The response reads well. It names authors, cites three studies, and includes statistics. Two of the citations are fabricated. The statistics have no traceable source.

The response sounds confident and looks professionally formatted. Those qualities can make AI useful for drafting and brainstorming. They can also make an unsupported answer feel ready to use. As Mollick has observed, evaluating AI output increasingly requires real knowledge of the subject.

This chapter introduces ABC-R, this course’s compact framework for evaluating AI outputs and interactions. It teaches the first three checks: Accuracy and Evidence, Bias and Perspective, and Context and Relevance. The R stands for Responsible Use, which Responsible Use in Education develops in more detail.

The ABC-R Framework

The U.S. Department of Labor’s AI Literacy Framework treats evaluating AI output and using AI responsibly as separate areas of AI literacy. Evaluation asks whether an output is factually supported and fit for its task. Responsible Use examines the conditions around using it, including privacy and accountability. The Association of College and Research Libraries similarly separates analysis and evaluation from ethical considerations. The University of Virginia’s AI literacy chapter uses familiar categories such as accuracy, relevance, and bias.

ABC turns the evaluation work into three questions.

The ABC Checks for Evaluating AI Outputs and Interactions
Dimension Key Questions
Accuracy and Evidence
Claim check
What exact claim might I use, repeat, cite, or share? What evidence supports it? Does the source exist, and does it support the claim?
Bias and Perspective
Frame check
Whose perspectives are represented? Who is missing? What assumptions, stereotypes, prompt effects, or patterns of over-agreement shape the response?
Context and Relevance
Fit check
Does this fit the task, required sources, discipline, institution, audience, course level, and learning purpose?

You can use these checks on a chatbot response, a source-linked answer, an agent’s finished work, or AI-generated material someone else has shared. The three checks ask different questions. A response can pass one and fail another.

Accuracy and Evidence: Check the Claim

Begin with the claim you might carry forward. What sentence would you repeat in class, place in course materials, cite in a report, or use to make a decision? Check that sentence before trying to judge the whole response.

The familiar term hallucination is often used for plausible but false AI-generated content. The NIST Generative AI Profile describes the problem as confidently presented false or erroneous content. NIST also warns that a model can generate false logic or citations that appear to justify an incorrect answer.

For routine checking, two distinctions cover most of the immediate work.

Fabrication occurs when the AI invents something. A citation does not exist. A quotation never appeared in the source. A statistic has no traceable origin. Walters and Wilder’s 2023 study of AI-generated bibliographic citations documented both fabricated sources and incorrect citation details in the model versions they tested. Those historical results do not establish a current error rate, but they show why polished citation formatting is not evidence that a source is real.

Misgrounding occurs when the source is real but does not support the claim attached to it. A chatbot might link to an actual institutional policy and then state a rule that never appears on the page. Stanford HAI uses this term in its review of AI legal research tools. A working link passes the source-existence check. It does not pass the source-support check.

Other errors can be caught through the same claim-checking routine. A response may conflate two real studies by combining their findings, or it may reverse an attribution by assigning a real quotation or action to the wrong person. In each case, separate the claim into parts and compare each part with its source.

Quantitative claims need the same attention. Suppose an AI response says that a rate increased from 20 percent to 30 percent and calls this a 10 percent increase. The change is 10 percentage points, but the relative increase is 50 percent. The numbers are real; the description of the change is wrong. Rework the denominator, units, formula, or comparison before using the result.

Red flags to watch for:

  • Exact dates, statistics, quotations, or names that lack a traceable source
  • Citations missing standard publication details or stable links
  • Quotations that align unusually well with the wording of the prompt
  • Information that conflicts with reliable sources you already know
  • Details that are individually plausible but combined in unfamiliar ways

Claim-check methods:

  • Exact-title check: Search for the exact title of a cited article, policy, report, or webpage.
  • Source-support check: Open the source and compare it with the exact claim.
  • Current-source check: Determine whether the source is recent enough for the claim and situation.
  • Attribution check: Separate what the source says from what the AI inferred.
  • Calculation check: Rework percentages, units, formulas, and quantitative comparisons.
  • Corroboration check: Leave the response and look for confirmation in another trustworthy source.

Source grounding gives the AI a document, webpage, policy, dataset, or source collection to work from. It can reduce generic responses and give you a source trail to inspect. Grounding is not verification. A grounded response can still misread a source, combine unrelated details, or make a claim the source does not support.

You may ask the AI for a concise rationale or a source trail. Treat that material as another claim to inspect. The AI’s explanation does not verify its answer. Open the source, rework the calculation, or corroborate the claim independently.

Bias and Perspective: Check the Frame

Bias and Perspective examines the frame that shaped a response. Who appears? Who disappears? What does the answer treat as normal? Which interpretation did the tool accept before it began answering?

Two kinds of bias are especially useful to check. Default bias appears when a response reproduces patterns and assumptions from its training data, model design, human feedback, or instructions built into the tool. Interaction bias appears when the response bends toward the user’s wording, assumptions, mood, or desired answer.

A bias check can also surface disciplinary defaults. A response may treat one research method, theoretical school, professional standard, or genre as the field’s neutral center. It may present a disputed interpretation as settled or assume that every field values the same kind of evidence. Detecting those defaults requires enough disciplinary knowledge to recognize which alternatives the response has ignored.

Bias does not always appear as an explicit stereotype. It can appear in the people chosen for examples, the resources assumed to be available, the evidence treated as authoritative, or the language presented as standard. A response can describe one plausible experience and still represent it as the ordinary experience.

In a study of one language model, Durmus et al. (2023) found that its responses were more similar to opinions from the United States, Canada, Australia, and several European and South American countries than to opinions from some other countries. The study does not describe every model or every kind of bias. It illustrates how a tool can present some cultural perspectives as ordinary while representing others less fully.

The same pattern can appear in teaching suggestions. An AI asked to design a group project for an Idaho college might assume that students own laptops, have reliable broadband, live near campus, can meet outside class, and are available during the day. The suggestion may work for some students while quietly treating their circumstances as universal.

Two prompting guides make this kind of frame check more concrete. Wiley’s guidance for authors recommends checking for assumptions about access to resources or technology, generalizations about populations or regions, and examples drawn from a limited cultural perspective. Jagadish Paudel’s chapter on linguistic and cultural bias (PDF) in the WAC Clearinghouse’s Writing and Rhetoric Studies in the Loop prompt library asks users to specify the audience, place, time, community, and cultural or linguistic conditions behind a request. It also asks what must be represented and what must not be generalized or misrepresented.

These prompting moves can expose assumptions and produce a more specific answer. A more detailed prompt does not make the answer unbiased. In Paudel’s comparison, the models still introduced incorrect cultural details, confused linguistic terms, and omitted important practices after receiving detailed constraints. Someone with relevant cultural and linguistic knowledge still had to evaluate the results.

Interaction Bias: When the AI Agrees Too Much

Sycophancy occurs when a chatbot agrees with the user or mirrors the user’s position even when the reasoning is weak. Anthropic’s 2023 research found this pattern across five assistants and four text-generation tasks. Responses that matched a user’s views were more likely to be preferred in the human-feedback data the researchers examined. OpenAI discussed a similar problem after rolling back a 2025 GPT-4o update that had become overly flattering and agreeable.

The practical effect is confirmation. The tool may accept a false premise or praise an idea that needs critique. Imagine asking, “Why are online students less engaged than students on campus?” The prompt treats lower engagement as established and asks the tool to explain it. A frame check begins by questioning the premise.

A Frame Check in Practice

Begin by separating what you observed from the explanation built into the prompt. The original question does not say how engagement was measured, which students were compared, or whether the difference exists. A revised prompt can make those gaps visible:

Before answering, identify the assumptions embedded in this question. Offer at least three alternate explanations for any observed difference in participation. Consider course design, work and caregiving responsibilities, reliable internet and technology access, commuting distance, disability, prior experience, and differences in how engagement is measured. Distinguish evidence from speculation. Do not generalize across all online students.

The revised prompt asks the tool to inspect the premise, consider conditions that the original wording excluded, and mark the limits of its evidence. The resulting answer may be more useful. It may also reproduce different assumptions or invent plausible explanations without evidence.

Complete the frame check outside the chatbot. Compare the response with relevant research, local course information, and the perspectives of the people represented. For the online-engagement question, that might include student feedback, participation data, course design, and conversations with online students. The chatbot can generate contrasts. It cannot determine whether it has represented those students fairly.

Use one or two methods that fit the response:

  • Default check: What kind of student, institution, culture, schedule, language, or discipline does the response treat as normal?
  • Representation check: Who appears, who disappears, and whose knowledge or experience counts as authoritative?
  • Prompt-frame check: What premise or desired answer did the wording supply?
  • Contrast check: How does the response change when you request alternate explanations or specify who and what must be represented?
  • Outside check: Which sources, local evidence, colleagues, students, or other people affected can test the response’s frame?

Context and Relevance: Check the Fit

Bias asks whose perspective shaped the answer. Context asks whether the answer fits the situation in front of you.

AI models often produce a plausible response for a generalized audience. That response may not fit your students, institution, discipline, or assignment. An explanation can be too advanced for an introductory course. A writing guide can default to generic academic prose while missing a discipline’s genre and citation practices. A polished rubric can reward qualities that the course outcome does not assess.

Stolpe, Larsson, and Johansson Falck (2026) describe critical AI engagement as a situated disciplinary practice. People need enough field knowledge to formulate a meaningful problem and judge whether AI contributes to the learning process. Correct terminology alone does not establish that fit. A response can sound appropriate while treating weak evidence as decisive, skipping a required form of reasoning, or applying the wrong professional standard.

Required sources create another kind of fit. A nursing instructor might ask for a lab activity based on the institution’s current safety protocol. The AI returns a sensible activity supported by a real general-health webpage, but it never uses the required local protocol. The answer may contain useful information and still fail both the source requirement and the teaching task. In a small pharmacy-education pilot, the FLUF Test asked students to examine credibility and fit for an audience or situation. Some participants reported that the framework needed clearer guidance for evaluating specificity and depth.

Use one or two methods that fit the response:

  • Task-fit check: Compare the response with the actual assignment, project, or decision.
  • Source-requirement check: Check whether it uses the required reading, policy, data, or disciplinary source.
  • Local-fit check: Ask whether it fits your institution, students, course format, class size, and available resources.
  • Audience-fit check: Ask whether the level, vocabulary, examples, and format fit the intended readers.
  • Tool-context check: Determine what sources or files the tool could actually access.
  • Learning-purpose check: Decide whether the response helps students practice the intended skill or performs that skill for them.

The learning-purpose check remains an evaluation of the output’s fit. Responsible Use in Education takes up the broader Responsible Use question: should AI perform this part of the task under these conditions?

What to Keep

  • Accuracy and Evidence checks one consequential claim and its support.
  • Bias and Perspective checks defaults, omissions, prompt effects, and over-agreement.
  • Context and Relevance checks whether the response fits the task, sources, discipline, institution, audience, course level, and learning purpose.
  • Source grounding creates a source trail. It does not verify the answer.
  • A short ABC note should end with a decision to use, revise, reject, or verify further.

Looking Ahead

This chapter covered the evaluation side of ABC-R. You checked the claim, frame, and fit of an AI output and made a decision about what to do with that output.

Responsible Use in Education asks whether the AI use itself is responsible. It examines human judgment and effects on learning. It also covers student-record privacy under the Family Educational Rights and Privacy Act (FERPA), institutional approval, and accountability. Faculty consider who else is affected and what disclosure the situation requires.

References

Alexander, K., Feild, C., Egan, C., & Parker, J. (2025). Application of the FLUF Test framework for evaluating AI-generated outputs in pharmacy education. The Journal of Applied Instructional Design, 14(3). DOI record

Anthropic. (2023, October 23). Towards understanding sycophancy in language models. Source link

Association of College and Research Libraries. (2025). AI competencies for academic library workers. Source link

Durmus, E., Nguyen, K., et al. (2023). Towards measuring the representation of subjective global opinions in language models. arXiv. Source link

Mollick, E. (2024, December 19). What just happened. One Useful Thing. Source link

National Institute of Standards and Technology. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). DOI record

OpenAI. (2025, April 29). Sycophancy in GPT-4o: What happened and what we’re doing about it. Source link

Paudel, J. (2026). Mitigating AI bias through tripartite prompting: A rhetorical-linguistic-translingual (RLT) approach. In A. Gupta & B. Gogan (Eds.), Writing and rhetoric studies in the loop: A GenAI prompt library (pp. 365–381). The WAC Clearinghouse. PDF source

Stanford Institute for Human-Centered Artificial Intelligence. (2024, May 23). AI on trial: Legal models hallucinate in 1 out of 6 or more benchmarking queries. Source link

Stolpe, K., Larsson, A., & Johansson Falck, M. (2026). Discipline-Specific AI literacy (DiSAIL): A theoretical framework for situated engagement with generative AI in education. International Journal of Technology and Design Education, 36, 1917–1932. DOI record

U.S. Department of Labor, Employment and Training Administration. (2026). Training and Employment Notice No. 07-25: Artificial Intelligence Literacy Framework. Source link

University of Virginia Library. (n.d.). Critical evaluation of AI outputs. AI Literacy in the Age of Generative AI. Source link

Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045. Source link

Wiley. (n.d.). AI guidelines: A guide for book authors. Source link

Willison, S. (2026, May 10). A quote from New York Times editors’ note. Simon Willison’s Weblog. Source link

Further Reading

  • D2L. (2026). AI literacies in practice. PDF source
  • Furze, L. (2025, September 15). Digital plastic as a framework for critical AI literacy. Source link
  • Mills, A. (2025, August). Ongoing sources of AI professional development. Anna Mills’ Substack. Source link

License

Icon for the Creative Commons Attribution-NonCommercial 4.0 International License

A Guide to Teaching and Learning with Artificial Intelligence Copyright © by Jason Blomquist; Liza Long; and Joel Gladd is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License, except where otherwise noted.