← Back to the overview

Feedback and Grading Analysis

Examine the reliability, consistency, and pedagogical value of AI-generated feedback and grading.

Academic Practice & IntegrityLevel 1Complexity: IntermediateDuration: 30-60 min

Recommended for: Writing-intensive courses across disciplines

Procedure

  1. Assign students to write a short essay (500 words) on a relevant topic (e.g., "Analyze the influence of social media on student well-being, including reflection on your own experience")
  2. Each essay receives feedback and a grade from at least three peer reviewers; do not provide a grading rubric; students must justify their assigned grades
  3. Students do not yet view the peer feedback they have received
  4. Students submit their essay to a generative AI system with a request for constructive feedback
  5. Students revise their essay based on AI feedback (they may use AI assistance or work independently)
  6. Students resubmit the revised essay to the same AI system and request feedback again
  7. Students analyze and document why the AI continues to provide feedback even after they have incorporated its previous suggestions
  8. Students request numerical grades from the AI system: first without specifications, then explicitly requesting a grade on a 1-10 scale with written justification
  9. Students verify whether the AI-assigned grade remains consistent when the same essay is submitted multiple times
  10. Instructor provides the official grading scheme or rubric
  11. Students compare peer feedback, AI feedback, and criteria-based evaluation
  12. Students exchange their peer-generated feedback and systematically compare all sources of evaluation
  13. Conduct comprehensive class discussion synthesizing findings across all comparison points

Extension options

  • Submit identical essays to different AI systems (ChatGPT, Claude, Gemini) and conduct comparative analysis of feedback quality and consistency
  • Compare feedback generated by minimal prompts versus highly detailed prompt instructions specifying strictness regarding grammar, structure, style, and argumentation
  • Provide the AI system with the official grading scheme and compare its evaluation to human peer assessment and instructor evaluation

Discussion questions

  • Why does AI continue to suggest revisions even after students have implemented its previous recommendations?
  • How does AI grading consistency compare to inter-rater reliability among human evaluators?
  • What does this exercise reveal about AI's understanding of textual quality, improvement, and educational standards?
  • How might students strategically game or manipulate AI feedback systems?
  • What are the pedagogical implications of students relying on AI feedback rather than developing internal standards of quality?

Assessment opportunity

Students produce a comparative analysis report documenting patterns across all feedback sources, evaluating the strengths and limitations of each approach.

Reference

  • Geysendorpher (2025) provides additional empirical analysis of AI feedback reliability in educational contexts.