Feedback and Grading Analysis
Examine the reliability, consistency, and pedagogical value of AI-generated feedback and grading.
Recommended for: Writing-intensive courses across disciplines
Procedure
- Assign students to write a short essay (500 words) on a relevant topic (e.g., "Analyze the influence of social media on student well-being, including reflection on your own experience")
- Each essay receives feedback and a grade from at least three peer reviewers; do not provide a grading rubric; students must justify their assigned grades
- Students do not yet view the peer feedback they have received
- Students submit their essay to a generative AI system with a request for constructive feedback
- Students revise their essay based on AI feedback (they may use AI assistance or work independently)
- Students resubmit the revised essay to the same AI system and request feedback again
- Students analyze and document why the AI continues to provide feedback even after they have incorporated its previous suggestions
- Students request numerical grades from the AI system: first without specifications, then explicitly requesting a grade on a 1-10 scale with written justification
- Students verify whether the AI-assigned grade remains consistent when the same essay is submitted multiple times
- Instructor provides the official grading scheme or rubric
- Students compare peer feedback, AI feedback, and criteria-based evaluation
- Students exchange their peer-generated feedback and systematically compare all sources of evaluation
- Conduct comprehensive class discussion synthesizing findings across all comparison points
Extension options
- Submit identical essays to different AI systems (ChatGPT, Claude, Gemini) and conduct comparative analysis of feedback quality and consistency
- Compare feedback generated by minimal prompts versus highly detailed prompt instructions specifying strictness regarding grammar, structure, style, and argumentation
- Provide the AI system with the official grading scheme and compare its evaluation to human peer assessment and instructor evaluation
Discussion questions
- Why does AI continue to suggest revisions even after students have implemented its previous recommendations?
- How does AI grading consistency compare to inter-rater reliability among human evaluators?
- What does this exercise reveal about AI's understanding of textual quality, improvement, and educational standards?
- How might students strategically game or manipulate AI feedback systems?
- What are the pedagogical implications of students relying on AI feedback rather than developing internal standards of quality?
Assessment opportunity
Students produce a comparative analysis report documenting patterns across all feedback sources, evaluating the strengths and limitations of each approach.
Reference
- Geysendorpher (2025) provides additional empirical analysis of AI feedback reliability in educational contexts.