← Back to the overview

AI-Generated Hypothesis and Experimental Design Evaluation

Develop students' ability to critically assess research design proposals, identifying methodological flaws and strengthening scientific rigor.

Disciplinary ExpertiseLevel 2Complexity: AdvancedDuration: 30-60 min

Recommended for: Research methods courses, laboratory courses, or capstone seminars

Procedure

  1. Instruct students to request that an AI system generate a testable hypothesis related to current course content or their research interests
  2. Have students ask the AI to propose a complete experimental design to empirically test this hypothesis
  3. Students systematically assess the AI's research proposal across multiple evaluative dimensions:
    • Feasibility: Can this experiment realistically be conducted with available resources, equipment, and techniques?
    • Controls: Are appropriate positive and negative controls included? What essential controls are missing?
    • Variables: Are independent and dependent variables clearly defined and appropriately operationalized? Are confounding variables identified and controlled?
    • Statistical Power: Is the proposed sample size adequate for detecting meaningful effects?
    • Logic: Does the experimental design actually test the stated hypothesis or address a different question?
    • Pitfalls: What could go wrong methodologically? What assumptions remain unexamined?
    • Ethics: Are there ethical considerations the AI overlooked (human subjects, animal welfare, environmental impact)?
    • Alternative explanations: Could the predicted results be explained by factors other than the hypothesis?
  4. Students revise and improve the experimental design to address all identified weaknesses
  5. Conduct small group discussions comparing findings, revisions, and alternative approaches

Extension options

  • Students compare experimental designs for the same research question generated by different AI-powered scientific search engines (e.g., Consensus, Elicit, Scite, SciSpace, Semantic Scholar)
  • Evaluate systematic differences in quality, comprehensiveness, and methodological rigor across AI platforms
  • Have students compare AI-generated designs to published experimental designs from peer-reviewed literature

Discussion questions

  • What categories of methodological errors or omissions were most common in AI-generated research designs?
  • What depth of domain knowledge and methodological training was necessary to identify these problems?
  • How might uncritical reliance on AI-generated experimental designs affect the quality and rigor of scientific research?
  • Under what circumstances might AI provide useful input in experimental design versus when is human expertise essential?

Assessment opportunity

Students produce a methodological critique of the AI-generated design and a revised design that addresses all identified limitations, with justification for each modification.