Skip to main navigation Skip to search Skip to main content

Generative AI Scoring of Constructed Responses: Psychometric Evidence for Research Applications

  • University of Wisconsin, La Crosse

Research output: Contribution to conferencePresentationpeer-review

Abstract

Constructed response assessments provide rich evidence of transfer and deep learning but require resource-intensive human scoring that creates practical barriers for management researchers. This study examines the reliability and validity of generative AI scoring as an alternative to traditional human scoring for research applications. Using an EvidenceCentered Design framework, we evaluated three commercial GenAI systems (GPT-4o, Claude 3.7 Sonnet, Gemini 2.5 Flash) scoring constructed responses assessing transfer of learning in motivation and leadership content domains (N = 243). Results demonstrated that GenAI achieved interrater reliability comparable to trained human raters, excellent test-retest reliability (ICC > .90) across all items, strong convergent validity with both research assistant and subject matter expert scores, and equivalent construct validity in predicting retention performance. Item-level analyses revealed that scoring challenges stemmed from item characteristics rather than rater type. These findings provide substantial psychometric evidence supporting GenAI as a reliable, valid, and accessible alternative to human scoring for constructed response assessments in management research contexts.
Original languageAmerican English
StatePublished - 2026
EventSociety for Industrial and Organizational Psychology Annual Conference - New Orleans, United States
Duration: Apr 29 2026May 2 2026

Conference

ConferenceSociety for Industrial and Organizational Psychology Annual Conference
Abbreviated titleSIOP2026
Country/TerritoryUnited States
CityNew Orleans
Period4/29/265/2/26

Keywords

  • Generative AI
  • Automated Scoring
  • Constructed Response
  • Validity
  • Reliability
  • Evidence-centered Design

Cite this