Validating LLM-based alternative uses test scoring across ages
Summary
This study tests whether general-purpose large language models can score Alternative Uses Test flexibility and originality reliably across childhood, adolescence, and adulthood. It compares automated scores with expert judgments across multiple datasets and model families, examining criterion validity, developmental patterns, and potential gender bias.
Topics
Creativity assessment · divergent thinking · Alternative Uses Test · flexibility · originality · large language models · validity
Citation
Hadas, E., Avital-Lev, B., & Hershkovitz, A. (2026). Validating LLM-based alternative uses test scoring across ages. Thinking Skills and Creativity, 60, Article 102066. https://doi.org/10.1016/j.tsc.2025.102066
Full text
The publisher version is linked above. No local full-text file is provided here unless a rights-cleared author manuscript is available.