Validating LLM-based alternative uses test scoring across ages

Eran Hadas, Ben Avital-Lev, and Arnon Hershkovitz · Thinking Skills and Creativity, 60, Article 102066 · 2026

Summary

This study tests whether general-purpose large language models can score Alternative Uses Test flexibility and originality reliably across childhood, adolescence, and adulthood. It compares automated scores with expert judgments across multiple datasets and model families, examining criterion validity, developmental patterns, and potential gender bias.

Topics

Creativity assessment · divergent thinking · Alternative Uses Test · flexibility · originality · large language models · validity

Citation

Hadas, E., Avital-Lev, B., & Hershkovitz, A. (2026). Validating LLM-based alternative uses test scoring across ages. Thinking Skills and Creativity, 60, Article 102066. https://doi.org/10.1016/j.tsc.2025.102066

Full text

The publisher version is linked above. No local full-text file is provided here unless a rights-cleared author manuscript is available.