I build AI systems for learning, and I measure whether they work. My background combines psychometric measurement (validity, fairness, item response theory), applied NLP research, and the engineering to ship usable + maintainable implementations.

I'm co-founder and CTO of Edumonia, the startup behind the iTELL learning platform, and I'm completing a PhD in cognitive science at Vanderbilt, where my dissertation introduces and evaluates LLM-surprisal as a measure of second-language proficiency.

Projects & Systems

iTELL · co-founder & CTO

A platform that turns course materials into interactive, AI-enhanced texts with comprehension checks, adaptive feedback, and content-aware chat. Deployed with institutional partners as an LTI 1.3 tool inside their learning management systems, running on Kubernetes; learning gains evaluated in a randomized controlled trial. Started as NSF-funded research, now commercialized.

RCT at Learning@Scale '25

PIILO · creator

An open-source system for deidentifying student writing that makes it easy to use the "hiding in plain sight" obfuscation strategy. As PI on a $30,000 award from The Learning Agency, I built the dataset behind Kaggle's PII Data Detection competition and helped to package the winning submissions into a lightweight, extensible Python package.

PIILO paper · CRAPII paper · code

Word predictability · dissertation

LLM surprisal as a measure of second-language proficiency. High-proficiency learners make more predictable word choices, and a predictability metric outperformed conventional measures of lexical sophistication in explaining TOEFL score variance. Published in Language Learning.

Language Learning paper · code

entroprisal · author

A Python package for information-theoretic n-gram statistics — entropy and surprisal — as measures of text readability. Paper forthcoming.

code

NAEP automated math scoring · grand prize winner

First place and the $30,000 grand prize in the NAEP Automated Math Scoring Challenge. LLM-based scoring of constructed-response items from the Nation's Report Card, matching human raters closely enough for operational use. Methods published in the International Journal of Artificial Intelligence in Education.

IJAIED paper

I also designed and maintain my lab's GPU compute cluster, fine-tune speech recognition models for children's voices, and help teach the graduate NLP course at Vanderbilt's Data Science Institute.

Selected Publications

Word Predictability as a Measure of Second Language Proficiency Holmes, Crossley, Morris, & Choi · Language Learning · 2026
Assessing Fairness in Finetuned Scoring Models with Demographically Restricted Training Data Holmes, Morris, Crossley, & Choi · Assessing Writing · 2026
Automated Scoring of Constructed Response Items in Math Assessment Using Large Language Models Morris, Holmes, Choi, & Crossley · Int. J. of Artificial Intelligence in Education · 2024 NAEP grand prize methods
iScore: Visual Analytics for Interpreting How Language Models Automatically Score Summaries Coscia, Holmes, Morris, Choi, Crossley, & Endert · ACM IUI · 2024
PIILO: An Open-Source System for Personally Identifiable Information Labeling and Obfuscation Holmes, Crossley, Sikka, & Morris · Information and Learning Sciences · 2023
The Cleaned Repository of Annotated Personally Identifiable Information Holmes, Crossley, Wang, & Zhang · Educational Data Mining · 2024

All publications →

Writing

Identifying Limitations and Bias in ChatGPT Essay Scores
The Cutting Ed · March 2025 · with L Burleigh, Ulrich Boser, Kennedy Smith, & Perpetual Baffour
Scrubbing Personally Identifiable Information from Large-Scale Datasets
The Cutting Ed · July 2024 · with Jules King