
New research explores why people often breeze through some sentences in a book or article but have to reread others to comprehend their meaning. A team of linguists and data scientists found a partial answer in AI, with some language processing parallels between humans and neural-network-based large language models (LLMs).
The study, conducted by researchers from New York University and the University of Massachusetts Amherst, shows that humans and AI process language in similar ways during the earliest moments of reading, relying on next-word predictions.
Humans’ processing differs from that of AI as reading continues and passages become more complex.
According to the report, LLMs can explain how long it takes people to recognize words when their eyes move smoothly forward through a text, but they fail to capture cases where people have difficulty integrating a word into the larger context of a sentence, often accompanied by rereading.
The authors note that despite the remaining uncertainty on how humans read, the findings offer a potential roadmap for improving language learning and addressing reading-related afflictions.
William Timkey, a linguistics doctoral student at NYU and the lead author of the paper, explains that LLMs develop their language understanding capabilities by being trained to predict the next word in a blood test of language.
No.
Related: CMS seeks feedback on breakthrough device coverage
The researchers used eye-tracking technology to analyze 368 adult readers, focusing on how long participants spent reading and rereading each word of carefully designed sentences, including syntactically challenging sentences known as garden-path sentences.
These sentences, such as “The old man the boat,” are grammatically correct but start in a way that a reader’s initial interpretation will likely be incorrect, making them good candidates for understanding how we process complex passages.
The results showed that AI models’ next word predictions can explain the first step of processing each word of a sentence, but they can’t explain the next step of integrating that word into the larger meaning of the sentence.
Brian Dillon, a professor of linguistics at UMass Amherst and the paper’s senior author, says that the findings offer a potential roadmap for improving language learning and addressing reading-related afflictions.
The research was supported by grants from the National Science Foundation. The study’s authors hope that their work will help create computational models that more closely match the human mind and can help understand in detail how it operates.
As Tal Linzen, an associate professor of linguistics and data science at NYU, notes, the human mind does not always work like standard AI systems, and they now have their work cut out for them to create models that can explain the complexities of human language processing.
They will use this knowledge to improve language learning.