Scaling predictability in Morphological paradigms
Scaling predictability in Morphological paradigms
Project
Scaling predictability in Morphological paradigms
Project members:
Sacha Beniamine (PI)
Period of award
Nov 2026 -- Dec 2030
Funder:
UKRI FLF
Humans are constantly faced with incomplete information about the world, and our minds have developed strong predictive abilities to fill in the gaps. These abilities are especially important in language, in linguistic systems of words called inflection, where a single word can take different forms to express grammatical information. Yet we do not know how these predictions work.
languages, inflectional systems can vary dramatically in size. In English, for example, nouns inflect only for number (singular "cat", vs. plural "cats"), while languages like Latin have a system of cases to mark the role of each word in a sentence. Some languages, like Vietnamese, do not inflect verbs at all, while others, like Archi (spoken in Russia), use thousands of verb forms to indicate different tenses, moods, aspects, person, number, and case.
Speakers only ever hear a very limited fraction of the inflectional systems they use. This means that they often have to produce words they have never heard before, and they excel at this task. For example, English speakers can guess that the past tense of a new verb like "glorp" might be "glorped", following the same pattern as "jump" / "jumped". This raises fundamental questions about how languages work: How hard it is to make these predictions? What properties of inflectional systems support the relevant inferences? Are they easier in some languages? How difficult can they be? Do languages become more predictable over time?
This Future Leaders Fellowship aims to reveal how the predictability of inflected word forms operates and varies across the world’s languages. I tackle a key challenge: while linguists know that these systems vary widely, they currently lack the tools and data to compare them effectively. Current methods are outdated and do not provide the detail needed to study this variation or draw meaningful comparisons. Moreover, while grammars describe the inflection of many languages, very few languages have the detailed digital datasets required for modern analysis.
To move beyond current limitations, this project will create high-quality datasets for a sample of under-documented languages. We will develop new computational tools to measure how predictable a language's word forms are. We will apply these tools to the datasets to explore both the inner variation within a language or language family (micro-variation) and the broader patterns of variation across languages (macro-variation). We will release all our data sets and tools as open-access resources for researchers and educators.
Many of the languages on which the project focuses are endangered. We will support the preservation and teaching of these languages by deriving pedagogical books from the datasets we create, in collaboration with language communities. In this way, we will produce pedagogical tools that are informed by cutting-edge research. This will ensure a lasting positive impact on language education, preservation, and revitalization. By transforming our understanding of the predictive capacities that make language possible, this project will offer insights that will benefit linguistics, cognitive science, language technology, and education for years to come.