A massive, free semantic database — English words, their meanings, and how they connect.
Linguabase is a dozen or so data tables — associations, senses, antonyms, families, definitions, clues — that together map English. We look at the language holistically: how common and familiar each word is, what people associate it with, and what it means.
It’s built on an amalgam of handcrafted lexicographic work and over 130 million frontier LLM inferences, run in 2025 and 2026 to tap the models’ implicit knowledge of language — so “hiking” carries flavors of both exercise and nature. But raw model output isn’t enough on its own. We did what no single prompt can — held the whole language in view at once, ranking associations, separating senses, and keeping it consistent across the whole vocabulary, from the most common words to the rare tail.
We want to boost public literacy by supporting all kinds of projects — nonprofit and commercial alike — built on English and its connections: creative work, communications, learning, research. For example, word-association games are having a moment. NYT Connections proved the mechanic, and this data can power a new generation of them:
Game makers can use Linguabase — plus minimal additional agentic coding — to build thousands of non-repeating levels, overcoming two core limitations of LLMs:
English is full of fascinating rabbit holes and edge cases, and Linguabase solves most (not all) of these problems — giving you ready-to-use data to drive your games, or any other project built on English and its interconnections.