Vocabulary & Frequency

Common 4-Letter Words:
The Ancient Core of English

Short words are not simple words. They are the load-bearing architecture of every sentence you speak — and most of them have been with us for over a thousand years.

~200
4-letter words in top-1000
90%+
of 4-letter commons are Old English
65%
of speech covered by 300 core words
1000+
years old — most common 4-letter words

The Architecture of Everyday Language

Count the words in the last ten sentences you spoke aloud. Chances are, more than half of them were four letters or fewer. This is not a coincidence — it is Zipf's Law made flesh. The linguist George Kingsley Zipf observed in the 1930s that the most frequently used words in any language are systematically the shortest, and that word length and word frequency follow a remarkably precise inverse relationship.

Four-letter words occupy a sweet spot in the English lexicon. They are long enough to carry distinct phonological identity (unlike a, I, of, in), short enough to have survived centuries of phonological erosion, and grammatically varied enough to include function words, verbs, adjectives, and nouns in roughly equal proportions. When linguists analyze large corpora like the British National Corpus or the Corpus of Contemporary American English (COCA), the 4-letter tier consistently produces the densest concentration of high-frequency items across all word classes.

What makes this linguistically fascinating is not just the frequency itself, but what that frequency tells us about the history of the language. The words that survived long enough to become this common are the words that were most indispensable — and that indispensability has kept them phonologically stable even as the rest of English grammar was being transformed by Norman French, Renaissance Latin borrowing, and the global spread of the modern era.

Zipf's Law and the Economy of Language

Zipf proposed that language is governed by a "principle of least effort": speakers try to communicate with minimal energy, and the most frequently needed items get compressed toward the shortest possible form. Conversely, rare words can afford to be long because the cognitive effort of retrieving them is amortized over fewer uses.

Zipf's Law in Practice

The most common word in English (the) occurs roughly twice as often as the second most common word (of), three times as often as the third (and), and so on. This power-law distribution means that knowing the top 300 words gives you coverage of about 65% of any typical English text — and the vast majority of those 300 words are under five letters long.

Among the top 200 English words ranked by corpus frequency, the average word length is approximately 3.6 letters. Four-letter words represent the longest category that still participates heavily in the very top frequency tiers.

Part-of-Speech Distribution Among Top 4-Letter Words

One of the most revealing features of the 4-letter frequency tier is how evenly it distributes across grammatical categories — unlike the very shortest words (1–2 letters), which are overwhelmingly function words, or longer words, which skew heavily toward content vocabulary.

Function Words
38%
that, with, this, from, they, your, what, when, than, then, into, some, also, been, just, only, over, such, very, even
Verbs
28%
have, will, were, been, know, make, come, give, take, find, keep, tell, seem, help, show, play, move, live, turn, feel
Nouns
22%
time, year, work, hand, part, face, home, life, back, area, room, word, name, kind, case, form, side, fact, city, body
Adjectives
12%
good, long, last, many, next, same, high, real, most, open, free, hard, full, dark, deep, near, late, huge, wide, true

The Most Common 4-Letter Words by Category

The following lists are organized by grammatical function rather than raw frequency rank, which makes them more useful for language learners and educators. Origin labels reflect the primary etymological source; most Old English words entered the language before 1100 CE.

Determiners, Pronouns & Conjunctions

These are the most frequent 4-letter words overall — the grammatical glue that holds sentences together. Native speakers typically use them unconsciously, which is precisely why learners of English as a second language find them so difficult to master: the rules governing that as determiner vs. relative pronoun vs. conjunction require years of exposure to internalize.

Top function words (4 letters)
1thatdet./conj./rel.pron.Old Englishmost common 4-letter word by a significant margin
2withprepositionOld Englishaccompaniment, means, and instrument
3thisdeterminer/pronounOld Englishproximal demonstrative; contrasts with "that" (distal)
4theypronoun (3pl)Old Norseone of the few common words of Norse origin; replaced OE hie
5fromprepositionOld Englishsource, origin, separation
6yourpronoun (2sg poss.)Old Englishoriginally second-person plural possessive; now both sg/pl
7whatinter./rel.pron.Old Englishinterrogative and relative pronoun
8whenconj./adverbOld Englishtemporal conjunction and adverb
9thanconjunctionOld Englishcomparative conjunction; historically same word as "then"
10thenadverbOld Englishtemporal: "at that time"; also logical: "in that case"
11intoprepositionOld Englishdirectional; distinguishes direction from location (in)
12somedeterminerOld Englishindefinite quantity determiner; also suffix (-some)
13alsoadverbOld Englishadditive adverb; cognate with German also (= "thus")
14onlyadverb/adjectiveOld Englishfocus particle; placement changes meaning subtly
15evenadverb/adjectiveOld Englishscalar particle and adjective; complex pragmatic functions
16veryadverbOld French verairare Latin-derived word in the top function-word tier
17justadverb/adjectiveLatin justusas adverb means "exactly/recently/simply"; as adj. means "fair"
18overprep./adverb/prefixOld Englishspatial, temporal, and aspectual functions
19suchdeterminerOld Englishdegree determiner: "of that kind/degree"
20likeprep./conj./verbOld Englishcomparison particle, verb, and now discourse marker

High-Frequency Verbs

Core 4-letter verbs
1haveauxiliary/lexicalOld Englishpossession + perfect aspect auxiliary; most irregular 4-letter verb
2willmodal auxiliaryOld Englishfuture and volition modal; historically a full lexical verb (to want)
3werepast tense (be)Old Englishpast plural + past subjunctive of be
4beenpast participle (be)Old Englishfrom a different root than be/is/was/were — suppletive paradigm
5knowstative verbOld Englishcognate with German kennen/wissen; English merges both
6makecausative verbOld Englishcreation and causation; 200+ phrasal/idiomatic uses
7comemotion verbOld Englishirregular: come/came/come; deictic movement toward speaker
8giveditransitiveOld Norsereplaced OE giefan; irregular: give/gave/given
9takecausative/motionOld Norseanother Norse replacement; antonym of give
10findperception verbOld Englishdiscovery + evaluative (find it difficult)
11keepcontinuativeOld Englishretention/continuation: keep doing, keep safe
12tellcommunicationOld Englishirregular: tell/told/told; cognate with German zählen (count)
13seemcopularOld Norseevidential copula: "appear to be"
14showperception/displayOld Englishcause to see; also intransitive (show up)
15movemotion/causativeOld Frenchone of the earlier French borrowings, well-integrated
16turndirectionalOld English/Frenchrotation and change; dozens of phrasal uses
17feelperceptionOld Englishtactile and emotional perception
18needmodal/lexicalOld Englishobligation and necessity; behaves as both modal and lexical
19playactivity verbOld Englishgames, music, theater, and light behavior
20openstate/actionOld Englishstate (the door is open) and action (open the door)

High-Frequency Nouns

time year work hand part face home life back area room word name line case kind head side body city fact form role view door mind type road book team land fire note plan data

High-Frequency Adjectives

good long last many next same high real open free hard full dark deep near late wide true huge cold bold fair flat fast warm safe dull able calm glad

The Old English Foundation

The dominance of Old English words in the 4-letter frequency tier reflects the general pattern of English vocabulary: the most common, most grammatically essential words are the oldest. When the Normans invaded in 1066 and French became the language of court, law, and prestige, the vernacular Germanic vocabulary did not disappear — it retreated into the everyday registers where it remains to this day.

Old English had a rich inflectional morphology: nouns declined for case (nominative, accusative, genitive, dative), verbs conjugated by person, number, tense, and mood, and adjectives agreed with their nouns in case, number, and gender. Over the Middle English period, this morphology was dramatically simplified, but the root words themselves survived — often with phonological changes that shortened them. The Old English word hwæt became what; þonne became than/then; þæt became that.

Proto-Germanic *þatą
Old English þæt
Middle English that
Modern English that (det./conj./pron.)
Proto-Germanic *habjaną
Old English habban
Middle English haven
Modern English have
Old Norse þeir
Northern ME thei
Modern English they
replaced OE hie

The Norse Interlopers

Several of the most common 4-letter words are not Old English but Old Norse — the language of the Viking settlers who occupied much of northern and eastern England from the 9th century onward. The pronouns they, them, their are all Norse, as are the verbs give, take, seem, call, and the noun skin. The Norse contribution is disproportionately concentrated in function words and basic vocabulary because Norse speakers and Old English speakers were in such close daily contact that grammatical items were borrowed alongside content words — a linguistic event rarely seen in recorded history.

Linguist Anatoly Liberman has argued that English is the only major language where the third-person plural pronoun system is entirely borrowed — they/them/their all come from Old Norse þeir/þeim/þeira. The original Old English forms (hie/him/hiera) became too similar to the singular masculine pronoun (he/him/his) as phonological change reduced final syllables, and Norse equivalents filled the gap.

The 4-Letter Words That Came From French

While Old English and Old Norse dominate the 4-letter frequency tier, some French-derived words have achieved sufficient frequency to join the highest ranks. These tend to be words that entered English early (before 1300) and referred to concepts with no close Old English equivalent, or that expressed abstract relationships in ways that were useful across many contexts.

also very just move real form role note turn face line type

How Other Languages Handle the Same Concepts

The grammatical machinery encoded in English's common 4-letter words is universal across languages — every language needs a way to express possession, location, temporal relationships, and basic actions. But the formal solutions differ dramatically. Comparing how different languages package these concepts illuminates both the universals of human cognition and the arbitrary accidents of linguistic history.

English
Spanish
French
German
Japanese
possession aux.have
tener
avoir
haben
ある/いる
proximity dem.this
este
ce/cet
dieser
この
temporal conj.when
cuando
quand
wenn/als
とき
creation verbmake
hacer
faire
machen
作る
3pl pronounthey
ellos/ellas
ils/elles
sie
彼ら
accompanimentwith
con
avec
mit
と/で

Notice how German mit parallels English with almost exactly — both derive from the same Proto-Germanic root *miþ. Meanwhile, French avec and Spanish con take completely different paths (Latin apud and cum respectively). This reflects the Germanic-Romance divide in European linguistics: English and German share a deep substrate of basic vocabulary, while French and Spanish diverge sharply from them in function-word territory.

Frequency Distribution: How Common Is Common?

Raw frequency numbers can be illuminating. In a corpus of one million words of typical written English, the top 4-letter words appear with stunning regularity. These figures are approximate but broadly consistent across major corpora including COCA, BNC, and the Google Ngram corpus.

that (det./conj.)
~26,000
with (prep.)
~19,000
have (aux./v.)
~18,000
this (det./pron.)
~17,000
from (prep.)
~15,000
they (pron.)
~14,300
will (modal)
~13,500
your (pron.)
~12,500
what (pron.)
~11,400
were (be-past)
~10,400
when (conj.)
~9,600
time (noun)
~7,800
know (verb)
~7,000
good (adj.)
~5,700
year (noun)
~4,900

Figures represent approximate occurrences per million words in mixed-register English text (COCA). Spoken corpora would show significantly higher frequencies for function words.

Register Variation: Spoken vs. Written

One of the most striking findings from corpus linguistics is how dramatically the frequency profile of 4-letter function words changes across registers. In spontaneous spoken conversation, words like that, with, just, like, know, yeah appear with far higher frequency than in formal written prose. Meanwhile, 4-letter nouns like time, work, life, case, form dominate academic and journalistic writing.

The word like presents a particularly fascinating case. In formal written corpora it ranks as a fairly common preposition and verb. But in spoken corpora — especially among younger speakers — it appears with extraordinary frequency as a discourse marker, quotative ("She was like, I can't believe it"), approximator ("There were like fifty people"), and hedge. This grammaticalization process — where a content word gradually takes on purely pragmatic functions — is one of the most productive mechanisms in language change, and like may be the fastest-moving example in contemporary English.

Mastering 4-Letter Words as a Language Learner

Whether you are learning English as a second language or teaching it, a sophisticated understanding of the common 4-letter word tier provides disproportionate returns. These words are not simply "easy vocabulary" to be dispatched quickly — they are cognitively and grammatically complex items whose correct use is often the last thing advanced learners achieve.

Consider even: a four-letter word that can mean "flat", "exactly", "surprisingly", "despite the fact that", or serve as an intensifier. Or just: "recently" (I just arrived), "exactly" (just right), "simply" (just do it), "only" (just one more), or "absolutely" (just beautiful). These semantic range differences — common for the most frequent words in any language — are what distinguishes native-like fluency from proficient-but-foreign performance.

Collocation First
Learn 4-letter words in their collocations, not in isolation. Don't learn make — learn make a decision, make progress, make sense. The British National Corpus online tool lets you search for the 50 most common right-hand collocates of any word.
Semantic Range Mapping
For each high-frequency 4-letter word, build a concept map of its different meanings. This is particularly critical for polysemous prepositions (over, with, from) and pragmatic particles (even, just, only).
Register Awareness
Note which 4-letter words shift meaning or frequency across registers. Like as a discourse marker is spoken-register only. Form (noun) is primarily formal-written. Corpus.byu.edu allows register comparison.
Pattern Practice
Function words like that, when, than, into are best learned through grammatical pattern drilling, not vocabulary lists. Sentence-combining exercises that require correct placement of these words build automatic control.
Phrasal Verb Focus
The most common 4-letter verbs (make, take, come, give, keep, find, show, turn) are the base of hundreds of phrasal verbs. Mastering phrasal verb families from these roots is more efficient than learning phrasal verbs alphabetically.
Spoken Frequency Immersion
Since function words appear with higher frequency in spoken than written English, podcast listening and conversation practice are more effective than reading alone for internalizing the natural rhythm of these words in connected speech.

The Schwa Problem

One of the most common pronunciation difficulties for non-native speakers involves the reduction of 4-letter function words to their weak forms in connected speech. In natural spoken English, that becomes /ðət/, were becomes /wə/, from becomes /frəm/, your becomes /jər/ — all reduced to a central vowel (the schwa /ə/). Learners who have learned these words from spelling often try to produce the "full" vowel in connected speech, which sounds unnatural and can interfere with comprehension. Training the ear to recognize reduced forms, and training the mouth to produce them, is one of the most impactful pronunciation interventions available at intermediate level.

Short Words in Other Languages: A Universal Pattern

The concentration of grammatical machinery in short, high-frequency forms is not a peculiarity of English — it is a linguistic universal observed across language families as diverse as Indo-European, Sino-Tibetan, Afroasiatic, and Austronesian. The reasons appear to be both cognitive (short words are processed faster, stored more efficiently) and communicative (high-frequency items reduce to minimum acoustic effort through phonological erosion).

Research published in the journal PNAS (Piantadosi, Tily, & Gibson, 2011, doi:10.1073/pnas.1012551108) analyzed 10 languages and found that word length inversely correlates with both frequency and contextual predictability — confirming that Zipf's observation is not an English peculiarity but a deep property of human language as a communication system optimized for efficient information transmission.

In Mandarin Chinese, many of the most common words are single morphemes of one or two syllables: 的 (de), 了 (le), 在 (zài), 是 (shì), 有 (yǒu). In Arabic, grammatical particles and common verbs tend toward the triliteral root pattern, but the most frequent items still show systematic phonological compression. Even in agglutinative languages like Finnish or Turkish — where words can grow to great length through affixation — the stems of the most common lexical items remain short.

What makes English distinctive is not the existence of short high-frequency words, but the historical layering visible within them: the Germanic stratum, the Norse interpolation, the French overlay, and the Latin-and-Greek academic register sitting on top. The 4-letter word tier sits at the interface of the two most fundamental layers — the Germanic vernacular and the Norse of the Danelaw — making it a particularly rich window into the demographic history of the language.

Explore Further

External Resources

Llexi Word of the DayA beautiful word, its story, and how to use it — daily.
Free forever · unsubscribe anytime · all 14 newsletters
That email did not go through — please check it and try again.