Which Languages Are Closest to Turkish?
- Seda
- 51 minutes ago
- 13 min read

For readers in a hurry: Turkish belongs to the Oghuz branch of the Turkic family. Its closest relatives sit inside that branch: Azerbaijani and Gagauz to the west, Turkmen further east. Which one counts as "closest" depends on what is being measured, shared ancestry, shared structure, or how much a Turkish speaker can actually follow without study, and controlled research on actual comprehension puts real, lower-than-expected numbers on that last question.
A Turkish speaker who turns on Azerbaijani television for the first time usually has the same experience. Words arrive that sound almost like their own, sometimes exactly like their own, and for a sentence or two the meaning seems to be right there. Then a verb ending twists in an unfamiliar direction, or a word that looked completely familiar turns out to mean something else, and the thread breaks. The listener is left with the sensation of nearly understanding a language they have never studied.
That is a reasonable place to start. "Which language is closest to Turkish" sounds like a question with one answer. It usually splits into three: which language shares the deepest common ancestry with Turkish, which language shares the most structure with it, and which language a Turkish speaker could actually follow in conversation. For most of the Turkic family, these answers only partly line up.
Where Turkish sits in the family
Turkish belongs to the Turkic language family and within that family to the Oghuz branch, the group that also contains Azerbaijani, Turkmen, and Gagauz. The Turkologist Lars Johanson's classification, a standard reference in the field, divides Oghuz into western and eastern groupings. A 2020 phylogenetic study published in the Journal of Language Evolution, built from lexical data across thirty-two Turkic languages, supports that division with quantitative modeling: Azeri, Turkish, and Gagauz cluster together as a West Oghuz node, while Turkmen appears as an earlier offshoot representing East Oghuz. Classifications like this vary somewhat between scholars and frameworks, and this study's authors work within Johanson's terminology specifically, so the West/East Oghuz split should be read as one well-supported model rather than a universally uncontested boundary.
Beyond Oghuz, the wider Turkic family holds several other branches within Proto-Turkic's larger family tree: Kipchak (Kazakh, Kyrgyz, Tatar, Bashkir), Karluk (Uzbek, Uyghur), Siberian Turkic (Yakut, Tuvan), and Oghuric, the earliest branch to split off.
Oghuric stands apart from the rest of the family, commonly grouped together as Common Turkic, and its only living member is Chuvash. Chuvash descends from the same Proto-Turkic ancestor as Turkish, further back, but sits outside the Common Turkic grouping that Oghuz, Kipchak, Karluk, and Siberian Turkic belong to.
Azerbaijani: a familiar Oghuz neighbor
Azerbaijani is consistently named as one of Turkish's closest relatives, and its place in West Oghuz reflects that relationship. The two languages share a large amount of core vocabulary, similar case systems, and the same basic suffix-based structure. For a Turkish speaker, many words and sentence patterns can feel familiar almost immediately.
That familiarity does not always turn into full understanding. Research on Turkish and Azerbaijani gives us some idea of the gap.
In one study, thirty Turkish university students who had never studied Azerbaijani were asked to understand both spoken and written Azerbaijani. They understood a fair amount, but much less than claims of “almost complete mutual understanding” would suggest. Later academic work puts their overall comprehension at around 42 percent.
Another study approached the question from the other direction. Mohammad Salehi and Aydin Neysani tested forty Azerbaijani speakers in Iran using short clips from Turkish television. On average, they understood about 56 percent of the spoken Turkish.
The two percentages cannot be compared directly. The studies involved different groups of people and used different tests, so they do not show that Azerbaijani speakers generally understand Turkish better than Turkish speakers understand Azerbaijani. They do, however, challenge the idea that the two languages are automatically understandable simply because they are closely related.
The Iranian Azerbaijani participants struggled more with unfamiliar vocabulary and false friends than with pronunciation differences. Many also reported that regular exposure to Turkish television had improved their understanding over time. This suggests that familiarity with another language can grow through repeated exposure, even when the languages are already closely related.
Personally, I understand Azerbaijani quite well, although there are moments when I lose the thread or come across words I don't recognize. It feels close enough that I suspect a few days surrounded by Azerbaijani would make a noticeable difference. I can already follow a good deal of it, and with that kind of daily exposure, I imagine communicating comfortably would come quite quickly. That is simply my experience as a native Turkish speaker, rather than a measure of mutual intelligibility.
Why the vocabulary drifted apart
Part of the gap has a specific historical cause. Türkiye and Azerbaijan standardized their national languages under different twentieth-century political pressures, and the vocabulary each ended up with reflects that difference.
In Türkiye, the language reform beginning in the 1930s, carried out through the newly founded Turkish Language Association, worked to replace much of the Ottoman-era Persian and Arabic vocabulary with Turkic-derived alternatives, some revived from older Anatolian and Central Asian sources, some newly coined. The process was neither complete nor entirely consistent. Large amounts of Persian and Arabic vocabulary remain in modern Turkish, but the reform did shift the language's core vocabulary away from its Ottoman predecessor.
Azerbaijani followed a different trajectory. Under Soviet rule, the language moved through repeated script changes, from a modified Arabic script to Latin, then to Cyrillic in 1939, then back to Latin after independence in 1991, and each shift brought its own wave of vocabulary standardization. Russian and Russian-mediated international vocabulary entered Azerbaijani heavily during the Soviet period, especially in technical, administrative, and scientific registers, in a way that never happened in Türkiye. Iranian Azerbaijani, spoken across the border with no Soviet history at all, followed a third path, retaining more of the Persian and Arabic vocabulary that standard Turkish moved away from.
These different standardization histories, Turkish reform, Soviet-era Azerbaijani policy, and Iranian Azerbaijani's separate path, contributed to the lexical distance speakers encounter today.
Gagauz, a West Oghuz sibling
Gagauz belongs to the same West Oghuz group as Turkish and Azerbaijani. Turkologist Astrid Menz describes it as especially closely related to Turkish, and the similarities can be striking.
This does not give us a simple answer to whether Gagauz or Azerbaijani is “closer” to Turkish. Azerbaijani has been tested for mutual intelligibility, while no comparable study exists for Gagauz. We know that Gagauz is very closely related to Turkish, but we do not have the same kind of evidence showing how much Turkish and Gagauz speakers actually understand each other.
I cannot call my own experience research, of course, but as a native Turkish speaker I find Gagauz the easiest of these languages to understand. After a minute or two of adjusting to the sound, it often feels less like listening to a different language and more like listening to a regional variety of Turkish. The distance feels comparable to what I hear when moving between Istanbul Turkish and some of the Turkish spoken in the Aegean, Central Anatolia, the Black Sea region, or eastern Anatolia. That is a personal impression rather than a measure of mutual intelligibility, but it gives some sense of how close Gagauz can sound to a Turkish ear.
Gagauz is spoken by a small Orthodox Christian Turkic community centered in Gagauzia, an autonomous region of Moldova, with related communities in Ukraine and the Balkans.
Its history makes Gagauz even more interesting. Gagauz speakers have been Orthodox Christian for centuries, probably since before the Ottoman conquest of the Balkans. Yet historians still disagree about where the community originally came from. Different theories connect them to the Turkic-speaking Bulgars of the Balkans, Kipchak-speaking Cuman groups, or Seljuk Turkic settlers who moved into the region.
And Gagauz did not develop in isolation. For centuries, its speakers lived alongside Bulgarian, Romanian, and Greek communities. Linguist Astrid Menz has documented how this long contact left its mark on Gagauz grammar. Linguists place the language within the Balkan Sprachbund, a group of neighboring languages that gradually came to share certain features even though they belong to different language families.
That gives Gagauz an unusual combination. Its foundations are clearly Turkic and remarkably close to Turkish, while centuries of life in the Balkans have also shaped the way the language works. Specialists still debate some of the finer details of these changes, especially exactly which features developed through contact and how far that influence reaches.
Turkmen, the more distant Oghuz relative
Turkmen is another close relative of Turkish, but it can sound much less familiar at first. It belongs to East Oghuz, which separated earlier from the branch that later gave us Turkish, Azerbaijani, and Gagauz. Some very old features survived in Turkmen while disappearing from its western relatives.
One of the most interesting is vowel length. In early Turkic, holding a vowel for a little longer could change the meaning of a word. Linguists reconstruct, for example, at for “horse” and āt for “name.” Turkish and Azerbaijani eventually lost this distinction, while Turkmen preserved many of these old long vowels. Historical linguist Anna Dybo has traced this pattern across the Turkic languages. Smaller traces also appear in Gagauz and Azerbaijani, while the more distant Yakut preserves the distinction much more extensively.
This means that something in modern Turkmen can preserve a sound difference that goes all the way back to Proto-Turkic, even though modern Turkish has lost it. It is a small detail, but it gives a glimpse of how differently two closely related languages can carry their shared past.
Turkmen still has the basic features that make its relationship with Turkish easy to see. It builds words with suffixes, uses vowel harmony, and shares a large amount of inherited vocabulary with other Oghuz languages. At the same time, its sound system developed differently, and Turkmen spent centuries in contact with Persian, Russian, and other Central Asian languages.
All of this makes Turkmen feel more distant from Turkish than its place in the family tree might suggest. We should be careful with that impression, though. Unlike Azerbaijani, Turkmen has not been tested against Turkish in a comparable controlled study, so there is no reliable percentage telling us how much the two languages understand each other.
My own experience with Turkmen is quite different. I can catch individual words and sometimes pieces of a sentence, but I cannot follow it nearly as easily as Gagauz or Azerbaijani. To my ear, it sometimes sounds strangely somewhere between Turkish and Russian. I can hear the Turkic connection underneath, yet the overall sound feels much more distant. Again, that is simply my impression as a native Turkish speaker, rather than a measure of how mutually intelligible the two languages are.
Familiar shape, foreign substance: the wider Turkic family
Move outside Oghuz entirely and the pattern shifts again. Kazakh and Kyrgyz belong to the Kipchak branch. Uzbek and Uyghur belong to Karluk. All of them, along with Tatar and Bashkir, still descend from the same Proto-Turkic ancestor as Turkish. The same grammatical pattern runs through all of them in recognizable form: agglutination, a fixed sequence of suffixes attached to a stable root, subject-object-verb order.
Vowel harmony itself is not applied evenly across the family today, worth noting since it is often presented as a universal Turkic constant. It belongs to the family's historical structural profile, but individual modern languages preserve or reorganize it to different degrees. Standard Uzbek, in particular, has undergone significant reduction of the older vowel harmony system, a development linked in Johanson's overview of the Turkic languages to Uzbek's long contact with Persian and Tajik and to the specific dialect base its modern literary standard draws on.
What does not carry over automatically across the family is comprehension. Centuries of separate branch-specific sound changes, independent vocabulary development, and heavy borrowing from Russian, Persian, Arabic, or Chinese through different historical channels than Turkish took, have widened the linguistic distance between Turkish and languages like Kazakh, Uzbek, or Uyghur well beyond what shared grammar alone would suggest. No controlled study comparable to the Azerbaijani research currently exists for these language pairs, and estimated percentages circulating for them online should be treated with real caution until traced to an original source with a stated method and sample. What can be said with more confidence is qualitative: related vocabulary and grammatical patterns remain visible to a Turkish speaker encountering Kazakh or Uzbek, scattered cognates, some suffixes, the general shape of a sentence, although their presence does not establish conversational intelligibility.
Branch membership is also not the whole story of distance. Tatar's centuries of close contact with Oghuz varieties, for instance, or Uyghur's very different modern contact history with Chinese, produce different profiles of closeness even within the same broader branch classification. Every non-Oghuz Turkic language is not equally far from Turkish.
My own experience becomes much more limited once I move beyond Oghuz. With Kazakh, I occasionally catch a familiar word, but I cannot really follow the language. Uzbek is even harder for me, and I understand very little Kyrgyz either.
Interestingly, all three sound much more Russian to my ear than Gagauz, Azerbaijani, or even Turkmen. I hear combinations of sounds that my ear associates with Russian, especially the harder consonants and “sh/zh”-like sounds. Of course, that is a personal impression of their sound, rather than a linguistic relationship. Kazakh, Kyrgyz, and Uzbek are Turkic languages, but as a native Turkish speaker listening without preparation, their Turkic roots are much less immediately audible to me.
Words that carry their history
A small set of cognates makes the relationship concrete. Turkish ev and Azerbaijani ev both mean “house,” and Turkmen uses öý for the same word, the same root with a shift in spelling and pronunciation. Turkish su, Azerbaijani su, and Turkmen suw all mean “water.” Turkish baş, Azerbaijani baş, and Turkmen baş all mean “head,” essentially unchanged across three branches of Oghuz. Gagauz, another West Oghuz language, preserves the same roots for all of these basic terms, part of what makes it feel immediately recognizable to Turkish speakers who encounter it.
Two questions that come up but are not about Turkic
Two other languages often come up when people search for languages related to Turkish: Finnish and Hungarian. The connection comes from the old Ural-Altaic hypothesis, developed in the nineteenth century. It proposed a distant relationship between Turkic, Mongolic, and Uralic languages, the family that includes Finnish and Hungarian.
It is easy to see why the idea once seemed convincing. Turkish, Finnish, and Hungarian share some noticeable features. They build words by adding suffixes, and all three have forms of vowel harmony. But languages can develop similar structures without coming from the same ancestor. To establish a real family relationship, linguists look for regular sound correspondences that can be traced back through time and used to reconstruct an earlier common language. That evidence has never been established between Turkic and Uralic. Today, Turkish is classified as Turkic, while Finnish and Hungarian belong to the Uralic family.
Persian is a different story. Here, the connection people notice is real, but it comes from centuries of contact. Turkish contains a substantial amount of vocabulary borrowed through Persian and Arabic, the result of a long history in which these languages were used alongside one another across the Ottoman and wider Islamic world.
Persian itself belongs to the Indo-Iranian branch of the Indo-European family, while Turkish belongs to the Turkic family. Their shared vocabulary tells us about historical contact rather than a shared linguistic ancestor. In a way, that makes the connection more interesting: the familiar words in Turkish carry traces of centuries in which Turkish and Persian speakers lived, wrote, traded, and governed within overlapping cultural worlds.
Why this matters
There is another thought I cannot quite leave alone. The ancestors of Anatolian Turks began moving westward from Central Asia roughly a thousand years ago, yet Turkish still shares some of its deepest grammatical patterns with languages spoken thousands of kilometres away. The vocabulary has changed, sounds have shifted, and the societies themselves have lived through very different histories. Still, something very old remains in the way these languages put a sentence together.
It makes me wonder about something we cannot easily measure. Language and thought can influence one another in complex ways. If generations of people grow up organizing experience through languages built on some of the same underlying patterns, are there also small similarities in how these communities make sense of the world?
Once we separate out religion, geography, political history, and all the other forces that shape a culture, is there anything left that travels with the language itself?
I do not know the answer. I am not sure we currently have a way to separate all those influences cleanly enough to find one. But when I hear an unfamiliar Turkic language and suddenly recognize the structure underneath it, I find myself wishing we could.
Vocabulary
Türkçe – the Turkish language.
Türk dilleri – the Turkic languages, the wider family Turkish belongs to.
Oğuz – the branch of the Turkic family that includes Turkish, Azerbaijani, Turkmen, and Gagauz.
akraba – relative, kin; used here for languages that share common ancestry.
ses uyumu – vowel harmony, the rule requiring a word's suffix vowels to match the vowels already in the word.
ek – suffix, the building block of Turkish's agglutinative grammar.
anlamak – to understand; the word at the center of every intelligibility question this article raises.
dil – language.
lehçe – dialect; a term whose use can vary in discussions of Turkic languages.
Add Learn Turkish with Seda as a Preferred Source on Google for more entries on Turkish language, history, and culture.
Frequently Asked Questions (FAQ)
Q: What language is closest to Turkish?
A: Within the Oghuz branch of Turkic, Azerbaijani and Gagauz are Turkish's closest West Oghuz relatives, with Turkmen representing a more easterly Oghuz development. Which one counts as closest depends on whether the question concerns shared ancestry, shared grammatical structure, or measured comprehension between speakers, and these do not always point to the same answer.
Q: Can Turkish speakers understand Azerbaijani?
A: Partially. A study by Sağın-Şimşek and König tested Turkish university students on their comprehension of spoken and written Azerbaijani. Their original paper found the widely assumed high mutual intelligibility was not borne out among Turkish speakers, and a 2022 doctoral dissertation drawing on that data states the resulting score as approximately 42 percent, below the popular assumption of near-total mutual understanding despite the languages' close structural relationship.
Q: Is Gagauz closer to Turkish than Azerbaijani?
A: Gagauz and Azerbaijani both belong to the same West Oghuz grouping as Turkish, and a specialist description by Astrid Menz calls Gagauz most closely related to Turkish. No controlled intelligibility study comparable to the Azerbaijani research exists for Gagauz, so this description and the measured Azerbaijani figures answer different kinds of questions and cannot be ranked directly against each other.
Q: Can Turkish speakers understand Kazakh or Uzbek?
A: Only partially and inconsistently. These languages share deep Turkic architecture, agglutination and suffix-based grammar among them, but belong to different branches (Kipchak and Karluk) with centuries of separate vocabulary development. Related vocabulary and grammatical patterns remain visible to a Turkish speaker, but their presence does not establish conversational intelligibility, and no controlled study comparable to the Azerbaijani research currently exists for these pairs.
Q: Are Turkish and Hungarian related?
A: No. Hungarian belongs to the Uralic language family and Turkish to Turkic, two separate families. The idea that they share ancestry comes from the nineteenth-century Ural-Altaic hypothesis, which has not been supported by the systematic sound correspondences that establishing genuine language relationship requires.
The television is still on when the segment ends and the anchor moves to another story. The words keep arriving the way they did at the start, some landing right away, some slipping past before they land at all.
Sources
Menz, Astrid. "Gagauz," in The Turkic Languages, 2nd edition, edited by Lars Johanson and Éva Ágnes Csató. Taylor & Francis. https://www.taylorfrancis.com/chapters/edit/10.4324/9781003243809-16/gagauz-astrid-menz
Savelyev, Alexander, and Martine Robbeets. "Bayesian phylolinguistics infers the internal structure and the time-depth of the Turkic language family." Journal of Language Evolution, vol. 5, no. 1, 2020, pp. 39–53. https://doi.org/10.1093/jole/lzz010
Sağın-Şimşek, Çiğdem, and Wolf König. "Receptive multilingualism and language understanding: Intelligibility of Azerbaijani to Turkish speakers." International Journal of Bilingualism, vol. 16, no. 3, 2012, pp. 315–331. https://doi.org/10.1177/1367006911426449
Salehi, Mohammad, and Aydin Neysani. "Receptive intelligibility of Turkish to Iranian-Azerbaijani speakers." Cogent Education, vol. 4, no. 1, 2017. https://doi.org/10.1080/2331186X.2017.1326653
Dybo, Anna V. "On the Reflexes of Proto-Turkic Vowel Length in the Turkic Languages." Studia Linguistica Universitatis Iagellonicae Cracoviensis, vol. 132, 2015, pp. 121–134. https://doi.org/10.4467/20834624SL.15.013.3934 (Full text: https://ejournals.eu/pliki_artykulu_czasopisma/pelny_tekst/4223ff8b-e0fa-4d78-af18-f3aa3870ecbb/pobierz)
Perry, John R. "Turkic-Iranian Contacts i. Linguistic Contacts." Encyclopaedia Iranica, published August 15, 2006. https://www.iranicaonline.org/articles/turkic-iranian-contacts-i-linguistic/
Afshar, Naeimeh. "Production and Perceptual Representation of American English Vowel Sounds by Monolingual Persian and Early Bilingual Azerbaijani-Persian Adolescents." PhD dissertation, University of Pannonia, Veszprém, 2022. https://doi.org/10.18136/PE.2022.825 (Cited for the 42 percent intelligibility figure, given in its literature review, derived from Sağın-Şimşek & König's 2012 data.)



Comments