search this blog

Saturday, April 30, 2016

Y-hg J2 cannot be a Proto-Indo-European marker


The claim that the Proto-Indo-Europeans came from West Asia and largely belonged to Y-haplogroup J2 seems to be popular online nowadays. I won't discuss here in detail the reasons why, but suffice to say it has a lot do with aggressive lobbying on several online forums and blogs by a few people of Southern European extraction, like Dienekes Pontikos.

It was always a shaky proposition, but difficult to debunk thoroughly. Until now.

Thanks to recent advances in both modern and ancient DNA research, we can now safely say that Y-haplogroup J2 was not involved in any rapid, large scale population expansions during the Late Neolithic/Early Bronze Age (LN/EBA), the generally accepted Proto-Indo-European time frame.

It thus fails to meet even the most basic criteria of a Proto-Indo-European diagnostic marker. The Proto-Indo-Europeans, after all, were surely highly patriarchal and patrilineal, and therefore expected to have left a clear signal of their migrations in the Y-chromosomes of many present-day Indo-European speakers.

For instance, an analysis of data from the deep sequencing of human Y-chromosomes as part of the 1000 Genomes Project suggests that not a single major subclade of J2 began expanding even roughly close to the LN/EBA. See here.

In the plot above three lineages jump out at you. E1b, R1a, and R1b. The first is associated with the Bantu expansion, that occurred over the last 4,000 years. The second two are likely associated with Indo-Europeans in both Asia and Europe, respectively. The timescale is on the order of 4 to 5,000 years in the past. The association between culture and genes, or the genetic lineages of males, is rather clear, in these cases. In other instances the growth was more gradual. For example, the lineages likely associated with the first Neolithic pulses, J and G.

Moreover, not a single instance of J2 has been reported from remains classified as belonging to the Andronovo, Battle-Axe, Corded Ware, Khvalynsk, Poltavka, Potapovka, Sintashta, Srubnaya and Yamnaya archaeological cultures. In other words, Kurgan and Kurgan-derived groups generally accepted to be early Indo-European, whch is a view that now has very strong support from ancient genomics. See here and here.

To date, most of these samples have probably come from elite burials. So at some point, when many more non-elite samples are sequenced, we are likely to see J2 among a few supposedly early Indo-European individuals. But so what?

There might be a couple of ways to salvage the Proto-Indo-Europeans = J2 theory. We'd have to argue that...

- the Proto-Indo-European time frame was actually the early Neolithic

and/or

- the Proto-Indo-Europeans were a small group that Indo-Europeanized the steppe Kurgan people, perhaps mainly via female migrations, and then did not partake in the main early Indo-European expansions

But the former is not particularly clever when viewed in the context of historical linguistics data. See here.

For instance, almost all IE language branches testify to a word designating ‘wool’. Since archaeological evidence suggests that wool sheep did not exist until the beginning of the fourth millennium BCE, the existence of the word in PIE would indicate that the disintegration of the proto-language could not have taken place before this date. Similarly, words for concepts such as ‘wheel’, ‘yoke’, ‘honey bee’ and ‘horse’ may be correlated directly with concrete, datable archaeological evidence.

And the latter isn't very parsimonious, and to me looks like special pleading. Why even bother?

Monday, April 25, 2016

Signals of ancient population explosions in our Y-chromosomes


Nature Genetics has a massive new paper on human Y-chromosomes based on the latest 1000 Genomes data. I'm still getting my head around the details, but at first glance it looks like a very capable effort. This part basically reads like some of my blog entries in recent years. The emphasis is mine.

In South Asia, we detected eight lineage expansions dating to ~4.0–7.3 kya and involving haplogroups H1-M52, L-M11, and R1a-Z93 (Supplementary Fig. 14b,d,e). The most striking were expansions within R1a-Z93, occurring 4.0–4.5 kya. This time predates by a few centuries the collapse of the Indus Valley Civilization, associated by some with the historical migration of Indo-European speakers from the Western Steppe into the Indian subcontinent 27. There is a notable parallel with events in Europe, and future aDNA evidence may prove to be as informative as it has been in Europe.

Poznik et al., Punctuated bursts in human male demography inferred from 1,244 worldwide Y-chromosome sequences, Nature Genetics, Published online 25 April 2016; doi:10.1038/ng.3559

See also...

The Poltavka outlier


Sunday, April 17, 2016

Estimating Basal Eurasian ancestry?


Basal Eurasians (BE) are a hypothetical ghost population that apparently split from other Eurasians no later than 45,000 years ago. If they actually existed, they had a significant impact on the ancestry of early Neolithic farmers, and thus all present-day West Eurasians.

Testing ancestry proportions from ghost populations isn't easy. However, Haak et al. 2015 made use of an f4 equation that seemingly gave an accurate estimate of BE admixture in LBK farmer Stuttgart: f4(Stuttgart,Loschbour;Onge,MA1)/f4(Mbuti,MA1;Onge,Loschbour) = 44%. The other LBK farmers scored an average of 40% BE, which also made sense.

Unfortunately, this equation doesn't appear to work too well for Caucasus Hunter-Gatherers (CHG) Kotias and Satsurblia. They both score around 25% BE, which, as far as I can see, seems way too low. Perhaps using MA1 in the equation is messing things up because CHG harbor significant MA1-related ancestry?

I tinkered around with Haak's equation and came up with this: f4(X,Iberia_Mesolithic;Dai,Karelia_HG)/f4(Mbuti,Karelia_HG;Dai,Iberia_Mesolithic). The results look solid, at least in relative terms (see image below). But is the equation actually valid?

My main worry is using both Iberia Mesolithic and Karelia HG. They share a lot of drift, much more than Loschbour and MA1. Also, even though both Dai and Onge belong to the so called Eastern non-African (ENA) clade, they're quite distinct, with Dai a lot less basal in the context of ENA diversity. Any thoughts? Suggestions?


Update 04/18/2016: Interestingly, my f4 equation essentially fails for most post-Neolithic Europeans, particularly those with relatively high ratios of Karelia HG-related ancestry. For instance, Yamnaya Kalmykia scores just 2.9% BE, which can't be right. Yamnaya Samara shows -2.2%, which is obviously wrong.

But I tried several combinations of reference samples and found that by replacing Karelia HG with Hungary HG and Dai with Ust-Ishim I was able to obtain coherent results for a wider range of groups, including Yamnaya.


To be honest, I still don't know what the hell I'm testing here exactly. The results appear to reflect the existence of two components within West Eurasia; one representing ancient hunter-gatherers from Europe and probably surrounding areas of the Near East, and another closely related to present-day Near Eastern populations. The latter might well be a signal of the so called Basal Eurasians, or perhaps a number of as yet unsampled meta populations from the ancient Near East?

Monday, March 28, 2016

PCA/nMonte open thread


Below are a few nMonte models of ancient individuals based on 25 principal components (PCs). The relevant datasheet and nMonte R script can be downloaded here and here, respectively.


Many of the outcomes are basically perfect. Others could certainly be better. But they all make sense.

The more complex the ancestry, the more difficult it is to model. Also, deamination, low coverage and missing markers are probably skewing things to some degree for most of these samples. So although time consuming, it might be a good idea to use population averages minus the most obvious outliers.

Are there any other ways to improve the analysis? Is 25 dimensions too much or too little? Let's run plenty of tests and see where this takes us.

I can update the datasheet with many more populations and dimensions later this week. Feel free to post your requests in the comments and I'll run them if I have them. Also, if anyone's wondering, I don't know yet which commercial genotype files I can run in this test, if any. I'll check.

Update 04/04/2016: A modified datasheet with 50 dimensions and many more samples is available here. It should be more useful in modeling South Central Asians, especially the Kalash. However, as far as I can tell, using just 9 dimensions, like in the version here, is faster and produces more accurate results.

Wednesday, March 16, 2016

Sintashta, BMAC and the Indo-Iranians


I'm perusing the online archives of Harvard Sanskrit Professor Michael Witzel. The links below are worth checking out for some background info on the prehistory of Eastern Europe and Central Asia. There's a very cool map on page 6 of the second PDF.

Sintashta, BMAC and the Indo-Iranians. A query.

Linguistic Evidence for Cultural Exchange in Prehistoric Western Central Asia.

The Home of the Aryans.

Autochthonous Aryans? The Evidence from Old Indian and Iranian Texts.

Looking back, these old school linguistics articles make a lot more sense than most of the supposedly cutting-edge population genetics papers coming out at around the same time dealing with the Indo-Aryan question.

Many population geneticists back then took the view that the ancestors of the Indo-Aryans could not have spread from the European steppes to India because Y-chromosome haplogroup R1a apparently showed the greatest haplotype diversity in the Indus Valley. Well, what a load of crock that turned out to be.

See also...

The Poltavka outlier

Friday, March 11, 2016

D-stats/nMonte open thread


I'll start the ball rolling with a 9-way mixture analysis of 93 European, Near Eastern and Central Asian present-day and ancient populations. The relevant datasheet and R script are available here and here.


Below is a simple tree/cluster analysis based on the results, using the freely available Past3 software. Makes perfect sense, I'd say.


It's important to understand that these sorts of tests are basically designed to estimate ancient ancestry proportions, rather than calculate minor admixtures. With that in mind, here are a few observations:

- Karasuk outlier RISE497 (the most eastern Karasuk individual) is surprisingly important for Near Eastern populations

- Nordic LNBA and Sintashta look very similar in terms of overall ancestry proportions, suggesting that they perhaps derive from the same ancestral population

- The effects of postmortem deaminantion or DNA damage appear to be expressed in many of the non-UDG treated ancient samples as minor Sub-Saharan admixture

Can anyone put together a better model for West Eurasians? Also, I'd really like to see a well thought out D-stats/nMonte analysis of South Central Asia.

See also...

Yamnaya = Khvalynsk + extra CHG + maybe something else

D-stats/nMonte open thread #2

Sunday, March 6, 2016

D-stats/4mix tour of ancient Eurasia


This 4mix experiment is based on a series of statistics of the form D(Chimp,Reference_pop/Test_pop)(Mbuti,X), where X represents one of 9 ancient and present-day outgroups. The input data is available here. Feel free to try it yourself and post your models in the comments below.


Here's a Principal Component Analysis (PCA) based on the D-stats. As far as I can see, it makes very good sense. Click to enlarge.

See also...

Yamnaya = Khvalynsk + extra CHG + maybe something else

PC/nMonte open thread

Thursday, March 3, 2016

Irano-Turko-Slavic roots of Ashkenazi Jews?


As far as I've been able to discern, Ashkenazi Jews are very similar to Sephardic Jews, except with minor admixture from Central and Eastern Europe, and perhaps Central Asia (via the Silk Road). So the hypothesis presented in this new paper at Genome Biology and Evolution doesn't work for me:

The Yiddish language is over one thousand years old and incorporates German, Slavic, and Hebrew elements. The prevalent view claims Yiddish has a German origin, whereas the opposing view posits a Slavic origin with strong Iranian and weak Turkic substrata. One of the major difficulties in deciding between these hypotheses is the unknown geographical origin of Yiddish speaking Ashkenazic Jews (AJs). An analysis of 393 Ashkenazic, Iranian, and mountain Jews and over 600 non-Jewish genomes demonstrated that Greeks, Romans, Iranians, and Turks exhibit the highest genetic similarity with AJs. The Geographic Population Structure (GPS) analysis localized most AJs along major primeval trade routes in northeastern Turkey adjacent to primeval villages with names that may be derived from "Ashkenaz." Iranian and mountain Jews were localized along trade routes on the Turkey's eastern border. Loss of maternal haplogroups was evident in non-Yiddish speaking AJs. Our results suggest that AJs originated from a Slavo-Iranian confederation, which the Jews call "Ashkenazic" (i.e., "Scythian"), though these Jews probably spoke Persian and/or Ossete. This is compatible with linguistic evidence suggesting that Yiddish is a Slavic language created by Irano-Turko-Slavic Jewish merchants along the Silk Roads as a cryptic trade language, spoken only by its originators to gain an advantage in trade. Later, in the 9th century, Yiddish underwent relexification by adopting a new vocabulary that consists of a minority of German and Hebrew and a majority of newly coined Germanoid and Hebroid elements that replaced most of the original Eastern Slavic and Sorbian vocabularies, while keeping the original grammars intact.

Das et al., Localizing Ashkenazic Jews to primeval villages in the ancient Iranian lands of Ashkenaz, Genome Biol Evol (2016), doi: 10.1093/gbe/evw046

See also...

Khazar shmazar

Khazar shmazar #2

Wednesday, February 24, 2016

Ancient DNA from early Medieval Muslim graves in France


A new paper at PLoS ONE reveals that three individuals from early Medieval burials in southern France belong to Y-chromosome haplogroup E1b1b1b-M81 and mtDNA haplogroups H1, K1a4a and L1c3a, and thus were probably of North African origin.

Abstract: The rapid Arab-Islamic conquest during the early Middle Ages led to major political and cultural changes in the Mediterranean world. Although the early medieval Muslim presence in the Iberian Peninsula is now well documented, based in the evaluation of archeological and historical sources, the Muslim expansion in the area north of the Pyrenees has only been documented so far through textual sources or rare archaeological data. Our study provides the first archaeo-anthropological testimony of the Muslim establishment in South of France through the multidisciplinary analysis of three graves excavated at Nimes. First, we argue in favor of burials that followed Islamic rites and then note the presence of a community practicing Muslim traditions in Nimes. Second, the radiometric dates obtained from all three human skeletons (between the 7th and the 9th centuries AD) echo historical sources documenting an early Muslim presence in southern Gaul (i.e., the first half of 8th century AD). Finally, palaeogenomic analyses conducted on the human remains provide arguments in favor of a North African ancestry of the three individuals, at least considering the paternal lineages. Given all of these data, we propose that the skeletons from the Nimes burials belonged to Berbers integrated into the Umayyad army during the Arab expansion in North Africa. Our discovery not only discusses the first anthropological and genetic data concerning the Muslim occupation of the Visigothic territory of Septimania but also highlights the complexity of the relationship between the two communities during this period.



Gleize Y, Mendisco F, Pemonge M-H, Hubert C, Groppi A, Houix B, et al. (2016) Early Medieval Muslim Graves in France: First Archaeological, Anthropological and Palaeogenomic Evidence. PLoS ONE 11(2): e0148583. doi:10.1371/journal.pone.0148583

Tuesday, February 9, 2016

CHG admixture in early western Anatolian farmers


Anatolian Neolithic farmer I0708 from the Mathieson et al. 2015 dataset belongs to Y-haplogroup J2a and is the most Caucasus-shifted of the early Anatolian farmers in my Principal Component Analysis (PCA) of West Eurasia (see below). This is unlikely to be a coincidence and provides strong evidence that at least some Neolithic farmers in western Anatolia harbored Caucasus Hunter-Gatherer (CHG) ancestry.


Note that the two CHG genomes sequenced to date courtesy of Jones et al. 2015, Kotias and Satsurblia, belonged to Y-haplogroups J and J2a. Moreover, J2 today shows peaks in frequency and diversity in and around the Caucasus. In other words, Y-haplogroup J, and in particular J2, appear to represent paternal signals of CHG admixture.

Unfortunately, it's not yet possible to demonstrate with formal tests beyond any doubt that I0708 has CHG admixture.

For instance, the D-stats below, in which a couple of the least Caucasus-shifted Anatolian farmers are Anatolia Neolithic1, while I0708 is Anatolia Neolithic2, fail to reach significance (Z=3). Please note, I ran the stats with the Amerindian and Siberian samples to test for Ancient North Eurasian (ANE) admixture, which appears to be a feature of CHG.

However, the results are all clearly positive, and might reach significance with higher quality data and/or a better reference than Anatolia Neolithic1.


Indeed, the subtle difference in ANE affinity between Anatolia Neolithic1 and Anatolia Neolithic2 is underlined by the D-stats below. Note that here Kotias shows significant signals of admixture when paired with Anatolia Neolithic1, but not when paired with Anatolia Neolithic2. This is despite the fact that Anatolia Neolithic2 is a higher coverage sequence (6.95x vs 2.66x) and offers more markers.


I0708 is unlikely to be the only early western Anatolian farmer with CHG/ANE admixture. The PCA above show that a couple of others are also pulling strongly towards the Caucasus. Indeed, all of the Anatolian and European Neolithic samples might harbor low levels of CHG ancestry. The problem with testing this idea at present is a lack of more basal Near Eastern ancient genomes from core areas of the Near East, like, say, the Levant.

Hopefully they're on their way, but in any case, it's almost certain now that CHG was already expanding west, and in all likelihood east, during the early Neolithic. This probably has some important implications for the peopling of West Eurasia and their linguistic affinities. Feel free to post what these might be in the comments.

Update 11/02/2016: I came up with new Anatolia Neolithic1 and Anatolia Neolithic2 sets using D-stats (by comparing each of the Anatolians to Kotias versus sample I0708). For a breakdown see here. Anatolia Neolithic2 now shows significant signals of admixture from Kotias, Dai, Surui and Han. This implies that it not only harbors CHG ancestry, but also ANE and East Asian-related admixtures.