search this blog

Showing posts with label qpAdm. Show all posts
Showing posts with label qpAdm. Show all posts

Thursday, February 22, 2024

Berkeley, we have a problem


A new preprint at bioRxiv by Kerdoncuff et al. makes the following, somewhat surprising, claim:

One of the individuals, referred to Sarazm_EN_1 (I4290) described above that was discovered with shell bangles showing affiliation with South Asia, has significant amount AHG-related ancestry, while a model without AHG-related ancestry provides the best fit for Sarazm_EN_2 (I4210) (Table S4.5).

First of all, the authors are actually referring to sample ID I4910 not I4210.

The aforementioned table, based on qpAdm output, shows that I4290 has 15.9% AHG-related ancestry and basically no Anatolian farmer-related ancestry. It also shows that I4910 has no AHG-related ancestry but 17.9% Anatolian farmer-related ancestry.

AHG stands for Andaman hunter-gatherer. The authors are using it as a proxy for South Asian hunter-gatherer ancestry.

However, I've looked at I4290 and I4910 in great detail over the years using ADMIXTURE, Principal Component Analysis (PCA), and qpAdm. And I'm quite certain that they do not show any obvious, above noise level South Asian ancestry. Indeed, I'd say that if they do have some minor South Asian ancestry, then I4910 probably has more of it than I4290.

Kerdoncuff et al. used the following "right pops" or outgroups: Ethiopia_4500BP.SG, WEHG, EEHG, ESHG, Dai.DG, Russia_Ust_Ishim_HG.DG, Iran_Mesolithic_BeltCave and Israel_Natufian.

This means they mixed data that were generated in very different ways (DG, SG and capture) and included some poor quality samples. For instance, the highest coverage version of Iran_Mesolithic_BeltCave offers just ~50K SNPs.

Mixing different types of data and relying on low coverage samples, even in part, often has negative consequences when using qpAdm. So I suspect that the above mentioned mixture results for I4290 are skewed by a poor choice of outgroups.

When I run qpAdm I try to stick to one type of data and avoid low quality singletons in the outgroups. This is the best qpAdm model that I can find for Sarazm_EN:

right pops:
Cameroon_SMA
Morocco_Iberomaurusian
Israel_Natufian
Levant_N
Iran_GanjDareh_N
Turkey_N
Russia_Karelia_HG
Russia_WestSiberia_HG
Mongolia_North_N
Brazil_LapaDoSanto_9600BP

Sarazm_EN
Kazakhstan_Botai_Eneolithic 0.113±0.017
Turkmenistan_C_Geoksyur_subset 0.887±0.017
P-value 0.06392

Sarazm_EN_1 (I4290)
Kazakhstan_Botai_Eneolithic 0.129±0.021
Turkmenistan_C_Geoksyur_subset 0.871±0.021
P-value 0.11019

Sarazm_EN_2 (I4910)
Kazakhstan_Botai_Eneolithic 0.104±0.021
Turkmenistan_C_Geoksyur_subset 0.896±0.021
P-value 0.07427

Also...

Sarazm_EN
Andaman_hunter-gatherer -0.018±0.020
Kazakhstan_Botai_Eneolithic 0.123±0.019
Turkmenistan_C_Geoksyur_subset 0.895±0.020
P-value 0.0298403
(Infeasible model)

Please note that Turkmenistan_C_Geoksyur_subset is made up of just three relatively high quality individuals: I8504, I12483 and I12487. That's because it's not possible to model the ancestry of Sarazm_EN using the full Geoksyur set, probably due to subtle genetic substructures within the latter.

Below is a PCA plot that, more or less, reflects my qpAdm model. I4290 and I4910 are sitting right next to each other in a cluster of ancient Central and Western Asians, and it's actually I4910 that is shifted slightly towards the South Asian pole of the PCA. Indeed, I can confidently say that there's no way to design a PCA in which I4290 is shifted significantly towards South Asia relative to I4910.

Citation...

Kerdoncuff et al., 50,000 years of Evolutionary History of India: Insights from ∼2,700 Whole Genome Sequences, bioRxiv, posted February 20, 2024, doi: https://doi.org/10.1101/2024.02.15.580575

See also...

The Nalchik surprise

A comedy of errors

Sunday, July 23, 2023

Dear Sandra, Wolfgang...a problem


In their recent paper, titled Early contact between late farming and pastoralist societies in southeastern Europe, Penske et al. make the following claim:

By contrast, Yamnaya Caucasus individuals from the southern steppe can be modelled as a two-way model of around 76% Steppe Eneolithic and 26% Caucasus Eneolithic/Maykop, confirming the findings of Lazaridis and colleagues 47. This two-way mix (40% + 60%, respectively) also provides a well-fit model (P = 0.09) for the Ozera outlier individual, consistent with the position in PCA and corroborating an influence from the Caucasus.

Err, nope.

The Ozera Yamnaya outlier, a female dated to 3096-2913 calBCE, is, in fact, a ~50/50 mix between standard Yamnaya and Late Maykop. It's a result that is totally unambiguous.

There are a number of ways to demonstrate this fact. For example, with the qpAdm software that was also used by Penske et al., except with different outgroups or right pops. Please note that in my dataset the Ozera outlier is labeled Ukraine_Ozera_EBA_Yamnaya_o.

right pops:
Cameroon_SMA
Levant_N
Iran_GanjDareh_N
Iran_C_SehGabi
Georgia_HG
Turkey_N
Serbia_IronGates_Mesolithic
Russia_WestSiberia_HG
Russia_Karelia_HG
Latvia_HG
Russia_Boisman_MN
Brazil_LapaDoSanto_9600BP

Ukraine_Ozera_EBA_Yamnaya_o
Russia_Caucasus_EneolithicMaykop 0.554±0.031
Russia_Steppe_Eneolithic 0.446±0.031
P-value 0.00109868 (FAIL)


Ukraine_Ozera_EBA_Yamnaya_o
Russia_LateMaykop 0.512±0.035
Russia_Samara_EBA_Yamnaya 0.488±0.035
P-value 0.462447 (PASS)

I can also do it with the Global25/Vahaduo method. And you, dear reader, can too, by putting the Target and Source Global25 coords from the text file here into the relevant fields here.

Target: Ukraine_Ozera_EBA_Yamnaya_o
Distance: 2.9292% / 0.02929202
50.6 Russia_Samara_EBA_Yamnaya
49.4 Russia_Caucasus_LateMaykop
0.0 Russia_Caucasus_EneolithicMaykop
0.0 Russia_Steppe_Eneolithic

Moreover, here's a self-explanatory Principal Component Analysis (PCA) plot that illustrates why my Late Maykop/Samara Yamnaya combo is much better than the reference populations used by Penske and colleagues. It was done with the PCA tools here.
I'm pointing this out for two main reasons. First of all, this is a fairly obvious mistake that should've been avoided, especially considering the level of expertise and experience among the authors (such as Wolfgang Haak and Johannes Krause).

Secondly, it's important to understand that the Ozera outlier comes out almost exactly 50% Samara Yamnaya because the standard Yamnaya genotype already existed well before she was alive, and thus she cannot be used to corroborate any sort of influence from the Caucasus in the formation of the mainstream Yamnaya population.


As for the Yamnaya Caucasus individuals, I don't know why Penske et al. attempted to model their ancestry as a group, because they don't form a coherent genetic cluster. RK1001 and ZO2002 are fairly similar to standard Yamnaya samples, while RK1007 and SA6010 resemble Eneolithic steppe samples from the Progress burial site. This is what happens when I try to reproduce the Penske et al. model with my outgroups.

Russia_Caucasus_EBA_Yamnaya
Russia_Caucasus_EneolithicMaykop 0.187±0.019
Russia_Steppe_Eneolithic 0.813±0.019
P-value 4.15842e-06 (HARD FAIL)

Oh, and Penske et al. modeled the ancestry of mainstream Yamnaya as a three-way mixture with Steppe Eneolithic, Caucasus Eneolithic/Maykop and Ukraine Neolithic (or Ukraine N). They succeeded, but with my outgroups it's another hard fail.

Russia_Samara_EBA_Yamnaya
Russia_Caucasus_EneolithicMaykop 0.177±0.017
Russia_Steppe_Eneolithic 0.706±0.026
Ukraine_N 0.116±0.014
P-value 4.73919e-07 (HARD FAIL)

Admittedly, proximal models aren't easy to get right. And if you throw enough outgroups into a model, a large proportion of plausible models will fail. But I'm somewhat taken aback by these poor statistical fits.

In my opinion, mainstream Yamnaya doesn't harbor any Caucasus ancestry that wasn't already present on the Pontic-Caspian steppe during the Eneolithic or even much earlier (see here). But ultimately this problem can only be solved with direct evidence from ancient DNA, so let's now wait patiently for the right samples.

Citation...

Penske et al., Early contact between late farming and pastoralist societies in southeastern Europe, Nature, https://doi.org/10.1038/s41586-023-06334-8

See also...

Understanding the Eneolithic steppe

Wednesday, August 19, 2020

Yamnaya-related ancestry proportions in present-day Poles


Modeling ancient ancestry proportions in present-day Europeans with the qpAdm software is now a lot more difficult. The reasons for this are updates to qpAdm as well as the availabiity of more useuful outgroups or right pops.

This isn't necessarily a bad thing, because users are forced to work harder to find successful models, which is likely to lead to some interesting discoveries. But it can be very frustrating.

I don't think that settling for poor statistical fits or using a small number of outrgoups are acceptable short cuts. Perhaps sequencing modern-day samples in exactly the same way as the ancient samples, and thus increasing the compatability between them, might help?

Limiting qpAdm runs to higher quality SNPs from transversion sites does help, but perhaps largely because of the significant reduction in markers?

In any case, I've now given up on running such analyses, at least until I see some serious pointers on the topic from Harvard's qpAdm experts. But before I put this project to bed for the time being, I'd like to share some new results for Poles from eastern and western Poland, respectively.

right pops:

CMR_Shum_Laka_8000BP
MAR_Taforalt
IRN_Ganj_Dareh_N
Levant_PPNB
GEO_CHG
TUR_Barcin_N
RUS_Piedmont_En
SRB_Iron_Gates_HG
WHG
RUS_Karelia_HG
MNG_North_N
RUS_Ust_Kyakhta

left pops:

Polish_East
CWC_Baltic_early 0.572±0.024
SWE_TRB 0.428±0.024
chisq 11.776
tail prob 0.300296
Full output

Polish_West
CWC_Baltic_early 0.587±0.021
SWE_TRB 0.413±0.021
chisq 11.165
tail prob 0.34478
Full output


Even using transversion sites, this is one of the very few combinations of ancient reference samples that works for the Poles with these right pops. That is, the combination of early Corded Ware samples from the East Baltic (CWC_Baltic_early) and Funnel Beaker samples from Scandinavia (SWE_TRB). The former are obviously the proxy here for Yamnaya-related ancestry.

Adding any sort of hunter-gatherer population to this model doesn't help or even makes things worse (for instance, see here and here). It is possible to add Baltic hunter-gatherers to a similar model after dropping CWC_Baltic_early in favor of closely related samples from the Early to Middle Bronze Age Pontic-Caspian steppe. Note, however, that the statistical fits are somewhat poorer.

Polish_East
Baltic_LTU_Narva 0.032±0.014
PC_steppe_EMBA 0.483±0.019
SWE_TRB 0.485±0.019
chisq 17.143
tail prob 0.0465198
Full output

Polish_West
Baltic_LTU_Narva 0.031±0.011
PC_steppe_EMBA 0.491±0.015
SWE_TRB 0.477±0.016
chisq 22.444
tail prob 0.00757421
Full output


Interestingly, but not surprisingly, the ancestry of many present-day Northwestern European populations can be modeled in basically the same way. That's because ancient ancestry proportions are more closely correlated with latitude than longitude across much of the European continent.

English_Kent
CWC_Baltic_early 0.527±0.024
SWE_TRB 0.473±0.024
chisq 13.042
tail prob 0.221357
Full output

Icelandic
CWC_Baltic_early 0.586±0.023
SWE_TRB 0.414±0.023
chisq 16.517
tail prob 0.085751
Full output

Scottish
CWC_Baltic_early 0.583±0.021
SWE_TRB 0.417±0.021
chisq 12.144
tail prob 0.275536
Full output


A zip file with the qpAdm output from this analysis and a list of the most relevant ancients is available here. I might try to run a few more populations over the next few days, but probably only from the northern half of Europe, so please check the zip file in a week or so to see what else is in there.

If anyone wants to challenge my results, note that these and very similar samples are freely available to the public via Harvard University here and here.

Update 22/08/2020: From Nick Patterson (Broad) in the comments:
My general advice for qpAdm is 1) Work on the right hand set. Don't include irrelevant population (except for one population as an outgroup); picking the best RHS can dramatically reduce s. errors on the admixture weights. 2) If qpAdm gives a very low p-value try and understand why, sometimes it is telling you that the target is not a mixture of the sources but sometimes the assumptions are violated, for example recent gene-flow from left pops -> right.

See also...

Ancient ancestry proportions in present-day Europeans

Monday, July 27, 2020

Ancient ancestry proportions in present-day Europeans (to be continued)


This year has already been massive in all sorts of ways, including for new data and software releases. So I'm thinking it might be time to update many of the analyses that were featured at this blog a while ago.

Let's start with the classic hunter vs farmer vs herder mixture model for present-day European populations. The rules of the game are as follows:


- run the latest version of qpAdm using qpfstats output

- use transversion sites and 1240K capture data

- pick a set of diverse and chronologically sound outgroups

- for a model to be successful the p-value must reach 0.01

- tweak the left pops in models that are clearly underperforming

- follow high end scientific literature, logic and common sense


Obviously, the reason that I decided to limit my analysis to markers from transversion sites is to mitigate problems associated with modeling the ancestry of modern, high quality samples with relatively low quality ancients. One of these problems appears to be qpAdm assigning faux East Asian/Siberian admixture to present-day Europeans (for instance, see figure 4 here).

My starting reference populations and outgroups are listed below. In qpAdm terminology the former are known as the "left pops", while the latter as the "right pops". Most of these samples are freely available at the David Reich Lab website here.

left pops:
HUN_Koros_N_HG
TUR_Barcin_N
UKR_Yamnaya

right pops:
CMR_Shum_Laka_8000BP
MAR_Taforalt
Levant_Natufian
IRN_Ganj_Dareh_N
Levant_PPNB
CZE_Vestonice16
BEL_GoyetQ116-1
Iberia_ElMiron
RUS_Karelia_HG
RUS_West_Siberia_HG
MNG_North_N
RUS_Ust_Kyakhta

As you can see, I picked a wide variety of right pops. But I chose most of them specifically to be able to differentiate the three streams of ancestry - from ancient hunters, farmers and herders - that are the focus of my analysis. I also intentionally avoided using samples in the right pops that may have experienced gene flow, including cryptic gene flow, from the populations in the left pops.

I somewhat speculatively earmarked HUN_Koros_N_HG, from the Early Neolithic Carpathian Basin, and UKR_Yamnaya, from the Early Bronze Age North Pontic steppe in what is now Ukraine, to represent the hunter-gatherer and pastoralist streams of ancestry, respectively.

That's because I expected HUN_Koros_N_HG to be the best proxy for the hunter-gatherer ancestry that was initially absorbed by the early farmers who fanned out from the Aegean region across much of the European continent, and of course it made sense to choose a steppe pastoralist population that was located close to Central Europe where such groups first made the biggest impact outside of the steppe.

Interestingly, HUN_Koros_N_HG and UKR_Yamnaya did prove to be among most effective choices for the types of ancestries that they represented. For instance, UKR_Yamnaya generally produced much stronger statistical fits than a very similar set of Yamnaya samples from the Caspian steppe (more precisely, from the Samara region in Russia). However, this might well be an artifact, due to very specific characteristics of these few ancient individuals. Larger sample sets would be welcome, especially from Yamnaya sites in Ukraine.

Below, dear audience, is a spreadsheet featuring the preliminary results. Click on the image to view and/or download the spreadsheet. The general rule is that the higher the tail prob, or p-value, the more likely it is that the ancestry proportions are close to the truth (a tail prob of well below 0.05 is usually a strong indication that something isn't right). For a detailed look at each of the qpAdm runs, feel free to consult the zip file here.


Note, however, that many of the European groups in my burgeoning genotype dataset are yet to make an appearance in the spreadsheet. That's because their models with the standard left pops showed p-values well under 0.01, which essentially meant that they failed, and I'm still trying to make them work.

But round one has certainly revealed some fascinating stuff. For instance, except for Hungarians and Estonians, none of the Uralic-speaking groups can be modeled successfully in the standard three-way model.

However, I managed to significantly improve the statistical fits in their models by adding a Siberian population, RUS_Baikal_BA, to the left pops. This is unlikely to be a coincidence, because the Proto-Uralic homeland was almost certainly located in or very near Siberia. Iain Mathieson please take note.

Saami
HUN_Koros_N_HG 0.134±0.043
RUS_Baikal_BA 0.270±0.015
TUR_Barcin_N 0.081±0.026
UKR_Yamnaya 0.515±0.058
chisq 19.865
tail prob 0.0108571

See also...


Saturday, June 27, 2020

Major updates to ADMIXTOOLS


An important message from Nick Patterson:

Dear Eurogenes bloggers,

Many of you use ADMIXTOOLS and you might like to know that there is a new release on github [LINK] with some important enhancements.

From the README

*** NEW ***

1)

Version 7.0 has numerous upgrades.

a) Two new executables --qpfstats qpfmv allow precomputation of f-statistic basis. This can greatly reduce computation costs.
b) qpAdm, qpWave, qpGraph support qpfstats output as input.
*** This is a much improved way of running with allsnps: YES. ***
c) A new experimental feature of qpGraph (halfscore: YES) allows comparison of 2 phylogenies + a (weak) goodness of fit score. Be careful if running with a large number of populations and consider reducing block size say blgsize: .005

2)

Note that several of the new ideas implemented in version 7.0 were developed collaboratively with Robert Maier, who has implemented them along with the great majority of other ADMIXTOOLS functionality in R: See https://github.com/uqrmaie1/admixtools
Executables run fast, and it has features not available in this C version, such as interactive exploration of graph phylogenies.
A manuscript describing the algorithmic ideas and providing documentation of the methods is in preparation.

qpfstats is the most important new executable. This estimates f-statistics and covariance on a basis.

a) This can be passed into other programs of the package without having to reaccess the genotype files, greatly speeding the computations.
b) In allsnps: YES mode a new computation is carried out (explained in qpfs.pdf) that is much more logical when there is a lot of missing data. Sometimes standard errors are greatly reduced.
qpfstats can be used with up to 30 populations. Much beyond that the output files become large.

As usual there may be bugs...

Nick Patterson 6/27/2020

Update 29/06/2020: As pointed out above, qpfstats is the most important new executable. Indeed, Nick Patterson now recommendeds that qpAdm analyses run with the allsnps: YES flag should be based on qpfstats output.

Several of my recent blog posts featured qpAdm models run with the allsnps: YES flag, but they were based on genotype data because obviously I didn't know anything about qpfstats at the time.

So I went back and ran some of these models again, just to make sure that they were still relevant. Below are three examples which you can compare to the original analyses here, here and here, respectively.

TUR_Arslantepe_LC_Maykop
RUS_Maykop_Novosvobodnaya 0.281±0.042
TUR_Arslantepe_LC 0.719±0.042
chisq 10.923
tail prob 0.449752
Full output

TUR_Barcin_C
RUS_Vonyuchka_En 0.137±0.031
TUR_Buyukkaya_EC 0.863±0.031
chisq 15.074
tail prob 0.0889099
Full output

UKR_N_admixed
RUS_Progress_En 0.083±0.020
UKR_N 0.917±0.020
chisq 6.825
tail prob 0.65538
Full output

As far as I can tell, they're very similar to the original runs, which is a relief, because it means that the conclusions in my blog posts still make sense.

Thursday, May 16, 2019

Fresh off the sledge


As things stand, the closest individual to a Proto-Uralic speaker in the ancient DNA record is arguably 0LS10 from an Iron Age tarand grave in what is now Estonia. I say that because:

- isotopic data suggest that 0LS10 wasn't born where he died, and considering his elevated Siberian ancestry relative to earlier and most contemporaneous Baltic ancients, he was very likely a migrant to the Baltic region from the east

- the tarand grave tradition appears to be specifically a Finnic (west Uralic) phenomenon that probably spread from the Volga-Oka region, which is just west of where most people place the Proto-Uralic homeland

- 0LS10 belongs to Y-chromosome haplogroup N-L1026, a paternal marker that is especially closely associated with Uralic-speaking populations and probably only appeared in the East Baltic region during the transition from the Bronze Age to the Iron Age

You can find more background info about 0LS10 and other relevant samples in Saag et al. 2019 (see here). This is where he sits in my Principal Component Analyses (PCA) focusing on fine scale Northern European genetic diversity. The relevant datasheets are available here and here, respectively.
Note that 0LS10 doesn't cluster strongly with any ancient or modern populations. To investigate this in more detail I ran a series of two-way qpAdm analyses, testing tens of ancient individuals and populations as potential admixture sources. These two models stood out above the rest in terms of their statistical fits, chronology and overall plausibility.

Baltic_EST_IA_0LS10
Baltic_EST_BA 0.826±0.045
RUS_Sintashta_MLBA_o1 0.174±0.045

chisq 12.527
tail prob 0.564048
Full output

Baltic_EST_IA_0LS10
Baltic_EST_BA 0.683±0.102
RUS_Mezhovskaya 0.317±0.102

chisq 13.811
tail prob 0.463864
Full output

Please note that RUS_Sintashta_MLBA_o1 isn't representative of the Sintashta culture population as a whole. It's a group of the most extreme genetic outliers among the Sintashta samples, and they may or may not have been Uralic speakers (see here). Interestingly, the Mezhovskaya culture population is generally associated with the Ugric branch of the Uralic language family.

I was also able to closely replicate these results with the Global25/nMonte method; down to almost one per cent. However, the statistical fits (distances) are poor, probably because the reference populations aren't the real mixture sources. This is in line with the fact that their Y-haplogroups are Q1a, R1a and R1b, rather than any type of N.

Baltic_EST_IA:0LS10
Baltic_EST_BA,83.8
RUS_Sintashta_MLBA_o1,16.2

distance%=4.7955

Baltic_EST_IA:0LS10
Baltic_EST_BA,69.8
RUS_Mezhovskaya,30.2

distance%=3.5783

I do realize that two Bronze Age samples from Bolshoy Oleni Ostrov, Kola Peninsula, belong to N-L1026, but adding them to my mixture models doesn't help. Little wonder, because the Kola Peninsula lies within the Arctic Circle, and I'm pretty sure that 0LS10 and his N-L1026 came from somewhere just north of the mixture cline marked on the map below. Unfortunately, I can't test this directly yet due to the scarcity of ancient samples from this region.



Monday, April 22, 2019

R1b-M269 in the Bronze Age Levant


The new Harvard genotype datasets that I blogged about recently include a couple of potentially very useful samples from the Levant dated to 1400-1100 BCE. Search for IDs I2062 and I1934 in the anno files here. They're both from an archeological paper about a Late Bronze Age (LBA) burial site in what is now Israel that was published back in 2017 (see here).

Surprisingly, individual I2062 is listed in the anno files as belonging to Y-haplogroup R1b1a1a2, which is also known as R1b-M269. The reason that this is a surprise to me is because R1b-M269 is closely associated with the Bronze Age expansions of pastoralists from the Pontic-Caspian steppe in Eastern Europe, and these expansions didn't impact the Levant in any direct or significant way.

The Y-haplogroup assignment may or may not be correct. Sometimes the Y-haplogroups in these sorts of datasheets are indeed wrong. Unfortunately, as far as I know, the BAM file for I2062 isn't available anywhere online, so I can't check whether he does really belong to R1b-M269. But, intriguingly, his autosomes do show a subtle signal of Yamnaya-related ancestry from the Pontic-Caspian steppe that is missing in earlier ancients from the Levant.

To characterize his genome-wide ancestry, I first ran a series of unsupervised and supervised analyses with the Global25/nMonte3 method (using this datasheet). For the sake of simplicity, I narrowed things down to the mixture models below based on three reference populations each. Levant_ISR_C is made up of Chalcolithic samples from Israel. The identities of the other reference sets should be obvious to most readers. If confused, feel free to ask for more details in the comments below.

Levant_ISR_MLBA:I2062
Levant_ISR_C,66.8
IRN_Seh_Gabi_C,27
Yamnaya_RUS_Samara,6.2

[1] distance%=1.8905

Levant_ISR_MLBA:I2062
Levant_ISR_C,66.2
Kura-Araxes_ARM_Kaps,30.2
Yamnaya_RUS_Samara,3.6

[1] distance%=2.0856

Levant_ISR_MLBA:I2062
Levant_ISR_C,67.8
Kura-Araxes_RUS_Velikent,31.8
Yamnaya_RUS_Samara,0.4

[1] distance%=2.1738

To further confirm the reliability of my models, I tested them with the formal statistics-based qpAdm software. As far as I can tell, the output from qpAdm looks very solid across the board.

Levant_ISR_MLBA_I2062
IRN_Seh_Gabi_C 0.193±0.052
Levant_ISR_C 0.710±0.038
Yamnaya_RUS_Samara 0.098±0.026

chisq 9.304
tail prob 0.67676
Full output

Levant_ISR_MLBA_I2062
Kura-Araxes_ARM_Kaps 0.249±0.076
Levant_ISR_C 0.681±0.051
Yamnaya_RUS_Samara 0.071±0.035

chisq 11.101
tail prob 0.52032
Full output

Levant_ISR_MLBA_I2062
Levant_ISR_C 0.661±0.042
Kura-Araxes_RUS_Velikent 0.339±0.042

chisq 7.979
tail prob 0.844942
Full output

Admittedly, even though I2062 can be modeled with Yamnaya-related admixture, he doesn't need to be. Indeed, his ratio of this type of ancestry varies significantly between the models, from around 10% to nothing. This appears to be dependent on the geography of the non-Levant and non-Yamnaya reference populations; the closer they are to the Pontic-Caspian steppe, the smaller the ratio of Yamnaya-related ancestry in I2062. I'd describe this as an artifact of the isolation-by-distance phenomenon, and it totally makese sense, but it prevents me from confirming beyond any doubt that I2062 does harbor genome-wide steppe ancestry. Unfortunately, individual I1934 doesn't offer enough data to be analyzed with the same methods.

Samples associated with the Kura-Araxes or Early Transcaucasian culture are particularly strong references for the eastern ancestry in I2062. This probably isn't a coincidence, and it might also explain his Y-haplogroup, because, at its maximum extent, the territory occupied by the Kura-Araxes culture stretched all the way from the Pontic-Caspian steppe to the southern Levant. The map below is from Wilkinson 2014.

See also...

Downloadable genotypes of present-day and ancient DNA data

Early chariot riders of Transcaucasia came from...

R-V1636: Eneolithic steppe > Kura-Araxes?

Friday, April 12, 2019

Armenians vs Georgians


Armenians and Georgians are ethnic groups that live side by side in the south Caucasus, or Transcaucasia. By all accounts, they've both been there since prehistoric times and they're very similar in terms of overall genetic structure.

However, they speak languages from totally unrelated families: Indo-European and Kartvelian, respectively. How did this happen and might the answer lie in the small genetic differences that do exist between them?

To investigate this issue, I ran a series of qpAdm formal mixture models of present-day Armenians and Georgians using tens of ancient reference populations. To come up with as straightforward and meaningful results as possible, I constrained myself to two-way models. I then discarded the runs that produced "tail probs" under 0.1 and retained less than 400K SNPs. Only a handful of models passed muster, including these two:

Armenian
Mycenaeans_&_Empuries2 0.233±0.041
Kura-Araxes_Kaps 0.767±0.041

chisq 18.422
tail prob 0.142151
Full output

Georgian
Globular_Amphora 0.071±0.025
Kura-Araxes_Kaps 0.929±0.025

chisq 18.419
tail prob 0.142266
Full output

At the most basic level, the results suggest that both Armenians and Georgians are overwhelmingly derived from populations of Bronze Age Transcaucasia associated with the Kura-Araxes archeological culture, albeit with minor ancestries from somewhat different sources from the west. As far as I can see, when using more than 400K SNPs and a wide range and large number of outgroups (or right pops), neither Armenians nor Georgians can pass perfectly for any one ancient population in my dataset.

The best proxies for the minor but significant western ancestry in Armenians are Mycenaeans of the Bronze Age Aegean region and Greek colonists from Iron Age Iberia (Empuries2). Obviously, and perhaps importantly, these are both attested Indo-European-speaking groups. On the other hand, the very minor western ancestry in Georgians is best characterized as gene flow from Middle to Late Neolithic European farmers rich in indigenous European forager ancestry. It's practically impossible to say what language or languages these farmers spoke. How about something Kartvelian?

In any case, for me, the perplexing thing about present-day Armenians is that they harbor very little steppe ancestry. By and large, no more than a few per cent. Compare that to the currently available samples from what is now Armenia dating to the Middle to Late Bronze Age, which show ratios of steppe ancestry of up to 25%. For now, I'm guessing that what we're dealing with here is the classic bounce back of older ancestry layers that has been documented for different parts and periods of prehistoric Europe.

See also...

Early chariot drivers of Transcaucasia came from...

Catacomb > Armenia_MLBA

Late PIE ground zero now obvious; location of PIE homeland still uncertain, but...

Monday, March 25, 2019

Celtic probably not from the west


The term "Celtic from the west" is the catchphrase for a working theory, offered in a couple of recent books, positing that the earliest speakers of Celtic languages lived in Atlantic Europe during the Bronze Age or even earlier. It'll be interesting to see how this theory holds up against increasing numbers of ancient samples from attested early Celtic-speaking populations.

More popular and long-standing theories postulate that the Proto-Celts are associated with the Urnfield and/or Hallstatt archeological cultures of Late Bronze Age and Iron Age Central Europe. I'm inclined to agree with these more mainstream views when looking at my qpAdm mixture models below of three Celtiberians from what is now La Hoya, northern Spain, from the recent Olalde et al. paper on the genomic history of Iberia.

Celtiberian_LaHoya
Halberstadt_LBA 0.207±0.077
Pre-Celtiberian_LaHoya 0.793±0.077

chisq 15.031
tail prob 0.522396
Full output

Celtiberian_LaHoya
Halberstadt_LBA 0.196±0.074
Non-Celtic_Iberian 0.804±0.074

chisq 17.366
tail prob 0.362297
Full output

The Celtiberians show a stronger signal of (Urnfield-related?) ancestry from the northeast than their Bronze Age predecessors in northern Iberia (Pre-Celtic_LaHoya) as well as their Iron Age contemporaries from eastern Iberia (Non-Celtic_Iberian). The latter group very likely spoke the non-Indo-European Iberian language. It's not clear what the Bronze Age northern Iberians spoke, but it may have been a language related to Basque, which is also non-Indo-European.

Of course, the fact that the Celtiberians harbored more northern Bell Beaker-related ancestry than basically all earlier Iberian groups was already reported in the Olalde et al. paper (on page 2), but I just wanted to see if I could flesh out some more details in regards to this observation by using chronologically and archeologically more proximate reference populations.

See also...

Open thread: What are the linguistic implications of Olalde et al. 2019?

An exceptional burial indeed, but not that of an Indo-European

Late PIE ground zero now obvious; location of PIE homeland still uncertain, but...

Monday, March 4, 2019

An exceptional burial indeed, but not that of an Indo-European


Not too many people have been buried sitting on wagons. The most famous case is that of an Early Bronze Age man who, considering his injuries, may have died in a high-speed crash - high-speed for its time anyway - on the Pontic-Caspian steppe in Eastern Europe.

It's likely that this guy was one of the very first wagon-drivers in human history, because his four-wheeled wooden model is dated to 3336-3105 calBCE, which makes it the oldest wagon discovered thus far. His genotype data, under the label Steppe Maykop SA6004, were published recently along with Wang et al. 2019.

Early wagons are very important for a couple of reasons: they revolutionized human transport and warfare, and they're often closely associated with the prehistoric expansions of Indo-European languages.

So I'm pretty sure that many of you must be thinking right now that wagon-driver SA6004 was an early Indo-European, or even a Proto-Indo-European! I bet that's what Wang et al. thought too, considering the conclusion in their paper. But, alas, the chances of this are slim to none.

Steppe Maykop samples show rather peculiar genetic structure considering their geographic origin, with a large proportion of their ancestry deriving from a source closely related to western Siberian hunter-gatherers (aka West_Siberia_N in the ancient DNA record). Indeed, SA6004 basically looks like a 50/50 mix between West_Siberia_N and Piedmont_Eneolithic. Here's a map with all of the relevant details.


Thus, clearly, the Steppe Maykop population wasn't ancestral or even directly related to the steppe and steppe-derived groups generally regarded to have been Indo-European speaking, such as those associated with the Yamnaya, Corded Ware, and Bell Beaker cultures. That's because these groups lack any discernible West_Siberia_N-related ancestry.

It also wasn't ancestral or directly related to any present-day or currently sampled ancient Indo-European speaking populations, again because these populations basically lack West_Siberia_N-related ancestry.

On the other hand, Yamnaya, Corded Ware and other closely related groups show an exceptionally strong genetic relationship with Indo-European speakers, especially those from across Northern Europe, which experienced massive migrations from the Pontic-Caspian steppe during the late Neolithic period, and hardly anything from elsewhere since then.

Case in point, the samples from Wang et al. labeled Yamnaya Caucasus were recovered from the same area of the Pontic-Caspian as their Steppe Maykop samples, and yet, take a look at this linear model based on outgroup f3-statistics. Steppe Maykop does show high genetic affinity to Indo-European speakers (no doubt mediated via its Piedmont_Eneolithic-related ancestry), but, unlike Yamnaya Caucasus, it also shows unusually high affinity for a West Eurasian population to Native Americans and Siberians. The relevant datasheet is available here.
So the only way that the Steppe Maykop population was Indo-European-speaking, was if it inherited its Indo-European speech from its Piedmont_Eneolithic-related ancestors. And even if it was Indo-European-speaking, it probably spoke an extinct Indo-European language not closely related to any extant Indo-European languages. In other words, the possibility that Steppe Maykop passed on its language to Yamnaya, along with its wagons, is close to zero. More likely, Yamnaya stole a few wagons from Steppe Maykop, and the rest is history.

See also...

The Steppe Maykop enigma

On Maykop ancestry in Yamnaya

Late PIE ground zero now obvious; location of PIE homeland still uncertain, but...

Wednesday, February 27, 2019

The Steppe Maykop enigma


Who were the Steppe Maykop people exactly? Their ancestry must surely rank as one of the biggest surprises served up by ancient DNA to date.

I always thought that they'd turn out roughly like a mixture between populations associated with the Kura-Araxes and Yamnaya cultures (mostly because their territory was located sort of in between them). Nope, that wasn't even close. This is where they cluster compared to Kura-Araxes and Yamnaya samples in my Principal Component Analysis (PCA) of world-wide genetic variation: the Global25.
To explore the ancestry of the Steppe Maykop people in more detail I ran a couple of unsupervised Global25/nMonte tests, using basically every ancient population in the (scaled) Global25 datasheet that seemed chronologically sensible and even remotely relevant. I narrowed things down to these two mixture models.

Steppe_Maykop
Geoksiur_Eneolithic,11.2
Piedmont_Eneolithic,44.4
West_Siberia_N,44.4
distance%=1.5161

Steppe_Maykop
Piedmont_Eneolithic,46.6
Sarazm_Eneolithic,10.4
West_Siberia_N,43
distance%=1.6408

But, you might say, Global25/nMonte isn't a published analytical method and it doesn't run on formal statistics, the meat and potatoes of ancient DNA papers. OK then, let's try the same models with the qpAdm software, which is a published method and does run on formal statistics, using exactly the same samples.

Steppe_Maykop
Geoksiur_Eneolithic 0.100±0.032
Piedmont_Eneolithic 0.433±0.053
West_Siberia_N 0.467±0.028
chisq 19.155
tail prob 0.159096
Full output

Steppe_Maykop
Piedmont_Eneolithic 0.429±0.051
Sarazm_Eneolithic 0.119±0.033
West_Siberia_N 0.452±0.026
chisq 18.090
tail prob 0.202699
Full output

They're basically identical. Importantly, my models must reflect reality at some level, because otherwise I wouldn't be able to produce a pair of essentially identical results using such vastly different statistical methods. So the pertinent question is what do these results actually mean?

It seems unlikely to me that we're dealing here with a highly complex three-way mixture process, involving populations from such far flung locations as western Siberia and southern Central Asia. Rather, I suspect that Steppe Maykop was the result of a two-way mixture between Piedmont_Eneolithic (the population that lived before it on the steppe north of the Caucasus) and someone just a little bit more easterly. I'm guessing that the latter was the (as yet unsampled) population associated with the Kelteminar archeological culture.


By the way, please note that Piedmont_Eneolithic is made up of samples from two different locations on the Piedmont steppe, and I occasionally treat them as separate populations labeled Progress_Eneolithic and Vonyuchka_Eneolithic (for instance, see here).

Update 28/02/2019: Below is a PCA focusing on West Eurasian genetic variation. Overall, the position of Steppe Maykop relative to Geoksiur_Eneolithic, Piedmont_Eneolithic and West_Siberia_N appears to reflect my nMonte and qpAdm models. However, as per our discussion in the comments, one of the Steppe Maykop individuals (the most southerly one in the PCA) probably also has recent ancestry from the Caucasus.

See also...

An exceptional burial indeed, but not that of an Indo-European

Maykop: a multi-ethnic layer cake?

Late PIE ground zero now obvious; location of PIE homeland still uncertain, but...