search this blog

Showing posts with label nMonte. Show all posts
Showing posts with label nMonte. Show all posts

Monday, April 22, 2019

R1b-M269 in the Bronze Age Levant


The new Harvard genotype datasets that I blogged about recently include a couple of potentially very useful samples from the Levant dated to 1400-1100 BCE. Search for IDs I2062 and I1934 in the anno files here. They're both from an archeological paper about a Late Bronze Age (LBA) burial site in what is now Israel that was published back in 2017 (see here).

Surprisingly, individual I2062 is listed in the anno files as belonging to Y-haplogroup R1b1a1a2, which is also known as R1b-M269. The reason that this is a surprise to me is because R1b-M269 is closely associated with the Bronze Age expansions of pastoralists from the Pontic-Caspian steppe in Eastern Europe, and these expansions didn't impact the Levant in any direct or significant way.

The Y-haplogroup assignment may or may not be correct. Sometimes the Y-haplogroups in these sorts of datasheets are indeed wrong. Unfortunately, as far as I know, the BAM file for I2062 isn't available anywhere online, so I can't check whether he does really belong to R1b-M269. But, intriguingly, his autosomes do show a subtle signal of Yamnaya-related ancestry from the Pontic-Caspian steppe that is missing in earlier ancients from the Levant.

To characterize his genome-wide ancestry, I first ran a series of unsupervised and supervised analyses with the Global25/nMonte3 method (using this datasheet). For the sake of simplicity, I narrowed things down to the mixture models below based on three reference populations each. Levant_ISR_C is made up of Chalcolithic samples from Israel. The identities of the other reference sets should be obvious to most readers. If confused, feel free to ask for more details in the comments below.

Levant_ISR_MLBA:I2062
Levant_ISR_C,66.8
IRN_Seh_Gabi_C,27
Yamnaya_RUS_Samara,6.2

[1] distance%=1.8905

Levant_ISR_MLBA:I2062
Levant_ISR_C,66.2
Kura-Araxes_ARM_Kaps,30.2
Yamnaya_RUS_Samara,3.6

[1] distance%=2.0856

Levant_ISR_MLBA:I2062
Levant_ISR_C,67.8
Kura-Araxes_RUS_Velikent,31.8
Yamnaya_RUS_Samara,0.4

[1] distance%=2.1738

To further confirm the reliability of my models, I tested them with the formal statistics-based qpAdm software. As far as I can tell, the output from qpAdm looks very solid across the board.

Levant_ISR_MLBA_I2062
IRN_Seh_Gabi_C 0.193±0.052
Levant_ISR_C 0.710±0.038
Yamnaya_RUS_Samara 0.098±0.026

chisq 9.304
tail prob 0.67676
Full output

Levant_ISR_MLBA_I2062
Kura-Araxes_ARM_Kaps 0.249±0.076
Levant_ISR_C 0.681±0.051
Yamnaya_RUS_Samara 0.071±0.035

chisq 11.101
tail prob 0.52032
Full output

Levant_ISR_MLBA_I2062
Levant_ISR_C 0.661±0.042
Kura-Araxes_RUS_Velikent 0.339±0.042

chisq 7.979
tail prob 0.844942
Full output

Admittedly, even though I2062 can be modeled with Yamnaya-related admixture, he doesn't need to be. Indeed, his ratio of this type of ancestry varies significantly between the models, from around 10% to nothing. This appears to be dependent on the geography of the non-Levant and non-Yamnaya reference populations; the closer they are to the Pontic-Caspian steppe, the smaller the ratio of Yamnaya-related ancestry in I2062. I'd describe this as an artifact of the isolation-by-distance phenomenon, and it totally makese sense, but it prevents me from confirming beyond any doubt that I2062 does harbor genome-wide steppe ancestry. Unfortunately, individual I1934 doesn't offer enough data to be analyzed with the same methods.

Samples associated with the Kura-Araxes or Early Transcaucasian culture are particularly strong references for the eastern ancestry in I2062. This probably isn't a coincidence, and it might also explain his Y-haplogroup, because, at its maximum extent, the territory occupied by the Kura-Araxes culture stretched all the way from the Pontic-Caspian steppe to the southern Levant. The map below is from Wilkinson 2014.

See also...

Downloadable genotypes of present-day and ancient DNA data

Early chariot riders of Transcaucasia came from...

R-V1636: Eneolithic steppe > Kura-Araxes?

Thursday, February 15, 2018

Modeling genetic ancestry with Davidski: step by step


There are many different ways to model your genetic ancestry but I prefer the Global25/nMonte method. This is a step by step guide to modeling ancient ancestry proportions with this simple but powerful method using my own genome.


As far as I know, the vast majority of my recent ancestors came from the northern half of Europe. This may or may not be correct, but it gives me somewhere to start, so that I can come up with a coherent model. If you don't have this sort of information, because, perhaps, you were adopted, then just look in the mirror, and work from there. Like I say, it's not imperative that you know anything whatsoever about your ancestry, because your genetic data will do the talking, but you do need a model when modeling.

In scientific literature nowadays Northern Europeans are often described as a three-way mixture between Yamnaya-related pastoralists, Anatolian-derived early farmers, and Western European Hunter-Gatherers (WHG). So let's see if this model works for me. Obviously, if it does, then it'll confirm the information that I have about my origins, but it might also reveal finer details that I'm not aware of. The datasheet that I'm using for this model is available here.

[1] distance%=6.9025 / distance=0.069025

Davidski

Yamnaya_Samara 53.9
Barcin_N 30.75
Rochedane 15.35
Tepecik_Ciftlik_N 0

Yep, the model does work, with a fairly reasonable distance of almost 7%. The ancestry proportions more or less match those from scientific literature and the plethora of analyses that I've featured at this blog on the topic. Please note that I've kept things very simple, using only four reference populations and individuals as proxies for four distinct streams of ancestry. But I've put my own twist on this Neolithic/Bronze Age model by including two populations from Neolithic Anatolia (Barcin_N and Tepecik_Ciftlik_N), just to see what would happen. The WHG proxy is Rochedane.

Admittedly, though, my Yamnaya cut of ancestry appears somewhat bloated at over 53%, and the model's distance is a little higher than what I normally see for really strong models. So let's check if I can get a better fitting and more sensible result by adding a slightly more easterly forager proxy than Rochedane: Narva_Lithuania.

[1] distance%=5.9331 / distance=0.059331

Davidski

Yamnaya_Samara 45.75
Barcin_N 31.45
Narva_Lithuania 22.8
Rochedane 0
Tepecik_Ciftlik_N 0

The statistical fit does improve, and when given a choice between Rochedane and Narva_Lithuania, the algorithm picks the latter as the only source of extra forager input in my genome.

What could this mean? It might mean that a large part of my ancestry derives from the Baltic region. Actually, I know for a fact that this is true. But even if I had no idea about my genealogy, this result would be a very strong hint about my genetic origins. Indeed, let's follow this trail and try to further improve the fit of the model by adding a more relevant Yamnaya-related proxy, such as early Baltic Corded Ware (CWC_Baltic_early).

[1] distance%=5.444 / distance=0.05444

Davidski

CWC_Baltic_early 54.95
Barcin_N 26.7
Narva_Lithuania 18.35
Rochedane 0
Tepecik_Ciftlik_N 0
Yamnaya_Samara 0

Holy shit! To be honest, I wasn't expecting this sort of resolution and accuracy, and I can't promise that everyone using the Global25/nMonte method will see such incredibly nuanced outcomes, but this isn't a fluke. It can't be, because it gels so well with everything that I know about my ancestry. Please note also that I belong to Y-chromosome haplogroup R1a-M417, which is a lineage intimately associated with the Corded Ware expansion across Northern Europe (for instance, see here).

But of course, the Baltic and nearby regions haven't been isolated from migrations and invasions since the Corded Ware times. For instance, at some point, probably during the Bronze Age, Uralic-speaking groups moved west across the forest zone of Northeastern Europe and into the East Baltic and northern Scandinavia. It's generally accepted that they brought Siberian admixture with them (see here). Moreover, from the Iron Age to the Middle Ages, East Central Europe was under intense pressure from a wide range of nomadic steppe groups with complex ancestry, such as the Sarmatians, Avars, Huns, and Mongolians. Did any of these peoples leave their mark on my genome? At the risk of overfitting the model, let's explore this possibility by adding a few more reference populations.

[1] distance%=5.444 / distance=0.05444

Davidski

CWC_Baltic_early 54.95
Barcin_N 26.7
Narva_Lithuania 18.35
Han 0
Mongolian 0
Nganassan 0
Rochedane 0
Sarmatian_Pokrovka 0
Tepecik_Ciftlik_N 0
Yamnaya_Samara 0

Nothing changes when I add the Han Chinese, Mongolians, Nganassans (a Uralic group from Siberia), and Sarmatians to the model. But what about if I throw in the only ancient Slav in my datasheet?

[1] distance%=2.9904 / distance=0.029904

Davidski

Slav_Bohemia 85.9
CWC_Baltic_early 7.7
Narva_Lithuania 6.4
Barcin_N 0
Rochedane 0
Tepecik_Ciftlik_N 0
Yamnaya_Samara 0

Considering that the vast majority of my recent ancestors were Poles, thus a Slavic-speaking people from near the Baltic, this outcome makes perfect sense. And check out the new distance! But the problem now is that I'm overfitting the model by using two very similar and probably very closely related references, CWC_Baltic_early and Slav_Bohemia. And overfitting should be avoided at all costs. So it might be useful to break up this effort into two models: one focusing on the Neolithic and Bronze Age, and the other on the Iron Age and Middle Ages. I'll do that soon, but not just yet, because there are still too few Iron Age and Medieval samples available from the Baltic region and surrounds for meaningful analyses of this type.

See also...

Genetic ancestry online store (to be updated regularly)