References
Suggest a paperAncient genomes from Ladakh reveal 2800-year-old admixture between Tibetans and South Asians
Reconstructing population history is harder in South Asia than in many other world regions due to a paucity of ancient DNA. We report genome-wide data for 10 individuals from Old Lady Spider Cave, which lies 4000 meters above sea level in the Himalayan region of Ladakh and dates to around 1500 years before the present (B.P.). These individuals were genetically homogeneous and had an ancestry signature rare in South Asians today: admixed in roughly 50-50% proportions between a population well proxied by present-day North Indians and another genetically similar to ancient Tibetans. By analyzing the typical sizes of segments of DNA inherited from each of these ancestral populations, we find that admixture of these groups began at least 50 generations before the date of the individuals, that is, by around 2800 B.P.
From prehistory to present-day: How isolation shaped the distinct genomic makeup of Italian Alpine valleys
Open access
ABSTRACT Historically, the eastern Italian Alps have provided crucial geographic corridors for cultural and genetic exchange between the Mediterranean region and Central and Northern Europe. Although recent archaeogenomic data suggest a strong regional persistence of Anatolian Neolithic ancestry followed by the arrival of Yamnaya-related components during the Bronze Age, the genomic landscape of modern Alpine populations remains largely unmapped. In this study, we verified the persistence of these signals in the present-day Rendena and Ledro Valleys by integrating novel complete mitochondrial genomes (N=185) and genome-wide SNP data (N=96), the latter combined with a novel Italian genomic dataset (N=139). Our findings show that the Rendena and Ledro populations form a distinct “modern Alpine” genomic group, which retains a significant proportion of early European Neolithic ancestry, aligning with ancient eastern Italian Alpine individuals and modern Sardinians. More recently, geographic isolation and localized genetic drift appear to have shaped the two valleys differently. Demographic reconstructions reveal asynchronous population declines over the past two millennia, followed by a sharp, synchronized bottleneck 200-300 years ago, which coincided with historical plague outbreaks. Remarkably, this structured drift persisted at an extremely fine microgeographic scale within the valleys, resulting in internal genetic subclusters that directly correlate with local topography. This micro-differentiation was likely maintained by steep geographic barriers and/or local endogamous practices. This study ultimately underscores how geographic barriers and isolation can preserve ancient genomic components and shape highly localized genetic structures over centuries even to the modern day.
Analysis of Y-Chromosomal STR Variation Reveals High Paternal Genetic Diversity in Albanian Populations
Open access
The Albanian population occupies an important position within the genetic landscape of the Balkan Peninsula. The present study investigates the Y-chromosomal genetic diversity and forensic characteristics of Albanian male populations using Y-chromosomal Short Tandem Repeat (Y-STR) markers. A total of 170 unrelated male individuals from different geographic regions of Albania were analysed using a multiplex Y-STR panel to assess allele variation, haplotype diversity, forensic efficiency, and internal population structure. The analysis identified 151 different haplotypes, including 136 unique haplotypes, indicating substantial paternal lineage heterogeneity. The overall haplotype diversity was 0.9979 ± 0.0018, demonstrating high Y-chromosomal variability within the studied population. The discrimination capacity was 0.8779, while the match probability was 0.0079. Combined forensic parameters showed near-complete discriminatory performance, with combined power of discrimination and combined power of exclusion approaching 1.000. The analysed loci DYS385, DYS458, DYS456, and DYS635 exhibited the highest polymorphism and forensic informativeness, whereas DYS392 and DYS391 showed lower allelic diversity. Concerning the haplogroup distribution the population was characterized by the predominance of E1b1b-V13, followed by J2b2-M241, I2a, R1b, and R1a, representing the principal paternal lineages of the Albanian population. These haplogroups are widely recognized as characteristic components of the western Balkan genetic landscape and provide evidence of long-term paternal lineage continuity in the region. Comparative analysis of haplogroup frequencies and multivariate clustering suggested regional variation between Gegë and Toskë population groups, indicating the presence of internal paternal genetic structuring within Albania. These findings provide new reference data for Albanian Y-chromosomal diversity and highlight the high forensic utility of Y-STR markers. The study contributes to the understanding of paternal genetic structure in Albania and offers valuable data for forensic genetics, population genetics, and comparative Balkan genetic research. Received: 25 March 2026 │ Accepted: 12 June 2026 │ Published: 23 July 2026
High-coverage ancient genomes reveal divergent population histories and prehistoric starch-related genetic variation in Japan
Open access
Prehistoric population movements played a central role in shaping the genetic diversity across Eurasia. As the eastern terminus of the Eurasian continent, the Japanese archipelago provides a pivotal context for reconstructing the genetic interplay between prehistoric indigenous lineages and continental migrations. However, high-resolution genomic evidence from mainland Japan is limited due to poor DNA preservation. In the present study, we report two low-contamination, high-coverage ancient genomes from mainland Japan, namely an Initial Jomon individual (>67-fold) and a Middle Yayoi individual (>46-fold). The genomes enabled diploid genotyping and individual-level demographic inferences at unparalleled resolutions. Demographic reconstructions revealed that while both lineages underwent population contractions during the Last Glacial Maximum (LGM), their subsequent trajectories diverged; the Jomon lineage was characterized by long-term postglacial stability with no discernible expansion, whereas the Yayoi-related ancestral population exhibited gradual and sustained growth. Our analysis further identified a distinct genetic continuity between the Middle Yayoi and present-day mainland Japanese, refining the evolutionary trajectory of modern genetic profiles. Furthermore, the genomic profile of the Yayoi individual reflected a complex ancestral background shaped by continental interactions before their arrival on the archipelago. Notably, the higher salivary amylase ( AMY1 ) copy number observed in the Initial Jomon individual suggests that AMY1 copy number variation, potentially relevant to starch-rich diets, was already present before the large-scale expansion of rice farming in the Japanese Archipelago. Collectively, the findings demonstrate that high-coverage ancient genomic approaches can provide a powerful framework for deciphering the contrasting population histories and evolutionary trajectories of human lifestyles.
Paleogenomic evidence for genetic heterogeneity and prior admixture in Gothic-associated communities of late antique Bulgaria
Open access
We report genome-wide ancient DNA from 37 individuals retained after conservative archaeological and genomic reassessment of 53 screened samples from two Gothic-period mortuary contexts in present-day Bulgaria: the Aquae Calidae necropolis in Thrace (n = 20; c. 320 CE–375 CE) and the Aul of Khan Omurtag (AKO) in Moesia Secunda (n = 17; c. 350 CE–489 CE). The filtering process excluded individuals with uncertain or probably later-medieval horizon assignment, including a C5-area component linked by READv2 kinship to suspected later-horizon burials. We evaluate three non-exclusive models for Balkan Gothic formation, namely, migration-continuity from northern and Pontic Gothic horizons, a Wenskus–Wolfram–Pohl tradition-bearing cultural–political formation model across changing demographic substrates, and the Roman frontier formation. The two assemblages share Gothic-associated material culture and Christian east–west burial orientation, but they are not a single genome-wide population. Proximal qpAdm models distinguish a southern Anatolian/Marmara-related and northern/Pontic structure at Aquae Calidae from a simpler Chernyakhov-related and Balkan Late Antique structure at AKO. f 4 statistics further show that the Anatolian-related component at Aquae Calidae cannot be explained solely as a Chernyakhov-carried signal. DATES places north–south admixture at approximately 12.44 ± 2.36 generations before burial (Z = 5.26), consistent with a pre-burial ancestry mixture rather than a single intact Gothic migration. Kinship is confined within sites. Together, the data reject a single biological Gothic population moving unchanged into the Balkans and instead support a formation process combining migration, multiethnic coalition, and frontier incorporation.
Genomic insights into the Iron Age Saka of Boz-Barmak, Kyrgyzstan
Open access
The nomadic cultures of the Iron Age played an important role in shaping the genetic and cultural landscape of Eurasian populations. Yet despite its key geographical location, the Central Eurasian region remains underrepresented in ancient DNA studies of humans. We address this gap through genomic analysis of 12 individuals from the Boz-Barmak burial site in Kyrgyzstan associated with Saka pastoralists (4th-2nd centuries BCE), 9 of which yielded low-coverage genomes (on average 0.7-fold coverage). Genetic clustering analysis placed these individuals within the genetic variation of ancient and modern Central Eurasian and Siberian populations. We found no evidence of first-degree relatives in a kinship analysis, however a network of second- and third-degree relationships seems to be present. Notably, all male individuals share the same Y-chromosomal haplotype, common in present-day Kyrgyz groups, while mitochondrial DNA showed comparably high diversity, with distinct haplogroups observed across the analysed individuals. These findings are in line with archaeological and ethnographic evidence of patrilocality in Early Iron Age Saka, where male lineages remained stable across generations, while female mobility contributed to genetic diversity. Our study complements our understanding of the interplay between kinship, social organisation and population history in nomadic cultures.
Haplogroup Dynamics in Afghanistan: A Genetic History from the Neolithic to the Steppe Migrations
Open access
This paper presents a comprehensive analysis of Y-chromosomal haplogroup distributions among the five major ethnic groups of Afghanistan—Hazara (Āzrah), Pashtun, Tajik, Uzbek, and Turkmen—based on peer-reviewed genetic studies published in PLOS ONE (Haber et al. 2012; Di Cristofaro et al. 2013). The data reveals that all Afghan populations share a common Neolithic ancestral substrate that emerged during the agricultural revolution (10,000–7,000 years ago). Inter-ethnic differentiation began during the Bronze Age (circa 4,700 years ago), driven by the formation of regional civilizations including the Bactria-Margiana Archaeological Complex (BMAC) and the Indus Valley Civilization. Key findings:- The Hazara (Āzrah) exhibit the highest frequency of Neolithic Iranian haplogroups (J2a1-Page55 at 13%) and the highest frequency of East Asian-related haplogroup C3-M217 (33.3%), making them the most genetically layered of all Afghan groups.- The Pashtun and Tajik populations are characterized by high frequencies of R1a1a-M17 (51% and 30.4%, respectively), traditionally associated with steppe migrations.- The Uzbek population shows a high frequency of C3-M217 (41.18%), consistent with their documented Turco-Mongol heritage.- The Turkmen, as a Turkic-speaking group, share similarities with the Uzbek but with their own distinctive admixture patterns. The paper argues that the genetic architecture of Afghanistan is best understood as a palimpsest of overlapping layers, where no single ethnic group is "pure" or "original," but each carries a unique mosaic of the region's deep history. Keywords: Population Genetics, Y-Chromosomal Haplogroups, Afghanistan, Hazara, Āzrah, Pashtun, Tajik, Uzbek, Turkmen, J2a1-Page55, R1a1a-M17, C3-M217, Neolithic Iranian Substrate, Steppe Migrations, Bronze Age, Central Asian Genetics, Helmand Civilization, PLOS ONE Studies
Long-read sequencing reveals novel mitochondrial genome variants undetected by short-read sequencing in Korean population
Open access
While long-read sequencing technologies (e.g., PacBio Revio, ONT) have revolutionized high-quality genome assembly for the human pangenome, mitochondrial genome (mtDNA) analysis still largely relies on short-read and Sanger sequencing. However, short-read sequencing often lacks the resolution required to resolve complex variations due to the unique features of mtDNA, such as high mutation rates and repetitive homopolymeric regions, which frequently lead to alignment artifacts and mapping ambiguities. To address this, we evaluated whether applying long-read sequencing to mtDNA improves analytical quality in empirical data. Through comprehensive bioinformatics analyses, we compared the performance of long-read sequencing against short-read sequencing and microarrays. Our results revealed that long-read sequencing detected the highest number of variants (n = 533), significantly outperforming both short-read sequencing (n = 525) and microarrays (n = 49). Notably, both sequencing methods provided significantly higher resolution in haplogroup assignment compared to microarrays in terms of phylogenetic depth (p < 0.05). Long-read sequencing demonstrated superior detection power, particularly for InDels. We identified two novel non-synonymous variants, including a unique InDel detected exclusively by long-read sequencing. Protein modeling and stability analysis validated that this InDel causes structural instability (RMSD > 2.0 Å, 𝛥𝛥G = -45.21 kcal/mol). Furthermore, we confirmed that this novel InDel is shared among haplogroup A samples in both the 1000 Genomes Project ONT dataset and the Korean population, highlighting the practical implications of long-read sequencing for molecular biology and population genetics.
Ancestry, admixture, and pathogens in contemporaneous Neolithic farmers and foragers on the Island of Gotland
Open access
Two archaeological cultural complexes; the Neolithic Funnelbeaker culture (FBC) and the Pitted ware culture (PWC), coexisted on Gotland for over 500 years, between ~3300 and 2800 calBCE. The ancestry of the FBC farmers and PWC marine foragers largely aligns with European Neolithic Farmers and European Mesolithic foragers, respectively, but the direct interactions between the groups on Gotland is not understood. We present a Middle Neolithic (MN) high-coverage genome and a Late Neolithic (LN) low-coverage genome from the Ansarve FBC dolmen. We investigate ancestry, admixture, and pathogens among these MN farmers (n = 6), foragers (n = 19), and the LN individual. We find that recent gene-flow between farmers and foragers could have taken place, although most gene-flow happened prior to their coexistence on the island. We also find evidence of different Yersinia pestis strains in the three cultural groups, showing that the pestis was widespread among groups with different subsistence strategies.
Whole‐Genome Sequencing Pilot of the Central Asian Genomic Diversity Project Reveals Distinct Histories, Adaptation, and Introgression
Open access
Central Asians are underrepresented in genomic research, limiting insights into their genetic history and disease risk. We established the Central Asian Genomic Diversity Project and sequenced whole genomes from 166 individuals across 20 Central Asian and Afghan Hazara groups. We identify marked differentiation driven by varying West/East Eurasian ancestry; Tajiks align with West Eurasians, Dungans with East Asians, and we report four geographically structured Turkic-related clusters, two Indo-European clines, and long-range migration events, including Siberian links in Hazaras and Sino-Tibetan ties in Dungans. Admixture dates cluster ∼650-1000 years ago, coinciding with the Song-Yuan era and Mongol expansion. We characterize distinct distributions of medically relevant variants and population-specific adaptation signatures across metabolic, immune, and neurological pathways, and illuminate shifts in subsistence practices correlated with trait-associated variation. We also detect Neanderthal-like and Denisovan-like segments that show group-specific associations with immunity, psychiatric risk, drug metabolism, and diabetes, underscoring the scientific imperative for a broader characterization of Central Asian evolutionary history and informing precision medicine.
Ancient DNA reveals elite dynastic rule among Iron Age Eurasian Steppe nomads
Open access
The Eurasian Steppe in the first millennium BCE saw the rise of the Scytho-Siberian archaeological horizon, which would come to stretch from the Altai Mountains in the east to the Black Sea in the west. We examined the genetic profiles of Iron Age Scythians to explore how social status shaped biological relatedness and ancestry patterns. We present genome-wide data from 85 individuals (38 elite and 47 non-elite), including 45 newly sequenced individuals and the first genome-wide data for the Scythian "Golden Man." We identify consanguineous unions, a reduced effective population size, and identity-by-descent links among the elites. Dynastic rule is supported by elite grandparent-grandchild relationships across cemeteries. While ancestries are heterogeneous, elite Iron Age Scythians show lower variation and no detectable patrilocal or matrilocal signal. These findings highlight hereditary status transmission and the emergence of social stratification in ancient nomadic societies.
Genetic diversity of late Neanderthals in northwestern Europe
Open access
Abstract Archaeological, osteological and genetic evidence suggests that Neanderthals lived in small groups 1,2 ; however, less is known about whether these groups were part of isolated communities or belonged to larger, well-connected populations 3 . The dense concentration of broadly contemporaneous Neanderthal sites in the Meuse Basin, Belgium 4 , provides a rare opportunity to study regional populations at high resolution. Here we generated genetic data from 27 Neanderthals who lived less than approximately 52,500 years ago from ten archaeological sites in Belgium and France, including a high-coverage genome from a 45,000-year-old individual from Goyet, Belgium. We show that most of these individuals are more closely related to one another than to other contemporaneous late Neanderthals in Europe. Further, some of these individuals carry DNA from a Neanderthal lineage predating the split of late Neanderthals. Although these Neanderthals overlapped temporally with early modern humans in northwestern Europe from around 47,000 years ago, we find no evidence of recent gene flow from modern humans. They also do not show the genetic signatures of mating among close relatives found in Altai Neanderthals, suggesting that they lived in larger or better-connected groups. Moreover, genetic load did not accumulate over time, arguing against progressive genetic deterioration as a driver of Neanderthal extinction.
Lethal plague outbreaks in Lake Baikal hunter-gatherers 5,500 years ago
Open access
Abstract Plague is among the most devastating diseases in human history 1 . However, early strains of the plague-causing bacterium Yersinia pestis lacked virulence factors that are required for the bubonic form until around 3,800 years ago 2,3 . Consequently, the morbidity and mortality of early plague strains remain unclear. Here we describe early plague strains that are associated with two phases of outbreaks among mid-Holocene hunter-gatherers near Lake Baikal in southeast Siberia, beginning from about 5,500 years ago. These outbreaks occur across four hunter-gatherer cemeteries, with a 39% detection rate for plague infection. By reconstructing kinship pedigrees, we show that small familial groups were affected, consistent with human-to-human spread of disease, and that the first outbreak occurred within a single generation. The infections appear to have resulted in acute mortality, especially among children (aged 8 to 11 years). We further note functional differences, including in the ypm superantigen locus, which is also present in present day Yersinia pseudotuberculosis . The new strains diverge ancestrally to known Y. pestis and constrain the timing of its emergence, indicating that this happened before approximately 5,700 years ago. These findings show that plague outbreaks happened earlier than previously thought and were indeed lethal. We contend that the occurrence of outbreaks among mid-Holocene hunter-gatherer communities well outside the sphere of Late Neolithic Europe challenges the notion that higher population densities and lifestyle changes during the Neolithic agricultural transition were prerequisites for plague epidemics.
Unveiling the complexity of post-Roman polity formation in Pannonia using ancient DNA
Open access
The transformation of the Roman world [fourth to ninth centuries common era (CE)], culminating in the Western Roman Empire’s fall, marked a fundamental transition in European history. Key questions persist regarding the regionally specific nature of this transformation. We generated a paleogenomic dataset to reconstruct post-Roman organizations in the Little Hungarian Plain at microregional resolution. Genetic and archaeological analyses of two Roman ( n = 68) and five post-Roman ( n = 246) sites reveal a rise in Northern European ancestry, reflecting large-scale population movements into this region. Moreover, despite post-Roman sites sharing similar genetic profiles, material culture, and burial practices, they show distinct social structures, especially regarding the role played by biological relatedness. These findings highlight local hierarchies and reveal the making of a post-Roman polity.
Population-scale Y chromosome assemblies reveal recurrent remodeling within constrained architectures
Open access
The human Y chromosome is among the most structurally dynamic chromosomes in the human genome, yet much of its diversity remains unresolved because of extensive palindromes, ampliconic gene families, satellite-rich heterochromatin and large segmental duplications. What remained unclear was how these diverse forms of variation fit together across the full chromosome, how often similar structures recur in different lineages, and which aspects of organization remain constrained despite rapid sequence turnover. Here, we generated and analyzed 142 nearly complete human Y chromosome assemblies from 17 major haplogroups spanning approximately 180,000 years of evolution, creating a population-scale resource for studying Y chromosome biology and diversity. These assemblies show that structural change on the Y chromosome is recurrent but constrained, even in its most repetitive regions. In the fertility-associated azoospermia factor c (AZFc) region, recurrent inversions, deletions, and complex rearrangements generate a limited repertoire of structural haplotypes. Multicopy ampliconic gene families follow distinct evolutionary paths: DAZ paralogues differ in structural constraint, RBMY evolves within a modular array, and TSPY copy number varies mainly through local expansion and contraction. The centromere and Yq12 heterochromatin vary greatly in size but retain a stable higher-order organization, including a single hypomethylated centromeric core and conserved Yq12 repeat composition and orientation. Methylation across palindromic and ampliconic regions is likewise structured by repeat class, copy order and local architecture. Together, these results provide a population-scale resource for the human Y chromosome and show that its rapid structural evolution is repeatedly funneled into a limited set of architectural outcomes.
A comprehensive analysis of Y-chromosomal diversity in Colombian populations
Open access
Abstract An intense gene flow between Native, European, and African groups, since the colonial period, has led to complex admixture patterns throughout Colombia. To investigate this genetic structure, we analyzed 23 Y-STRs in 975 males from five major regions of the country: Andes, Pacific, Caribbean, Orinoquía, and Amazon. To explore the paternal lineages and to reconstruct the underlying historical and demographic processes, a subset of 175 individuals was genotyped for 859 Y-SNPs using massively parallel sequencing. A high haplotypic diversity was found for the 23 Y-STRs (0.9998), compared to other Latin American populations studied. Pairwise genetic distance analyses ( F ST and R ST ) showed a clearer differentiation among Colombian regions, when increasing the number of Y-STRs from 17 to 23. In population comparisons, Caribbean and Orinoquía regions clustered closer to European and Admixed populations, while the Pacific region moves towards African populations. The Amazon clustered with Peru and Ecuador that have high Native American ancestry. The predominant paternal ancestry in the Caribbean and Andes was European (79% and 87.9%, respectively), with macrohaplogroup R being the most frequent (39% and 64%, respectively). In the Caribbean region, sub-Saharan African lineages are the second most prevalent (16%), while in the Andean region the Native American lineages are the second most represented (7.6%), followed by a smaller proportion of sub-Saharan African lineages (4.5%). The results attest to the extensive regional genetic diversity within the country. The Y-STRs showed good intra and inter-population discrimination capabilities and support the establishment of specific regional databases for forensic purposes.
Y-STR Mutation Patterns in North Indian Male Relatives at 16 Loci: A Preliminary Study
Open access
Y-chromosomal short tandem repeats (Y-STRs) are passed from father to son through the male line only. They change very little over generations and are widely used in forensic science and for tracing paternal ancestry. However, mutations at Y-STR loci can complicate relationship analyses, and may result in inaccurate exclusions during father-son or extended paternal lineage testing. . In this preliminary study, we looked for possible Y-STR mutations in 292 cases from the North Indian population. These cases included pairs of siblings, uncles and nephews, first cousins, and second cousins. We found a total of 29 mutations across 11 Y-STR markers out of 16 markers. All mutations were single-step changes consistent with the stepwise mutation model. The highest mutation rates were at DYS458 (0.006) and DYS385 (0.005), while loci such as DYS19, DYS390, and DYS389II had lower rates (0.001). These results align with previous findings on mutation rates at different loci. They also support the observation that Y-STRs with rapid mutations provide better information for distinguishing closely related paternal lineages. Our findings give baseline estimates of mutation rates for the North Indian populace, highlighting the need for population-specific data in forensic work and genealogical research.
Genetic Ancestry and Population Structure Across Ecuador
Open access
Background: Ecuador is a genetically diverse population setting shaped by long-term interactions among Native American, European, and African populations across distinct ecological regions. Although multiple studies have examined ancestry patterns in Ecuadorian populations, the available evidence remains fragmented and methodologically heterogeneous. Objective: To systematically identify, critically appraise, and synthesize published studies on genetic ancestry and population structure in Ecuador. Methods: A systematic review was conducted in accordance with PRISMA 2020. Searches were performed in PubMed/MEDLINE, Scopus, Web of Science Core Collection, SciELO, and Google Scholar through 31 January 2026. Eligible studies reported extractable ancestry-related data from Ecuadorian populations using autosomal, mitochondrial DNA, Y-chromosomal, or other ancestry-relevant genetic markers. Methodological quality was assessed using an adapted Joanna Briggs Institute framework. Owing to substantial heterogeneity across marker systems, sampling strategies, and ancestry inference methods, findings were synthesized qualitatively rather than by meta-analysis. Results: Of 1243 records identified, 12 studies met the inclusion criteria. Across marker systems, the evidence consistently supported a three-way admixture framework involving Native American, European, and African ancestry components, together with substantial regional and population-specific heterogeneity. Autosomal studies generally showed higher Native American ancestry in Highland and Native American populations, whereas African ancestry was more prominent in Afro-Ecuadorian and some Coastal populations. Uniparental markers further supported persistent sex-biased admixture, with predominant Native American maternal lineages and comparatively greater European or African paternal contributions depending on region and population history. Conclusions: Ecuadorian populations share a broad three-way admixture framework, but with marked internal heterogeneity across regions and population groups. These findings highlight the importance of geographic and demographic context in ancestry interpretation and the need for larger, more balanced, and methodologically standardized genomic studies in Ecuador.
Novel Y-STRs with elevated mutation rates further improve male relative differentiation
Open access
Y-chromosomal short tandem repeats (Y-STRs) with elevated mutation rates are valuable markers for distinguishing male suspects from their paternal male relatives - something that is typically not possible with standard Y-STRs. However, while the 26 rapidly mutating Y-STRs (RM Y-STRs) we identified in our two previous screens substantially improve male relative differentiation compared to standard Y-STRs, many close relatives cannot be separated with these markers. Aiming to further enhance the discrimination power of male relatives, particularly closely related ones, we performed a new chromosome-wide search for Y-STRs with elevated mutation rates by integrating in-silico marker discovery with experimental marker verification. Relative to previous screens, three major advancements were applied: (1) use of the Y-chromosome sequence from the telomere-to-telomere reference genome and other genomes, (2) consideration of all repeat motifs from homopolymers to hexanucleotides, and (3) use of targeted massively parallel sequencing to genotype male relatives for marker verification. To ensure robust allele calling and mutation detection for dinucleotide repeats prone to PCR slippage, we developed and applied a novel curve-fitting approach that accounts for all observed signals: true alleles and stutter products. Overall, we identified 14 novel Y-STRs with previously unreported elevated mutation rates, most of which were dinucleotide repeats. Relative to the 30-marker set of the RMplex tool, this novel set increased the empirical differentiation rates of close relatives separated by 1-4 meioses by 20.0%, 15.6%, 9.9% and 4.1%, respectively. The combined set of 44 novel and previous markers empirically differentiated 46.9%, 80.4%, 86.4%, and 84.8% of close relatives separated by 1-4 meioses and 96.4-100% of distant relatives separated by 5-15 meioses. The differentiation capacities of this 44-marker set, estimated from locus-specific mutation rates, were 50.2%, 75.2%, 87.6%, and 93.8% for these close relatives and 96.9-100% for the distant ones. Provided the development and successful forensic validation of a targeted genotyping tool, we anticipate that this expanded set of 44 Y-STRs with elevated mutation rates will enhance the ability to distinguish a male suspect from his paternal male relatives. This will benefit solving criminal cases where an autosomal STR profile of the male perpetrator cannot be generated and where standard Y-STR profiling yields a haplotype match between the suspect and the trace as well as haplotype sharing between the suspect and his male relative(s).
Paternal genetic structure and Y-chromosomal haplogroup prediction in the Tujia and Bai ethnic groups of Guizhou, Western China
Open access
BACKGROUND: Y chromosome genetic markers, with strict paternal inheritance and lack of recombination, are particularly valuable tools for tracing male lineages. They complement autosomal analyses in forensic applications and anthropological inference by increasing resolution for patrilineal structure. METHODS: We genotyped 382 unrelated male individuals from two Tibeto-Burman-speaking populations in Guizhou (Tujia, n = 220; Bai, n = 162) using the Goldeneye DNA Identification System Y Plus kit comprising 44 Y-markers. We calculated haplotype-level forensic indices and assessed inter-population structure via Rst-based multidimensional scaling (MDS) and a neighbor-joining (NJ) tree, based on genetic distances with 47 reference groups. Y-chromosomal haplogroups were predicted from Y-STR profiles to characterize paternal lineages. RESULTS: In the Tujia population, 338 alleles and 219 haplotypes were detected, with allelic frequencies ranging from 0.0045 to 0.9364. The haplotype diversity (HD), haplotype match probability (HMP), and discrimination capacity (DC) were 0.9999, 0.0046, and 0.9955, respectively. In the Bai population, 309 alleles and 141 haplotypes were detected, with allelic frequencies ranging from 0.0062 to 0.9691, with HD = 0.9979, HMP = 0.0083, and DC = 0.8704. Population genetic analysis revealed that the Guizhou Tujia and Bai groups share closer genetic affinity with Southern Han than with Northern Han and cluster with Tibeto-Burman-speaking groups, including the Sichuan and Guizhou Yi populations. Similarly, the Y-STR haplogroup prediction results revealed a multilayered paternal structure dominated by haplogroup O2a2, accompanied by contributions from indigenous East Asian lineages and minor inputs from West Eurasia and other regions. CONCLUSIONS: Our study provides valuable Y-STR data and forensic parameters for Tibeto-Burman-speaking ethnic groups in China, as well as population genetics evidence in patrilineal history. The Tujia and Bai of Guizhou exhibit a complex paternal genetic structure, offering insights into the demographic dynamics of Southwest China. The 43 Y-marker system exhibits high polymorphism and strong discriminatory power, supporting its utility as a powerful supplementary tool for forensic investigations, particularly for male lineage inference and suspect screening.
Select a publication to see its samples.