Biomarkers for Jomon ancestry testing

The identification of 132 SNPs as biomarkers enables accurate Jomon ancestry testing, addressing the limitations in understanding Japanese genetic ancestry and phenotypic variations.

JP2026047651APending Publication Date: 2026-03-16OSAKA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2026-03-16

AI Technical Summary

Technical Problem

The existing understanding of the genetic ancestry of the Japanese population, particularly the Jomon ancestors, is limited due to the scarcity of ancient DNA data and the unclear applicability of the ternary ancestral structure across the archipelago, necessitating comprehensive population-level genomic data for accurate modeling.

Method used

Identification of 132 SNPs as biomarkers for Jomon ancestry testing, including specific SNPs such as rs28729170(A), rs13017060(C), and others, to determine the presence and proportion of Jomon ancestors in a subject's genomic DNA.

Benefits of technology

Provides a reliable method to test for and estimate the proportion of Jomon ancestors, contributing to a clearer understanding of the genetic structure of the Japanese population and its phenotypic variations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026047651000008
    Figure 2026047651000008
  • Figure 2026047651000009
    Figure 2026047651000009
  • Figure 2026047651000010
    Figure 2026047651000010
Patent Text Reader

Abstract

To provide biomarkers for Jomon ancestry testing. [Solution] A biomarker for Jomon ancestor testing, comprising at least one SNP selected from a group including the group consisting of (x)rs28729170(A), rs13017060(C), rs2645158(G), rs6446239(C), rs1026980(T), rs2057165(A), rs4302225(T), rs4981864(T), rs629577(G), and rs6097031(C).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to biomarkers for Jomon ancestry testing, etc. [Background technology]

[0002] Anatomically modern humans, originating in Africa, began a global dispersal 50,000 to 60,000 years ago, undergoing repeated migration, settlement, and interbreeding. Their arrival in East Asia can be traced back at least 40,000 to 50,000 years ago. A crucial event in human history was the encounter between indigenous hunter-gatherers and migrant farmers, which subsequently led to a significant change in lifestyles. While this transition to agriculture occurred globally, its timing and process varied by region. The agricultural revolution in East Eurasia dates back approximately 10,000 years.

[0003] Archaeological evidence suggests that humans inhabited the Japanese archipelago, an island region in East Eurasia, as early as 38,000 years ago during the Paleolithic period. However, due to the scarcity of ancient DNA data, our understanding of the connection between their ancestors and modern humans is limited. One of the ancestral groups well-studied in Japan is the Jomon people. The Jomon people are known for their pioneering use of pottery, which is one of the oldest examples in the world. The Jomon period lasted until approximately 3,000 years ago, but immigrants from the continent introduced rice cultivation during the Yayoi period, from 3,000 to 1,700 years ago. This agricultural revolution spurred socio-political development, leading to the establishment of the Japanese state during the Kofun period, which lasted for 200-300 years starting around 1,700 years ago.

[0004] The long-standing model for the origin of the Japanese population is a dual ancestral structure. This model posits that modern Japanese are a hybrid of Jomon hunter-gatherers from Southeast Asia and agriculturalists from Northeast Asia. However, recent ancient DNA research provides compelling evidence that the genetic origin of the Japanese population is composed of three distinct ancestors (i.e., a ternary ancestral structure): (1) the ancient hunter-gatherers, the Jomon people; (2) Northeast Asian ancestors introduced during the Yayoi period, the agricultural era; and (3) East Asian ancestors introduced during the Kofun period, the state-forming era. This ternary structure, established during the Kofun period, persists in the modern population. The Japanese population is well characterized by a north-south gradient of genomic variation. However, due to the limitations of the modern human samples and geographical representativeness used in previous studies, the applicability of the ternary model and the variability of the three distinct ancestral components across the archipelago remain unclear. Therefore, it is crucial to model this ancestral structure using comprehensive population-level genomic data. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Marnetto, D. et al. Ancestral genomic contributions to complex traits in contemporary Europeans. Curr. Biol. 32, 1412-1419.e3 (2022). [Overview of the project] [Problems that the invention aims to solve]

[0006] Recent advances in ancient genomics have not only made it possible to identify diverse genetic ancestries regionally and globally, but also, for example, the exacerbation and resistance to coronavirus infection due to genes introduced by Neanderthals. Genetic predisposition to multiple sclerosis brought about by the ancestors of the Steppes. However, the extent to which the past of humanity has shaped the phenotypic variations today, especially outside of Europe, is not yet fully understood (Non-Patent Document 1).

[0007] The present invention focuses on the Jomon people, the ancestors of the Japanese, and aims to provide biomarkers for Jomon ancestry testing.

Means for Solving the Problems

[0008] As a result of intensive research in view of the above problems, the present inventor succeeded in identifying 132 SNPs as biomarkers for Jomon ancestry testing. Based on this finding, the present inventor further advanced the research and completed the present invention. That is, the present invention includes the following aspects.

[0009] Item 1. (x) The group consisting of rs28729170(A), rs13017060(C), rs2645158(G), rs6446239(C), rs1026980(T), rs2057165(A), rs4302225(T), rs4981864(T), rs629577(G), and rs6097031(C), and (y)rs4298845(C)、rs7595986(A)、rs10764532(T)、rs111468015(A)、rs1172495(A)、rs11076581(C)、rs806469(A)、rs77125278(C)、rs13257488(C)、rs13253574(C)、rs6456790(T)、rs11030161(G)、rs3812526(G)、rs7832131(G)、rs388394(T)、rs35692656(T)、rs10497147(A)、rs13134059(A)、rs934213(C)、rs2486729(T)、rs12699444(T)、rs2722810(C)、rs12325527(T)、rs404725(A)、rs2570415(T)、rs9537153(A)、rs2172180(G)、rs7202291(C)、rs1478870(T)、rs34349793(T)、rs912758(G)、rs11960706(G)、rs6687335(G)、rs10759526(G)、rs12615487(A)、rs2248501(G)、rs12277503(C)、rs1716646(T)、rs712591(A)、rs7795435(C)、rs60799144(T)、rs2997881(T)、rs930592(G)、rs9496721(T)、rs957303(T)、rs437239(C)、rs1858730(G)、rs1358968(A)、rs2871660(G)、rs488165(A)、rs12675727(A)、rs359546(G)、rs11852766(G)、rs28495134(G)、rs3800577(T)、rs4902386(C)、rs6667615(G)、rs7980063(A)、rs6586660(A)、rs7030507(G)、rs78890485(T)、rs644713(C)、rs8030121(G)、rs7200600(G)、rs7069081(C)、rs10067255(G)、rs6703740(T)、rs1168551(A)、rs12498572(T)、rs647548(A)、rs2469954(C)、rs6105551(T)、rs11083432(G)、rs2100791(T)、rs12416497(A)、A group consisting of rs11769455(T), rs3954044(G), rs2505859(A), rs10923743(T), rs9499598(T), rs61776759(T), rs1237302(T), rs4716901(G), rs61849335(G), rs6823966(A), rs9294661(C), rs7140599(C), rs685218(A), rs4947605(C), rs10866961(G), rs11061561(A), rs76839556(T), rs10835830(A), rs10792719(C), rs10822062(A), rs1031026(G), rs11705906(A), rs6434910(G), rs6913029(C), rs6556892(A), rs10167446(G), rs3790272(T), rs536618(C), rs695872(A), rs10091796(A), rs9893926(T), rs2337247(G), rs6726814(C), rs1884784(G), rs307575(G), rs6075075(T), rs4698413(T), rs6534672(T), rs113114417(A), rs608873(C), rs7332756(C), rs1290900(A), rs2060179(T), rs2280973(A), rs881647(A), rs11779483(A), and rs4104180(A).[[ID=[]]] [[ID=[]]]A biomarker for Jomon ancestor testing, consisting of at least one SNP selected from the group consisting of [[ID=[]]] [[ID=[]]]

[0010] [[ID=[]]] [[ID=[]]]Item 6. The biomarker for Jomon ancestor testing according to Item 1, wherein the SNP contains at least one selected from the group (x).[[ID=[]]] [[ID=[]]]

[0011] [[ID=[]]] [[ID=[]]]Item 3. (1) In the genomic DNA of a biological sample collected from a subject, [[ID=[]]] (x) The group consisting of rs28729170(A), rs13017060(C), rs2645158(G), rs6446239(C), rs1026980(T), rs2057165(A), rs4302225(T), rs4981864(T), rs629577(G), and rs6097031(C), and (y)rs4298845(C)、rs7595986(A)、rs10764532(T)、rs111468015(A)、rs1172495(A)、rs11076581(C)、rs806469(A)、rs77125278(C)、rs13257488(C)、rs13253574(C)、rs6456790(T)、rs11030161(G)、rs3812526(G)、rs7832131(G)、rs388394(T)、rs35692656(T)、rs10497147(A)、rs13134059(A)、rs934213(C)、rs2486729(T)、rs12699444(T)、rs2722810(C)、rs12325527(T)、rs404725(A)、rs2570415(T)、rs9537153(A)、rs2172180(G)、rs7202291(C)、rs1478870(T)、rs34349793(T)、rs912758(G)、rs11960706(G)、rs6687335(G)、rs10759526(G)、rs12615487(A)、rs2248501(G)、rs12277503(C)、rs1716646(T)、rs712591(A)、rs7795435(C)、rs60799144(T)、rs2997881(T)、rs930592(G)、rs9496721(T)、rs957303(T)、rs437239(C)、rs1858730(G)、rs1358968(A)、rs2871660(G)、rs488165(A)、rs12675727(A)、rs359546(G)、rs11852766(G)、rs28495134(G)、rs3800577(T)、rs4902386(C)、rs6667615(G)、rs7980063(A)、rs6586660(A)、rs7030507(G)、rs78890485(T)、rs644713(C)、rs8030121(G)、rs7200600(G)、rs7069081(C)、rs10067255(G)、rs6703740(T)、rs1168551(A)、rs12498572(T)、rs647548(A)、rs2469954(C)、rs6105551(T)、rs11083432(G)、rs2100791(T)、rs12416497(A)、rs11769455(T), rs3954044(G), rs2505859(A), rs10923743(T), rs9499598(T), rs61776759( T), rs1237302(T), rs4716901(G), rs61849335(G), rs6823966(A), rs9294661(C), rs7140599( C), rs685218(A), rs4947605(C), rs10866961(G), rs11061561(A), rs76839556(T), rs108358 30(A), rs10792719(C), rs10822062(A), rs1031026(G), rs11705906(A), rs6434910(G), rs691 3029(C), rs6556892(A), rs10167446(G), rs3790272(T), rs536618(C), rs695872(A), rs1009 1796(A), rs9893926(T), rs2337247(G), rs6726814(C), rs1884784(G), rs307575(G), rs60750 The group consisting of 75(T), rs4698413(T), rs6534672(T), rs113114417(A), rs608873(C), rs7332756(C), rs1290900(A), rs2060179(T), rs2280973(A), rs881647(A), rs11779483(A), and rs4104180(A), A method for examining Jomon ancestors, comprising the step of checking for the presence or absence of at least one SNP selected from a group consisting of the following.

[0012] Item 4. The testing method described in Item 3, which is a method for testing the presence or absence of Jomon ancestors and / or the proportion of Jomon ancestors.

[0013] Item 5. Furthermore, (2) If at least one of the SNPs is detected in step (1), a step of determining whether the subject has Jomon ancestors or Japanese ancestors based on the detected SNP, or estimating the proportion of Jomon ancestors in the subject. The inspection method described in item 4, including the inspection method described in item 4.

[0014] Item 6. The testing method according to item 5, wherein in step (2b), the proportion of the subject to Jomon ancestors is estimated based on the detected SNPs and the prediction coefficients for each SNP.

[0015] Item 7. The group consisting of (x)rs28729170(A), rs13017060(C), rs2645158(G), rs6446239(C), rs1026980(T), rs2057165(A), rs4302225(T), rs4981864(T), rs629577(G), and rs6097031(C), and (y)rs4298845(C)、rs7595986(A)、rs10764532(T)、rs111468015(A)、rs1172495(A)、rs11076581(C)、rs806469(A)、rs77125278(C)、rs13257488(C)、rs13253574(C)、rs6456790(T)、rs11030161(G)、rs3812526(G)、rs7832131(G)、rs388394(T)、rs35692656(T)、rs10497147(A)、rs13134059(A)、rs934213(C)、rs2486729(T)、rs12699444(T)、rs2722810(C)、rs12325527(T)、rs404725(A)、rs2570415(T)、rs9537153(A)、rs2172180(G)、rs7202291(C)、rs1478870(T)、rs34349793(T)、rs912758(G)、rs11960706(G)、rs6687335(G)、rs10759526(G)、rs12615487(A)、rs2248501(G)、rs12277503(C)、rs1716646(T)、rs712591(A)、rs7795435(C)、rs60799144(T)、rs2997881(T)、rs930592(G)、rs9496721(T)、rs957303(T)、rs437239(C)、rs1858730(G)、rs1358968(A)、rs2871660(G)、rs488165(A)、rs12675727(A)、rs359546(G)、rs11852766(G)、rs28495134(G)、rs3800577(T)、rs4902386(C)、rs6667615(G)、rs7980063(A)、rs6586660(A)、rs7030507(G)、rs78890485(T)、rs644713(C)、rs8030121(G)、rs7200600(G)、rs7069081(C)、rs10067255(G)、rs6703740(T)、rs1168551(A)、rs12498572(T)、rs647548(A)、rs2469954(C)、rs6105551(T)、rs11083432(G)、rs2100791(T)、rs12416497(A)、rs11769455(T), rs3954044(G), rs2505859(A), rs10923743(T), rs9499598(T), rs61776759( T), rs1237302(T), rs4716901(G), rs61849335(G), rs6823966(A), rs9294661(C), rs7140599( C), rs685218(A), rs4947605(C), rs10866961(G), rs11061561(A), rs76839556(T), rs108358 30(A), rs10792719(C), rs10822062(A), rs1031026(G), rs11705906(A), rs6434910(G), rs691 3029(C), rs6556892(A), rs10167446(G), rs3790272(T), rs536618(C), rs695872(A), rs1009 1796(A), rs9893926(T), rs2337247(G), rs6726814(C), rs1884784(G), rs307575(G), rs60750 The group consisting of 75(T), rs4698413(T), rs6534672(T), rs113114417(A), rs608873(C), rs7332756(C), rs1290900(A), rs2060179(T), rs2280973(A), rs881647(A), rs11779483(A), and rs4104180(A), A diagnostic reagent for use in the testing method described in any of items 3 to 6, comprising at least one SNP detection agent selected from the group consisting of the following. [Effects of the Invention]

[0016] According to the present invention, a biomarker for Jomon ancestry testing can be provided. [Brief explanation of the drawing]

[0017] [Figure 1]This shows the population composition of Biobank Japan. a) Seven geographical regions are represented by different colors, indicating the location where participants were registered at local hospitals. b) Scatter plot of PCA of BBJ participants using 1KG EAS samples. c) Clustering results of BBJ participants based on PCA. d) Geographic location of the Ryukyu Islands. e) Scatter plot of PCA of BBJ participants in the Ryukyu Islands. Color-coded according to the location of the registered hospital. In a) and d), the map of Japan was drawn using the R package "jpndistrict" (https: / / github.com / uribo / jpndistrict). PC stands for Principal Components, 1KG for the 1000 Genomes Project, JPT for Japanese in Tokyo, CDX for Chinese Dai in Xishuangbanna, CHB for Han Chinese in Beijing, CHS for Han Chinese in the South, KHV for Kinh in Ho Chi Minh City, NEA for Northeast Asians, and EA for East Asians. [Figure 2] This graph shows the changes in the ternary ancestry structure among Biobank Japan participants. The bar graphs show the proportions of three different ancestors. a) The population is defined as the entire BBJ sample or by the region in which participants were registered. c) The BBJ sample is divided into five different island populations based on the PCA clusters shown in Figure 1. Error bars represent the standard error, and the value at the top of each bar indicates the tail probability of the ternary model for each group. NEA is Northeast Asian, EA is East Asian. [Figure 3] This shows the genetic legacy of Jomon ancestors remaining in the entire Japanese population. a) Projection of the Jomon proportion onto a PCA plot. b) Absolute value of the correlation coefficient between the Jomon proportion and PCs. c) Box plot showing the variation in the Jomon proportion by registration region of BBJ participants. d) Box plot showing the variation in the Jomon proportion by various genetic clusters defined by PCA. In c) and d), the boxes represent the interquartile range (IQR), the median is shown as a white horizontal bar; the whiskers extend up to 1.5 times the IQR; outliers are shown as individual points. PC is the principal component, EA is East Asian. [Figure 4]This study demonstrates the association between the Jomon people's ancestors and 80 composite traits. The associations were examined using a generalized linear model for a) all BBJ participants (n = 163,243) and b) only participants in the mainland cluster (n = 152,148). Quantitative traits were modeled using linear regression, and binary traits were analyzed using logistic regression. The direction of the triangles corresponds to the sign of the beta coefficient for each feature. The gray dashed line represents statistical significance based on Bonferroni correction (P < 0.05 / 80 = 6.3 × 10⁻⁴). The control group shows 10 dummy phenotypes. All labeled traits on the plot have nominal significance (P < 0.05). BMI, Body Mass Index; BW, Body Weight; LVM, Left Ventricular Mass; E / A, E / A Ratio; RBC, Red Blood Cell Count; LDLC, Low-Density Lipoprotein Cholesterol. [Figure 5] This paper presents the functional and genomic characterization of Jomon-related marker SNPs. a) Stratified LD score regression analysis was performed across 10 major cell type groups. The false detection rate (FDR) was calculated using the Benjamini-Hochberg method. b) The violin plot compares the length of the region with strong LD for each of the 132 Jomon-related SNPs (r² > 0.8) to that of 132 frequency-matched non-Jomon-related SNPs. The p-value (3.1 × 10⁻³²) is calculated using the Wilcoxon rank-sum test. c) The locus plot highlights the Jomon-related SNPs linked to the longest haplotype among the 132 variants (shown as purple diamonds in the figure above). The dashed line defines statistical significance as a p-value of 5 × 10⁻⁸. The lower plot represents the GENCODE-based protein-coding genes present in the highlighted regions. [Figure 6]This paper presents the results of testing the predictive power of the genetic heritage of Jomon ancestors and 132 Jomon-related variants using an independent cohort of BBJ-2nd. a) The PCA plot includes all individuals of BBJ-2nd (n = 68,632) and overlays their respective Jomon proportions. The box plot shows the distribution of Jomon proportions observed at each decile of the predicted score. R² represents the squared Pearson correlation coefficient between the predicted score and the residuals of the Jomon proportions regressed on the 10 PCs. b) In this case, the boxes represent the interquartile range (IQR), the median is shown as a white horizontal bar; whiskers extend up to 1.5 times the IQR; and outliers are shown as individual points. [Figure 7] This paper presents the estimation of Jomon ancestry and the reproduction of phenotypic associations using UKB EAS. Jomon component prediction scores are divided into decimal numbers. Box plots show the distribution of Jomon proportions observed in certain deciles of prediction scores, based on a) all UKB EAS participants (n = 566), b) EG6 (self-reported multi-ethnic group; n = 200), and c) EG5 (self-reported Chinese group; n = 308). d) Forest plots show the influence of Jomon ancestry on BMI in UKB EAS, EG5, and EG6. Error bars indicate 95% confidence intervals. Asterisks represent P < 0.05. In a), b), and c), the boxes represent the interquartile range (IQR), and the median is shown by a white horizontal bar. [Modes for carrying out the invention]

[0018] In this specification, the terms “contains” and “includes” include the concepts of “contains,” “includes,” “substantially consist of,” and “consist solely of.”

[0019] In one embodiment, the present invention relates to a biomarker for Jomon ancestry testing (which may be referred to herein as "the biomarker of the present invention") comprising at least one SNP (single nucleotide polymorphism) selected from the group consisting of group (x) and group (y).

[0020] Group (x) consists of rs28729170(A), rs13017060(C), rs2645158(G), rs6446239(C), rs1026980(T), rs2057165(A), rs4302225(T), rs4981864(T), rs629577(G), and rs6097031(C).

[0021] (y) group is rs4298845 (C), rs7595986 (A), rs10764532 (T), rs111468015 (A), rs1172495 (A), rs11076581 (C), rs806469 (A), rs77125278 (C), rs13257488 (C), rs13253574 (C), rs6456790 (T), rs11030161 (G), rs3812526 (G), rs7832131 (G), rs388394 (T), rs35692656 (T), rs10497147 (A), rs13134059 (A), rs934213 (C), rs2486729 (T), rs12699444 (T), rs2722810 (C), rs123rs11769455(T), rs3954044(G), rs2505859(A), rs10923743(T), rs9499598(T), rs61776759( T), rs1237302(T), rs4716901(G), rs61849335(G), rs6823966(A), rs9294661(C), rs7140599( C), rs685218(A), rs4947605(C), rs10866961(G), rs11061561(A), rs76839556(T), rs1083583 0(A), rs10792719(C), rs10822062(A), rs1031026(G), rs11705906(A), rs6434910(G), rs6913 029(C), rs6556892(A), rs10167446(G), rs3790272(T), rs536618(C), rs695872(A), rs100917 96(A), rs9893926(T), rs2337247(G), rs6726814(C), rs1884784(G), rs307575(G), rs6075075 This group consists of (T), rs4698413(T), rs6534672(T), rs113114417(A), rs608873(C), rs7332756(C), rs1290900(A), rs2060179(T), rs2280973(A), rs881647(A), rs11779483(A), and rs4104180(A).

[0022] Group (x) consists of biomarkers that have a greater effect (Effect Size: Tables 2-5) on determining the presence and proportion of Jomon ancestors. From this perspective, among group (x), rs28729170(A), rs13017060(C), rs2645158(G), rs6446239(C), and rs1026980(T) are preferred, more preferably rs28729170(A) and rs13017060(C), and particularly preferably rs28729170(A).

[0023] Each biomarker is indicated by its registration number (rs number) in the NCBI SNP Database (http: / / www.ncbi.nlm.nih.gov / snp / ), with the single-letter base of each biomarker shown in parentheses. For example, rs28729170(A) indicates that the base of the SNP represented by rs28729170 is adenine (A).

[0024] By examining the presence or absence of the biomarker of the present invention, it is possible to test for Jomon ancestors. From this perspective, the present invention relates in one embodiment to a method for testing for Jomon ancestors (which may also be referred to as "the testing method of the present invention" in this specification), comprising the step of examining the presence or absence of at least one SNP selected from the group consisting of group (x) and group (y) in the genomic DNA of a biological sample taken from a subject.

[0025] The biomarker of this invention is a single nucleotide polymorphism in a human gene. Therefore, the subject is human.

[0026] The subjects' place of residence is not particularly limited and could include, for example, Asia, Europe, Africa, North America, South America, Oceania, etc. Of these, Asia is preferred, and Japan is particularly preferred.

[0027] The number of subjects to whom the testing method of the present invention is applied is not particularly limited, but for example, it may be 10 or more, 50 or more, 100 or more, or 1000 or more. The upper limit of this number is not particularly limited, but for example, it may be 100,000, 10,000, or 5,000.

[0028] The subjects may be either specimens for which information about the Jomon ancestors is unknown, or specimens for which information about the Jomon ancestors is known.

[0029] The biological sample is not particularly limited as long as it contains chromosomal genomic DNA contained in the nucleus of the subject's cells. Examples of biological samples include body fluids, skin, mucous membranes, and internal tissues. Among these, body fluids, skin, and mucous membranes are preferred, and more preferably, body fluids, from the viewpoint of ease of collection and minimal invasiveness.

[0030] Examples of bodily fluids include blood, follicular fluid, menstrual blood, saliva, cerebrospinal fluid, synovial fluid, urine, tissue fluid, sweat, and tears. Examples of mucous membranes include the oral mucosa and nasal mucosa.

[0031] The biological sample may be the sample taken directly from the living organism, or it may be a sample obtained by concentrating and purifying genomic DNA.

[0032] Biological samples may be used individually or in combination of two or more types.

[0033] Biological samples can be collected from a subject by methods known to those skilled in the art. For example, whole blood can be collected by blood collection using a syringe or the like. It is preferable that blood collection be performed by a medical professional such as a doctor or nurse. Serum is the portion of blood from which blood cells and certain blood clotting factors have been removed, and can be obtained, for example, as the supernatant after blood has coagulated. Plasma is the portion of blood from which blood cells have been removed, and can be obtained, for example, as the supernatant after centrifugation under conditions that do not cause blood to coagulate.

[0034] Step (1) involves examining the genomic DNA of the subject for the presence or absence of the biomarker of the present invention. The presence of the biomarker of the present invention serves as an indicator that the subject has Jomon ancestors or Japanese ancestors, and / or estimates the proportion of Jomon ancestors in the subject.

[0035] The method for determining the presence or absence of the biomarker of the present invention is not particularly limited as long as it is a method capable of specifically detecting DNA of a specific base sequence, and various known methods or methods similar thereto can be employed. Examples of such methods include PCR (e.g., real-time PCR), DNA sequencing, DNA microarrays, and Southern hybridization. More specifically, methods that use primers and probes specific to polymorphisms and detect amplification by PCR and polymorphism of the amplified product by PCR by luminescence include the TaqMan-PCR method, Invader method, method using FRET, ASP-PCR method, MALDI-TOF / MS method using primer extension method, RCA method, method using DNA chips or microarrays, and DigiTag2 method.

[0036] In one embodiment of the above method, DNA-binding molecules (e.g., primers, probes, etc.) can be used. Primer pairs and probes can be designed and synthesized based on the base sequence of the biomarker of the present invention. The base lengths of the primers and probes are not particularly limited. The base length of the primer can be, for example, 10 to 50 nucleotides, preferably 15 to 30. The base length of the probe can be, for example, 10 to 5000 nucleotides, preferably 10 to 1000, more preferably 20 to 150.

[0037] Primer pairs and probes can be made from natural nucleic acids such as RNA and DNA, or, if necessary, from natural nucleic acids to chemically modified nucleic acids or pseudo-nucleic acids. Examples of chemically modified nucleic acids and pseudo-nucleic acids include PNA (Peptide Nucleic Acid), LNA (Locked Nucleic Acid; registered trademark), methylphosphonate-type DNA, phosphorothioate-type DNA, and 2'-O-methyl-type RNA. Furthermore, primers and probes may contain fluorescent substances and / or quencher substances, or radioisotopes (e.g., 32 P, 33 P, 35The material may be labeled or modified using a labeling substance such as S), or a modifying substance such as biotin, (strept)avidin, or magnetic beads.

[0038] The labeling substance is not limited and commercially available substances can be used. For example, fluorescent substances such as FITC, Texas, Cy3, Cy5, Cy7, Cyanine3, Cyanine5, Cyanine7, FAM, HEX, VIC, fluorescein and its derivatives, and rhodamine and its derivatives can be used. Quencher substances such as AMRA, DABCYL, BHQ-1, BHQ-2, or BHQ-3 can be used. The labeling position of the labeling substance on the primer and probe may be determined appropriately according to the characteristics of the modifying substance and the intended use. Generally, modification is often performed at the 5' or 3' end. Furthermore, a single primer and probe molecule may be labeled with one or more types of labeling substances. The design of primer and probe nucleotide sequences and the selection of labeling substances are well-known and disclosed in molecular biology experimental protocol manuals such as Molecular Cloning: A Laboratory Manual (3rd ed., Cold Spring Harbor Laboratory Press, 2001) by Sambrook, J and Russell, DW.

[0039] The number of biomarkers of the present invention examined in step (1) is preferably large from the viewpoint of judgment and estimation accuracy, for example, 2 or more, 3 or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, 120 or more, 130 or more, or 132.

[0040] Step (1) includes checking whether each single nucleotide polymorphism is present in both chromosomes of the pair, and checking whether it is present in one of the chromosomes of the pair.

[0041] According to the inspection method of the present invention, which includes step (1), it is possible to provide information regarding the presence or absence of the biomarker of the present invention, which is an indicator for the examination of Jomon ancestors, thereby assisting in the examination of Jomon ancestors.

[0042] The inspection method of the present invention, in one embodiment, moreover, (2) If at least one of the SNPs is detected in step (1), a step of determining whether the subject has Jomon ancestors or Japanese ancestors based on the detected SNP, or estimating the proportion of Jomon ancestors in the subject. It is preferable that it includes.

[0043] Since the Jomon people are the ancestors of the Japanese people, it can be determined that having Jomon ancestors means having Japanese ancestors.

[0044] The phrase "based on detected SNPs" is not particularly limited as long as it uses information about the detected SNPs that are biomarkers of the present invention. Examples include the presence or absence of the SNPs, the number of SNPs, and the discriminant formula obtained from the SNPs. From the viewpoint of judgment / estimation accuracy, it is preferable to use the discriminant formula obtained from the detected SNPs that are biomarkers of the present invention.

[0045] The discriminant formula can be created by performing, for example, regression analysis based on known statistical analysis, etc., on a subject population in which Jomon ancestry information (for example, Jomon ancestry information obtained by the analysis method of the examples described later) is clear, such as the presence or absence of Jomon ancestors and the proportion of DNA derived from Jomon ancestors in the total genomic DNA, and on the biomarker information of the present invention for each subject in the said population.

[0046] In one embodiment, the discriminant formula is as follows: For each biomarker of the present invention, a coefficient determined from the viewpoint of judgment / estimation accuracy is used, and for each biomarker of the present invention detected in the subject, the coefficient is multiplied by the coefficient, and the value obtained by summing the products can be used as the judgment score. For example, if the judgment score is above a standard value or within a certain range, it can be determined that the subject has Jomon ancestors or Japanese ancestors, and the proportion of Jomon ancestors in the subject can be estimated (for example, multiple levels (e.g., three levels: high, medium, and low) are set for the proportion of Jomon ancestors and it is determined which one applies, or the value of the proportion is estimated).

[0047] For each biomarker of the present invention, it is preferable to use the Effect Size as the coefficient. As the coefficient, for example, a value of ±0.010 of the Effect Size shown in Tables 2 to 5 of the examples described below can be used.

[0048] Furthermore, weighting can be applied depending on whether the biomarker of the present invention is heterozygous or homozygous. For example, the coefficient can be multiplied by 1 if the biomarker of the present invention is heterozygous, and by 2 if it is homozygous.

[0049] The biomarker detection agent of the present invention can be used as a diagnostic reagent for use in the testing method of the present invention. In this view, the present invention relates in one embodiment to a diagnostic reagent for use in the testing method of the present invention (which may also be referred to herein as "the diagnostic reagent of the present invention") comprising a detection agent for at least one SNP selected from the group consisting of group (x) and group (y).

[0050] The detection agent is not particularly limited as long as it can detect the biomarker of the present invention. Examples of such detection agents include primers, probes, etc., for the biomarker of the present invention.

[0051] The detection agent may be modified, provided that its function is not significantly impaired. Examples of modifications include the addition of labels such as fluorescent dyes, enzymes, proteins, radioisotopes, chemiluminescent substances, and biotin.

[0052] Suitable fluorescent dyes used in the present invention are those generally used to label nucleotides for the detection and quantification of nucleic acids. Examples include, but are not limited to, HEX (4,7,2',4',5',7'-hexachloro-6-carboxylfluorescein, green fluorescent dye), fluorescein, NED (trade name, Applied Biosystems, yellow fluorescent dye), or 6-FAM (trade name, Applied Biosystems, yellow-green fluorescent dye), rhodamin or its derivatives (e.g., tetramethylrhodamin (TMR)). Any suitable known labeling method can be used to label nucleotides with a fluorescent dye (see Nature Biotechnology, 14, 303-308 (1996)). Commercially available fluorescent labeling kits can also be used (e.g., Amersham Pharmacia's Oligonucleotide ECL 3'-Oligolabeling System).

[0053] The detection agent can also be immobilized on any solid phase for use. Therefore, the detection agent of the present invention can be provided in the form of a substrate on which the detection agent is immobilized (for example, a microarray chip on which a probe is immobilized).

[0054] The solid phase used for immobilization is not particularly limited as long as it can immobilize polynucleotides, etc., and examples include glass plates, nylon membranes, microbeads, silicon chips, capillaries, or other substrates. The method of immobilizing the detection agent onto the solid phase is not particularly limited. The immobilization method is well known in the art, depending on the type of immobilized probe, for example, using a commercially available spotter (such as one from Amersham) in the case of a microarray [e.g., in situ synthesis of oligonucleotides using photolithographic technology (Affymetrix), inkjet technology (Rosetta Inpharmatics), etc.].

[0055] Primers, probes, etc., are not particularly limited as long as they selectively (specifically) recognize the biomarker of the present invention. Here, "selectively (specifically) recognized" means, for example, that the biomarker of the present invention is specifically amplified in the PCR method, but is not limited to that; it is sufficient if a person skilled in the art can determine the presence or absence of the biomarker of the present invention from the detected substance, amplified substance, or the presence or absence thereof.

[0056] Specific examples of primers and probes include the polynucleotides listed in (a) below and the polynucleotides listed in (b) below: (a) A polynucleotide having at least 15 consecutive bases in the base sequence of a gene having the biomarker of the present invention and / or a polynucleotide complementary to said polynucleotide, (b) A polynucleotide having at least 15 bases that hybridizes under stringent conditions to the base sequence of the gene having the biomarker of the present invention or a base sequence complementary thereto. At least one selected from the group consisting of the following is mentioned.

[0057] A complementary polynucleotide or complementary base sequence (complementary strand, reverse strand) refers to a polynucleotide or base sequence that is nucleotide-complementary to the full-length polynucleotide sequence consisting of the base sequence of the gene having the biomarker of the present invention, or a partial sequence having at least 15 consecutive bases in length (for convenience, these are also referred to here as the "forward strand"). This complementary relationship is based on base pairings such as A:T and G:C. However, such a complementary strand is not limited to forming a perfectly complementary sequence with the base sequence of the target forward strand; it may also have a complementary relationship to the target forward strand to the extent that it can hybridize under stringent conditions. Here, stringent conditions can be determined based on the melting temperature (Tm) of the nucleic acid to which the complex or probe is bound, as taught in Berger and Kimmel (1987, Guide to Molecular Cloning Techniques Methods in Enzymology, Vol. 152, Academic Press, San Diego CA). For example, typical washing conditions after hybridization include conditions of approximately "1×SSC, 0.1%SDS, 37°C". It is preferable that the complementary strand maintains its hybridized state with the target positive strand even after washing under such conditions. While not particularly limited, more stringent hybridization conditions include washing conditions of approximately "0.5×SSC, 0.1%SDS, 42°C", and even more stringent hybridization conditions include washing conditions of approximately "0.1×SSC, 0.1%SDS, 65°C". Specifically, examples of such complementary strands include strands consisting of base sequences that are completely complementary to the base sequence of the target positive strand, and strands consisting of base sequences that have at least 90%, preferably 95%, more preferably 98% or more, identity with the positive strand.

[0058] Primers, probes, etc., can be designed, for example, based on the base sequence of the biomarker of the present invention, using various design programs.

[0059] The base length of primers, probes, etc., is not particularly limited as long as it has a length of at least 15 consecutive bases, as described above, and can be set appropriately depending on the application. For example, when used as a primer, the base length can be 15 to 35 bases, and when used as a probe, the base length can be 15 to 35 bases.

[0060] The diagnostic reagent of the present invention may also be in the form of a composition. The composition may optionally contain other components. Examples of other components include bases, carriers, solvents, dispersants, emulsifiers, buffers, stabilizers, excipients, binders, disintegrants, lubricants, thickeners, humectants, colorants, fragrances, chelating agents, and the like.

[0061] The diagnostic reagent of the present invention may be in the form of a kit. In addition to the above-mentioned detection agent or a composition containing the same, the kit may also contain other materials that can be used to detect the biomarker of the present invention in the genomic DNA of a biological sample of a subject. Specific examples of such materials include various reagents (e.g., buffer solutions), and equipment (e.g., equipment for collecting, purifying, and separating biological samples). [Examples]

[0062] The present invention will be described in detail below based on examples, but the present invention is not limited to these examples.

[0063] (1) Method (1-1) Biobank Japan (BBJ) Dataset We used 171,287 participants from the initial BBJ cohort, which enrolled participants between 2003 and 2007. BBJ is a hospital-based genomic cohort, enrolling participants with at least one of 47 diseases from 12 medical institutions in 7 regions of Japan. All participants provided written informed consent with the approval of the Ethics Committee of the Institute of Medical Science, University of Tokyo. This study was approved by the Ethics Committees of the Graduate School of Medicine, Osaka University and the Graduate School of Medicine, University of Tokyo.

[0064] (i) Individuals with a low call rate (<99%), (ii) closely related individuals with a genetic relevance of 0.178 or higher calculated from the Genetic Relevance Matrix (GRM) using GCTA (version 1.93.3β2), and (iii) individuals located far from the Japanese cluster defined within the 1000 Genomes Project (1KG) dataset in the PCA plot using PLINK2 (v2.00a2.3) were excluded. Based on the location where participants registered, seven geographically defined populations were defined for the Japanese archipelago (i.e., Hokkaido, Tohoku, Kanto-Koshinetsu, Chubu-Hokuriku, Kinki, Kyushu, Okinawa (from northeast to southwest)) and eight populations for the Ryukyu Islands (Yakushima, Amami, Kikai, Okinoerabu, Tokunoshima, Yoron, Okinawa, Miyako). Five subpopulations were identified by visual inspection of the PCA plots (i.e., Mainland, Ryukyu, Ryukyu_admix, EA_admix, Hokkaido_sub).

[0065] BBJ GWAS data were genotyped using either the Illumina HumanOmniExpressExome BeadChip or a combination of the Illumina HumanOmniExpress BeadChip and HumanExome BeadChip. Genotype quality control is described in other literature. Briefly, variants meeting the following criteria were excluded: (i) call rate < 99%, (ii) Hardy-Weinberg equilibrium (HWE) p-value < 1.0 × 10⁻⁶. -6 (iii) number of heterozygotes < 5, (iv) concordance rate < 99.5%, or non-reference concordance rate between GWAS array and whole genome sequence. Genotype data were complemented by Eagle v2 and imputed with WGS merged from 1000 Genomes Project Phase3v5 (n = 2,504) and BBJ1K WGS (n = 1,037) using Mimimac3 software (2.0.1).

[0066] (1-2) Data merging between modern and ancient genome data We used ancient genomes from the Japanese archipelago and East Eurasia, integrated into the Simons Genome Diversity Project (SGDP) panel (SGDP_Ancient). This SGDP_Ancinet dataset contains 14 ancient Japanese individuals, all of whom were shotgun-sequenced. Under strict quality control (including ancient DNA damage and low-coverage data), the ancient genomes (n = 22) totaled 3,867,366 sites, consisted only of translocations, and were pseudo-diploid, with a minor allele frequency of 1%. Furthermore, to integrate the ancient genome data with BBJ genotype data, we used PLINK2 to convert BBJ's accurate doser data (Rsq≧0.7) into genotype data. Next, using PLINK (v1.90b4.4), we extracted sites present in both the SGDP_Ancient and BBJ genotype data, obtaining a final merged dataset containing a total of 2,038,260 sites.

[0067] (1-3) Modeling of mixtures and f 4 test We applied qpAdm from AdmixTools (version 7.0.2). In our previous study, we used only translocation sites with a global minor allele frequency of 1%, and set them to "allsnps:NO". In qpAdm, we set the left and right populations as the source and reference populations, respectively. We selected nine individuals as the right population: Sardinians (n=3), Kusundans (n=2), Papuans (n=14), Dai (n=4), Amis (n=2), Naxi (n=3), Tien Yuan (n=1), Chokhopani (n=1), and Maltese (n=1). Three ancient groups were set up as the left-hand group: Jomon people (n = 12), Northeast Asians (n ​​= 2; WLR_BA_o and HMMH_MN, individuals from the Bronze Age and Middle Neolithic periods in the Western Liao River basin), Han Chinese (n = 3; SGI), and Han Chinese (n = 3; SGDP panel).

[0068] We evaluated whether the ternary structure fits the data better at the population level than other possible scenarios, based on the p-values ​​of nested models with a cutoff of 0.05. These alternative models include a dual-structure model with any combination of Jomon, Northeast Asian, and East Asian ancestors, and a single-ancestor model. When performing interbreeding modeling at the individual level, we considered the ternary model to fit if (i) the tail probability is greater than or equal to 0.05 and / or (ii) the estimate of the interbreeding rate is viable (i.e., between 0.0 and less than 1.0). The correlation of the interbreeding rates of three different ancestors was calculated using Pearson's method. For individuals that did not support the ternary model because the tail probability was less than 5%, we attempted to identify the alternative dual-ancestor model with the highest tail probability, and then confirmed it with a nested p-value > 0.05 by comparing the ternary model and the dual-structure model. If the single-ancestor model was valid, we further evaluated whether the individual could be adequately explained by a specific ancestor alone using nested p-values.

[0069] The f4 statistic was measured using qpDstat with the f4 mode of AdmixTools, in the form f4(Mbuti, Jomon; Han, X). The target populations were either BBJ and 1KG populations (n ​​= 2,504) or SGDP populations (in the form X).

[0070] (1-4) Phenotype curation at Biobank Japan BBJ collected clinical status, laboratory data, and behavioral information from all participants through interviews using standardized questionnaires and review of medical records. As a result, 80 items were extracted (3 anthropometric items, 55 biomarker items, 2 behavioral items, 2 reproductive items, and 18 disease items). Data from individuals aged 18 and over was used, but only alcohol consumption and smoking traits from individuals aged 20 and over were included. For quantitative biomarkers, the same processing and quality control methods as previously reported were applied. In short, laboratory values ​​measured when participants first visited the recruitment center were used, and the values ​​were adjusted based on the type of medication used. Next, rank-based inverse normal transformation was applied to normalize the biomarker traits. Behavioral characteristics, including drinking history and smoking history (with and without drinking history, with and without smoking history), were analyzed as binary phenotypes. Reproductive traits were coded as age of menarche and age of menopause. Cases where the age of menarche was under 10 or over 20, or where the age of menopause was under 40 or over 60, were excluded. Patients with myocardial infarction, stable angina, and unstable angina were reclassified as having coronary artery disease (CAD). Eighteen diseases were selected from the BBJ's target disease group, with sufficient numbers of cases and controls (arrhythmia, asthma, cataracts, CAD, dyslipidemia, ischemic stroke, congestive heart failure, osteoporosis, glaucoma, chronic hepatitis C, colorectal cancer, gastric cancer, hay fever, urolithiasis, rheumatoid arthritis, prostate cancer, breast cancer, and type 2 diabetes). Individuals without any of the specific diseases studied were treated as controls.

[0071] Additionally, a dummy phenotype was set as a negative control. Using the GCTA GWAS simulation method, a predefined heritability (h) was determined from 10,000 causal variants randomly sampled from the BBJ GWAS data. 2 Ten phenotypes with = 0.5 were simulated. The values ​​were normalized by applying rank-based inverse normalization.

[0072] (1-5) Phenotypic influence of Jomon ancestors on the Japanese population The association between Jomon ancestors and traits was tested using the glm() function implemented in R software (version 4.1.0). Linear regression models were applied to quantitative traits, and logistic regression models were applied to diseases and behavioral habits, with adjustments made for covariates. The proportion of Jomon ancestors was normalized using rank-based inverse normal transformation. Covariates included sex, age, age squared, top 20 PCs, 45 disease statuses, geographical region, PCA clusters, and trait-specific covariates. For BMI, the association with Jomon ancestors was further stratified by sex and age, with an age threshold of 65 years, reflecting the mean age of BBJ participants.

[0073] To address the problem of multicollinearity between the Jomon ancestors and PC, we adopted a different approach. First, we performed regression analysis of quantitative trait measurements for all covariates, including PC. Then, we used a linear regression model to test the association between the Jomon lineage and the residuals obtained from this regression.

[0074] (1-6) Evaluation of genome-wide estimation and polygene prediction based on the presence or absence of Jomon people's ancestors We investigated the effect of Jomon ancestry proportion on the accuracy of BMI GWAS and BMI polygene score (PGS). We performed a BMI GWAS on all BBJ participants using a generalized mixed linear model (MLM) approach with GCTA-fastGWA. We employed a sparse genetic relation matrix (GRM) constructed with variants subject to minimal LD ​​pruning. The covariates in the original BMI GWAS were sex, age, age squared, top 20 PCs, and 45 disease statuses. To assess the effect of Jomon ancestry, we introduced Jomon ancestry proportion as an additional covariate and compared it with the results obtained from the original GWAS.

[0075] To calculate the PGS, since there was no independent external reference for GWAS or genotype data with Jomon ancestry, the 5-fold leave-one-group-out GWAS method was adopted. Briefly, first, BBJ was randomly divided into five subsets. Then, using GCTA-fastGWA, a BMI GWAS was performed on the samples excluding the subset under investigation (i.e., the target subset). To estimate the SNP posterior effects from the GWAS summary data, PRS-CS-auto (version June 4, 2021) was utilized with the 1KG EAS HapMap3 LD reference panel. These posterior estimates of the effect sizes were estimated from both the original GWAS and the GWAS with the Jomon pedigree as covariates. Finally, the PLINK2 score function was applied to calculate the PGS for individuals within the target subset and compare the scores with and without adding the Jomon pedigree as a covariate.

[0076] The increase in the predictive performance of Jomon ancestry was quantified as follows: 2 :

[0077]

Number

[0078] Here, R 2 PGS is the R of the PGS modeled by trait ~ PGS + covariates (sex, age, age squared, top 20 PCs, 45 disease statuses, geographical region, PCA cluster), 2 and R 2 PGS + Jomon is the R of the same PGS model 2 but with the proportion of Jomon people additionally included.

[0079] (1-7) Identification of Jomon-related variants To detect genetic variations associated with Jomon lineage, genome-wide association studies were performed between the proportion of Jomon individuals and 6,861,976 double-stranded variants (MAF≧1%, Rsq≧0.7). (i) MLM-based approach using covariate-adjusted GCTA-fastGWA: age, age squared, sex, top 20 PCs, 45 disease statuses, geographical region, PCA cluster; (ii) fixed-effects meta-analysis using METAL (version 2020-05-05) of mainland summary data including populations from mainland and EA_admix clusters (n = 152,148), and Ryukyu summary data including populations from Ryukyu, Ryukyu admix, and Hokkaido_sub clusters (n = 10,142); and (iii) dual genome control correction method using METAL. Next, the Z-score for each variant was calculated, taking into account the sign of the beta coefficient and the associated p-value.

[0080] To identify independent variants, BBJ1K and 1KG EAS were used as reference populations, and variants with positive Z-scores were subjected to LD clamping with the following PLINK parameters: p1=1, p2=1, r2=0.01, kb=2000. Jomon-related variants (132 variants in total) were independent, not located at HLA loci, and showed genome-wide significance (P < 5.0 × 10⁻¹⁰). -8 ) was defined as satisfying the following conditions. For Jomon-related variants between 1KG EAS groups, Hudson's F implemented in PLINK2 was used. ST Using F STValues ​​were calculated. Variants with a p-value greater than 0.05 were classified as non-Jomon-related variants. Next, using nearest neighbor matching, 132 non-Jomon-related variants consistent with the allele frequencies of Jomon-related variants were identified, and Jomon-related and non-Jomon-related variants were compared by examining frequency differences between Japanese and East Asian populations, and the length of haplotypes with an LD greater than 0.8. The enrichment of selection signals in Jomon-related variants was validated based on the Z-scores of singleton density scores (SDS) estimated in previous studies. The sum of squares of rank-based normalized Z-scores was compared to a chi-squared distribution where the degrees of freedom are equal to the number of available variants.

[0081] Stratified LD score regression was applied to summary data of Jomon proportions from a meta-analysis, and the enrichment of cell populations was estimated using the recommended baseline LD model. For LD score regression, HapMap3 SNPs excluding SNPs within HLA regions were used, and pre-calculated LD scores between 1KG EAS populations obtained from the LDSC software website were used. Pleiotropic effects of Jomon-related variants were investigated using BioBank Japan PheWeb (https: / / pheweb.jp / ) for Japanese populations and Open Targets Genetics (https: / / genetics.opentargets.org / ) for European populations. The eQTL effects of variants were also evaluated using GTEx (https: / / gtexportal.org / home / ).

[0082] (1-8) Verification using an independent Japanese cohort The Nagahama Cohort Study and the 2nd BBJ Cohort (BBJ-2nd) were used as independent replication cohorts. The Nagahama Cohort Study was community-based, recruiting participants from Nagahama City, Shiga Prefecture. Genotyping of this cohort was determined using six different genotyping arrays. Next, two platforms with large sample sizes were selected (Nagahama A; Illumina Human610-Quad Beadchip, Nagahama B; Illumina HumanOmni2.5-4v1 Beadchip). From the EAS population (n = 1,591 for Nagahama A, n = 1,444 for Nagahama B), individuals with low call rates, high heterozygosity rates, closely related individuals, and PCA abnormalities were excluded. In addition, (i) call rate < 0.98, (ii) MAF < 1%, (iii) HWE P-value < 1.0 × 10⁻⁶ -6 The variants were excluded. Genotype data were phase-transformed using Eagle v2 and imputed using Mimimac3 with reference panels from 1000 Genomes Project Phase3v5 and BBJ1K. As described above, high-quality imputed dose data were converted to genotype data and integrated with ancient genome data. As a result, 1,982,989 shared variants were obtained in Nagahama A and 2,109,225 in Nagahama B.

[0083] BBJ-2nd is an additional cohort of participants independent of the BBJ 1st cohort. BBJ-2nd consists of approximately 80,000 participants collected between 2013 and 2018. Participants were genotyped using Illumina Asian Screening Array Chips. As with the 1st cohort, strict QC filtering was applied to both participants and SNPs, as described elsewhere. Briefly, individuals with low call rates (< 0.98), closely related individuals (King's kinship index ≥ 0.0884), and outliers from the EAS cluster in PCA with samples from the HapMap3 project (n = 72,695) were excluded. Additionally, call rates less than 0.99, minor allele counts less than 5, and HWE P-values ​​of 1.0 × 10⁻⁶ were excluded. -10Variants with an allele frequency difference of less than 0.05 compared to the Japanese WGS reference panel were excluded. Genotype data were complemented using SHAPEIT (version 4.2.1) and imputed using Mimimac4 (version 1.0.1) with the 1000 Genomes Project Phase 3v5 and BBJ1K reference panels. High-quality input dosage data (Rsq ≥ 0.7) were converted to genotype data and merged with ancient genome data to obtain 1,940,657 shared variants.

[0084] Similar to the first cohort of BBJ, the proportion of Jomon ancestors was estimated for individuals in each cohort using qpAdm. Using the PLINK2 score option, a score predicting Jomon ancestors was derived from 132 independent Jomon-related variants, and this was defined as the Jomon component prediction score. After regression on the Jomon proportion with 10 PCs as covariates, the variance explained by the scaled Jomon component prediction score was estimated using Pearson correlation. In each cohort, the principal components were calculated using projections on the PCs of the first cohort of BBJ. Clamping and thresholding approaches were employed to evaluate the power of the Jomon component prediction score. The Jomon component score for BBJ-2nd individuals was constructed using gene mutations that satisfied the following p-value threshold: 5 × 10 -8 , 5×10 -7 , 1×10 -6 , 1×10 -5 , 1×10 -4 , 1×10 -3 , 0.01, 0.05, 0.1. To assess the proportion of variance explained by the Jomon component score, an adjusted R-squared model was used from a full model including the score and all covariates, compared to a null model that does not include the score. 2 I calculated it.

[0085] (1-9) Reproduction analysis using East Asians from the UK Biobank As an independent source of East Asian (EAS) populations, we focused on the EAS population in the UK Biobank (UKB). EAS individuals were extracted by visual inspection of PCA plots including a 1KG population as a reference for EAS ancestors (n = 2,286). UKB individuals were genotyped using either the Applied Biosystems UK BiLEVE Axiom Array or the Applied Biosystems UKB Axiom Array. Genotypes were imputed using a 1000 Genomes Phase 3 reference panel with the Haplotype Reference Consortium, UK10K, and IMPUTE4. 72 Detailed characteristics of the cohort are described in the following literature. For UKB EAS individuals, variants with an INFO score of 0.8 or higher were converted into genotype data, yielding 3,610,183 sites shared with the ancient genome data.

[0086] The genetic affinity between the Jomon people and the UKB EAS, or between the Jomon people and each ethnic background group within the UKB EAS, was tested using the f4 statistic of form f4(Mbuti, Jomon; Han, X). The breakdown of ethnic backgrounds self-reported by participants at their first visit (Data-Field 21000) is as follows: Data coding 1 "Caucasian" (n=1), Data coding 1001 "British" (n=1), Data coding 2003 "Caucasian and Asian" (n=2), Data coding 2004 "Other mixed race" (n=3), Data coding 3 "Asian or Asian British" (n=2), Data coding 3004 "Other Asian" (n=300), Data coding 5 "Chinese" (n=1,375), Data coding 6 "Other EG" (n=541).

[0087] The proportion of Jomon individuals in each cohort was estimated using qpAdm. Where an individual did not support a tripartite ancestry structure, the same procedure used in BBJ was applied to identify a plausible model of bilateral interbreeding in which the Jomon were one of the primary ancestors. Jomon component prediction scores based on 132 independent Jomon-related variants were calculated using PLINK2.

[0088] To evaluate the association between Jomon ancestry and BMI in the UKB EAS, the mean BMI of participants measured two or three times was calculated. Association tests were adjusted for sex, age, age squared, confirmation center information, batch information, and ethnic background. This method was also applied to EG5 (i.e., self-reported Chinese, n=1,375) and EG6 (i.e., self-reported other ethnic groups, n=541) with the same covariates, excluding ethnic background.

[0089] (2) Results (2-1) Estimation of the ternary ancestry structure in the Biobank Japan dataset To evaluate the fit of the triplicate model to modern populations, we used BBJ GWAS data (Figure 1a). The total number of participants was 171,287, distributed across the archipelago from northeast to southwest as follows: Hokkaido = 7,955, Tohoku = 11,013, Kanto-Koshinetsu = 94,981, Chubu-Hokuriku = 9,489, Kinki = 25,200, Kyushu = 15,962, and Okinawa = 5,804. Our PCA clearly defined clusters (Figure 1b), separating the Ryukyu cluster, mainly consisting of the Okinawa population, from the other populations, as reported in previous studies. Regional clusters were also observed in Tohoku, Kanto-Koshinetsu, Kinki, and Kyushu. Based on the PCA results, five distinct genetic clusters were defined within the Japanese population: EA_admix (n = 1,019), Mainland (n = 159,642), Ryukyu_admix (n = 640), Ryukyu (n = 9,847), and Hokkaido_sub (n = 139) (Figure 1c). Within the Ryukyu Islands, the stratification of populations is even more pronounced, reflecting the geographical affinity within these islands (Figures 1d and e, Yakushima, n=431, Amami, n=1,531, Kikai, n=561, Okinoerabu, n=845, Tokunoshima, n=476, Yoron, n=167, Okinawa, n=4,795, Miyako, n=827).

[0090] This diverse Japanese population was then integrated with ancient genome data from Japan and the Eurasian continent. To find regions present in both the array-typing genome data of BBJ and the pseudo-diploid genome data of ancient humans, high-precision imputation-surge data (Rsq≧0.7) was converted to genotype data. This conversion yielded 2,038,260 shared variants (n = 171,287 in BBJ and n = 22 in ancient genomes). Subsequently, AdmixTools' qpAdm was applied to evaluate the goodness of fit of the hybridization model and to estimate the hybridization ratio at both the population and individual levels of the biobank. To comprehensively capture the geographical and genetic diversity of the Japanese population, ternary models were fitted to geographically defined populations across the Japanese and Ryukyu Islands (Figures 1b and 1e), five genetically defined populations (Figure 1c), and the entire BBJ dataset. This analysis showed that the triplicate structure provides a better fit for all populations, both at broad and local scales, compared to all possible dual-family structure models. The only exception is EA_admix, a population that can be adequately explained by bidirectional interbreeding between Northeast Asian and East Asian ancestors.

[0091] Across the entire BBJ dataset, the proportions of the three distinct ancestral components are largely consistent with those reported in previous studies (Jomon: 12.4%, Northeast Asia: 21.2%, East Asia: 66.4%). However, the proportion of Jomon ancestors varied regionally, being 9.8% in Kinki and 26.1% in Okinawa (Figure 2a). The proportion of Jomon ancestors was high in the Ryukyu Islands, with the highest proportion on Yoron Island (Figure 2b). In Hokkaido_sub (31.6%, Figure 2c), one of the genetically defined populations, the proportion of Jomon ancestors was even higher. In contrast, EA_admix is ​​likely composed of continental East Asian individuals and contains very few Jomon ancestors. The mainland cluster reflects the proportions of the entire BBJ, as it contains the majority of the samples in the data (159,642 out of 171,287 individuals). Even when individuals are separated from this cluster based on geographical origin (i.e., where samples were collected), these proportions remain relatively consistent across different regions. The Ryukyu cluster represents the ancestral composition originating from Okinawa, and the ratio of the Ryukyu_admix cluster is positioned between that of the mainland cluster and the Ryukyu cluster.

[0092] Next, we investigated whether the ancestors of the Jomon people exist only in Japanese populations or are also observed in continental populations using f4 statistics (Mbuti, Jomon; Han, X). The populations included were the Simon Genome Diversity Project (SGDP) panel, the 1000 Genomes Project (1KG), and subpopulations within BBJ. However, the Ulchi people of SGDP and East Asians (EAS) of 1KG were exceptions. In the 1KG EAS population, only Japanese people in Tokyo (JPT) showed a significant affinity to the Jomon people. Among BBJ participants, the Hokkaido_sub and Ryukyu subpopulations showed extremely strong affinity to the Jomon people, which is consistent with the high rate of Jomon ancestry in mixed-race modeling (Figure 1e).

[0093] Overall, our analysis provides a detailed map of regional differences in the ternary ancestry structure across the entire archipelago.

[0094] (2-2) Individual differences in the three-part structure The hybridization model revealed that the two subpopulations of Ryukyu and Hokkaido have a higher proportion of Jomon ancestors compared to other subpopulations (Figure 2). To visualize the genetic distance between ancient and modern populations, ancient or modern individuals representing three different ancestors that underlie the ternary structure of modern Japanese populations were projected onto PCA plots, along with individuals of even older Japanese populations (i.e., Yayoi and Kofun people). Kofun people are included within the variation of the current population, while Jomon people are clustered in a position extending from the Ryukyu and Hokkaido subpopulations. Two individuals from the Yayoi period are morphologically Jomon, but genetically hybridized with Jomon. Two individuals from the Yayoi period are morphologically considered to be Jomon, but genetically hybridized with Jomon and continental people. Two individuals from the Yayoi period are morphologically considered to be Jomon, but genetically hybridized with Jomon and continental ancestors.

[0095] Since population-based hybridization modeling only represents average patterns within a population, a ternary model was then applied to each BBJ participant. This model fit 154,339 out of 171,287 individuals (90.1%), with varying proportions of the three ancestral components. These three ancestral proportions were negatively correlated with each other, supporting a previously proposed scenario that two continental ancestors, Northeast Asians and East Asians, likely arrived in the archipelago independently. Approximately 5% of individuals (8,932 individuals) did not fit the ternary model, and a bilateral hybridization including either Jomon and East Asian ancestors, or Northeast Asian and East Asian ancestors, was shown to be a better fit. However, as indicated by the nested p-value < 0.05, the bilateral hybridization model did not provide sufficient support compared to the ternary model, and 28 individuals were excluded from further analysis. There were also a few exceptions, with 10 individuals preferring the East Asian ancestor-only model over the dual-structure model. It is important to note that none of the models adequately fit the remaining individuals, who accounted for nearly 5% (7,962 people). This is likely because a 5% tail probability cutoff was used to account for the inherent variability in the data, as a predetermined proportion of all tests were expected to fall outside the tested model.

[0096] When the proportion of Jomon ancestors was incorporated into the PCA plot, a significant slope was observed along the first and second principal components in the proportion of Jomon individuals (Figure 3a). In fact, the proportion of Jomon individuals showed a remarkably strong correlation with the first and second principal components (Figure 3b; |R| = 0.61, 0.09, PC1 and PC2, respectively). Similar correlation patterns were observed between Northeast Asian or East Asian ancestors and PCs, but the strength of the correlation was not as pronounced as with the Jomon individuals (|R| = 0.14 for Northeast Asia and PC1; |R| = 0.26 for East Asia and PC1). The correlation between Jomon ancestors and PCs remains clear even when focusing only on individuals from the mainland or Ryukyu clusters. These results strongly suggest that Jomon ancestors play a crucial role in shaping Japanese genomic variation at the PCA level. Furthermore, these individual-based estimates not only reflect population-based patterns in their mean values, but also highlight significant variability in Jomon ancestors across the Japanese archipelago (Figures 3c and 3d).

[0097] (2-3) The influence of Jomon ancestors on phenotypic variation in the Japanese population Next, we investigated whether the ancestors of the Jomon people phenotypically influenced the current population. Of the 163,243 individuals in the BBJ (British Behavioral Journal) whose genetic ancestry was successfully modeled, the average proportion of Jomon ancestry was 12.5 ± 6.3% (mean ± SD). There were no significant differences in the proportion of Jomon ancestry between age groups or genders.

[0098] Next, robust adjustments were made for genetic and geographical subpopulations to examine the association between the Jomon people's ancestors and 80 composite traits. To account for type I error, the statistical significance threshold was set to P < 0.05 / 80 = 6.3 × 10⁻⁶. -4 The threshold was set based on Bonferroni correction. This threshold was confirmed to be calibrated by simulating 10 dummy hereditary phenotypes as negative controls. In the analysis of all BBJ participants, a significant association with increased body mass index (BMI) was found (Figure 4a; β=0.012, standard error [SE]=0.003, P=3.0×10⁻⁶). -5). However, it is important to note that regional disparities, such as Okinawa's higher BMI compared to other regions, can complicate this association. However, even after adjusting for geographical factors as a covariate, it is important to note that regional disparities, such as Okinawa's higher BMI compared to other regions, can still complicate this association. To address this concern, we limited the analysis to individuals included in the mainland cluster (n = 152,148; see Figure 1c). The association with BMI remained statistically significant (Figure 4b; Beta = 0.012, SE = 0.003, P = 7.9×10⁻⁶). -5 Furthermore, a significant association with BMI was confirmed regardless of gender or age.

[0099] While our approach may be considered conservative, we prioritized addressing the potential inflation of association signals due to population stratification. Therefore, it was crucial to account for the effects of stratification stemming from differences in Jomon ancestry among individuals by incorporating PC as a covariate. To mitigate this issue, we also employed a method of correcting for phenotypes by regression against all covariates, including PC, before testing associations with Jomon ancestry. As a result, a significant association was found for BMI, but no statistically significant differences were observed for any of the other traits. Taken together, the robustness of the BMI signal was consistently demonstrated regardless of the method used to correct for population stratification (Figure 4).

[0100] To assess the impact of Jomon ancestry on BMI, we incorporated the proportion of Jomon ancestry as a covariate into a genome-wide association study (GWAS) of BMI. While most read SNPs were consistent, a reduction in the influence of Jomon ancestry on the polygene score (PGS) of BMI was observed. This effect persists to some extent, and it is noteworthy that it indicates a functional correlation between Jomon ancestry and BMI-PGS. Nevertheless, our findings suggest that Jomon ancestry acts as a confounding factor when estimating the effect size of BMI. Indeed, when Jomon ancestry is not considered in the GWAS, PGS tends to be inflated, and the degree of inflated is weakly correlated with the proportion of Jomon ancestry.

[0101] Our analysis includes height, which has been shown to exhibit a north-south gradient in the Japanese population. The association between Jomon ancestry and decreased height is observable only when PC is not considered in the test, suggesting that this association may be confused by population stratification. Overall, these results suggest that the genetic legacy of the ancient hunter-gatherers, the Jomon people, significantly influences BMI across the population today, regardless of geographical differences, and consequently contributes to the increased risk of obesity. Furthermore, we examined the impact of using the proportion of Jomon people on the predictive power of the polygene score (PGS) for BMI. The result showed an incremental predictive performance of -2.8 × 10⁻⁶. -3 It was limited by its size.

[0102] (2-4) Identification of the underlying genetic mutations in the ancestors of the Jomon people We investigated genome-wide variants associated with individual differences in Jomon ancestors. In this analysis, the proportion of an individual's Jomon ancestor was considered a surrogate phenotype. However, conventional null hypotheses, as used in standard genotype-phenotype association studies, do not directly apply to this phenotype. To ensure robust correction for genomes and geographical subpopulations and to control genomic inflation of test statistics, we employed a dual genome control correction method and a mixed linear model approach including PC as a covariate. Furthermore, because the proportion of Jomon individuals differed significantly among these populations, we performed meta-analyses of association tests individually for individuals within the mainland and Ryukyu populations (Figure 3d). To gain biological insights into the genetic signals enriched in Jomon ancestors, we performed stratified linkage disequilibrium score regression (S-LDSC) on the genome-wide meta-analysis results. S-LDSC results identified significant enrichment of heritability across major cell populations in skeletal muscle cells (Figure 5a).

[0103] Based on our association analysis, a positive Z-score indicates genome-wide significance (P < 5.0 × 10⁻¹⁰) after strict control measures. -8 Variants that reached 10¹⁶ (P = 3.1 × 10¹⁶) were classified as Jomon-related variants. As a result, 132 independent variants were identified from LD clamping (Tables 1-5). To explore the evolutionary background of these Jomon-related variants, the haplotype structure of the genomic region containing the Jomon-related variants was examined. When these haplotype structures were compared with the haplotype structures of non-Jomon-related mutants with matched allele frequencies (i.e., mutants with P > 0.05), it was found that Jomon-related mutants exhibited significantly longer haplotypes than non-Jomon-related mutants (P = 3.1 × 10¹⁶). -32(Figures 5b, 5c). This strongly supports the Jomon origin of these long haplotypes, and it is thought that strong linkage disequilibrium occurs across the genome, probably due to the high genetic homogeneity of the Jomon population and the small effective population size (~1,000 people). Furthermore, the persistence of these long Jomon-derived haplotypes in modern populations is thought to be due to the relatively recent interbreeding with continental ancestors. This haplotype is thought to be due to the relatively recent interbreeding with continental ancestors and the lack of sufficient recombination time to divide the haplotype. We utilized the results of selection scans based on singleton density scores (SDS). Furthermore, we observed enrichment of selection signals in Jomon-derived haplotypes (P = 0.008). These results support the idea that Jomon-associated variants function as markers for quantifying Jomon ancestry, tagging Jomon-derived haplotypes, and potentially subjecting them to recent selection pressures.

[0104] These Jomon-related variants are significantly more frequent in the JPT compared to other East Asian populations in the 1KG dataset (Table 1). Within the BBJ population, the Ryukyu population has a higher frequency of these variants than the mainland population, supporting a strong link to the Jomon-derived segment. To assess the specificity in the Japanese population, a fixation index (F) was used for 132 Jomon-related variants. ST ) was measured. F among East Asian populations ST There isn't a huge difference, but rs536618 in 4p12 and rs2871660 in 1q31 are relatively high F values ​​in the Japanese population. ST The values ​​were shown. Notably, the top mutation, rs13017060, is located in the intron region of the NBAS gene and exhibits multifaceted effects on BMI and weight gain in the modern Japanese population. Furthermore, as is evident from the GTEx data, it also functions as an eQTL (equivalent quantity limit) of NBAS expression in the cardiomyocyte. In addition, it is associated with increased leg fat mass in the UK Biobank. These findings provide further evidence supporting a significant association between the Jomon ancestors and BMI.

[0105] In Tables 1-5, Ref (REF) represents the wild-type base of the SNP, and Alt (ALT) represents the mutant base of the SNP. In Tables 2-5, Effect Size can be used to calculate the Jomon prediction score. That is, the following formula:

[0106]

number

[0107] Based on this, the sum of the products of the effect size and the individual's genotype can be calculated as the Jomon prediction score using 132 Jomon proportion-related variants. In the above formula, Xi is 0 if allele i is wild-type homozygous, 1 if the variant is heterozygous, and 2 if the variant is homozygous.

[0108] [Table 1]

[0109] [Table 2]

[0110] [Table 3]

[0111] [Table 4]

[0112] [Table 5]

[0113] (2-5) A triplicate model using an independent Japanese cohort and verification of Jomon-related variations. To validate the mixed-race model of the tripartite mixed-race structure, we used two independent Japanese cohorts: the Nagahama Cohort and the BBJ Second Cohort (BBJ-2nd). The Nagahama Cohort includes participants from Nagahama City, Shiga Prefecture, in the Kinki region of Japan (consisting of Nagahama A with n=1,549 and Nagahama B with n=1,444). BBJ-2nd is an independent additional population distinct from the BBJ First population, representing various regions of the Japanese archipelago (n = 72,695).

[0114] Specifically, this involved (i) estimating the proportion of Jomon ancestors at the individual level, and (ii) deriving a Jomon ancestor prediction score as a form of PGS. In the BBJ-2nd PCA, a significant correlation was confirmed between the first two PCs (Figure 6a) and Jomon ancestors (Figure 3a). Next, using 132 Jomon-related variants identified from the first BBJ cohort, a Jomon ancestor prediction score was derived. Under robust adjustment for population stratification, the variance of Jomon ancestors explained by the prediction score was as follows: 0.14 for Nagahama A, 0.12 for Nagahama B, and 0.13 for BBJ-2nd. PN Nagahama A = 2.5 × 10 -50 、PN Nagahama B = 1.3 × 10 -40 , PBBJ -2nd < 1.0 × 10 -300 Notably, when the predicted scores are divided into decile groups, the observed proportion of Jomon people gradually increases along the predicted scores across different cohorts (Figure 6b).

[0115] p-value cutoff 5 × 10 -8 From 1 x 10 -3 When the p-cutoff was relaxed, the predictive power of the Jomon-related variant increased, resulting in variances of Jomon ancestors explained by genetic variants of 0.27 (Nagahama A), 0.26 (Nagahama B), and 0.24 (BBJ-2nd), respectively. However, it should be noted that this clamping and thresholding method may overfit the data, especially when the p-cutoff is relaxed. Therefore, in subsequent analyses, the p-cutoff for predicting Jomon ancestors was set to 5 × 10⁻¹⁰.-8 That's what I decided.

[0116] (2-6) Reconstruction of the relationship between Jomon people's ancestry and BMI using East Asians from the UK Biobank We attempted to replicate our findings on the association between Jomon ancestry and increased BMI using a completely independent cohort. The subjects were East Asians within the UK Biobank (UKB EAS). First, we selected individuals from the UKB EAS based on PCA plots (n = 2,286). An f4-test in the form f4(Mbuti, Jomon; Han, UKB EAS) showed a symmetric relationship between the Jomon and both the Han and UKB EAS (Z = 0.93). Next, we divided the UKB EAS into different groups based on self-reported ethnic background and performed f4-tests individually for each group. This analysis revealed that ethnic group (EG) 6 (i.e., self-reported other ethnic groups, n = 541) showed a higher genetic affinity to the Jomon than to the Han Chinese, with a Z score of 5.4. None of the other EGs, including EG5 consisting of self-reported Chinese (n = 1,375), supported a significant affinity to the Jomon. These results suggest that the EG6 participants had Jomon ancestry and represent a suitable subgroup for our reconstruction analysis.

[0117] By fitting the mixed-race model to the UKB EAS cohort (n = 2,286), we were able to quantify the ancestry of 566 Jomon individuals. The predictive power of the score based on 132 Jomon-related variants accounted for approximately 2% of the total variance of the observed proportion of Jomon individuals (R 2 = 0.02 and P = 2.7 × 10 -4 (Figure 7a). Focusing only on EG 6 within the UKB EAS, this force increases significantly (R 2 = 0.06 and P = 3.8 × 10 -4 (Figure 7b). In contrast, this prediction is less effective in EG5 (R 2 = 0.03, P = 0.002 (Figure 7c), which is consistent with the fact that this group clearly has few ancestors of the Jomon people.

[0118] Finally, we examined the influence of Jomon ancestry on BMI in the UKB EAS. No significant association was found in the UKB EAS as a whole or in EG5 of the UKB EAS (Figure 7d; β=0.11, SE=0.58, P=0.86 for the UKB EAS as a whole; β=-0.83, SE=1.30, P=0.52 for EG5). However, when focusing only on EG6, where the genetic influence of Jomon is stronger, a significant association was found between Jomon ancestry and increased BMI (Beta = 2.2, SE = 0.99, P = 0.03).

[0119] Overall, these results highlight the potential influence of Jomon ancestors on the obesity risk of modern populations, regardless of differences in living environments, such as those observed in Britain and Japan.

Claims

1. (x) The group consisting of rs28729170(A), rs13017060(C), rs2645158(G), rs6446239(C), rs1026980(T), rs2057165(A), rs4302225(T), rs4981864(T), rs629577(G), and rs6097031(C), and (y)rs4298845(C)、rs7595986(A)、rs10764532(T)、rs111468015(A)、rs1172495(A)、rs11076581(C)、rs806469(A)、rs77125278(C)、rs13257488(C)、rs13253574(C)、rs6456790(T)、rs11030161(G)、rs3812526(G)、rs7832131(G)、rs388394(T)、rs35692656(T)、rs10497147(A)、rs13134059(A)、rs934213(C)、rs2486729(T)、rs12699444(T)、rs2722810(C)、rs12325527(T)、rs404725(A)、rs2570415(T)、rs9537153(A)、rs2172180(G)、rs7202291(C)、rs1478870(T)、rs34349793(T)、rs912758(G)、rs11960706(G)、rs6687335(G)、rs10759526(G)、rs12615487(A)、rs2248501(G)、rs12277503(C)、rs1716646(T)、rs712591(A)、rs7795435(C)、rs60799144(T)、rs2997881(T)、rs930592(G)、rs9496721(T)、rs957303(T)、rs437239(C)、rs1858730(G)、rs1358968(A)、rs2871660(G)、rs488165(A)、rs12675727(A)、rs359546(G)、rs11852766(G)、rs28495134(G)、rs3800577(T)、rs4902386(C)、rs6667615(G)、rs7980063(A)、rs6586660(A)、rs7030507(G)、rs78890485(T)、rs644713(C)、rs8030121(G)、rs7200600(G)、rs7069081(C)、rs10067255(G)、rs6703740(T)、rs1168551(A)、rs12498572(T)、rs647548(A)、rs2469954(C)、rs6105551(T)、rs11083432(G)、rs2100791(T)、rs12416497(A)、The group consisting of rs11769455 (T), rs3954044 (G), rs2505859 (A), rs10923743 (T), rs9499598 (T), rs61776759 (T), rs1237302 (T), rs4716901 (G), rs61849335 (G), rs6823966 (A), rs9294661 (C), rs7140599 (C), rs685218 (A), rs4947605 (C), rs10866961 (G), rs11061561 (A), rs76839556 (T), rs10835830 (A), rs10792719 (C), rs10822062 (A), rs1031026 (G), rs11705906 (A), rs6434910 (G), rs6913029 (C), rs6556892 (A), rs10167446 (G), rs3790272 (T), rs536618 (C), rs695872 (A), rs10091796 (A), rs9893926 (T), rs2337247 (G), rs6726814 (C), rs1884784 (G), rs307575 (G), rs6075075 (T), rs4698413 (T), rs6534672 (T), rs113114417 (A), rs608873 (C), rs7332756 (C), rs1290900 (A), rs2060179 (T), rs2280973 (A), rs881647 (A), rs11779483 (A), and rs4104180 (A). A biomarker for Jomon ancestry testing, comprising at least one SNP selected from the group consisting of the following.

2. The biomarker for Jomon ancestry testing according to claim 1, wherein the SNP comprises at least one selected from group (x).

3. (1) In the genomic DNA of a biological sample taken from a subject, (x) The group consisting of rs28729170(A), rs13017060(C), rs2645158(G), rs6446239(C), rs1026980(T), rs2057165(A), rs4302225(T), rs4981864(T), rs629577(G), and rs6097031(C), and (y)rs4298845(C)、rs7595986(A)、rs10764532(T)、rs111468015(A)、rs1172495(A)、rs11076581(C)、rs806469(A)、rs77125278(C)、rs13257488(C)、rs13253574(C)、rs6456790(T)、rs11030161(G)、rs3812526(G)、rs7832131(G)、rs388394(T)、rs35692656(T)、rs10497147(A)、rs13134059(A)、rs934213(C)、rs2486729(T)、rs12699444(T)、rs2722810(C)、rs12325527(T)、rs404725(A)、rs2570415(T)、rs9537153(A)、rs2172180(G)、rs7202291(C)、rs1478870(T)、rs34349793(T)、rs912758(G)、rs11960706(G)、rs6687335(G)、rs10759526(G)、rs12615487(A)、rs2248501(G)、rs12277503(C)、rs1716646(T)、rs712591(A)、rs7795435(C)、rs60799144(T)、rs2997881(T)、rs930592(G)、rs9496721(T)、rs957303(T)、rs437239(C)、rs1858730(G)、rs1358968(A)、rs2871660(G)、rs488165(A)、rs12675727(A)、rs359546(G)、rs11852766(G)、rs28495134(G)、rs3800577(T)、rs4902386(C)、rs6667615(G)、rs7980063(A)、rs6586660(A)、rs7030507(G)、rs78890485(T)、rs644713(C)、rs8030121(G)、rs7200600(G)、rs7069081(C)、rs10067255(G)、rs6703740(T)、rs1168551(A)、rs12498572(T)、rs647548(A)、rs2469954(C)、rs6105551(T)、rs11083432(G)、rs2100791(T)、rs12416497(A)、The group consisting of rs11769455 (T), rs3954044 (G), rs2505859 (A), rs10923743 (T), rs9499598 (T), rs61776759 (T), rs1237302 (T), rs4716901 (G), rs61849335 (G), rs6823966 (A), rs9294661 (C), rs7140599 (C), rs685218 (A), rs4947605 (C), rs10866961 (G), rs11061561 (A), rs76839556 (T), rs10835830 (A), rs10792719 (C), rs10822062 (A), rs1031026 (G), rs11705906 (A), rs6434910 (G), rs6913029 (C), rs6556892 (A), rs10167446 (G), rs3790272 (T), rs536618 (C), rs695872 (A), rs10091796 (A), rs9893926 (T), rs2337247 (G), rs6726814 (C), rs1884784 (G), rs307575 (G), rs6075075 (T), rs4698413 (T), rs6534672 (T), rs113114417 (A), rs608873 (C), rs7332756 (C), rs1290900 (A), rs2060179 (T), rs2280973 (A), rs881647 (A), rs11779483 (A), and rs4104180 (A). A method for examining Jomon ancestors, comprising the step of checking for the presence or absence of at least one SNP selected from a group consisting of the following.

4. The inspection method according to claim 3, which is a method for examining the presence or absence of Jomon ancestors and / or the proportion of Jomon ancestors.

5. moreover, (2) If at least one of the SNPs is detected in step (1), a step of determining whether the subject has Jomon ancestors or Japanese ancestors based on the detected SNP, or estimating the proportion of the subject to be Jomon ancestors. The inspection method according to claim 4, including the method described in claim 4.

6. The testing method according to claim 5, wherein in step (2b), the proportion of the subject to Jomon ancestors is estimated based on the detected SNPs and the prediction coefficients for each SNP.

7. (x) The group consisting of rs28729170(A), rs13017060(C), rs2645158(G), rs6446239(C), rs1026980(T), rs2057165(A), rs4302225(T), rs4981864(T), rs629577(G), and rs6097031(C), and (y)rs4298845(C)、rs7595986(A)、rs10764532(T)、rs111468015(A)、rs1172495(A)、rs11076581(C)、rs806469(A)、rs77125278(C)、rs13257488(C)、rs13253574(C)、rs6456790(T)、rs11030161(G)、rs3812526(G)、rs7832131(G)、rs388394(T)、rs35692656(T)、rs10497147(A)、rs13134059(A)、rs934213(C)、rs2486729(T)、rs12699444(T)、rs2722810(C)、rs12325527(T)、rs404725(A)、rs2570415(T)、rs9537153(A)、rs2172180(G)、rs7202291(C)、rs1478870(T)、rs34349793(T)、rs912758(G)、rs11960706(G)、rs6687335(G)、rs10759526(G)、rs12615487(A)、rs2248501(G)、rs12277503(C)、rs1716646(T)、rs712591(A)、rs7795435(C)、rs60799144(T)、rs2997881(T)、rs930592(G)、rs9496721(T)、rs957303(T)、rs437239(C)、rs1858730(G)、rs1358968(A)、rs2871660(G)、rs488165(A)、rs12675727(A)、rs359546(G)、rs11852766(G)、rs28495134(G)、rs3800577(T)、rs4902386(C)、rs6667615(G)、rs7980063(A)、rs6586660(A)、rs7030507(G)、rs78890485(T)、rs644713(C)、rs8030121(G)、rs7200600(G)、rs7069081(C)、rs10067255(G)、rs6703740(T)、rs1168551(A)、rs12498572(T)、rs647548(A)、rs2469954(C)、rs6105551(T)、rs11083432(G)、rs2100791(T)、rs12416497(A)、The group consisting of rs11769455 (T), rs3954044 (G), rs2505859 (A), rs10923743 (T), rs9499598 (T), rs61776759 (T), rs1237302 (T), rs4716901 (G), rs61849335 (G), rs6823966 (A), rs9294661 (C), rs7140599 (C), rs685218 (A), rs4947605 (C), rs10866961 (G), rs11061561 (A), rs76839556 (T), rs10835830 (A), rs10792719 (C), rs10822062 (A), rs1031026 (G), rs11705906 (A), rs6434910 (G), rs6913029 (C), rs6556892 (A), rs10167446 (G), rs3790272 (T), rs536618 (C), rs695872 (A), rs10091796 (A), rs9893926 (T), rs2337247 (G), rs6726814 (C), rs1884784 (G), rs307575 (G), rs6075075 (T), rs4698413 (T), rs6534672 (T), rs113114417 (A), rs608873 (C), rs7332756 (C), rs1290900 (A), rs2060179 (T), rs2280973 (A), rs881647 (A), rs11779483 (A), and rs4104180 (A). A diagnostic reagent for use in the testing method according to any one of claims 3 to 6, comprising a detection agent for at least one SNP selected from the group consisting of the following.