Composite amplification detection system containing 73 polymorphic dip loci and use thereof
By constructing a composite amplification detection system containing 73 polymorphic DIP sites and combining machine learning algorithms with six-color fluorescent labeling technology, the problem of insufficient biogeographical tracing efficiency of internal populations in East Asia has been solved, and efficient biogeographical ancestral tracing has been achieved. It is suitable for samples of various tissue types and has a high accuracy rate.
Patent Information
- Application Number
- PCT/CN2024/121972
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2024-09-27
- Publication Date
- 2025-10-16
AI Technical Summary
In existing technologies, the biogeographical traceability of populations within East Asia is insufficient, and the existing composite amplification system has low efficiency in distinguishing populations within East Asia, making it difficult to provide more biogeographical origin information. Moreover, as the number of loci increases, the balance control of amplification conditions becomes more difficult, and machine learning methods have not been fully developed in genetic marker selection and evidence analysis.
A composite amplification detection system containing 73 polymorphic DIP sites was constructed. High-resolution DIP sites were screened using the data from the International 1000 Genomes Project. Combining machine learning algorithms and six-color fluorescent labeling technology, primer combinations were designed for same-tube composite amplification. Detection was performed using a capillary electrophoresis platform, which is suitable for samples from various tissue types.
It has achieved efficient biogeographic ancestral tracing of major intercontinental groups and groups within East Asia, can subdivide the population within East Asia, and improves the applicability and feasibility in forensic practice, with an accuracy rate of over 98%.
Smart Images

Figure CN2024121972_16102025_PF_FP_ABST
Abstract
Description
Compound amplification detection system comprising 73 polymorphic DIP loci and application thereof TECHNICAL FIELD
[0001] The present application relates to the technical field of nucleic acid detection, in particular to a compound amplification detection system comprising 73 polymorphic DIP loci and application thereof. BACKGROUND
[0002] In forensic practice, in addition to conventional biological samples, various types of difficult biological samples such as old and degraded samples are also common. The detection and analysis of such samples is difficult, and if short tandem repeat sequences are used for conventional forensic DNA analysis, DNA typing failure may occur, which brings difficulties and challenges to the work of testing and identification. In addition, the number of short tandem repeat loci that can be accommodated in the same reaction system is limited, resulting in low cumulative identification efficiency of the system, and additional loci are needed to provide more information about the identity of unknown individuals in the case.
[0003] Deletion / Insertion Polymorphism (DIP) genetic markers are an alternative form of genetic variation, which is a DNA polymorphism formed by the insertion or deletion of DNA fragments. It is widely distributed in the human genome, has a low mutation rate (about 10-8), has racial and population genetic polymorphism, and has short amplification fragments, no artificial stutter peaks, and simple allele typing. It has advantages in the analysis of difficult biological samples such as old and degraded samples, and therefore has good forensic application prospects. It is also one of the commonly used ancestral information genetic markers. The detection of DIP genetic markers is suitable for second-generation sequencing platforms or capillary electrophoresis platforms. The main advantage of the former is high throughput, and different types of genetic markers can be detected, but the technical requirements for experimenters are higher, and the data processing process is relatively complex. In addition, factors such as sequencing errors and primer binding region mutations can easily cause misjudgment of results, so second-generation sequencing has not been routinely applied to forensic actual case testing. The traditional PCR-CE detection method has high accuracy, low cost, short time consumption, and simple operation, and is a reliable, economical and efficient detection platform suitable for forensic biological geographical ancestry tracing. Since the six-color fluorescence detection scheme was introduced in 2014, the number of sites that can be simultaneously separated by capillary has been greatly improved, and the corresponding system efficiency has also been improved, making it possible to trace the ancestry of multiple regions and multiple levels, which helps to promote the application of the DIP system for biological geographical ancestry tracing in grassroots forensic DNA laboratories.
[0004] In the past five years, a series of multiplex compound amplification systems based on DIP markers with ancestral information have been developed and verified, such as a 30-DIP-loci-based system, a 45-DIP-loci-based system, a 60-DIP-loci-based system, a 70-DIP-loci-based system, and a 73-DIP-loci-based system. DIPplex system, AGCU InDel 60 kit system containing 60 DIP sites, and 39-fold di-allelic DIP and 41-fold multi-InDel composite amplification system independently developed by Zhu et al. The above systems basically achieve accurate differentiation of the three major ancestral components (Africa, Europe and East Asia), but the differentiation efficiency between European, South Asian and South American individuals still needs to be improved. In order to further improve the biogeographic tracing efficiency of the AI-DIP composite amplification system in Asian populations, Sun et al. used 15 multi-InDel sites to preliminarily achieve the three classification of East Asian population, Southeast Asian population and Northeast Asian Russian Adyghe. Zhang et al. used 21 high-efficiency AI-DIP sites to achieve the three classification of 10 Asian populations, and the cross-validation results showed that the average correct rate was 81.70%. Although the above studies preliminarily explored the tracing strategy for the populations in the pan-Asian region, the number of selected sites and the biogeographic tracing efficiency still have deficiencies, and therefore cannot provide more beneficial biogeographic information in the forensic practice of East Asian internal populations. Although East Asian populations have migrated, exchanged and formed a state of multi-ethnic, multi-cultural and multi-lingual in the long history, the molecular biological results of previous phylogenetic and population genetic studies have shown that there is still a high genetic structure homogeneity between the internal populations of East Asia. Therefore, constructing a biogeographic ancestral tracing system of East Asian internal populations containing more high-efficiency AI-DIP sites is one of the difficulties of ancestry inference, which puts higher requirements on the development technology of composite amplification system, the selection method of genetic markers and the evidence analysis ability. The development of six-color fluorescence labeling technology enables the composite amplification system to accommodate more number or type of loci, so that the system can provide more genetic information and higher forensic identification efficiency. However, with the increase of the number of loci in the composite amplification system, the relative balance control of each locus becomes more difficult due to competition, and the adaptability of amplification conditions becomes higher. Therefore, it is necessary to repeatedly verify the amplification parameters, continuously adjust the primer concentration and ratio, and improve the balance of site amplification. In addition, machine learning algorithms have unique advantages in genetic marker selection and evidence strength analysis, and their applicability to high-dimensional and sparse data helps to mine sites with potential for ancestral information inference from massive whole-genome data, but the machine learning methodology suitable for ancestral information DIP system and its biogeographic ancestral tracing efficiency have not been developed and verified.
[0005] SUMMARY
[0006] The present application aims to disclose a composite amplification detection system containing 73 polymorphic DIP sites and its application, to solve one or more technical problems existing in the prior art, and to provide at least one beneficial option or create conditions.
[0007] The first aspect of the present application provides a composite amplification detection system.
[0008] The second aspect of the present application provides a kit containing the composite amplification detection system according to the first aspect of the present application.
[0009] The third aspect of the present application provides the use of the composite amplification detection system according to the first aspect of the present application or the kit according to the second aspect of the present application in biogeographic ancestry tracing.
[0010] The composite amplification detection system in the first aspect of the application comprises a primer combination and a composite amplification premix, wherein the primer combination targets 73 DIP sites, respectively, and the DIP sites are as follows: rs73611618, rs28741387, rs141511864, rs71377077, rs10660476, rs879841278, rs55681325, rs140698686, rs200216987, rs71879919, rs10531408, rs59369367, rs56120126, rs71097946, rs77514652, rs561904853, rs5780349, rs2067285, rs5789056, rs5789729, rs141160384, rs139988800, rs141928144, rs56968651, rs59127488, rs1347535145, rs551883542, rs3994057, rs1342356747, rs3044086, rs35450593, rs72104851, rs59005026, rs71712626, rs138600078, rs879662430, rs10564190, rs140202531, rs10573591, rs10630253, rs77624782, rs141471313, rs10600917, rs71408252, rs1404627509, rs74816196, rs112473811, rs57051438, rs766586871, rs10628367, rs58227077, rs143267128, rs140671911, rs138465422, rs200935491, rs141613931, rs66462883, rs71110898, rs35991174, rs35880452, rs59377169, rs778835021, rs79710335, rs59605350, rs112524265, rs5886296, rs59218555, rs67579111, rs71004215, rs56358449, rs140200174, rs56783915, and rs141047228.The complex amplification detection system is based on the international 1000 Genome Project phase III (1KGP), and the whole genome data of 2598 individuals in the extended data set is evaluated according to a series of site evaluation criteria. 27,000 second-allele DIP sites with biogeographical ancestry tracing potential for major continental populations and East Asian populations are selected, and then 157 candidate DIP sites are selected by combining different dimension reduction visualization methods and machine learning feature selection algorithms. Through site primer design, feature importance performance evaluation, 73 high-resolution and better detection effect autosomal DIP sites are further selected as the candidate development site set of the complex system. And according to the 73 DIP sites obtained, a same-tube complex amplification system is designed, which can detect 73 DIP sites at the same time through one reaction. The specific technical route is shown in Figure 1.
[0011] In order to realize the same-tube complex amplification, the application further provides a corresponding primer combination in some application embodiments of the first aspect of the application, which contains 73 pairs of primer pairs respectively targeting the above-mentioned 73 DIP sites. The nucleotide sequences of the specific primers are shown in Table 1.
[0012] In some application embodiments of the first aspect of the application, in order to improve the accuracy of the complex amplification detection system, the final concentration of each pair of primers in the primer combination can be further limited, as shown in Table 1.
[0013] In some application embodiments of the first aspect of the application, at least one primer in each primer pair is labeled with a fluorescent dye at the 5' end, and the fluorescent dye is selected from FAM, HEX, SUM, LYN, PUR, A514, TAMRA, ROX, VIC, A555, PET, NED, TAZ, A488, SF488 or A568. By using different fluorescent dyes, the number of similar-length amplification products is converted into fluorescent signal output, thereby realizing the same-tube complex amplification detection of multiple DIP sites.
[0014] In some application embodiments of the first aspect of the application, the site arrangement is designed according to the length range of the amplification product, and the primer pairs with similar or small differences in amplification product length are grouped into different groups. One fluorescent dye is selected for each group, and the fluorescent dyes selected for each group are different from each other. Through the allocation of different groups, the types of fluorescent dyes used can be greatly reduced, and the difficulty of detection and analysis is reduced.
[0015] In some embodiments of the first aspect of the present application, the fluorescent dye selected from the first group is FAM, the fluorescent dye selected from the second group is HEX, the fluorescent dye selected from the third group is SUM, the fluorescent dye selected from the fourth group is LYN, and the fluorescent dye selected from the fifth group is PUR, as shown in Table 1.
[0016] Table 1. Primer sequences and final concentrations of primers for 73 DIP loci
[0017] In some embodiments of the first aspect of the present application, the components of the master mix include dNTPs, Taq enzyme, Tris-HCl buffer, KCl, MgCl2, and bovine serum albumin (BSA).
[0018] The human biological sample can be genomic DNA extracted from human body fluids / tissues (e.g., bloodstains, blood, saliva, buccal swabs, hair follicle-containing roots, semen, muscle tissue, exfoliated cells, etc.), and can be extracted and quantified by Chelex-100 method, phenol-chloroform method, magnetic bead method, etc. Alternatively, the sample can be directly amplified without extraction, such as blood filter paper, blood gauze, FTA card, saliva card, exfoliated cells, etc.
[0019] The kit of the second aspect of the present application comprises the composite amplification detection system of the first aspect of the present application.
[0020] In some embodiments of the second aspect of the present application, molecular cloning techniques are used to prepare an allelic ladder for analysis, i.e., a positive quality control. The positive quality control is composed of a mixture of products amplified from each locus corresponding to the allelic fragments.
[0021] In some embodiments of the second aspect of the present application, nuclease-free water is used as a negative quality control.
[0022] In some application embodiments of the second aspect of the present application, the amplification product is subjected to fluorescence detection on a capillary electrophoresis genetic analyzer. The PCR amplification product (1 μL), deionized formamide (9.5 μL), and SIZE-500 molecular weight internal standard (0.5 μL) are mixed, denatured at 95°C for 3 minutes, incubated on ice for 3 minutes, and then subjected to capillary electrophoresis detection on a genetic analyzer (including but not limited to 3100 series, 3130 series, 3500 series genetic analyzer). Subsequently, the typing results of 73 DIP sites are interpreted.
[0023] Based on the combination of more DIP genetic markers, six-color fluorescence labeling technology and machine learning algorithm, it is of great practical significance to successfully develop a biological geographical ancestry tracing system. The present application aims to deeply mine and systematically select DIP molecular genetic markers strongly related to different intercontinental and geographical regions in the whole genome range by combining machine learning algorithm, and to construct a DIP composite detection system containing more high-performance sites using capillary electrophoresis platform, thereby providing a new detection scheme and evidence analysis process for the accurate identification of biological geographical ancestry of intercontinental and East Asian populations.
[0024] The present application has the following advantages:
[0025] The present application provides a high-ancestral information DIP site composite amplification detection system suitable for major intercontinental populations and East Asian internal populations, which can simultaneously detect 73 DIP sites in one reaction and is suitable for various types of test materials. The composite amplification detection system and the corresponding kit can simultaneously achieve the accurate tracing of the biological geographical ancestry of the major five intercontinental populations including African, European, East Asian, South Asian and South American populations except for mixed American populations, and the East Asian internal population, and further subdivide the East Asian internal population into Han, Southeast Asian and Japanese populations, thereby making up for the deficiency of the previous system in the large range of biological geographical ancestry tracing, effectively improving the ancestral resolution of the East Asian internal population, and improving the applicability and feasibility of the system in forensic practice in China. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 is a technical roadmap for developing the 73-site composite amplification system described in Example 1.
[0027] Figure 2 is an arrangement diagram of the 73 DIP site amplification products of the composite amplification detection system described in Example 1.
[0028] Figure 3 is an Allelic Ladder electropherogram as described in Example 2.
[0029] Figure 4 is an electropherogram of a human DNA test sample as described in Example 3.
[0030] Figure 5 is a t-SNE dimensionality reduction visualization result of five intercontinental reference populations in Example 4.
[0031] Figure 6 is the phylogenetic reconstruction results of the five continental reference populations in Example 4;
[0032] Figure 7 is the ancestral composition analysis results of the five continental reference populations in Example 4;
[0033] Figure 8 is the ancestral composition analysis results of the East Asian reference populations in Example 4;
[0034] Figure 9 is the ten-fold cross-validation results of the biogeographical ancestry inference model described in Example 4, Figure 9A is the normalized ten-fold cross-validation confusion matrix of a series of biogeographical ancestry inference models constructed based on the five continental reference populations, and Figure 9B is the normalized ten-fold cross-validation confusion matrix of the biogeographical ancestry inference model constructed based on the East Asian reference populations. DETAILED DESCRIPTION
[0035] The following examples further illustrate the present application, but should not be construed as limiting the application. Modifications and adaptations of the methods, steps or conditions described may become apparent to those skilled in the art and can be made without departing from the spirit and scope of the application.
[0036] Unless otherwise specified, the technical means used in the examples are conventional means known to those skilled in the art.
[0037] The molecular biology test methods not specifically described in the following examples were performed according to the Molecular Cloning Laboratory Guide (3rd edition) or according to the reagent kit and product instructions; the biological materials of the reagent kit, if not specifically described, can be obtained from commercial channels.
[0038] Example 1: Screening of DIP sites and primer design in the multiplex amplification system
[0039] (1) Screening of 73 polymorphic DIP sites
[0040] Based on the whole genome data of 2598 individuals in the 1KGP expansion dataset, a systematic and comprehensive bio-geographical ancestral origin marker screening was carried out. According to the specific technical route shown in Figure 1, based on the above screening criteria, the final candidate site set used for system construction was determined by using data preprocessing, dimensionality reduction visualization (t-SNE), feature importance evaluation based on tree model and cross-validation method, in turn, rs73611618, rs28741387, rs141511864, rs71377077, rs10660476, rs879841278, rs55681325, rs140698686, rs200216987, rs71879919, rs10531408, rs59369367, rs56120126, rs71097946, rs77514652, rs561904853, rs5780349, rs2067285, rs5789056, rs5789729, rs141160384, rs139988800, rs141928144, rs56968651, rs59127488, rs1347535145, rs551883542, rs3994057, rs1342356747, rs3044086, rs35450593, rs72104851, rs59005026, rs71712626, rs138600078, rs879662430, rs10564190, rs140202531, rs10573591, rs10630253, rs77624782, rs141471313, rs10600917, rs71408252, rs1404627509, rs74816196, rs112473811, rs57051438, rs766586871, rs10628367, rs58227077, rs143267128, rs140671911, rs138465422, rs200935491, rs141613931, rs66462883, rs71110898, rs35991174, rs35880452, rs59377169, rs778835021, rs79710335, rs59605350, rs112524265, rs5886296, rs59218555, rs67579111, rs71004215, rs56358449, rs140200174, rs56783915 and rs141047228.
[0041] The primer combination is designed for the selected DIP site. Due to the large difference in the amplification product fragment size and the amplification efficiency of each polymorphic site, the primer pairs for all sites need to be repeatedly verified and screened by using a complex amplification system for detection and analysis of a variety of samples, and the primer concentration of each site is optimized and adjusted through a series of experiments to gradually improve the amplification balance of the primer combination, and finally the optimal concentration of each primer pair in the complex detection system is obtained. The final primer combination is shown in Table 1, and the arrangement of the DIP site amplification product is shown in Figure 2.
[0042] Example 2: Construction of a complex amplification detection system
[0043] On the basis of successfully establishing a single site amplification system, new site primers are gradually added to the amplification system for testing, and various parameters in the complex amplification system, including cycle parameters, annealing temperature, final extension time, enzyme amount, complex amplification reaction system volume, and template DNA amount, are determined through repeated experiments to achieve stable and balanced typing detection results of the complex system. The final optimal amplification reaction volume of this system is 10 μL, and the specific complex amplification detection system is shown in Table 2.
[0044] Table 2, reaction components and sample volume of the complex amplification detection system
[0045] The Master Mix contains dNTPs, Taq enzyme, Tris-HCl buffer, KCl, MgCl2, and BSA; the human biological sample to be tested can be genomic DNA extracted from human body fluids / tissues (such as blood marks, blood, saliva, oral swabs, hair root with hair follicles, semen, muscle tissue, exfoliated cells, etc.), which can be extracted and quantified by Chelex-100 method, phenol chloroform method, magnetic bead method, etc.; or direct amplification without extraction can be performed on various test materials (such as blood filter paper, blood gauze, FTA card, saliva card, exfoliated cells, etc.). The specific complex amplification program is shown in Table 3.
[0046] Table 3, complex amplification program
[0047] The amplification product is detected by fluorescence on a capillary electrophoresis genetic analyzer. The PCR amplification product (1 μL), deionized formamide (9.5 μL), and SIZE-500 molecular weight internal standard (0.5 μL) are mixed, denatured at 95°C for 3 minutes, incubated on ice for 3 minutes, and then subjected to capillary electrophoresis detection on a genetic analyzer (including but not limited to 3100 series, 3130 series, 3500 series genetic analyzer). After capillary electrophoresis, the ID-X software analyzes and processes data. First, the corresponding Bin file and Panel file are written according to the software format requirements, the insertion (Insertion) allele in the DIP marker is named I, and the deletion (Deletion) allele is named D, the electrophoresis analysis method of the system is created, then the capillary electrophoresis data is imported, the corresponding Panel, Bin, Analysis Method and Size Standard analysis parameters are selected, and the electrophoresis results of 73 composite amplification DIP sites are interpreted.
[0048] The application also uses molecular cloning technology to prepare an allelic typing standard (Allelic Ladder) for analysis, which is composed of a mixture of products corresponding to the allele fragments amplified by primers at each site, and Figure 3 is an electropherogram of the Allelic Ladder of the system.
[0049] Example 3: Practical application of the composite amplification detection system
[0050] According to the composite amplification detection system provided in Example 2, it is applied to the detection of actual human DNA samples. 1 ng of human blood card DNA sample extracted by Chelex-100 method is used for detection, and the specific detection steps are shown in Example 2, and the typing spectrum of the detection result is shown in Figure 4. The peaks of each site of the composite amplification detection system are normal, and the peak height balance between sites is good, indicating that the above system can realize effective detection of DNA samples.
[0051] Example 4: Evaluation of the biogeographical ancestry tracing efficiency of 73 DIP sites
[0052] The ancestry information inference efficiency of the selected DIP genetic markers is evaluated using the five intercontinental population data in the 1KGP extended data set, and the biogeographical ancestry tracing efficiency of the 73 DIP genetic markers is evaluated based on nonlinear dimension reduction algorithm (t-SNE), phylogenetic reconstruction, Structure and other population genetics methods.
[0053] The t-SNE dimension reduction visualization result is shown in Figure 5, and the 73 DIP sites selected in the system can basically distinguish the five intercontinental populations, and the preliminary clustering of Japanese, Han and Southeast Asian ethnic groups can be seen within the East Asian population. The phylogenetic reconstruction result based on the maximum likelihood ratio method is shown in Figure 6, which also clearly shows the evolutionary branches of the five intercontinental populations, and three small branches can be seen within the East Asian population. The ancestry composition analysis based on the Structure method shows that the system can display the genetic structure differences between different intercontinental populations at the level of five intercontinental populations, and the best K value is 4 (see Figure 7), and the genetic structure differences within the East Asian population can also be identified, and the best K value is 2 (see Figure 8).
[0054] Based on the 73-loci system and the 1KGP extended dataset, a series of biogeographic ancestry models were constructed using algorithms such as polynomial naive Bayes, support vector machine, random forest, and extreme gradient boosting (XGBoost). The ten-fold cross-validation confusion matrix results of the above models showed that the 73-DIP system described in the present application had an average classification accuracy of more than 98% in the five intercontinental populations [see Fig. 9(A)], and an average classification accuracy of more than 90% in the East Asian internal population [see Fig. 9B]. Table 4 shows the ten-fold cross-validation results of each machine learning model in the test set under different biogeographic ancestry classification tasks. As shown in Table 4, the 73-DIP loci have high classification performance and generalization ability in the five intercontinental (Africa, Europe, East Asia, South Asia, and America) and East Asian internal (Han, Japanese, and Southeast Asian populations) populations.
[0055] The above population genetics analysis results show that the 73-DIP loci in the present system have good resolution ability for individuals from five intercontinental populations, and can further subdivide the ancestral origins of East Asian internal populations.
[0056] Table 4, ten-fold cross-validation results of the 73-DIP loci in the test set for biogeographic ancestry tracing
[0057] Example 5: Application verification of the composite system described in the present application
[0058] The composite system described in Example 1 was used to genotype two different individuals with known biogeographic ancestry origins according to the method and procedure of Example 2, using standard sample 9948 DNA as a positive control. The biogeographic ancestry origins of these individuals were inferred using the naive Bayes method, and the biogeographic ancestry tracing performance of the system for real samples was evaluated.
[0059] The genotyping results of the two different individuals with known origins are shown in Table 5.
[0060] Table 5, genotyping results of two different individuals with known origins
[0061] The population matching probability and likelihood ratio results of the two samples at the 73-DIP loci are shown in Table 6.
[0062] Table 6, population matching probability and likelihood ratio results of two samples at the 73-DIP loci
[0063] The above results show that both of the two real samples can be completely detected for genotyping in all loci, and are correctly inferred as real biogeographical ancestral sources, indicating that the composite system has good detection ability and biogeographical ancestral tracing ability for real samples, and the system and method can be applied to effectively infer individual biogeographical ancestral sources.
[0064] It will be obvious to a person skilled in the art that, without departing from the scope of the present application, the application can be implemented in other particular forms apparent to the person skilled in the art. Therefore, the examples should be considered as exemplary and non-limiting, the scope of the application being defined by the claims appended hereto rather than the description set forth above, and all the changes falling within the meaning and the scope of the equivalent elements of the claims are intended to be embraced in the present application.
Claims
1. A composite amplification detection system, characterized in that: The invention comprises a primer combination and a composite amplification premix, wherein the primer combination targets 73 DIP sites, respectively, and the DIP sites are: rs73611618, rs28741387, rs141511864, rs71377077, rs10660476, rs879841278, rs55681325, rs140698686, rs200216987, rs71879919, rs10531408, rs59369367, rs56120126, rs71097946, rs77514652, rs5 61904853, rs5780349, rs2067285, rs5789056, rs5789729, rs141160384, rs139988800, rs141928144, rs56968651, rs59127488, rs 1347535145, rs551883542, rs3994057, rs1342356747, rs3044086, rs35450593, rs72104851, rs59005026, rs71712626, rs1386000 78. rs879662430, rs10564190, rs140202531, rs10573591, rs10630253, rs77624782, rs141471313, rs10600917, rs71408252, rs14 04627509, rs74816196, rs112473811, rs57051438, rs766586871, rs10628367, rs58227077, rs143267128, rs140671911, rs138465 422, rs200935491, rs141613931, rs66462883, rs71110898, rs35991174, rs35880452, rs59377169, rs778835021, rs79710335, rs59605350, rs112524265, rs5886296, rs59218555, rs67579111, rs71004215, rs56358449, rs140200174, rs56783915, and rs141047228.
2. The composite amplification detection system according to claim 1, characterized in that: The primer combination includes: a primer pair targeting rs73611618 as shown in SEQ ID No: 1 and SEQ ID No: 2; a primer pair targeting rs28741387 as shown in SEQ ID No: 3 and SEQ ID No: 4; a primer pair targeting rs141511864 as shown in SEQ ID No: 5 and SEQ ID No: 6; a primer pair targeting rs71377077 as shown in SEQ ID No: 7 and SEQ ID No: 8; a primer pair targeting rs10660476 as shown in SEQ ID No: 9 and SEQ ID No: 10; a primer pair targeting rs879841278 as shown in SEQ ID No: 11 and SEQ ID No: 12; a primer pair targeting rs55681325 as shown in SEQ ID No: 13 and SEQ ID No: 14; a primer pair targeting rs140698686 as shown in SEQ ID No: 15 and SEQ ID No: 16; the primer pair targeting rs200216987 is shown in SEQ ID No: 17 and SEQ ID No: 18; the primer pair targeting rs71879919 is shown in SEQ ID No: 19 and SEQ ID No: 20; the primer pair targeting rs10531408 is shown in SEQ ID No: 21 and SEQ ID No: 22; the primer pair targeting rs59369367 is shown in SEQ ID No: 23 and SEQ ID No: 24; the primer pair targeting rs56120126 is shown in SEQ ID No: 25 and SEQ ID No: 26; the primer pair targeting rs71097946 is shown in SEQ ID No: 27 and SEQ ID No: 28; the primer pair targeting rs77514652 is shown in SEQ ID No: 29 and SEQ ID No: 30; the primer pair targeting rs561904853 is shown in SEQ ID No: 31 and SEQ ID No: 32; the primer pair targeting rs5780349 is shown in SEQ ID No: 33 and SEQ ID No: 34; the primer pair targeting rs2067285 is shown in SEQ ID No: 35 and SEQ ID No: 36; the primer pair targeting rs5789056 is shown in SEQ ID No: 37 and SEQ ID No: 38; the primer pair targeting rs5789729 is shown in SEQ ID No: 39 and SEQ ID No: 40; the primer pair targeting rs141160384 is shown in SEQ ID No: 41 and SEQ ID No: 42; the primer pair targeting rs139988800 is shown in SEQ ID No: 43 and SEQ ID No: 44;The primer pair targeting rs141928144 is shown in SEQ ID No: 45 and SEQ ID No: 46; the primer pair targeting rs56968651 is shown in SEQ ID No: 47 and SEQ ID No: 48; the primer pair targeting rs59127488 is shown in SEQ ID No: 49 and SEQ ID No: 50; the primer pair targeting rs1347535145 is shown in SEQ ID No: 51 and SEQ ID No: 52; the primer pair targeting rs551883542 is shown in SEQ ID No: 53 and SEQ ID No: 54; the primer pair targeting rs3994057 is shown in SEQ ID No: 55 and SEQ ID No: 56; the primer pair targeting rs1342356747 is shown in SEQ ID No: 57 and SEQ ID No: 58; the primer pair targeting rs3044086 is shown in SEQ ID No: 59 and SEQ ID No: 60; the primer pair targeting rs35450593 is shown in SEQ ID No: 61 and SEQ ID No: 62; the primer pair targeting rs72104851 is shown in SEQ ID No: 63 and SEQ ID No: 64; the primer pair targeting rs59005026 is shown in SEQ ID No: 65 and SEQ ID No: 66; the primer pair targeting rs71712626 is shown in SEQ ID No: 67 and SEQ ID No: 68; the primer pair targeting rs138600078 is shown in SEQ ID No: 69 and SEQ ID No: 70; the primer pair targeting rs879662430 is shown in SEQ ID No: 71 and SEQ ID No: 72; the primer pair targeting rs10564190 is shown in SEQ ID No: 73 and SEQ ID No: 74; the primer pair targeting rs140202531 is shown in SEQ ID No: 75 and SEQ ID No: 76; the primer pair targeting rs10573591 is shown in SEQ ID No: 77 and SEQ ID No: 78; the primer pair targeting rs10630253 is shown in SEQ ID No: 79 and SEQ ID No: 80; the primer pair targeting rs77624782 is shown in SEQ ID No: 81 and SEQ ID No: 82; the primer pair targeting rs141471313 is shown in SEQ ID No: 83 and SEQ ID No: 84; the primer pair targeting rs10600917 is shown in SEQ ID No: 85 and SEQ ID No: 86; the primer pair targeting rs71408252 is shown in SEQ ID No: 87 and SEQ ID No: 88;The primer pair targeting rs1404627509 is shown in SEQ ID No: 89 and SEQ ID No: 90; the primer pair targeting rs74816196 is shown in SEQ ID No: 91 and SEQ ID No: 92; the primer pair targeting rs112473811 is shown in SEQ ID No: 93 and SEQ ID No: 94; the primer pair targeting rs57051438 is shown in SEQ ID No: 95 and SEQ ID No: 96; the primer pair targeting rs766586871 is shown in SEQ ID No: 97 and SEQ ID No: 98; the primer pair targeting rs10628367 is shown in SEQ ID No: 99 and SEQ ID No: 100; the primer pair targeting rs58227077 is shown in SEQ ID No: 101 and SEQ ID No: 102; the primer pair targeting rs143267128 is shown in SEQ ID No: 103 and SEQ ID No: 104; the primer pair targeting rs140671911 is shown in SEQ ID No: 105 and SEQ ID No: 106; the primer pair targeting rs138465422 is shown in SEQ ID No: 107 and SEQ ID No: 108; the primer pair targeting rs200935491 is shown in SEQ ID No: 109 and SEQ ID No: 110; the primer pair targeting rs141613931 is shown in SEQ ID No: 111 and SEQ ID No: 112; the primer pair targeting rs66462883 is shown in SEQ ID No: 113 and SEQ ID No: 114; the primer pair targeting rs71110898 is shown in SEQ ID No: 115 and SEQ ID No: 116; the primer pair targeting rs35991174 is shown in SEQ ID No: 117 and SEQ ID No:
118. No: 118; the primer pair targeting rs35880452 is shown in SEQ ID No: 119 and SEQ ID No: 120; the primer pair targeting rs59377169 is shown in SEQ ID No: 121 and SEQ ID No: 122; the primer pair targeting rs778835021 is shown in SEQ ID No: 123 and SEQ ID No: 124; the primer pair targeting rs79710335 is shown in SEQ ID No: 125 and SEQ ID No: 126; the primer pair targeting rs59605350 is shown in SEQ ID No: 127 and SEQ ID No: 128; the primer pair targeting rs112524265 is shown in SEQ ID No: 129 and SEQ ID No: 130;The primer pair targeting rs5886296 is shown in SEQ ID No: 131 and SEQ ID No: 132; the primer pair targeting rs59218555 is shown in SEQ ID No: 133 and SEQ ID No: 134; the primer pair targeting rs67579111 is shown in SEQ ID No: 135 and SEQ ID No: 136; the primer pair targeting rs71004215 is shown in SEQ ID No: 137 and SEQ ID No: 138; the primer pair targeting rs56358449 is shown in SEQ ID No: 139 and SEQ ID No: 140; the primer pair targeting rs140200174 is shown in SEQ ID No: 141 and SEQ ID No: 142; the primer pair targeting rs56783915 is shown in SEQ ID No: 143 and SEQ ID No: 144; the primer pair targeting rs141047228 is shown in SEQ ID No: 1 No: 145 and SEQ ID No:
146.
3. The composite amplification detection system according to claim 2, characterized in that: The final concentration of the primer pair targeting rs73611618 was 0.0476 μM; the final concentration of the primer pair targeting rs28741387 was 0.0311 μM; the final concentration of the primer pair targeting rs141511864 was 0.0385 μM; the final concentration of the primer pair targeting rs71377077 was 0.0476 μM; the final concentration of the primer pair targeting rs10660476 was 0.0458 μM; the final concentration of the primer pair targeting rs879841278 was 0.0641 μM; the final concentration of the primer pair targeting rs55681325 was 0.0348 μM; the final concentration of the primer pair targeting rs140698686 was 0.0660 μM; the final concentration of the primer pair targeting rs The final concentration of the primer pair targeting rs200216987 was 0.0513 μM; the final concentration of the primer pair targeting rs71879919 was 0.0861 μM; the final concentration of the primer pair targeting rs10531408 was 0.0531 μM; the final concentration of the primer pair targeting rs59369367 was 0.0586 μM; the final concentration of the primer pair targeting rs56120126 was 0.0531 μM; the final concentration of the primer pair targeting rs71097946 was 0.0257 μM; the final concentration of the primer pair targeting rs77514652 was 0.0403 μM; the final concentration of the primer pair targeting rs561904853 was 0.0660 μM; the final concentration of the primer pair targeting rs578034 The final concentration of the primer pair targeting rs141928144 was 0.0708 μM; the final concentration of the primer pair targeting rs56968651 was 0.0764 μM; the final concentration of the primer pair targeting rs59127488 was 0.0766 μM; the final concentration of the primer pair targeting rs2067285 was 0.0751 μM; the final concentration of the primer pair targeting rs5789056 was 0.0403 μM; the final concentration of the primer pair targeting rs5789729 was 0.0623 μM; the final concentration of the primer pair targeting rs141160384 was 0.0708 μM; the final concentration of the primer pair targeting rs139988800 was 0.0793 μM; the final concentration of the primer pair targeting rs141928144 was 0.0708 μM; the final concentration of the primer pair targeting rs56968651 was 0.0764 μM; the final concentration of the primer pair targeting rs59127488 was 0.0766 μM. The final concentration of the primer pair targeting rs1347535145 was 0.1132 μM; the final concentration of the primer pair targeting rs551883542 was 0.0594 μM; the final concentration of the primer pair targeting rs3994057 was 0.0906 μM; the final concentration of the primer pair targeting rs1342356747 was 0.0849 μM; the final concentration of the primer pair targeting rs3044086 was 0.0708 μM; the final concentration of the primer pair targeting rs35450593 was 0.0531 μM; the final concentration of the primer pair targeting rs72104851 was 0.0403 μM; the final concentration of the primer pair targeting rs59005026 was 0.The final concentration of the primer pair targeting rs71712626 was 0.0311 μM; the final concentration of the primer pair targeting rs138600078 was 0.0366 μM; the final concentration of the primer pair targeting rs879662430 was 0.1026 μM; the final concentration of the primer pair targeting rs10564190 was 0.0879 μM; the final concentration of the primer pair targeting rs140202531 was 0.0458 μM; the final concentration of the primer pair targeting rs10573591 was 0.0678 μM; the final concentration of the primer pair targeting rs10630253 was 0.0898 μM; the final concentration of the primer pair targeting rs77624782 was 0. The final concentration of the primer pair targeting rs141471313 was 0.1869 μM; the final concentration of the primer pair targeting rs10600917 was 0.0733 μM; the final concentration of the primer pair targeting rs71408252 was 0.0421 μM; the final concentration of the primer pair targeting rs1404627509 was 0.0293 μM; the final concentration of the primer pair targeting rs74816196 was 0.0421 μM; the final concentration of the primer pair targeting rs112473811 was 0.0825 μM; the final concentration of the primer pair targeting rs57051438 was 0.1154 μM; the final concentration of the primer pair targeting rs766586871 was 0 The final concentration of the primer pair targeting rs10628367 was 0.0679 μM; the final concentration of the primer pair targeting rs58227077 was 0.1274 μM; the final concentration of the primer pair targeting rs143267128 was 0.0764 μM; the final concentration of the primer pair targeting rs140671911 was 0.0849 μM; the final concentration of the primer pair targeting rs138465422 was 0.0340 μM; the final concentration of the primer pair targeting rs200935491 was 0.0679 μM; the final concentration of the primer pair targeting rs141613931 was 0.1076 μM; the final concentration of the primer pair targeting rs66462883 was The final concentration of the primer pair targeting rs71110898 was 0.1104 μM; the final concentration of the primer pair targeting rs35991174 was 0.0651 μM; the final concentration of the primer pair targeting rs35880452 was 0.0708 μM; the final concentration of the primer pair targeting rs59377169 was 0.0758 μM; the final concentration of the primer pair targeting rs778835021 was 0.0884 μM; the final concentration of the primer pair targeting rs79710335 was 0.0455 μM; the final concentration of the primer pair targeting rs59605350 was 0.0758 μM; the final concentration of the primer pair targeting rs112524265 was 0.The final concentration of the primer pair targeting rs5886296 was 0.0379 μM; the final concentration of the primer pair targeting rs5886296 was 0.0303 μM; the final concentration of the primer pair targeting rs59218555 was 0.0425 μM; the final concentration of the primer pair targeting rs67579111 was 0.0474 μM; the final concentration of the primer pair targeting rs71004215 was 0.0278 μM; the final concentration of the primer pair targeting rs56358449 was 0.0360 μM; the final concentration of the primer pair targeting rs140200174 was 0.0343 μM; the final concentration of the primer pair targeting rs56783915 was 0.0425 μM; and the final concentration of the primer pair targeting rs141047228 was 0.0409 μM.
4. The composite amplification detection system according to claim 2 or 3, characterized in that: The 5' end of at least one primer in each primer pair is labeled with a fluorescent dye, and the fluorescent dye is selected from FAM, HEX, SUM, LYN, PUR, A514, TAMRA, ROX, VIC, A555, PET, NED, TAZ, A488, SF488 or A568.
5. The composite amplification detection system according to claim 4, characterized in that: The primer combinations include primers having nucleotide sequences shown in SEQ ID No: 1 to SEQ ID No: 30 as a first group, primers having nucleotide sequences shown in SEQ ID No: 31 to SEQ ID No: 62 as a second group, primers having nucleotide sequences shown in SEQ ID No: 63 to SEQ ID No: 88 as a third group, primers having nucleotide sequences shown in SEQ ID No: 89 to SEQ ID No: 118 as a fourth group, and primers having nucleotide sequences shown in SEQ ID No: 119 to SEQ ID No: 146 as a fifth group. Each group uses a different fluorescent dye.
6. The composite amplification detection system according to claim 5, characterized in that: The fluorescent dye selected by the first group is FAM, the fluorescent dye selected by the second group is HEX, the fluorescent dye selected by the third group is SUM, the fluorescent dye selected by the fourth group is LYN, and the fluorescent dye selected by the fifth group is PUR.
7. The composite amplification detection system according to claim 6, characterized in that: The components of the composite amplification premix include: dNTPs, Taq enzyme, Tris-HCl buffer, KCl, MgCl2 and bovine serum albumin.
8. A kit, characterized in that It comprises the composite amplification detection system according to any one of claims 1 to 7.
9. The kit according to claim 8, characterized in that Also include 1 positive control and / or 1 negative control.
10. Use of the multiplex amplification detection system according to any one of claims 1 to 7 or the kit according to any one of claims 8 to 9 in forensic biogeographic ancestry inference and differentiation of African, European, East Asian, South Asian, American populations or Han Chinese, Southeast Asian, and Japanese populations.
Citation Information
Patent Citations
Method and system for carrying out African, East Asia and European group source analysis on individuals with unknown sources
CN112011622A
Composite amplification kit for simultaneously detecting 60 InDel genetic polymorphic sites and application thereof
CN114438173A
Composite amplification box for degrading biological geographic ancestors DIPs for sample inference and sex identification
CN116064842A
Composite amplification detection system containing 73 polymorphic DIP sites and application thereof
CN118028493A
Compositions and methods for inferring ancestry
WO2004016768A2