A molecular marker combination and primer combination and their application in tracing paternal ancestors

By designing the molecular marker combination and primer combination of the Y chromosome InDel and SNP sites, combined with PCR and capillary electrophoresis technology, the problem of difficult traceability of male ancestor information in complex areas is solved, and accurate ancestral inference of male samples in southwestern China is achieved.

CN118910278BActive Publication Date: 2025-08-26SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411140395.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2025-08-26
Estimated Expiration
2044-08-20

AI Technical Summary

Technical Problem

The prior art is difficult to accurately infer individual ancestral information through paternal genetic markers in areas with complex geographical environments and numerous ethnic groups, and is not effectively applied in male and female bodily fluid mixed spot samples.

Method used

A molecular marker combination was designed, including 11 Y chromosome InDel sites and 10 Y chromosome SNP sites, and a matching primer combination was used to trace the ancestral information of male samples through PCR amplification and single-base extension reaction, combined with capillary electrophoresis and machine learning models.

Benefits of technology

Accurate ancestral inference of male blood mark examination materials was achieved, especially men in southwestern China, and the inference accuracy and sensitivity of mixed samples in forensic science were improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118910278B_ABST
    Figure CN118910278B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of molecular genetics and forensic medicine technology, and specifically relates to a molecular marker combination and a primer combination and their application in tracing paternal ancestors. The molecular marker combination provided by the present invention includes 11 Y chromosome InDel sites and 10 Y chromosome SNP sites. Based on the site information of the molecular marker combination provided by the present invention, it is possible to accurately detect male DNA components in blood samples, and to perform ancestral inference of male bloodstain specimens, providing a powerful tool for forensic male ancestral inference applications. By utilizing a primer combination designed based on the molecular marker combination described in the present invention, only a simple PCR amplification and extension reaction is required, and the obtained extension product is introduced into the ancestral inference model to obtain the ancestral information of the sample. The method is simple, has strong specificity, and high sensitivity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of molecular genetics and forensic medicine, and in particular relates to a molecular marker combination and a primer combination and their application in tracing paternal ancestors. Background Art

[0002] In forensic work, inferring ancestral information from biological specimens left at the scene can provide crucial clues for investigation. However, in regions with complex geography and diverse ethnic groups, diverse cultures and unique regional characteristics exist. Numerous studies have examined the genetic structure of different ethnic groups in such regions, but most reported epigenetic ancestry prediction models are based on autosomal loci and cannot be used to screen paternal lines. Furthermore, mixed male and female body fluids are often encountered during actual investigations, and currently reported epigenetic ancestry prediction models are unable to produce satisfactory results in mixed male and female samples. Summary of the Invention

[0003] The purpose of the present invention is to make up for the shortcomings of the existing technology, provide a molecular marker combination and primer combination and their application in tracing paternal ancestors, accurately trace paternal ancestor information, and provide an effective means for solving family screening and inferring the ancestors of male individuals in mixed spot samples.

[0004] The present invention provides a molecular marker combination, which includes a first molecular marker, a second molecular marker, a third molecular marker, a fourth molecular marker, a fifth molecular marker, a sixth molecular marker, a seventh molecular marker, an eighth molecular marker, a ninth molecular marker, a tenth molecular marker, an eleventh molecular marker, a twelfth molecular marker, a thirteenth molecular marker, a fourteenth molecular marker, a fifteenth molecular marker, a sixteenth molecular marker, a seventeenth molecular marker, an eighteenth molecular marker, a nineteenth molecular marker, a twentieth molecular marker, and a twenty-first molecular marker;

[0005] The first molecular marker has a single nucleotide polymorphism site Y-InDel-1, located at the 13990180bp site on the human Y chromosome, and the polymorphism is C / CT;

[0006] The second molecular marker has a single nucleotide polymorphism site Y-SNP-1, located at the 22177026bp site on the human Y chromosome, and the polymorphism is A / C;

[0007] The third molecular marker has a single nucleotide polymorphism site Y-SNP-2, located at the 4954280bp site on the human Y chromosome, and the polymorphism is T / C;

[0008] The fourth molecular marker has a single nucleotide polymorphism site Y-InDel-2, located at the 8506135-8506136bp site on the human Y chromosome, and the polymorphism is GC / G;

[0009] The fifth molecular marker has a single nucleotide polymorphism site Y-InDel-3, located at the 9884493-9884494bp site on the human Y chromosome, and the polymorphism is CT / C;

[0010] The sixth molecular marker has a single nucleotide polymorphism site Y-SNP-3, located at the 2734854bp site on the human Y chromosome, and the polymorphism is C / T;

[0011] The seventh molecular marker has a single nucleotide polymorphism site Y-SNP-4, located at the 8492876bp site on the human Y chromosome, and the polymorphism is C / T;

[0012] The eighth molecular marker has a single nucleotide polymorphism site Y-InDel-4, located at the 15385547-15385548bp site on the human Y chromosome, and the polymorphism is AG / A;

[0013] The ninth molecular marker has a single nucleotide polymorphism site Y-InDel-5, located at the 16371765-16371768 bp site on the human Y chromosome, and the polymorphism is CAGT / C;

[0014] The tenth molecular marker has a single nucleotide polymorphism site Y-SNP-5, located at the 15469724bp site on the human Y chromosome, and the polymorphism is G / A;

[0015] The 11th molecular marker has a single nucleotide polymorphism site Y-SNP-6, located at the 21764674bp site on the human Y chromosome, and the polymorphism is A / G;

[0016] The twelfth molecular marker has a single nucleotide polymorphism site Y-SNP-7, located at the 14495243bp site on the human Y chromosome, and the polymorphism is T / C;

[0017] The 13th molecular marker has a single nucleotide polymorphism site Y-InDel-6, located at the 8107158bp site on the human Y chromosome, and the polymorphism is A / AC;

[0018] The fourteenth molecular marker has a single nucleotide polymorphism site Y-SNP-8, located at the 15581983bp site on the human Y chromosome, and the polymorphism is A / G;

[0019] The 15th molecular marker has a single nucleotide polymorphism site Y-InDel-7, located at the 15508699 to 15508704 bp site on the human Y chromosome, and the polymorphism is ACTTCT / A;

[0020] The sixteenth molecular marker has a single nucleotide polymorphism site Y-InDel-8, located at the 16269068bp site on the human Y chromosome, and the polymorphism is G / GA;

[0021] The 17th molecular marker has a single nucleotide polymorphism site Y-InDel-9, located at the 19470788-19470789 bp site on the human Y chromosome, and the polymorphism is GA / A;

[0022] The 18th molecular marker has a single nucleotide polymorphism site Y-InDel-10, located at the 18894227-18894228 bp site on the human Y chromosome, and the polymorphism is TA / T;

[0023] The 19th molecular marker has a single nucleotide polymorphism site Y-InDel-11, which is located at the 6661104-6661107bp site on the human Y chromosome, and the polymorphism is GACA / G;

[0024] The 20th molecular marker has a single nucleotide polymorphism site Y-SNP-9, located at the 22928068bp site on the human Y chromosome, and the polymorphism is G / A;

[0025] The 21st molecular marker has a single nucleotide polymorphism site Y-SNP-10, which is located at the 22749853bp site on the human Y chromosome, and the polymorphism is A / C.

[0026] The present invention also provides a primer combination for detecting the molecular marker combination described in the above technical solution, wherein the primer combination includes an amplification primer combination; the amplification primer combination includes a first amplification primer, a second amplification primer, a third amplification primer, a fourth amplification primer, a fifth amplification primer, a sixth amplification primer, a seventh amplification primer, an eighth amplification primer, a ninth amplification primer, a tenth amplification primer, an eleventh amplification primer, a twelfth amplification primer, a thirteenth amplification primer, a fourteenth amplification primer, a fifteenth amplification primer, a sixteenth amplification primer, a seventeenth amplification primer, an eighteenth amplification primer, a nineteenth amplification primer, a twentieth amplification primer, and a twenty-first amplification primer;

[0027] The first amplification primer includes an upstream primer 1 having a nucleotide sequence as shown in SEQ ID NO.1 and a downstream primer 1 having a nucleotide sequence as shown in SEQ ID NO.2;

[0028] The second amplification primers include an upstream primer 2 having a nucleotide sequence as shown in SEQ ID NO.3 and a downstream primer 2 having a nucleotide sequence as shown in SEQ ID NO.4;

[0029] The third amplification primer includes an upstream primer 3 whose nucleotide sequence is shown as SEQ ID NO.5 and a downstream primer 3 whose nucleotide sequence is shown as SEQ ID NO.6;

[0030] The fourth amplification primer includes an upstream primer 4 having a nucleotide sequence as shown in SEQ ID NO.7 and a downstream primer 4 having a nucleotide sequence as shown in SEQ ID NO.8;

[0031] The fifth amplification primer includes an upstream primer 5 having a nucleotide sequence as shown in SEQ ID NO.9 and a downstream primer 5 having a nucleotide sequence as shown in SEQ ID NO.10;

[0032] The sixth amplification primer includes an upstream primer 6 having a nucleotide sequence as shown in SEQ ID NO.11 and a downstream primer 6 having a nucleotide sequence as shown in SEQ ID NO.12;

[0033] The seventh amplification primer includes an upstream primer 7 whose nucleotide sequence is shown in SEQ ID NO.13 and a downstream primer 7 whose nucleotide sequence is shown in SEQ ID NO.14;

[0034] The eighth amplification primer includes an upstream primer 8 having a nucleotide sequence as shown in SEQ ID NO.15 and a downstream primer 8 having a nucleotide sequence as shown in SEQ ID NO.16;

[0035] The ninth amplification primer includes an upstream primer 9 having a nucleotide sequence as shown in SEQ ID NO.17 and a downstream primer 9 having a nucleotide sequence as shown in SEQ ID NO.18;

[0036] The tenth amplification primer includes an upstream primer 10 having a nucleotide sequence as shown in SEQ ID NO.19 and a downstream primer 10 having a nucleotide sequence as shown in SEQ ID NO.20;

[0037] The 11th amplification primer includes an upstream primer 11 having a nucleotide sequence as shown in SEQ ID NO.21 and a downstream primer 11 having a nucleotide sequence as shown in SEQ ID NO.22;

[0038] The 12th amplification primer includes an upstream primer 12 having a nucleotide sequence as shown in SEQ ID NO. 23 and a downstream primer 12 having a nucleotide sequence as shown in SEQ ID NO. 24;

[0039] The 13th amplification primer includes an upstream primer 13 having a nucleotide sequence as shown in SEQ ID NO.25 and a downstream primer 13 having a nucleotide sequence as shown in SEQ ID NO.26;

[0040] The 14th amplification primer includes an upstream primer 14 having a nucleotide sequence as shown in SEQ ID NO. 27 and a downstream primer 14 having a nucleotide sequence as shown in SEQ ID NO. 28;

[0041] The 15th amplification primer includes an upstream primer 15 having a nucleotide sequence as shown in SEQ ID NO.29 and a downstream primer 15 having a nucleotide sequence as shown in SEQ ID NO.30;

[0042] The 16th amplification primer includes an upstream primer 16 having a nucleotide sequence as shown in SEQ ID NO.31 and a downstream primer 16 having a nucleotide sequence as shown in SEQ ID NO.32;

[0043] The 17th amplification primer includes an upstream primer 17 having a nucleotide sequence as shown in SEQ ID NO.33 and a downstream primer 17 having a nucleotide sequence as shown in SEQ ID NO.34;

[0044] The 18th amplification primer includes an upstream primer 18 having a nucleotide sequence as shown in SEQ ID NO.35 and a downstream primer 18 having a nucleotide sequence as shown in SEQ ID NO.36;

[0045] The 19th amplification primer includes an upstream primer 19 having a nucleotide sequence as shown in SEQ ID NO.37 and a downstream primer 19 having a nucleotide sequence as shown in SEQ ID NO.38;

[0046] The 20th amplification primer includes an upstream primer 20 having a nucleotide sequence as shown in SEQ ID NO.39 and a downstream primer 20 having a nucleotide sequence as shown in SEQ ID NO.40;

[0047] The 21st amplification primer includes an upstream primer 21 having a nucleotide sequence as shown in SEQ ID NO.41 and a downstream primer 21 having a nucleotide sequence as shown in SEQ ID NO.42.

[0048] Preferably, the primer combination further comprises an extension primer combination;

[0049] The extension primer combination includes the first to twenty-first extension primers; the nucleotide sequences of the first to twenty-first extension primers are shown in SEQ ID NO.43 to SEQ ID NO.63, respectively.

[0050] The present invention also provides the use of the primer combination described in the above technical solution in the preparation of forensic identification products.

[0051] Preferably, the forensic identification products include forensic products for tracing paternal ancestors; and the males include males in southwest China.

[0052] The present invention also provides a kit for tracing paternal ancestors, which comprises the primer combination described in the above technical solution.

[0053] The present invention also provides the use of the molecular marker or primer combination or kit described in the above technical solution in tracing paternal ancestors.

[0054] The present invention also provides a method for tracing paternal ancestors, comprising the following steps:

[0055] Extracting genomic DNA from male samples to be tested;

[0056] Using the DNA as a template, PCR amplification is performed using the amplification primers in the primer combination described in the above technical solution to obtain an amplified product;

[0057] purifying the amplified product to obtain a purified amplified product;

[0058] Performing a single-base extension reaction on the purified amplification product using the extension primer combination in the primer combination described in the above technical solution to obtain an extension product;

[0059] performing capillary electrophoresis detection on the extension product to obtain a capillary electrophoresis detection result;

[0060] The capillary electrophoresis detection results are introduced into an ancestry inference model to obtain ancestry information of the male sample to be tested.

[0061] Preferably, the PCR amplification system is: QIAGEN Multiplex PCR Master Mix 2.5 μL, 1 μL of 100 μM amplification primer mixture, 0.125-10 ng of DNA template, and nuclease-free water to 5 μL;

[0062] The volume ratio of the first to twenty-first amplification primers in the amplification primer mixture is 5:5:1:10:5:20:5:1:10:2:1:20:30:2:20:30:2:50:5:2:3; the concentration ratio of the upstream primer to the downstream primer in each amplification primer is 1:1;

[0063] The PCR amplification program is as follows: 95°C for 15 min; 94°C for 30 s, 59°C for 90 s, 72°C for 60 s, 11 cycles, with each cycle decreasing by 1°C from the second cycle; 94°C for 30 s, 49°C for 90 s, 72°C for 60 s, 25 cycles; 60°C for 30 min; and storage at 4°C.

[0064] Preferably, the single base extension reaction system is: Platinum Multiplex Ready Reaction Mix 1.2 μL, 100 μM extension primer mixture 1.5 μL, purified amplification product 1 μL and nuclease-free water 1.5 μL;

[0065] The volume ratio of the first to twenty-first extension primers in the extension primer mixture is 5:5:1:10:5:20:5:1:10:2:1:20:30:2:20:30:2:50:5:2:3;

[0066] The procedure of the single base extension reaction is: 96° C. for 10 s, 50° C. for 5 s, 60° C. for 30 s, 26 cycles; and storage at 4° C.

[0067] Beneficial effects:

[0068] The molecular marker combination provided by the present invention includes 11 Y chromosome InDel sites and 10 Y chromosome SNP sites. Based on the site information of the molecular marker combination provided by the present invention, it can accurately detect the male DNA component in blood samples and can infer the ancestry of male bloodstain samples, providing a powerful tool for forensic male ancestry inference applications.

[0069] Furthermore, the present invention designs a primer combination based on the molecular marker combination, which only requires a simple PCR amplification and extension reaction. The obtained extension product is introduced into the ancestral inference model to obtain the ancestral information of the sample. The method is simple, specific and sensitive. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments are briefly introduced below.

[0071] Figure 1 It is the result of single-site detection of the Y chromosome molecular marker site;

[0072] Figure 2 The capillary electrophoresis results of step (3) in Example 4 are shown below:

[0073] Figure 3 The capillary electrophoresis results of the sensitivity test in Example 4 are as follows;

[0074] Figure 4The capillary electrophoresis results of the specific detection in Example 4 are as follows;

[0075] in, Figures 2-4 The horizontal axis value represents the DNA fragment length, and the vertical axis value represents the fluorescence intensity;

[0076] Figure 5 is the K-fold cross-validation result obtained by the random forest algorithm; A and B are the confusion matrices generated by random forest for the 1000 Genomes reference population and the three Chinese populations; C and D are the sensitivity, specificity, and balanced accuracy of the 1000 Genomes reference population and the three Chinese populations in the cross-validation. DETAILED DESCRIPTION

[0077] The present invention provides a molecular marker combination, which includes a first molecular marker, a second molecular marker, a third molecular marker, a fourth molecular marker, a fifth molecular marker, a sixth molecular marker, a seventh molecular marker, an eighth molecular marker, a ninth molecular marker, a tenth molecular marker, an eleventh molecular marker, a twelfth molecular marker, a thirteenth molecular marker, a fourteenth molecular marker, a fifteenth molecular marker, a sixteenth molecular marker, a seventeenth molecular marker, an eighteenth molecular marker, a nineteenth molecular marker, a twentieth molecular marker, and a twenty-first molecular marker;

[0078] The first molecular marker has a single nucleotide polymorphism site Y-InDel-1, located at the 13990180bp site on the human Y chromosome, and the polymorphism is C / CT;

[0079] The second molecular marker has a single nucleotide polymorphism site Y-SNP-1, located at the 22177026bp site on the human Y chromosome, and the polymorphism is A / C;

[0080] The third molecular marker has a single nucleotide polymorphism site Y-SNP-2, located at the 4954280bp site on the human Y chromosome, and the polymorphism is T / C;

[0081] The fourth molecular marker has a single nucleotide polymorphism site Y-InDel-2, located at the 8506135-8506136bp site on the human Y chromosome, and the polymorphism is GC / G;

[0082] The fifth molecular marker has a single nucleotide polymorphism site Y-InDel-3, located at the 9884493-9884494bp site on the human Y chromosome, and the polymorphism is CT / C;

[0083] The sixth molecular marker has a single nucleotide polymorphism site Y-SNP-3, located at the 2734854bp site on the human Y chromosome, and the polymorphism is C / T;

[0084] The seventh molecular marker has a single nucleotide polymorphism site Y-SNP-4, located at the 8492876bp site on the human Y chromosome, and the polymorphism is C / T;

[0085] The eighth molecular marker has a single nucleotide polymorphism site Y-InDel-4, located at the 15385547-15385548bp site on the human Y chromosome, and the polymorphism is AG / A;

[0086] The ninth molecular marker has a single nucleotide polymorphism site Y-InDel-5, located at the 16371765-16371768 bp site on the human Y chromosome, and the polymorphism is CAGT / C;

[0087] The tenth molecular marker has a single nucleotide polymorphism site Y-SNP-5, located at the 15469724bp site on the human Y chromosome, and the polymorphism is G / A;

[0088] The 11th molecular marker has a single nucleotide polymorphism site Y-SNP-6, located at the 21764674bp site on the human Y chromosome, and the polymorphism is A / G;

[0089] The twelfth molecular marker has a single nucleotide polymorphism site Y-SNP-7, located at the 14495243bp site on the human Y chromosome, and the polymorphism is T / C;

[0090] The 13th molecular marker has a single nucleotide polymorphism site Y-InDel-6, located at the 8107158bp site on the human Y chromosome, and the polymorphism is A / AC;

[0091] The fourteenth molecular marker has a single nucleotide polymorphism site Y-SNP-8, located at the 15581983bp site on the human Y chromosome, and the polymorphism is A / G;

[0092] The 15th molecular marker has a single nucleotide polymorphism site Y-InDel-7, located at the 15508699 to 15508704 bp site on the human Y chromosome, and the polymorphism is ACTTCT / A;

[0093] The sixteenth molecular marker has a single nucleotide polymorphism site Y-InDel-8, located at the 16269068bp site on the human Y chromosome, and the polymorphism is G / GA;

[0094] The 17th molecular marker has a single nucleotide polymorphism site Y-InDel-9, located at the 19470788-19470789 bp site on the human Y chromosome, and the polymorphism is GA / A;

[0095] The 18th molecular marker has a single nucleotide polymorphism site Y-InDel-10, located at the 18894227-18894228 bp site on the human Y chromosome, and the polymorphism is TA / T;

[0096] The 19th molecular marker has a single nucleotide polymorphism site Y-InDel-11, which is located at the 6661104-6661107bp site on the human Y chromosome, and the polymorphism is GACA / G;

[0097] The 20th molecular marker has a single nucleotide polymorphism site Y-SNP-9, located at the 22928068bp site on the human Y chromosome, and the polymorphism is G / A;

[0098] The 21st molecular marker has a single nucleotide polymorphism site Y-SNP-10, which is located at the 22749853bp site on the human Y chromosome, and the polymorphism is A / C.

[0099] Based on a data set obtained from the 1000 Genomes Database, the present invention analyzes ancestral information of all Y-InDel sites and screens out 11 Y-InDel sites with ancestral pre-correlation; and by comparing haplogroup sites that account for a relatively large and unique proportion in different populations, 10 Y-SNP sites with ancestral pre-correlation are screened out. The combination of the 11 Y-InDel sites and the 10 Y-SNP sites can infer the ancestral information of men, especially those in southwest China, providing a basis for the application of male ancestry inference in forensic medicine.

[0100] In view of the effects of the molecular marker combination provided by the present invention, the present invention also provides a primer combination for detecting the molecular marker combination described in the above technical solution, wherein the primer combination includes an amplification primer combination; the amplification primer combination includes a first amplification primer, a second amplification primer, a third amplification primer, a fourth amplification primer, a fifth amplification primer, a sixth amplification primer, a seventh amplification primer, an eighth amplification primer, a ninth amplification primer, a tenth amplification primer, an eleventh amplification primer, a twelfth amplification primer, a thirteenth amplification primer, a fourteenth amplification primer, a fifteenth amplification primer, a sixteenth amplification primer, a seventeenth amplification primer, an eighteenth amplification primer, a nineteenth amplification primer, a twentieth amplification primer, and a twenty-first amplification primer;

[0101] The first amplification primer includes an upstream primer 1 having a nucleotide sequence as shown in SEQ ID NO.1 and a downstream primer 1 having a nucleotide sequence as shown in SEQ ID NO.2;

[0102] The second amplification primers include an upstream primer 2 having a nucleotide sequence as shown in SEQ ID NO.3 and a downstream primer 2 having a nucleotide sequence as shown in SEQ ID NO.4;

[0103] The third amplification primer includes an upstream primer 3 whose nucleotide sequence is shown as SEQ ID NO.5 and a downstream primer 3 whose nucleotide sequence is shown as SEQ ID NO.6;

[0104] The fourth amplification primer includes an upstream primer 4 having a nucleotide sequence as shown in SEQ ID NO.7 and a downstream primer 4 having a nucleotide sequence as shown in SEQ ID NO.8;

[0105] The fifth amplification primer includes an upstream primer 5 having a nucleotide sequence as shown in SEQ ID NO.9 and a downstream primer 5 having a nucleotide sequence as shown in SEQ ID NO.10;

[0106] The sixth amplification primer includes an upstream primer 6 having a nucleotide sequence as shown in SEQ ID NO.11 and a downstream primer 6 having a nucleotide sequence as shown in SEQ ID NO.12;

[0107] The seventh amplification primer includes an upstream primer 7 whose nucleotide sequence is shown in SEQ ID NO.13 and a downstream primer 7 whose nucleotide sequence is shown in SEQ ID NO.14;

[0108] The eighth amplification primer includes an upstream primer 8 having a nucleotide sequence as shown in SEQ ID NO.15 and a downstream primer 8 having a nucleotide sequence as shown in SEQ ID NO.16;

[0109] The ninth amplification primer includes an upstream primer 9 having a nucleotide sequence as shown in SEQ ID NO.17 and a downstream primer 9 having a nucleotide sequence as shown in SEQ ID NO.18;

[0110] The tenth amplification primer includes an upstream primer 10 having a nucleotide sequence as shown in SEQ ID NO.19 and a downstream primer 10 having a nucleotide sequence as shown in SEQ ID NO.20;

[0111] The 11th amplification primer includes an upstream primer 11 having a nucleotide sequence as shown in SEQ ID NO.21 and a downstream primer 11 having a nucleotide sequence as shown in SEQ ID NO.22;

[0112] The 12th amplification primer includes an upstream primer 12 having a nucleotide sequence as shown in SEQ ID NO. 23 and a downstream primer 12 having a nucleotide sequence as shown in SEQ ID NO. 24;

[0113] The 13th amplification primer includes an upstream primer 13 having a nucleotide sequence as shown in SEQ ID NO.25 and a downstream primer 13 having a nucleotide sequence as shown in SEQ ID NO.26;

[0114] The 14th amplification primer includes an upstream primer 14 having a nucleotide sequence as shown in SEQ ID NO. 27 and a downstream primer 14 having a nucleotide sequence as shown in SEQ ID NO. 28;

[0115] The 15th amplification primer includes an upstream primer 15 having a nucleotide sequence as shown in SEQ ID NO.29 and a downstream primer 15 having a nucleotide sequence as shown in SEQ ID NO.30;

[0116] The 16th amplification primer includes an upstream primer 16 having a nucleotide sequence as shown in SEQ ID NO.31 and a downstream primer 16 having a nucleotide sequence as shown in SEQ ID NO.32;

[0117] The 17th amplification primer includes an upstream primer 17 having a nucleotide sequence as shown in SEQ ID NO.33 and a downstream primer 17 having a nucleotide sequence as shown in SEQ ID NO.34;

[0118] The 18th amplification primer includes an upstream primer 18 having a nucleotide sequence as shown in SEQ ID NO.35 and a downstream primer 18 having a nucleotide sequence as shown in SEQ ID NO.36;

[0119] The 19th amplification primer includes an upstream primer 19 having a nucleotide sequence as shown in SEQ ID NO.37 and a downstream primer 19 having a nucleotide sequence as shown in SEQ ID NO.38;

[0120] The 20th amplification primer includes an upstream primer 20 having a nucleotide sequence as shown in SEQ ID NO.39 and a downstream primer 20 having a nucleotide sequence as shown in SEQ ID NO.40;

[0121] The 21st amplification primer includes an upstream primer 21 having a nucleotide sequence as shown in SEQ ID NO.41 and a downstream primer 21 having a nucleotide sequence as shown in SEQ ID NO.42.

[0122] In the present invention, the primer combination preferably further includes an extension primer combination; the extension primer combination preferably includes the 1st to 21st extension primers; the nucleotide sequences of the 1st to 21st extension primers are shown in SEQ ID NO.43 to SEQ ID NO.63, respectively.

[0123] The present invention also provides the use of the primer combination described in the above technical solution in the preparation of a forensic identification product. In the present invention, the forensic identification product preferably includes a product for forensic tracing of paternal ancestry; the males include males from southwestern China, and more preferably include males from one or more of the Hui, Qiang, and Han ethnic groups in southwestern China. The product of the present invention preferably includes a kit, and more preferably includes an SNaPshot detection kit.

[0124] The present invention also provides a kit for tracing paternal ancestors, the kit comprising the primer combination described in the above technical solution. In the present invention, the kit preferably also comprises components of a conventional detection kit, such as QIAGEN Multiplex PCR Master Mix and / or nuclease-free water.

[0125] The present invention designs primer combinations (including amplification primers and extension primers) based on the molecular marker sites, which can simultaneously perform multiple site detections, and can trace the ancestral information of men, especially men from some groups in southwest China, based on the capillary electrophoresis platform detection and analysis commonly used in forensic laboratories, providing a scientific basis for the application of inferring male ancestors in mixed samples.

[0126] In view of the effects of the molecular marker or primer combination or kit provided by the present invention, the application of the molecular marker or primer combination or kit in tracing paternal ancestors also falls within the scope of protection of the present invention.

[0127] The present invention also provides a method for tracing paternal ancestors, comprising the following steps:

[0128] Extracting genomic DNA from male samples to be tested;

[0129] Using the DNA as a template, PCR amplification is performed using the amplification primers in the primer combination described in the above technical solution to obtain an amplified product;

[0130] purifying the amplified product to obtain a purified amplified product;

[0131] Performing a single-base extension reaction on the purified amplification product using the extension primer combination in the primer combination described in the above technical solution to obtain an extension product;

[0132] The extension product is introduced into an ancestry inference model to obtain ancestry information of the male sample to be tested.

[0133] The present invention has no strict requirements on the step of extracting genomic DNA from the male sample to be tested, and conventional steps in the art can be used.

[0134] After obtaining genomic DNA, the present invention uses the DNA as a template and performs PCR amplification using the amplification primers in the primer combination described in the above technical solution to obtain an amplified product. In the present invention, the PCR amplification system is preferably: 2.5 μL of QIAGEN Multiplex PCR Master Mix, 1 μL of 100 μM amplification primer mixture, 0.125-10 ng of DNA template, and nuclease-free water to 5 μL; the amount of the DNA template is preferably 0.25-5 ng, more preferably 0.1-0.5 ng, and more preferably 0.2 ng; the volume ratio of the first to 21st amplification primers in the amplification primer mixture is preferably 5:5:1:10:5:20:5:1:10:2:1:20:30:2:20:30:2:50:5:2:3; the concentration ratio of the upstream primer to the downstream primer in each amplification primer is preferably 1:1, respectively. The PCR amplification program of the present invention is preferably: 95°C for 15 min; 94°C for 30 s, 59°C for 90 s, 72°C for 60 s, 11 cycles, with the temperature of each cycle decreasing by 1°C from the second cycle; 94°C for 30 s, 49°C for 90 s, 72°C for 60 s, 25 cycles; 60°C for 30 min; and storage at 4°C.

[0135] After obtaining the amplification product, the present invention purifies the amplification product to obtain a purified amplification product. In the present invention, the purification preferably includes the step of removing excess nucleotides and primers in the amplification product; the method of removing excess nucleotides and primers in the amplification product is preferably: adding exonuclease I and shrimp alkaline phosphatase to the amplification product, digesting and inactivating it; the volume ratio of the amplification product, exonuclease I and shrimp alkaline phosphatase is preferably 5:1.1:2.5; the enzymatic activity of the exonuclease I is preferably 5U / mL; the enzymatic activity of the shrimp alkaline phosphatase is preferably 1U / mL. The temperature of the digestion treatment of the present invention is preferably 37°C, and the time is preferably 60 minutes; the temperature of the inactivation is preferably 80°C, and the time is preferably 20 minutes.

[0136] After obtaining the purified amplification product, the present invention uses the extension primer combination in the primer combination described in the above technical solution to perform a single base extension reaction on the purified amplification product to obtain an extension product. In the present invention, the system of the single base extension reaction is preferably: 1.2 μL of Platinum Multiplex Ready Reaction Mix, 1.5 μL of 100 μM extension primer mixture, 1 μL of purified amplification product, and 1.5 μL of nuclease-free water; the volume ratio of the 1st to 21st extension primers in the extension primer mixture is preferably 5:5:1:10:5:20:5:1:10:2:1:20:30:2:20:30:2:50:5:2:3; the procedure of the single base extension reaction is preferably: 96°C for 10 s, 50°C for 5 s, 60°C for 30 s, 26 cycles; and stored at 4°C.

[0137] After obtaining the extension product, the present invention performs capillary electrophoresis detection on the extension product to obtain the capillary electrophoresis detection result. The present invention has no strict requirements on the capillary electrophoresis detection step, and conventional operations can be performed.

[0138] After obtaining the capillary electrophoresis test results, the present invention imports the capillary electrophoresis test results into an ancestral inference model to obtain the ancestral information of the male sample to be tested. In the present invention, the ancestral inference model includes using the caret, nnet, rf, knn, and multinom software packages in R language v4.3.2, and a confusion matrix generated by a machine learning model; the machine learning model includes one or more of a random forest model, a K-nearest neighbor model, a neural network model, and a multinomial logistic regression model, and is further preferably a random forest model.

[0139] The present invention uses amplification primers and extension primers to detect genomic DNA, and combines the detection and analysis results of the capillary electrophoresis platform with the ancestral inference model to trace the ancestral information of men, especially those in some groups in southwest China. It has strong specificity and high sensitivity, and provides a scientific basis for the application of male ancestry inference in mixed samples.

[0140] To further illustrate the present invention, a molecular marker combination and primer combination provided by the present invention and their application in tracing paternal ancestors are described in detail below with reference to the accompanying drawings and examples, but they should not be construed as limiting the scope of protection of the present invention.

[0141] Example 1

[0142] 1. Sample Collection

[0143] Following informed consent, DNA samples were collected from 382 healthy, unrelated male participants, including 137 Han Chinese from Sichuan, 170 Qiang Chinese from Sichuan, 36 Hui Chinese from Sichuan, and 39 Hui Chinese from Yunnan. The samples were then frozen at -20°C for subsequent DNA extraction.

[0144] 2. Sample Genomic DNA Extraction

[0145] DNA was extracted using the Chelex-100 method. After DNA extraction, it was quantified using a Nanodrop™ 2000 spectrophotometer and stored in a -20°C freezer.

[0146] Example 2

[0147] 1. Y site screening

[0148] (1) F was calculated by comparing five reference populations in the 1000 Genome Database: Mixed American (AMR), European (EUR), East Asian (EAS), South Asian (SAS), and African (AFR). ST The sum of the values ​​is used to sort the InDel sites. ST The top 10 gene loci with the highest values.

[0149] (2) These loci were typed for 10 different individuals with significant genetic differences (Han from Sichuan, Han from Beijing, Qiang, Yi, Tibetan, Uyghur, Oroqen, Kirgiz, Hui, and Manchu). The total gene frequencies of these individuals were compared with the intercontinental F values ​​of the corresponding loci in the genomes of thousands of people by linear regression analysis. ST The values ​​were calculated and all InDel sites were evaluated according to the following formula (1) obtained by linear regression:

[0150] F=0.566×X1+1.053×X2-0.14 Formula (1)

[0151] Among them, F represents the predicted gene frequency, X1 represents the gene frequency of the corresponding site of the East Asian population in the 1000 Genomes, and X2 represents the average F of the reference population in the 1000 Genomes. ST .

[0152] Loci with a gene frequency exceeding 0.1 calculated based on the above formula were screened and considered as candidate genes.

[0153] (3) Eliminate loci from candidate genes and screen for Y-InDel sites based on the following criteria:

[0154] a. Number of alleles ≥ 3;

[0155] b. Loci for which suitable amplification and sequencing primers cannot be designed due to flanking sequence characteristics;

[0156] c. For two loci that are 10 kbp apart, select the locus with the higher predicted gene frequency;

[0157] d. Loci with insertions or deletions of alleles exceeding 20 base pairs;

[0158] e. The flanking regions of the loci contain SNPs, InDels, or other genetic polymorphisms;

[0159] f. Loci that were unsuccessfully amplified under complex conditions.

[0160] (4) Screen Y-SNP sites by screening haplogroup sites that are relatively large and unique in different populations.

[0161] In this example, a total of 11 Y-InDel sites and 10 Y-SNP sites were screened out, and the specific site information is shown in Table 1 below.

[0162] Table 1 Y chromosome molecular marker site information

[0163]

[0164]

[0165] Note: The polymorphism at 13990180bp in the table is an example of C-CT, which indicates that there is an insertion marker between the 13990180 and 13990181bp sites, and the inserted base type is T; the polymorphism at 8506135 to 8506136bp in the table is an example of CG-C, which indicates that there is a base deletion at the 8506136bp site.

[0166] 2. Primer Design

[0167] According to the location of the site screened in step 1 in the GRCh37 gene, sequence information was obtained through the UCSC (https: / / genome.ucsc.edu / ) database. Amplification primers were then designed using the Primer-BLAET online tool in NCBI (https: / / www.ncbi.nlm.nih.gov / ), and sequencing primers were designed using the PyroMarkAssayDesign 2.0 software. Primers that produce primer dimers and hairpin structures found in the software were excluded. After the design was completed, the primers were simulated and amplified using the UCSC In-Silico PCR tool to verify the specificity of the primers. Among them, for the Y-InDel site, the SNaPsho technology was also used for detection. After the sequence undergoes an InDel mutation, the base at the corresponding position will be replaced by an inserted sequence or filled with a sequence that lacks a suffix, so the base at this position will be different. The primers for the InDel site were designed using the same method. The primer sequences designed in this embodiment are shown in Tables 2 and 3.

[0168] Table 2 Primers for amplification of Y chromosome molecular marker sites

[0169]

[0170]

[0171] Table 3 Y chromosome molecular marker site extension primers

[0172] Molecular markers Physical location (bp) Sequence (5'-3') Sequence number Y-InDel-1 13990180 TCATGAAGAAAAATATATCT SEQ ID NO.43 Y-SNP-1 22177026 TCCACGGGACCATCTC SEQ ID NO.44 Y-SNP-2 4954280 TGCACCCCTCACTTCTGCACT SEQ ID NO.45 Y-InDel-2 8506135 CCAGGCCAGCTCTGC SEQ ID NO.46 Y-InDel-3 9884493 ACTGGCTTCTCTCCTTT SEQ ID NO.47 Y-SNP-3 2734854 GCAGGGCAATAAACCTTGGATTTC SEQ ID NO.48 Y-SNP-4 8492876 TCCACGGGACCATCTC SEQ ID NO.49 Y-InDel-4 15385547 GGTATTTGTTCCTGCC SEQ ID NO.50 Y-InDel-5 16371765 CCAAATGCACTTTAACTACT SEQ ID NO.51 Y-SNP-5 15469724 TCGATCTTTCCCCCAATT SEQ ID NO.52 Y-SNP-6 21764674 CTTTATTCAGATTTTCCCCTGAGAGC SEQ ID NO.53 Y-SNP-7 14495243 GGTTACATAAATAAGGTTTTTTTTTGGTTG SEQ ID NO.54 Y-InDel-6 8107158 CCAGTTGTACTGTTTCTTTT SEQ ID NO.55 Y-SNP-8 15581983 GGCAAATGTAAGTCAAGCAAGAAATTTA SEQ ID NO.56 Y-InDel-7 15508699 CATGCCTTCTCACTTCTC SEQ ID NO.57 Y-InDel-8 16269068 TCTTTGTAACCACTTTTTTT SEQ ID NO.58 Y-InDel-9 19470788 AGAGTTTGAGGGCAGA SEQ ID NO.59 Y-InDel-10 18894227 TTGATCCCCACCAAT SEQ ID NO.60 Y-InDel-11 6661104 TTTAGTGAGGAAGCTGAC SEQ ID NO.61 Y-SNP-9 22928068 TCTCGGGGATTCTAAAATGTTTCCAG SEQ ID NO.62 Y-SNP-10 22749853 CGTCTTATACCAAAATATCACCAGTTGT SEQ ID NO.63

[0173] Example 3

[0174] Single-site detection analysis

[0175] (1) Using the genomic DNA extracted in Example 1 as a template, PCR amplification was performed on a ProFlex™ PCR instrument using the Platinum Multiplex PCR Master Mix kit according to the following system ratios and procedures to obtain amplified products; each reaction system contained only one pair of amplification primers corresponding to one site;

[0176] The PCR amplification system was as follows: QIAGEN Multiplex PCR Master Mix 2.5 μL, primer mix 1 μL (upstream primer concentration was 100 μM, downstream primer concentration was 100 μM), DNA template 10 ng, and nuclease-free water was added to 5 μL;

[0177] The PCR amplification program was as follows: 95°C for 15 min; 94°C for 30 s, 59°C for 90 s, 72°C for 60 s, for 11 cycles, with the temperature of each cycle decreasing by 1°C from the second cycle; 94°C for 30 s, 49°C for 90 s, 72°C for 60 s, for 25 cycles; 60°C for 30 min; and storage at 4°C.

[0178] (2) Purify the amplified product obtained in step (1). Specifically, take 5 μL of the amplified product obtained in step 1, add 1.1 μL of EXOI (exonuclease I) with an enzyme activity of 5 U / mL and 2.5 μL of SAP (shrimp alkaline phosphatase) with an enzyme activity of 1 U / mL, shake and mix, centrifuge, and then digest in a ProFlexTM PCR instrument at 37°C for 60 min to eliminate excess nucleotides and primers, and then keep at 80°C for 20 min to inactivate the purified enzyme to obtain a purified product, which was stored at 4°C.

[0179] (3) Using the Platinum Multiplex Ready Reaction Mix, the purified product, and the extension primer, perform a single-base extension reaction according to the following system ratio and procedure. The extension reaction is a single-site extension, that is, each reaction system contains only one extension primer corresponding to one site;

[0180] Single-base extension reaction system: Platinum Multiplex Ready Reaction Mix 1.2 μL, extension primer 1.5 μL, purified amplification product 1 μL, and nuclease-free water 1.5 μL;

[0181] Single base extension reaction program: 96°C for 10 s, 50°C for 5 s, 60°C for 30 s, 26 cycles; stored at 4°C.

[0182] (4) After the single-base extension reaction in step (3) is completed, add 1 μL of SAP to the product to digest unincorporated ddNTPs. Digest at 37°C for 60 min and then hold at 80°C for 20 min before use in capillary electrophoresis. If overnight storage is required, store at -20°C.

[0183] Take 1 μL of the purified single base extension reaction product and 9 μL of HiDi Formamide internal standard mixture (add 3 μL of Genescan to 1 ml of HiDi Formamide) TM 120LIZ TM(Thermo Fisher Scientific) internal standard and mix well), mix well and centrifuge, perform capillary electrophoresis on 3130 genetic analyzer, use Data Collection Software V3.0 to collect electrophoresis data, and use Gene Mapper ID V3.2 to analyze the capillary electrophoresis results. The results show that each set of primers can amplify the correct result and the peak appears at the expected position. Among them, the capillary electrophoresis results of Y-InDel-1 molecular marker (Q1), Y-SNP-1 molecular marker (Q2), Y-SNP-2 molecular marker (Q3), Y-InDel-2 molecular marker (Q4), Y-InDel-3 molecular marker (Q5) and Y-SNP-3 molecular marker (Q6) are shown in Figure 1 shown.

[0184] Example 4

[0185] Establishment and testing of NaPshot composite system

[0186] (1) Composite amplification primers

[0187] Y-InDel-1 amplification primer, Y-SNP-1 amplification primer, Y-SNP-2 amplification primer, Y-InDel-2 amplification primer, Y-InDel-3 amplification primer, Y-SNP-3 amplification primer, Y-SNP-4 amplification primer, Y-InDel-4 amplification primer, Y-InDel-5 amplification primer, Y-SNP-5 amplification primer, Y-SNP-6 amplification primer, Y-SNP-7 amplification primer, Y-InDel-6 amplification primer, Y-SNP-8 amplification primer, Y-SNP-9 amplification primer, Y-SNP-10 amplification primer, Y-SNP-11 amplification primer, Y-SNP-12 amplification primer, Y-SNP-13 amplification primer, Y-SNP-14 amplification primer, Y-SNP-15 amplification primer, Y-SNP-16 amplification primer, Y-SNP-17 amplification primer, Y-SNP-18 amplification primer, Y-SNP-19 amplification primer, Y-SNP-20 amplification primer, Y-SNP-21 amplification primer, Y-SNP-22 amplification primer, Y-SNP-23 amplification primer, Y-SNP-24 amplification primer, Y-SNP-25 amplification primer, Y-SNP-26 amplification primer, Y-SNP-27 amplification primer, Y-SNP-28 amplification primer, NP-8 amplification primer, Y-InDel-7 amplification primer, Y-InDel-8 amplification primer, Y-InDel-9 amplification primer, Y-InDel-10 amplification primer, Y-InDel-11 amplification primer, Y-SNP-9 amplification primer and Y-SNP-10 amplification primer are mixed in a volume ratio of 5:5:1:10:5:20:5:1:10:2:1:20:30:2:20:30:2:50:5:2:3 to obtain a composite amplification primer;

[0188] The independent concentration of each amplification primer before mixing was 100 μM, and the concentration ratio of the upstream primer to the downstream primer in each amplification primer was 1:1.

[0189] (2) Composite extension primer

[0190] Y-InDel-1 extension primer, Y-SNP-1 extension primer, Y-SNP-2 extension primer, Y-InDel-2 extension primer, Y-InDel-3 extension primer, Y-SNP-3 extension primer, Y-SNP-4 extension primer, Y-InDel-4 extension primer, Y-InDel-5 extension primer, Y-SNP-5 extension primer, Y-SNP-6 extension primer, Y-SNP-7 extension primer, Y-InDel-6 extension primer, Y-SNP-8 extension primer, Y-SNP-9 extension primer, Y-SNP-10 extension primer, Y-SNP-11 extension primer, Y-SNP-12 extension primer, Y-SNP-13 extension primer, Y-SNP-14 extension primer, Y-SNP-15 extension primer, Y-SNP-16 extension primer, Y-SNP-17 extension primer, Y-SNP-18 extension primer, Y-SNP-19 extension primer, Y-SNP-20 extension primer, Y-SNP-21 extension primer, Y-SNP-22 extension primer, Y-SNP-23 extension primer, Y-SNP-24 extension primer, Y-SNP-25 extension primer, Y-SNP-26 extension primer, Y-SNP-27 extension primer, Y-SNP-28 extension primer, Y-SNP-29 extension primer, Y-SNP-30 extension primer, NP-8 extension primer, Y-InDel-7 extension primer, Y-InDel-8 extension primer, Y-InDel-9 extension primer, Y-InDel-10 extension primer, Y-InDel-11 extension primer, Y-SNP-9 extension primer and Y-SNP-10 extension primer are mixed at a volume ratio of 5:5:1:10:5:20:5:1:10:2:1:20:30:2:20:30:2:50:5:2:3 to obtain a composite extension primer;

[0191] The concentration of each extension primer before mixing was 100 μM.

[0192] (3) a. Using commercial standard DNA (2800M) as a template, PCR amplification was performed using the above composite amplification primers to obtain an amplified product;

[0193] The PCR amplification system was as follows: QIAGEN Multiplex PCRMasterMix 2.5 μL, primer mix 1 μL, DNA template 10 ng, and nuclease-free water to 5 μL.

[0194] The PCR amplification program was as follows: 95°C for 15 min; 94°C for 30 s, 59°C for 90 s, 72°C for 60 s, for 11 cycles, with the temperature of each cycle decreasing by 1°C from the second cycle; 94°C for 30 s, 49°C for 90 s, 72°C for 60 s, for 25 cycles; 60°C for 30 min; and storage at 4°C.

[0195] The PCR amplification program was as follows: 94°C for 8 min; 94°C for 40 s, 65°C for 25 s, 72°C for 20 s, for 15 cycles, with the temperature of each cycle decreasing by 1°C from the second cycle; 94°C for 40 s, 55°C for 25 s, 72°C for 20 s, for 20 cycles; 72°C for 8 min; and storage at 4°C.

[0196] b. After obtaining the amplified product, the amplified product was purified by referring to Example 3, and the purified amplified product was subjected to a single base extension reaction using the composite extension primer, and the capillary electrophoresis results were analyzed. The results were as follows. Figure 2 shown.

[0197] according to Figures 1-2 It can be seen that the SNaPshot composite system provided by the present invention can specifically amplify the molecular markers of the Y chromosome. According to the capillary electrophoresis results, the results of the test sample at all 21 Y chromosome sites can be clearly observed, which constitutes the ancestral information value of the sample and can be used to infer ancestors by model, proving that the Y chromosome molecular markers and SNaPshot composite system provided by the present invention can accurately detect biological samples (blood).

[0198] (4) Sensitivity detection

[0199] Referring to step (3), different masses of genomic DNA (10ng, 2ng, 1ng, 0.5ng, 0.25ng and 0.125ng) were detected and analyzed. The results are as follows: Figure 2 shown.

[0200] according to Figure 2 It can be seen that the peaks at each site can be completely detected when the DNA input amount ranges from 10 ng to 0.125 ng, which proves that the SNaPshot composite system of Y on chromosome Y of the present invention has good sensitivity.

[0201] (5) Specificity detection

[0202] Male standard DNA (2800M, purchased from Promega, USA) and standard DNA (F312, purchased from Promega, USA) were collected and mixed at certain mass ratios (1:1, 1:10, 1:50, 1:100, 0:1) to obtain DNA templates;

[0203] Referring to step (3), the DNA template is detected and analyzed, and the results are as follows: Figure 3 shown.

[0204] according to Figure 3 It can be seen that as the female DNA concentration increases, the peak height decreases, but the electrophoresis diagrams of all mixed samples can show complete peaks, which proves that the Y chromosome molecular markers and SNaPshot composite system provided by the present invention are suitable for the detection of male components in mixed blood samples of single males and females.

[0205] Example 5

[0206] (1) In order to evaluate the efficiency of the 21 molecular markers screened in Example 2 in assigning individuals to their correct biogeographical origins, a K-fold (10-fold) cross-validation analysis was used to evaluate the ability of this new system to accurately assign unknown individuals to their correct ancestral origins. Specifically, a cross-validation analysis was performed on the 1000 Genomes reference population and three Chinese populations (Hui (denoted as CCPW), Qiang (denoted as CQSC), and Han (denoted as CHSC)). The confusion matrix was generated using caret, nnet, rf, knn, and multinom packages in Rv4.3.2 using four methods (random forest, K-nearest neighbor, neural network, and multinomial logistic regression). It was found that random forest had the highest predictive power. Subsequently, the sensitivity-specificity index test was calculated and the results of the random forest were visualized, wherein the code used was as follows:

[0207] library(openxlsx)

[0208] library(caret)

[0209] data<-read.xlsx("your_data.xlsx")

[0210] colnames(data)<-c("Ethnicity",paste0("SNP",1:21))

[0211] control<-trainControl(method="cv",number=10)

[0212] model<-train(Ethnicity~.,data=data,method="rf",trControl=control)

[0213] #rf is random forest, which can be replaced by other methods

[0214] predicted<-predict(model,data)

[0215] conf_matrix<-table(Actual=data$Ethnicity,Predicted=predicted)

[0216] print(conf_matrix)

[0217] The K-fold cross validation results obtained using the random forest algorithm are as follows Figure 5 As shown. Figure 5It can be seen that random forests produced accurate results. Among the five intercontinental populations (1000 Genomes reference population), the predictions for Asian and African populations were more accurate. Specifically, the sensitivity of the 21 molecular markers for correctly predicting the ancestral origins of East Asian (EAS) and African (AFR) individuals was 98.20% and 88.92%, respectively. Since the AMR population provided was itself a mixed population, no one was classified into that group. Individuals from AMR were incorrectly assigned to the EUR and SAS populations. Among the three Chinese populations (Hui (CHuSW), Qiang (CQSC), Han (CHSC)), it showed better sensitivity and specificity because its clustering was significantly different from the other two populations, with a sensitivity of 100%. The Han and Qiang currently show a certain ability to distinguish.

[0218] Based on the above, it can be seen that the molecular marker combination provided by the present invention can accurately detect male components in blood samples and infer the ancestry of male samples. The method is simple, specific and sensitive.

[0219] Although the above embodiment provides a detailed description of the present invention, it is only a part of the embodiments of the present invention, not all of the embodiments. People can also obtain other embodiments based on this embodiment without creativity, and these embodiments all fall within the scope of protection of the present invention.

Claims

1. A molecular marker combination, characterized in that: The molecular marker combination includes a first molecular marker, a second molecular marker, a third molecular marker, a fourth molecular marker, a fifth molecular marker, a sixth molecular marker, a seventh molecular marker, an eighth molecular marker, a ninth molecular marker, a tenth molecular marker, an eleventh molecular marker, a twelfth molecular marker, a thirteenth molecular marker, a fourteenth molecular marker, a fifteenth molecular marker, a sixteenth molecular marker, a seventeenth molecular marker, an eighteenth molecular marker, a nineteenth molecular marker, a twentieth molecular marker, and a twenty-first molecular marker; The first molecular marker has a single nucleotide polymorphism site Y-InDel-1, located at the 13990180bp site on the human Y chromosome, and the polymorphism is C / CT; The second molecular marker has a single nucleotide polymorphism site Y-SNP-1, located at the 22177026bp site on the human Y chromosome, and the polymorphism is A / C; The third molecular marker has a single nucleotide polymorphism site Y-SNP-2, located at the 4954280bp site on the human Y chromosome, and the polymorphism is T / C; The fourth molecular marker has a single nucleotide polymorphism site Y-InDel-2, located at the 8506135-8506136bp site on the human Y chromosome, and the polymorphism is GC / G; The fifth molecular marker has a single nucleotide polymorphism site Y-InDel-3, located at the 9884493-9884494bp site on the human Y chromosome, and the polymorphism is CT / C; The sixth molecular marker has a single nucleotide polymorphism site Y-SNP-3, located at the 2734854bp site on the human Y chromosome, and the polymorphism is C / T; The seventh molecular marker has a single nucleotide polymorphism site Y-SNP-4, located at the 8492876bp site on the human Y chromosome, and the polymorphism is C / T; The eighth molecular marker has a single nucleotide polymorphism site Y-InDel-4, located at the 15385547-15385548bp site on the human Y chromosome, and the polymorphism is AG / A; The ninth molecular marker has a single nucleotide polymorphism site Y-InDel-5, located at the 16371765-16371768 bp site on the human Y chromosome, and the polymorphism is CAGT / C; The tenth molecular marker has a single nucleotide polymorphism site Y-SNP-5, located at the 15469724bp site on the human Y chromosome, and the polymorphism is G / A; The 11th molecular marker has a single nucleotide polymorphism site Y-SNP-6, located at the 21764674bp site on the human Y chromosome, and the polymorphism is A / G; The twelfth molecular marker has a single nucleotide polymorphism site Y-SNP-7, located at the 14495243bp site on the human Y chromosome, and the polymorphism is T / C; The 13th molecular marker has a single nucleotide polymorphism site Y-InDel-6, located at the 8107158bp site on the human Y chromosome, and the polymorphism is A / AC; The fourteenth molecular marker has a single nucleotide polymorphism site Y-SNP-8, located at the 15581983bp site on the human Y chromosome, and the polymorphism is A / G; The 15th molecular marker has a single nucleotide polymorphism site Y-InDel-7, located at the 15508699 to 15508704 bp site on the human Y chromosome, and the polymorphism is ACTTCT / A; The sixteenth molecular marker has a single nucleotide polymorphism site Y-In Del-8, located at the 16269068bp site on the human Y chromosome, and the polymorphism is G / GA; The 17th molecular marker has a single nucleotide polymorphism site Y-InDel-9, located at the 19470788-19470789 bp site on the human Y chromosome, and the polymorphism is GA / A; The 18th molecular marker has a single nucleotide polymorphism site Y-InDel-10, located at the 18894227-18894228 bp site on the human Y chromosome, and the polymorphism is TA / T; The 19th molecular marker has a single nucleotide polymorphism site Y-InDel-11, which is located at the 6661104-6661107bp site on the human Y chromosome, and the polymorphism is GACA / G; The 20th molecular marker has a single nucleotide polymorphism site Y-SNP-9, located at the 22928068bp site on the human Y chromosome, and the polymorphism is G / A; The 21st molecular marker has a single nucleotide polymorphism site Y-SNP-10, which is located at the 22749853bp site on the human Y chromosome, and the polymorphism is A / C.

2. A primer combination for detecting the molecular marker combination according to claim 1, characterized in that: The primer combination includes an amplification primer combination; the amplification primer combination includes a first amplification primer, a second amplification primer, a third amplification primer, a fourth amplification primer, a fifth amplification primer, a sixth amplification primer, a seventh amplification primer, an eighth amplification primer, a ninth amplification primer, a tenth amplification primer, an eleventh amplification primer, a twelfth amplification primer, a thirteenth amplification primer, a fourteenth amplification primer, a fifteenth amplification primer, a sixteenth amplification primer, a seventeenth amplification primer, an eighteenth amplification primer, a nineteenth amplification primer, a twentieth amplification primer, and a twenty-first amplification primer; The first amplification primer includes an upstream primer 1 having a nucleotide sequence as shown in SEQ ID NO.1 and a downstream primer 1 having a nucleotide sequence as shown in SEQ ID NO.2; The second amplification primers include an upstream primer 2 having a nucleotide sequence as shown in SEQ ID NO.3 and a downstream primer 2 having a nucleotide sequence as shown in SEQ ID NO.4; The third amplification primer includes an upstream primer 3 whose nucleotide sequence is shown as SEQ ID NO.5 and a downstream primer 3 whose nucleotide sequence is shown as SEQ ID NO.6; The fourth amplification primer includes an upstream primer 4 having a nucleotide sequence as shown in SEQ ID NO.7 and a downstream primer 4 having a nucleotide sequence as shown in SEQ ID NO.8; The fifth amplification primer includes an upstream primer 5 having a nucleotide sequence as shown in SEQ ID NO.9 and a downstream primer 5 having a nucleotide sequence as shown in SEQ ID NO.10; The sixth amplification primer includes an upstream primer 6 having a nucleotide sequence as shown in SEQ ID NO.11 and a downstream primer 6 having a nucleotide sequence as shown in SEQ ID NO.12; The seventh amplification primer includes an upstream primer 7 whose nucleotide sequence is shown in SEQ ID NO.13 and a downstream primer 7 whose nucleotide sequence is shown in SEQ ID NO.14; The eighth amplification primer includes an upstream primer 8 having a nucleotide sequence as shown in SEQ ID NO.15 and a downstream primer 8 having a nucleotide sequence as shown in SEQ ID NO.16; The ninth amplification primer includes an upstream primer 9 having a nucleotide sequence as shown in SEQ ID NO.17 and a downstream primer 9 having a nucleotide sequence as shown in SEQ ID NO.18; The tenth amplification primer includes an upstream primer 10 having a nucleotide sequence as shown in SEQ ID NO.19 and a downstream primer 10 having a nucleotide sequence as shown in SEQ ID NO.20; The 11th amplification primer includes an upstream primer 11 having a nucleotide sequence as shown in SEQ ID NO.21 and a downstream primer 11 having a nucleotide sequence as shown in SEQ ID NO.22; The 12th amplification primer includes an upstream primer 12 having a nucleotide sequence as shown in SEQ ID NO. 23 and a downstream primer 12 having a nucleotide sequence as shown in SEQ ID NO. 24; The 13th amplification primer includes an upstream primer 13 having a nucleotide sequence as shown in SEQ ID NO.25 and a downstream primer 13 having a nucleotide sequence as shown in SEQ ID NO.26; The 14th amplification primer includes an upstream primer 14 having a nucleotide sequence as shown in SEQ ID NO. 27 and a downstream primer 14 having a nucleotide sequence as shown in SEQ ID NO. 28; The 15th amplification primer includes an upstream primer 15 having a nucleotide sequence as shown in SEQ ID NO.29 and a downstream primer 15 having a nucleotide sequence as shown in SEQ ID NO.30; The 16th amplification primer includes an upstream primer 16 having a nucleotide sequence as shown in SEQ ID NO.31 and a downstream primer 16 having a nucleotide sequence as shown in SEQ ID NO.32; The 17th amplification primer includes an upstream primer 17 having a nucleotide sequence as shown in SEQ ID NO.33 and a downstream primer 17 having a nucleotide sequence as shown in SEQ ID NO.34; The 18th amplification primer includes an upstream primer 18 having a nucleotide sequence as shown in SEQ ID NO.35 and a downstream primer 18 having a nucleotide sequence as shown in SEQ ID NO.36; The 19th amplification primer includes an upstream primer 19 having a nucleotide sequence as shown in SEQ ID NO.37 and a downstream primer 19 having a nucleotide sequence as shown in SEQ ID NO.38; The 20th amplification primer includes an upstream primer 20 having a nucleotide sequence as shown in SEQ ID NO.39 and a downstream primer 20 having a nucleotide sequence as shown in SEQ ID NO.40; The 21st amplification primer includes an upstream primer 21 having a nucleotide sequence as shown in SEQ ID NO.41 and a downstream primer 21 having a nucleotide sequence as shown in SEQ ID NO.

42.

3. The primer combination according to claim 2, characterized in that The primer combination also includes an extension primer combination; The extension primer combination includes the first to twenty-first extension primers; the nucleotide sequences of the first to twenty-first extension primers are shown in SEQ ID NO.43 to SEQ ID NO.63, respectively.

4. Use of the primer combination according to claim 2 or 3 in the preparation of forensic identification products; The forensic medicine identification product is a product for forensic medicine tracing male ancestors; The males are East Asian and African male individuals; The males are male individuals of the Han and Qiang nationalities.

5. A kit for tracing male ancestors, characterized in that: The kit comprises the primer combination according to claim 2 or 3; The males are East Asian and African male individuals; The males are male individuals of the Han and Qiang nationalities.

6. Use of the molecular marker according to claim 1, the primer combination according to claim 2 or 3, or the kit according to claim 5 in family screening and / or ancestry inference of male individuals in mixed spot samples The males are East Asian and African male individuals; The males are male individuals of the Han and Qiang nationalities.

7. A method for tracing paternal ancestors, characterized in that: The steps include: Extracting genomic DNA from male samples to be tested; Using the genomic DNA as a template, PCR amplification is performed using the amplification primers in the primer combination of claim 2 or 3 to obtain an amplified product; purifying the amplified product to obtain a purified amplified product; performing a single-base extension reaction on the purified amplification product using the extension primer combination in the primer combination according to claim 3 to obtain an extension product; performing capillary electrophoresis detection on the extension product to obtain a capillary electrophoresis detection result; Importing the capillary electrophoresis test results into an ancestry inference model to obtain ancestry information of the male sample to be tested; The males are East Asian and African male individuals; The males are male individuals of the Han and Qiang nationalities.

8. The method according to claim 7, characterized in that The PCR amplification system is as follows: QIAGEN Multiplex PCR Master Mix 2.5 μL, 1 μL of 100 μM amplification primer mixture, 0.125-10 ng of DNA template, and nuclease-free water to 5 μL; The volume ratio of the first to twenty-first amplification primers in the amplification primer mixture is 5:5:1:10:5:20:5:1:10:2:1:20:30:2:20:30:2:50:5:2:3; the concentration ratio of the upstream primer to the downstream primer in each amplification primer is 1:1; The PCR amplification program is as follows: 95°C for 15 min; 94°C for 30 s, 59°C for 90 s, 72°C for 60 s, 11 cycles, with each cycle decreasing by 1°C from the second cycle; 94°C for 30 s, 49°C for 90 s, 72°C for 60 s, 25 cycles; 60°C for 30 min; and storage at 4°C.

9. The method according to claim 7, characterized in that The single base extension reaction system is: Platinum Multiplex Ready Reaction Mix 1.2 μL, 100 μM extension primer mixture 1.5 μL, purified amplification product 1 μL and nuclease-free water 1.5 μL; The volume ratio of the first to twenty-first extension primers in the extension primer mixture is 5:5:1:10:5:20:5:1:10:2:1:20:30:2:20:30:2:50:5:2:3; The procedure of the single base extension reaction is: 96° C. for 10 s, 50° C. for 5 s, 60° C. for 30 s, 26 cycles; and storage at 4° C.

Citation Information

Patent Citations

  • Forensic medicine composite examination kit based on 55 Y-chromosome SNP (Single Nucleotide Polymorphism) genetic markers

    CN108060237A

  • Primer combination for detecting 617 SNPs and InDel and application of primer combination in forensic identification and genetic relationship identification

    CN111500748A