A SNP genetic marker combination, detection method and application for human individual identification based on next-generation sequencing
Through CleanPlex ultra-high multiplex PCR technology based on second-generation sequencing, primer combinations and library building kits are designed for 30 SNP sites, which solves the problem of typing of trace samples and degraded samples in forensic science, and achieves efficient and accurate SNP typing and individual recognition.
Patent Information
- Application Number
- CN202311439213.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-11-01
AI Technical Summary
The existing capillary electrophoresis technology is difficult to effectively handle SNP typing of trace samples and degraded samples in forensic science, especially when identifying complex kinship, identifying identity of degraded samples or distinguishing mixed samples, the detection efficiency and accuracy are insufficient.
Using CleanPlex ultra-high multiplex PCR technology based on second-generation sequencing, a specific primer combination suitable for 30 SNP sites was designed and synthesized. Combined with the CleanPlex Targeted Ultra Library Kit for 2-Pool Panels library building kit, high-throughput sequencing and data analysis were performed to achieve accurate typing of a large number of SNP sites in micro samples.
It has achieved one-time detection of more genetic information in trace and degraded samples, improved the accuracy and repetition of the test, and is suitable for Chinese population, suitable for individual identification and successful classification of degraded samples in forensic science.
Smart Images

Figure CN117604119B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of gene detection, and specifically relates to a human individual identification SNP genetic marker combination based on second-generation sequencing, a detection method and an application. Background Art
[0002] Single nucleotide polymorphisms (SNPs) refer to polymorphisms in nucleic acid sequences caused by changes in a single nucleotide base. They have high distribution density, short fragments, and high genetic stability. The amplification product of a single SNP site can be controlled below 150bp, making it easy to achieve simultaneous amplification of multiple sites and conducive to the typing of degraded samples. Therefore, SNPs are also considered to be the third generation of genetic markers after STR, and have attracted great attention in the field of forensic medicine in terms of application.
[0003] At present, the most widely used DNA typing technology in the forensic industry is PCR-capillary electrophoresis (CE). Its principle is to use multi-color fluorescence composite amplification technology to type DNA by detecting the length of amplified fragments with different fluorescent groups. In order to place more sites in each color of fluorescence, some sites will retain longer flanking sequences when designing primers, resulting in long amplified fragments, which is not conducive to the typing of degraded samples. In addition, although CE technology can currently detect more than 40 STR sites in one system, the amount of sample required for one test is relatively large, and the detection of extremely small amounts of samples is limited. Therefore, the first-generation STR typing technology using capillary electrophoresis is difficult to handle such as complex kinship identification, identification of degraded samples, or distinguishing mixed samples.
[0004] The second generation sequencing technology is also called the next generation sequencing technology (NGS). In recent years, the NGS technology has gradually developed and matured and has many advantages in the field of forensics: it is not limited by fluorescent labels and can detect a large number of SNP sites in one system; it can more accurately analyze SNP typing and sequence information; it has a large detection throughput and can detect hundreds of samples in one experiment, and it is also more advantageous to analyze degraded samples.
[0005] Therefore, there is an urgent need for a system that can simultaneously detect a large number of SNPs in trace samples and a precise typing method to facilitate more accurate analysis of trace samples and degraded specimens in forensic medicine. Summary of the invention
[0006] One object of the present invention is to provide a method for detecting SNP sites based on high-throughput sequencing, which can amplify 30 SNP sites simultaneously. These specific primer pairs are synthesized using the CleanPlex ultra-high multiplex PCR technology (Paragon Genomics, USA). The locus numbers are shown in Table 1. The library construction kit used is the CleanPlex Targeted UltraLibrary Kit for 2-Pool Panels (Paragon Genomics).
[0007] Another object of the present invention is to provide the application of the above detection method for individual identification and improving the successful genotyping of degraded forensic samples; the main process of the detection method includes library construction of genomic DNA, sequencing on a machine, and obtaining genotype data of all sites through data analysis.
[0008] In order to achieve the above object of the invention, the technical solutions adopted by the present invention are as follows:
[0009] A SNP genetic marker combination for human individual identification based on next-generation sequencing, including 30 SNP sites, and the 30 SNP sites include: rs10783100, rs10839751, rs1166235, rs4773029, rs8013483, rs10852588, rs11655774, rs12605006, rs2348475, rs530913, rs4588273, rs5748311, rs9849233, rs2613019, rs271397, rs397728, rs2354159, rs1012515, rs7789598, rs6984007, rs1281350, rs3091244, rs2298556, rs3812847, rs3743842, rs941454, rs3816662, rs385780, rs356167, rs2307223.
[0010] The 30 SNP sites are respectively:
[0011]
[0012]
[0013] The present invention also provides a primer set for the SNP genetic marker combination for human individual identification based on next-generation sequencing. The primer set is used to detect the above 30 SNP sites, and the sequences of the primer set are shown as SEQ ID NO: 1-60. (Table 1):
[0014] Table 1: Primer Sequences of Human SNP Genetic Marker Combinations
[0015]
[0016]
[0017] The present invention also provides an amplification system for human individual identification, and the amplification system includes reagents for detecting at least 20 SNP sites out of the above-mentioned 30 SNP sites.
[0018] Furthermore, the reagents include detection primers and reagents required for PCR amplification; the detection primers are the above-mentioned primer sets.
[0019] Furthermore, the amplification system includes a first amplification system and a second amplification system; the first amplification system includes: 2 μL of reaction enzyme and buffer premix, 2 μL of primer mixture, a total of 6 μL of DNA sample and ddH2O; the second amplification system includes: 10 μL of TE Buffer eluate, 8 μL of secondary amplification enzyme, 2 μL of I5 index primer, 2 μL of I7 index primer, 18 μL of ddH2O.
[0020] Furthermore, the PCR reaction conditions for the first amplification are: pre-denaturation at 95°C for 10 minutes; denaturation at 98°C for 15 seconds, annealing at 60°C for 5 minutes, repeating 10 times; after the cycle ends, store at 10°C; the PCR reaction conditions for the second amplification are: pre-denaturation at 95°C for 10 minutes; denaturation at 98°C for 15 seconds, annealing at 60°C for 75 seconds, repeating 14 times; after the cycle ends, store at 10°C.
[0021] Furthermore, the product after the first amplification is purified using 1.3X AMPure XP magnetic beads and eluted with 10 μL of TE Buffer; the eluted product is digested in a metal bath at 37°C for 10 min using the digestion enzyme in the kit, and the digestion system includes: 10 μL of eluted product, 2 μL of digestion enzyme, 2 μL of digestion buffer, 6 μL of ddH2O; the digested product needs to be purified again, using 1.3X AMPure XP magnetic beads and eluted with 10 μL of TE Buffer; the product of the second amplification is purified again using 1.3X AMPure XP magnetic beads and eluted with 10 μL of ddH2O.
[0022] Furthermore, the concentration of the DNA sample used is 1 - 10 ng / ul.
[0023] Furthermore, the DNA sample is derived from blood cards, hair, nails or cigarette butts.
[0024] A method for human individual identification detection includes the following steps:
[0025] (1) Extract human DNA samples; (2) Specifically amplify DNA fragments containing 30 SNP sites, construct a DNA fragment library, prepare NGS sequencing templates, and perform NGS sequencing on the obtained DNA library; (3) Analyze the data to obtain the site depth information and genotype data of all sites.
[0026] The present invention also provides a method for human individual identification detection, including the following steps:
[0027] (1) DNA extraction
[0028] In this example, the QIAamp DNA Investigator kit (QIAGEN, Germany) was used to extract DNA samples of different types including blood cards, hair, nails, and cigarette butts. The Qubit fluorescence quantification method was used to detect the DNA concentration, and the DNA samples with higher or lower concentrations were diluted and concentrated. Finally, the concentrations of all DNA samples in this example were in the range of 1 - 10 ng / ul, which could meet the requirements of subsequent experiments.
[0029] (2) Multiplex PCR amplification and library construction
[0030] The extracted DNA samples were specifically PCR amplified using the SNP primer combination of the invention. A 10 μL multiplex amplification system was used, including: 2 μL of reaction enzyme and buffer premix, 2 μL of primer mixture, and a total of 6 μL of DNA sample + ddH2O. The PCR reaction conditions were: pre-denaturation at 95°C for 10 minutes; denaturation at 98°C for 15 seconds, annealing at 60°C for 5 minutes, repeated 10 times; after the cycle ended, stored at 10°C.
[0031] The amplified product was purified using 1.3X AMPure XP magnetic beads and eluted with 10 μL of TE Buffer.
[0032] The eluted product was digested in a metal bath at 37°C for 10 min using the digestion enzyme in the kit. The digestion system included: 10 μL of eluted product, 2 μL of digestion enzyme, 2 μL of digestion buffer, and 6 μL of ddH2O.
[0033] The digested product needed to be purified again, using 1.3X AMPure XP magnetic beads and eluted with 10 μL of TE Buffer.
[0034] The secondary elution product needs to be amplified a second time. The 40 μL amplification system includes: 10 μL of TE Buffer eluate, 8 μL of secondary amplification enzyme, 2 μL of I5 index primer, 2 μL of I7 index primer, and 18 μL of ddH2O. The PCR reaction conditions are: pre-denaturation at 95 °C for 10 minutes; denaturation at 98 °C for 15 seconds, annealing at 60 °C for 75 seconds, repeated 14 times; after the cycle ends, store at 10 °C.
[0035] The product of the second amplification is purified again with 1.3X AMPure XP magnetic beads and eluted with 10 μL of ddH2O. This product is the constructed DNA library.
[0036] Pipette 2 μL for Qubit concentration measurement and another 2 μL for 2% electrophoresis to detect the fragment size.
[0037] (3) NGS sequencing
[0038] Sequence the constructed library using Illumina Novaseq 6000 according to the instructions.
[0039] (4) Data analysis
[0040] Align the sequencing results after running off the machine with the reference genome sequence. Use the GATK software to call SNPs at the target sites of the panel for this project, and obtain the SNP typing of each site in the sample by whether a mutation is detected at the site. At the same time, use the sambamba software to count the depths of the four base types ATCG at the sites as an auxiliary reference.
[0041] The present invention also provides the application of the above primer set in the preparation of a human individual identification kit.
[0042] Beneficial effects
[0043] The present invention creatively provides a set of primer combinations and detection methods based on NGS suitable for detecting 30 autosomal SNPs. The SNP loci screened by the present invention are more suitable for the Chinese population. The detection method provided by the present invention can achieve the detection of more genetic information at one time in trace and degraded samples, providing a reliable guarantee for accurate typing, and having the advantages of good repeatability and high detection performance.
[0044] The present invention provides a SNP typing detection method suitable for the Chinese population, which can detect 30 high-resolution single nucleotide polymorphism genetic marker sites at one time. The amplified fragments are uniform and short, making the amplification efficiency of each SNP site as consistent as possible, and the amplification results are maintained in balance and effectiveness to the greatest extent. Through this method, degraded forensic samples can be identified to complete individual identification, with high sensitivity, good accuracy and repeatability, and strong identification ability. Brief Description of the Drawings
[0045] Figure 1 It is a locus detection map under different starting amounts of 10 ng, 5 ng, 1 ng, 0.5 ng, 0.25 ng, 0.125 ng, and 0.0625 ng in Example 2.
[0046] Figure 2 It is a repeatability result map among different instruments and different personnel in Example 3.
[0047] Figure 3 It is a locus detection map of degraded test samples with different starting amounts in Example 4.
[0048] Figure 4 It is a locus detection map of different test sample types in Example 5.
[0049] Figure 5 It is a locus detection map of species-specific testing in Example 6. Detailed Description of the Invention
[0050] The specific steps of the present invention are as follows:
[0051] 1) Locus screening and primer design:
[0052] Through a large amount of analysis of relevant literature and research on similar kits on the market, the polymorphism, mutation rate of SNP loci, and the corresponding relationship between some SNP loci and phenotypes were statistically analyzed, and finally 30 loci were screened. Primer pairs were synthesized using the CleanPlex ultra-high multiplex PCR technology (Paragon Genomics, USA). The locus numbers are shown in Table 1.
[0053] 2) Library construction:
[0054] Use the CleanPlex Targeted Ultra Library Kit for 2-Pool Panels (Paragon Genomics) for specific amplification and library construction.
[0055] 3) Sequencing on the machine:
[0056] The constructed library was sequenced using the Illumina Novaseq 6000 according to the instructions.
[0057] 4) Data analysis:
[0058] The sequencing results after sample extraction were aligned with the reference genome sequence. The GATK software was used to call SNPs at the target sites of the panel for this project, and the SNP genotyping of each site in the sample was obtained by detecting whether there were variations at the sites. At the same time, the sambamba software was used to count the depths of the four base types A, T, C, and G at the sites as an auxiliary reference.
[0059] To evaluate the performance of the constructed detection method, performance experiments such as accuracy, sensitivity, repeatability, analysis of degraded forensic samples, applicability of forensic sample types, and species specificity were conducted.
[0060] Example 1:
[0061] The specific genotyping of NA12878 and the reference genome at each SNP site in Example 1 is as follows:
[0062]
[0063]
[0064] Accuracy analysis
[0065] The test template was NA12878 with a starting amount of 0.5 ng of DNA, and each was repeated 3 times to test the accuracy of the kit;
[0066] For the sample NA12878, the repeatability among the 3 replicates was good. At the depth thresholds of 20X and 100X, the test results were consistent with the SNP genotyping of the existing reference genome, and the genotyping results among the 3 replicates were also consistent. The repeatability was 100% and the accuracy was 100%. The results are shown in Table 2.
[0067] Example 2: Sensitivity analysis
[0068] Using the forensic standard 9948 as the template, different gradients of starting amounts of DNA were set, namely 10 ng, 5 ng, 1 ng, 0.5 ng, 0.25 ng, 0.125 ng, and 0.0625 ng, and each was repeated 3 times.
[0069] As Figure 1 shown, at different starting amount gradients, as the starting amount decreased, the number of sites detected at each depth gradually decreased; at a data volume of 1G, when the starting amount was 0.0625 ng, all three replicates could still be detected at 100%; the higher the starting amount, the better the data quality and the site detection rate.
[0070] Example 3: Repeatability analysis
[0071] Using forensic standard 9948 as a template, starting with 0.5 ng of DNA, with 3 replicates for each; for instrument repeatability, one person used different models of PCR instruments commonly used in the laboratory to complete the library construction operation; for personnel repeatability, 2 experimental personnel used the same PCR instrument to complete the library construction operation.
[0072] As Figure 2 shown, the repeatability is good between different instruments and different personnel with CV < 20%, and the SNP site detection rate is 100% and the repeatability is 100%.
[0073] Example 4: Analysis of degraded forensic samples
[0074] To simulate real degraded forensic samples, the 9948 standard was randomly fragmented to about 300 bp using ultrasonic fragmentation method, with starting amounts of 1 ng, 0.5 ng, 0.25 ng, 0.125 ng, 0.0625 ng, and 3 replicates for each;
[0075] As Figure 3 shown, the 9948 standard was ultrasonically fragmented to simulate degraded forensic samples, with the main band at about 300 bp. At a data volume of 1 G, when the starting amount was 0.125 ng, all three replicates could be detected 100%; for different starting amounts of the fragmented samples, the number of detected sites still decreased with the decrease of the starting amount. Therefore, it is recommended to increase the starting amount for library construction when encountering degraded samples in the future.
[0076] Example 5: Analysis of the applicability of forensic sample types
[0077] Samples of different genders and different forensic sample types were tested, including hair, cigarette butts, blood cards, and nail samples from male and female, with a starting amount of 0.5 ng of DNA and 3 replicates for each;
[0078] As Figure 4 shown, the site detection rate of all types of forensic sample samples was 100%. Therefore, this product can be compatible with the detection of different types of forensic samples such as hair, cigarette butts, blood cards, and nails, and the repeatability between the same samples is 100%.
[0079] Example 6: Species specificity analysis
[0080] Using the DNA of common cats, dogs, cows, pigs, chickens, Escherichia coli, Staphylococcus aureus, and environmental soil in the detection environment as templates, with forensic standard 9948 as a control, starting with 0.5 ng of DNA and 3 replicates for each;
[0081] As Figure 5 shown, all sites could be detected for the 9948 standard, but for the DNA of non-human species, at a depth threshold of 20X, only 1 SNP was detected in 3 samples, and the rest were not detected. Therefore, this product has good specificity.
Claims
1. A SNP genetic marker combination for human individual identification based on next-generation sequencing, characterized in that, It includes 30 SNP sites, and the 30 SNP sites include: rs10783100, rs10839751, rs1166235, rs4773029, rs8013483, rs10852588, rs11655774, rs12605006, rs2348475, rs530913, rs4588273, rs5748311, rs9849233, rs2613019, rs271397, rs397728, rs2354159, rs1012515, rs7789598, rs6984007, rs1281350, rs3091244, rs2298556, rs3812847, rs3743842, rs941454, rs3816662, rs385780, rs356167 and rs2307223.
2. The primer set for the SNP genetic marker combination for human individual identification based on next-generation sequencing according to claim 1, wherein The primer set is used to detect the 30 SNP sites described in claim 1, and the sequences of the primer set are shown as SEQ ID NO: 1 to 60.
3. An amplification system for human individual identification, characterized in that, The amplification system includes reagents for detecting the 30 SNP sites described in claim 1.
4. The amplification system according to claim 3, wherein The reagents include detection primers and reagents required for PCR amplification; the detection primers are the primer set described in claim 2.
5. The amplification system according to claim 4, characterized in that, The amplification system includes a first amplification system and a second amplification system; the first amplification system includes: 2 μL of reaction enzyme and buffer premix, 2 μL of primer mixture, a total of 6 μL of DNA sample and ddH2O; the second amplification system includes: 10 μL of TE Buffer eluate, 8 μL of secondary amplification enzyme, 2 μL of I5 index primer, 2 μL of I7 index primer, 18 μL of ddH2O.
6. The amplification system according to claim 4, wherein The PCR reaction conditions for the first amplification are: pre-denaturation at 95°C for 10 minutes; denaturation at 98°C for 15 seconds, annealing at 60°C for 5 minutes, repeating 10 times; after the cycle ends, store at 10°C; the PCR reaction conditions for the second amplification are: pre-denaturation at 95°C for 10 minutes; denaturation at 98°C for 15 seconds, annealing at 60°C for 75 seconds, repeating 14 times; after the cycle ends, store at 10°C.
7. The amplification system according to claim 5, wherein The product after the first amplification is purified using 1.3X AMPure XP magnetic beads and eluted with 10 μL of TE Buffer; the eluted product is digested in a metal bath at 37°C for 10 min using the digestion enzyme in the kit, and the digestion system includes: 10 μL of eluted product, 2 μL of digestion enzyme, 2 μL of digestion buffer, 6 μL of ddH2O; the digested product needs to be purified again, using 1.3X AMPure XP magnetic beads for purification and eluting with 10 μL of TE Buffer; the product of the second amplification is purified again using 1.3X AMPure XP magnetic beads and eluted with 10 μL of ddH2O.
8. The amplification system according to claim 5, wherein The concentration of the DNA sample used is 1 - 10 ng / ul.
9. The amplification system according to claim 5, wherein The DNA sample is derived from blood cards, hair, nails or cigarette butts.
10. Use of the primer set according to claim 2 in the preparation of a human individual identification kit.