Genetic marker system containing 91 high-performance autosomal microhaplotypes as well as detection primer and application of genetic marker system

Through single-ended primer extension technology and flexible reading method, a genetic marker system containing 91 high-performance autosomal microhaplotypes was constructed, solving the typing problem of micro and degraded DNA samples, and achieving high sensitivity and accurate forensic analysis.

CN120442816APending Publication Date: 2025-08-08SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510709457.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

It is difficult to effectively analyze trace and degraded DNA samples in the prior art, especially in STR detection, where there are problems such as limited number of genetic markers, stutter interference and insufficient typing efficacy.

Method used

A genetic marker system is designed using single-ended primer extension technology, and specific primers are designed for each microhaplotype site. The general primers are bound at the breaking position of the target fragment, and amplification and typing are used to use a second-generation sequencing platform. Combined with a flexible reading method to maximize the use of sequencing data.

Benefits of technology

It has achieved stable amplification of 91 sites in trace and degraded DNA samples, improved the typing ability, reduced stutter interference, enhanced the accuracy and sensitivity of the detection results, and is suitable for individual identification and kinship identification in forensic science.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120442816A_ABST
    Figure CN120442816A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of forensic medicine identification, in particular to a genetic marker system containing 91 high-performance autosomal microhaplotypes as well as a detection primer and application of the genetic marker system. The invention provides a genetic marker system containing 91 high-performance autosomal microhaplotypes. According to the system, a single-ended primer extension technology is utilized for amplification, and the requirement for the integrity of DNA fragments is lower. The specific application of single-ended primer extension in the system is as follows: at least one specific primer is designed for each micro haplotype site, and the other end is a universal primer suitable for all sites in the composite system. The universal primer is combined at the fracture position of the target fragment, so that even if the target DNA fragment is fractured, partial amplification can still be performed and sequencing can be completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of forensic identification, and in particular to a genetic marker system comprising 91 high-efficiency autosomal microhaplotypes, detection primers thereof, and applications thereof. Background Art

[0002] Trace DNA samples are samples that contain only very small amounts of DNA. These samples may come from tiny substances at a crime scene or processed items, such as fingerprints, bloodstains, saliva stains, etc. Due to the extremely small amount of content, the extraction and analysis of these trace DNA samples usually require the use of highly sensitive DNA detection technology to ensure accurate analysis results. Degraded DNA samples are samples in which the DNA molecules have been partially or completely decomposed due to environmental factors or the passage of time. This degradation may be caused by factors such as light, temperature, humidity, or chemicals. Degraded DNA samples usually reduce the length and integrity of DNA molecules, which may pose challenges to DNA analysis.

[0003] Trace and degraded samples often contain low DNA content and may be partially degraded, making them insufficient for conventional STR analysis. Current STR-based approaches for detecting trace and degraded samples include screening for highly polymorphic STR loci, designing and optimizing specific primers, performing PCR amplification, separating the amplified products by capillary electrophoresis and performing fluorescent staining, and then analyzing the data and interpreting the results. This entire process requires strict quality control to ensure the accuracy and reliability of the results. MiniSTR technology is often used, which involves designing primers closer to the STR core repeat region to shorten the amplified fragments to account for the shorter degraded DNA fragments. However, due to the limited number of fluorescent signal pathways, the number of STR loci is limited, making it difficult to detect sufficient genetic markers in trace and degraded DNA. Therefore, when dealing with trace and degraded DNA samples, a detection system with higher sensitivity, more genetic information, and better typing capabilities is needed for analysis.

[0004] Microhaplotypes (MHs) are haplotypes formed by the combination of two or more closely linked SNPs. MHs are widely distributed across the genome, with low mutation rates and high polymorphism. They have been extensively studied for individual identification and kinship testing.

[0005] When analyzing trace and degraded DNA samples, stuttering can affect the accurate typing of STR. Stuttering is an inherent phenomenon that occurs during the PCR amplification process of STR. Specifically, it is an additional DNA fragment generated due to replication slippage, which is usually manifested as a reduction in STR repeat units. A weak signal peak can be observed before the allele peak in the STR spectrum. Researchers can minimize or correct typing errors caused by stuttering by optimizing PCR reaction conditions and using specialized data analysis software. However, these methods may result in peak loss when analyzing trace and degraded samples. MH, as a genetic marker composed of two or more SNPs, does not stutter during the typing process, and therefore can improve the accuracy and reliability of typing trace and degraded DNA samples.

[0006] When dealing with trace samples, the DNA copy number is low, and some STR loci may not be detected. At the same time, due to the limited number of STR loci compounded in the amplification system itself, insufficient typing efficiency is prone to occur. With the help of the second-generation sequencing platform, micro-haplotypes can compound a sufficient number of loci in the amplification system to compensate for the inability to detect all loci when testing trace samples, while still ensuring sufficient system efficiency. At the same time, the amplicon length of MH can be shorter than that of commonly used STR loci, which is more conducive to the analysis of trace and degraded samples. Therefore, MH has a clear advantage when dealing with trace and degraded DNA samples.

[0007] Conventional PCR is based on double-ended primer extension technology. Double-ended primer amplification uses a pair of primers, a forward primer and a reverse primer, to bind to each strand of the template DNA to amplify double-stranded DNA. This technique is suitable for use with large amounts of template DNA and can produce a large number of target fragments. However, when dealing with trace amounts of degraded samples, any breakage of the target strand can reduce amplification efficiency, resulting in fewer amplified products and lowering the detection capacity of the complex system. Summary of the Invention

[0008] The purpose of the present invention is to provide a genetic marker system comprising 91 high-efficiency autosomal microhaplotypes, its detection primers and applications. This system utilizes single-end primer extension technology for amplification, which has lower requirements for the integrity of DNA fragments. The specific application of single-end primer extension in this system is as follows: at least one specific primer is designed for each microhaplotype site, and the other end is a universal primer applicable to all sites in the composite system. The universal primer binds to the break position of the target fragment, so that even if the target DNA fragment is broken, it can still be partially amplified and sequenced (such as Figure 1 shown).

[0009] Since this system adopts single-end primer extension technology, a series of amplified fragments with different lengths but the same starting point will be generated accordingly. The MH genetic marker itself is composed of multiple SNPs. In theory, different SNP combinations can produce different MHs. Therefore, in the data analysis system, the present invention adopts a "flexible reading" method to comprehensively consider amplified fragments of different lengths for MH typing, maximizing the use of sequencing data (such as Figure 2 shown).

[0010] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:

[0011] The present invention provides a genetic marker system comprising 91 high-efficiency autosomal microhaplotypes, wherein the genetic marker system comprises loci as shown in Table 1:

[0012] Table 1

[0013]

[0014]

[0015]

[0016]

[0017]

[0018]

[0019]

[0020]

[0021] The present invention also provides a primer set for amplifying the genetic marker system, and the nucleotide sequences of the primer set are shown in SEQ ID NOs. 1 to 145.

[0022] The present invention also provides application of the primer set in forensic identification.

[0023] The present invention also provides a forensic medicine identification kit, comprising the primer set.

[0024] Preferably, the kit can be used for individual identification or kinship identification.

[0025] The present invention also provides a method for using the kit, which comprises extracting genomic DNA of a sample to be tested as a template, performing PCR amplification using the kit, performing a linker sequence PCR reaction on the obtained amplified product to obtain an amplified library, and performing quantification and second-generation sequencing detection and analysis on the amplified library to obtain the typing results of the microhaploid genome.

[0026] 1. The present invention constructs a new MH complex system that does not overlap with the currently reported MH system. After optimization, the MH system of the present invention can simultaneously amplify fragments of 91 sites in the system. The amplification primers do not interfere with each other, and typing of all sites can be stably obtained. The present invention compounds 91 MH sites in the same system, and the fragment length is within 200bp, with good polymorphism.

[0027] 2. The MH system of this invention is suitable for testing all types of forensic DNA samples, especially trace and degraded samples. It offers stable detection results and high sensitivity. Compared with conventional SNP systems, it offers better polymorphism, shorter fragment lengths, and the absence of stutter products, which is unaffected by stutter interference, thus improving typing capabilities for complex samples.

[0028] 3. The present method utilizes single-primer extension amplification technology to fully leverage the flexible and variable nature of MH, forming different MHs when different SNP numbers are detected, maximizing the genetic information provided by the specimen. It is capable of individual identification even with trace amounts of DNA, with an allele detection rate exceeding 80% even with input DNA as low as 0.25 ng. This method is suitable for testing trace amounts of forensic specimens and is also suitable for individual identification of trace amounts of degraded specimens.

[0029] 4. The single-primer extension method reduces the complexity of primer pair design and eliminates the nonspecific amplification caused by primer cross-reactions during multiplex amplification. The present invention designs multiple specific single primers for some of the 91 MH loci, which also makes the detection results more accurate and reliable. It is also suitable for massively parallel sequencing technology, with high single-shot detection efficiency and more genetic information obtained than conventional STR-CE platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a single primer extension library preparation method. GSP is a specific primer and UP is a universal primer.

[0031] Figure 2 Schematic diagram of "flexible read" analysis. Ref is the allele information of the "complete" MH, and mh1-12 are the allele types obtained from the valid sequence through "flexible read". DETAILED DESCRIPTION

[0032] The technical solution provided by the present invention is described in detail below with reference to embodiments.

[0033] Example 1 Screening sites and constructing the MH system, the specific steps are as follows: Step 1: SNPs were screened from the expanded 1000 Genomes Project (e1kGP) to ensure Hardy-Weinberg equilibrium (p > 0.05). MHs were then assembled, aligning all possible SNP combinations within a 350-bp range to form MHs. The effective allele number (Ae) was calculated based on the frequency of each MH allele in the Southern Chinese (CHS) population.

[0034] Step 2: Remove all sites containing indels from the assembled MHs. MH sites that meet the objectives of this study were further screened based on the following criteria: ① MH length ≤ 200 bp, Ae value ≥ 4 in the 1000 Genomes CHS population. ② If MHs overlapped, retain the site with the largest Ae and shortest fragment length among the overlapping MHs. ③ Hardy-Weinberg equilibrium and no linkage disequilibrium were achieved in the CHS population. ④ The sequence contained no poly or repeat structures, and the GC content was < 60%.

[0035] Step 3: Primers were designed based on the selected loci, and multiplex amplification and sequencing were performed. MH loci that could not be consistently detected, had low total sequencing reads, or had significant differences in heterozygous allele sequencing reads were eliminated. 91 loci that could be stably amplified and sequenced were retained to form the MH multiplex system. Detailed information on the 91 MH loci is shown in Table 2, and the corresponding primer sequences are shown in Table 3.

[0036] Table 2 Information of 91 MH sites screened

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045] Table 3. Primers corresponding to MH sites

[0046]

[0047]

[0048]

[0049]

[0050]

[0051]

[0052]

[0053] ("A, T, C, G, U" in the primer sequence represent the bases of nucleic acids, respectively. "A" represents adenine, "T" represents thymine, "C" represents cytosine, "G" represents guanine, and "M" represents "U", which represents uracil. "0" in the primer direction represents a forward primer, which is used to amplify the positive strand of the target sequence. "1" represents a reverse primer, which is used to amplify the antisense strand of the target sequence) as shown in SEQ ID NOs: 1 to 145.

[0054] Example 2

[0055] The specific detection method of the 91 MH sites of the present invention includes the following detailed steps:

[0056] 1. Reagents required for the method

[0057] Table 4

[0058]

[0059]

[0060] Step 1: DNA fragmentation and end repair: Prepare the Mix according to the table below. The DNA used is from Coriell Human Genomic DNA Standard (NA12878).

[0061] Table 5

[0062]

[0063] Use the PCR instrument to incubate according to the table below. When using, you can enable the hot cover and set the temperature to 65℃

[0064] Table 6

[0065]

[0066]

[0067] 1Pause when precooling to 4°C, add the prepared Mix to the PCR instrument and continue

[0068] Step 2: Connect the Primer-UMI related adapters and prepare the Mix on ice according to the table below.

[0069] Table 7

[0070]

[0071] Use the PCR instrument to incubate according to the table below. The heated cover can be enabled and the temperature can be set to 65°C.

[0072] Table 8

[0073]

[0074] 1 Pause when the mixture is just cooled to 4°C and add the prepared Mix to the PCR instrument before continuing.

[0075] Step 3: Ligation Cleanup Reagent: Add 2μl Ligation Cleanup Reagent to each sample and incubate using a PCR instrument according to the table below. Enable the heated cover when using, and the temperature can be set to 95°C.

[0076] Table 9

[0077]

[0078] 1 Pause when the mixture is just cooled to 4°C and add the prepared Mix to the PCR instrument before continuing.

[0079] Step 4: Target enrichment: Prepare the Mix on ice according to the table below.

[0080] Table 10

[0081]

[0082]

[0083] Use the PCR instrument to perform PCR as shown in the table below. The heated cover can be enabled and the temperature can be set to the default temperature (105°C)

[0084] Table 11

[0085]

[0086] Step 5: Clean up the TEPCR reagents and prepare the mix on ice according to the table below.

[0087] Table 12

[0088]

[0089] Use the PCR instrument to incubate according to the table below. Enable the hot cover when using, and the temperature can be set to the default temperature (105℃)

[0090] Table 13

[0091]

[0092] 1 Pause when the mixture is just cooled to 4°C and add the prepared Mix to the PCR instrument before continuing.

[0093] Step 6: For universal PCR, prepare the Mix on ice according to the table below.

[0094] Table 14

[0095]

[0096] Use the PCR instrument to perform PCR as shown in the table below. The heated cover can be enabled and the temperature can be set to the default temperature (105°C)

[0097] Table 15

[0098]

[0099] Step 7: Purify the PCR product using magnetic beads to remove unligated adapters, primers, and other impurities.

[0100] Step 8: Quantify the product based on Qubit to ensure that the quality and concentration of the library meet the sequencing requirements.

[0101] Step 9: Sequencing on the Illumina platform.

[0102] Step 10: The raw data files generated from Illumina sequencing were converted into short reads using base calling and recorded in FASTQ format, which contains sequence information and sequencing quality information. Fastp (version 0.23.1) was then used to perform quality control on the raw data to remove poor-quality data, ensuring the accuracy and reliability of subsequent analyses. This included removing adapter contamination, excessive numbers of uncertain bases, and data with a high proportion of low-quality bases. Finally, a Python script developed in the laboratory was used to extract the genotype for each MH locus from the data.

[0103] Example 3 Sensitivity Verification Results of 91 MH Composite Systems in the Present Invention

[0104] The NA12878 DNA standard samples with template amounts of 2 ng, 1 ng, 0.5 ng, 0.25 ng, 0.125 ng, 0.0625 ng, and 0.03125 ng were analyzed using the above-mentioned specific detection method. Each template amount was repeated three times. Comparison of the complete read and flexible read analysis methods:

[0105] The complete read results showed that when the sample template amount was 2 ng, the number of loci dropouts in the three replicates was 39, 44, and 35, respectively; the number of allele dropouts in the three replicates was 7, 14, and 7, respectively. When the sample template amount was 1 ng, the number of loci dropouts in the three replicates was 41, 52, and 42, respectively; the number of allele dropouts in the three replicates was 10, 15, and 10, respectively. The flexible read results showed that when the sample template amount was 2 ng, the number of loci dropouts in the three replicates was 0, 1, and 0, respectively; the number of allele dropouts in the three replicates was 2, 4, and 1, respectively. When the sample template amount was 1 ng, the number of loci dropouts in the three replicates was 0, 0, and 1, respectively; the number of allele dropouts in the three replicates was 2, 3, and 2, respectively. All loci experiencing allele dropout were heterozygous alleles. The results show that the flexible read strategy can improve the detection rate in samples with different template amounts. The detection rates for samples with different template amounts are shown in Table 16 below:

[0106] Table 16. Detection rate of MH system at different template amounts

[0107]

[0108] Detection rate = (total number of loci × 2 - total number of dropout alleles in three replicate samples at that concentration) / (total number of loci × 2)

[0109] Example 4 Detection results of DNA samples with different degradation degrees using the MH system of the present invention

[0110] NA12878 standard DNA was heated for different periods of time to simulate the degradation of forensic DNA samples. When the samples were heated for 60, 120, and 180 minutes, the degradation index increased to 2.98, 4.81, and 6.82, respectively. DNA samples with different degradation indices were typed using the aforementioned method, with three replicates for each degradation index. The allele detection rates for samples with different degradation levels are shown in Table 17 below.

[0111] Table 17. Detection rates of MH system at different degradation levels

[0112] Degradation index 6.82 4.81 2.98 Complete read call rate 0.55% 1.28% 1.65% Flexible read detection rate 36.08% 44.14% 53.66%

[0113] It can be found from the table that the larger the degradation index, the smaller the detection rate of complete read and flexible read, and the detection rate successfully reflects the degree of degradation of the fragment, which shows that the single primer extension strategy used by the MH system of the present invention is effective, and the number of alleles detected by templates of different fragment lengths shows differences. At the same time, under each degradation index, the detection rate of flexible read is higher than that of complete read, which shows that the single primer extension combined with flexible read sequencing used by the MH system of the present invention has played a role and maximized the detection ability of degraded samples. When the degradation index is 6.82, 36.08% of the alleles can still be detected using flexible read, which also shows that the MH system of the present invention has the ability to detect degraded samples and can be applied to forensic degraded samples. The above is only a preferred embodiment of the present invention.

Claims

1. A genetic marker system comprising 91 high-performance autosomal microhaplotypes, characterized in that: The genetic marker system includes sites as shown in Table 1: Table 1 2. A primer set for amplifying the genetic marker system according to claim 1, characterized in that: The nucleotide sequences of the primer set are shown in SEQ ID NOs. 1 to 145.

3. Use of the primer set according to claim 2 in forensic identification.

4. A kit for forensic identification, characterized in that: The method comprises the primer set according to claim 2.

5. The kit according to claim 4, characterized in that The kit can be used for individual identification or kinship identification.

6. The method for using the kit according to claim 4 or 5, characterized in that: The genomic DNA of the sample to be tested is extracted as a template, and PCR amplification is performed using the kit. The obtained amplified product is then subjected to a linker sequence PCR reaction to obtain an amplified library. The amplified library is quantified and analyzed by second-generation sequencing to obtain the typing results of the microhaplotype genome.