Human microhaplotype genetic marker composition, kit and application thereof
By developing a micro-haplotype genetic marker composition containing 358 nucleic acid fragments, the shortcomings of existing micro-haplotypes in forensic applications have been addressed, enabling efficient individual identification and kinship determination, and improving the sensitivity and accuracy of forensic testing.
Patent Information
- Application Number
- CN202410935194.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2026-01-13
AI Technical Summary
Existing microhaplotype genetic markers have several drawbacks in forensic applications, including a limited number of loci, poor polymorphism in the Chinese population, lack of commercialization, and limited practical applications, making it difficult to meet the needs of individual identification and kinship determination.
A human microhaplotype genetic marker composition was developed, containing 358 nucleic acid fragments ranging from 60 to 400 bp in length, covering 353 microhaplotypes and 5 sex identification loci. Sequencing libraries were constructed using multiplex PCR amplification and high-throughput sequencing technologies to achieve efficient individual identification and kinship determination.
It improves the identification rate and polymorphism of groups and individuals, enhances the sensitivity and accuracy of forensic testing, can accurately detect complex mixed DNA and trace samples, provides clues to difficult cases, and increases the possibility of solving forensic cases.
Smart Images

Figure CN121320547A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human gene detection technology, and in particular to a human microhaplotype genetic marker composition, kit, and its application. Background Technology
[0002] Forensic DNA technology is crucial in criminal investigations, identification of disaster victims, identification of missing persons, combating trafficking in women and children, and civil paternity testing. In forensic research, short tandem repeats (STRs) are the most widely used. STRs are a class of DNA repetitive sequences widely found in the human genome, with a core sequence of 2-6 base repeats. They offer advantages such as high practicality, low mutation rate, good genetic polymorphism, and ease of detection, and are widely used in individual identification, paternity testing, and population genetics research. However, STR-based detection techniques have certain limitations. The main limitation is that, due to their multiple tandem repeats, slippage can easily occur during amplification, leading to stuttering peaks that affect sample peak reading, especially in mixed sample analysis where interference is more pronounced.
[0003] Microhaplotypes (MHs) are novel genetic markers discovered in recent years and have attracted widespread attention in the international forensic genetics community. Specifically, they refer to combinations containing two or more SNPs within a relatively short fragment. The amplicon fragment length is generally tens to hundreds of base pairs. Because each microhaplotype locus contains multiple SNP sites, microhaplotypes are also known as multi-allelic genetic markers, containing richer genetic information. They combine the advantages of STRs and SNPs, effectively avoiding the influence of shadow peaks, and show great potential in mixed DNA typing, better assisting in individual identification, kinship determination, and the detection of complex and mixed cases.
[0004] There are few publicly available microhaplotypes and related studies, and they generally suffer from the following shortcomings: 1. The number of loci included is small, resulting in insufficient efficacy for forensic applications; 2. Existing loci do not show good polymorphism in the Chinese population, making it difficult to meet the needs of forensic applications; 3. Similar genetic markers are only in the scientific and technological research stage and have not been commercialized or made applicable, resulting in a lack of practical applications.
[0005] In conclusion, developing human micro-haplotype genetic markers with stronger individual identification, kinship identification, and forensic application value remains a crucial technical challenge that forensic medicine urgently needs to address. Summary of the Invention
[0006] The purpose of this application is to provide a novel genetic marker composition for human microhaplotypes, a genetic marker detection kit for human microhaplotypes based on this genetic marker composition, and its application.
[0007] To achieve the above objectives, this application adopts the following technical solution:
[0008] The first aspect of this application discloses a genetic marker composition for human microhaplotypes, which consists of 358 nucleic acid fragments with a length between 60 and 400 bp. The 358 nucleic acid fragments consist of 353 microhaplotypes on 21 autosomes, the Amel_X microhaplotype and the Amel_Y microhaplotype at the Amelogenin locus, and 3 microhaplotypes at the Yindel locus. The genetic marker information of the microhaplotypes is shown in Table 1.
[0009] It should be noted that the human microhaplotype genetic marker composition of this application can simultaneously analyze 353 microhaplotype genetic markers, exhibiting advantages such as high population-individual identification rate, rich polymorphism, and good stability, and can also be used for auxiliary sex identification. This human microhaplotype genetic marker composition can effectively assist in solving related forensic problems and is used in forensic examinations such as forensic identification, individual identification, and kinship identification. Compared with existing microhaplotype genetic markers and their detection technologies, this application has advantages such as more loci, higher sensitivity, more accurate detection results, and wider applicability. This human microhaplotype genetic marker composition can detect more loci at once, which is more helpful for individual identification of human biological samples. It can be better used for the analysis and degradation of complex mixed DNA or the detection of trace samples, better solving difficult and complex problems in the detection of human biological samples, providing clues for solving difficult cases, and assisting in the solving of forensic cases. In summary, the human microhaplotype genetic marker composition of this application, and the forensic identification, individual identification, or kinship identification based on the genetic marker composition, are far ahead of the existing technology level and bring more possibilities for technological innovation.
[0010] It should also be noted that the human microhaplotype genetic marker composition of this application can be in the form of 358 nucleic acid fragments in physical form, used as a positive control or standard for forensic identification, individual identification, or kinship identification. The human microhaplotype genetic marker composition of this application can also be the nucleic acid sequence of 358 nucleic acid fragments recorded in a computer-readable carrier or the nucleic acid sequence of 358 nucleic acid fragments existing in a database form, serving as standard reference data for comparing the sequencing results of the sample to be tested. In practical applications, the sequencing results of the sample to be tested do not need to be compared with the massive data of the human reference genome; they only need to be directly compared with the nucleic acid sequence of the 358 nucleic acid fragments of this application, greatly improving the efficiency and quality of sequencing result analysis. In a specific human microhaplotype genetic marker detection kit, the 358 nucleic acid fragments of the human microhaplotype genetic marker composition of this application can be contained as a positive control or standard, and simultaneously contain the nucleic acid sequence of the 358 nucleic acid fragments that can be recognized and used by the comparison software as standard reference data.
[0011] In one implementation of this application, the 358 nucleic acid fragments of the genetic marker composition are the fragments determined by the forward and reverse primers shown in Table 2.
[0012] It should be noted that the sequences of the PCR amplification products of the paired forward and reverse primers are determined; therefore, the 358 nucleic acid fragments of this application can be determined using forward and reverse primers. The key to the human microhaplotypes used as genetic marker combinations in this application lies in the presence of two or more single nucleotide polymorphism (SNP) sites with linkage disequilibrium. The fragments determined by the forward and reverse primers shown in Table 2 are only one specific microhaplotype used in one implementation of this application. In principle, any microhaplotype covering the genetic marker information shown in Table 1 of this application can be applied to this application, not limited to the fragments determined by the forward and reverse primers shown in Table 2. Alternatively, based on the fragments determined by the forward and reverse primers shown in Table 2, without changing the genetic marker information shown in Table 1, several bases can be added or removed at the 5' and / or 3' ends to form new human microhaplotype genetic marker combinations.
[0013] It should also be noted that the microhaplotype containing two or more SNP sites refers to the 21 autosomes; for the Yindel locus of this application, and Amel_X and Amel_Y of the Amelogenin locus, as long as they are specific and can be combined with the microhaplotype of this application and its amplification primers, they are acceptable.
[0014] The second aspect of this application discloses the use of the genetic marker composition of this application in the preparation of reagents for forensic identification, individual identification, or kinship identification.
[0015] It should be noted that the reagents for forensic identification, individual identification, or kinship identification based on the genetic marker composition of this application are actually primers and / or probes designed based on the 358 nucleic acid fragments of this application, capable of detecting the genetic marker information shown in Table 1 of this application. Forensic identification includes, for example, case investigation and identification of the source of a deceased person.
[0016] Therefore, in one implementation of this application, the reagents for forensic identification, individual identification, or kinship identification include primers and / or probes for detecting genetic marker information of 358 microhaplotypes.
[0017] Preferably, the reagents for forensic identification, individual identification, or kinship identification are the amplification primers for the 358 micro-haplotypes shown in Table 2.
[0018] It is understood that the amplification primers for the 358 microhaplotypes shown in Table 2 are only the amplification primers specifically used in one implementation of this application. In principle, any primer that can amplify the microhaplotypes of this application and whose amplified fragments cover the genetic marker information shown in Table 1 can be used in this application. Of course, it is best to also be able to achieve multiplex PCR amplification of the 358 microhaplotypes, or to divide the 358 microhaplotypes into two or more groups and perform multiplex amplification separately.
[0019] In one implementation of this application, kinship identification includes parentage identification, full sibling identification, and two-level kinship identification involving only two individuals.
[0020] In one implementation of this application, the human microhaplotype genetic marker composition of this application can be used in conjunction with STR genetic markers for kinship identification.
[0021] It should be noted that combining the human micro-haplotype genetic marker combination of this application with STR genetic markers for complex kinship determination can improve the ability to infer complex kinship and increase the accuracy of the inference results. The STR genetic markers can be, for example, combinations of 30 common autosomal STR genetic markers, such as the STR genetic marker combinations detected by the BGI Genomics IDentifier DNA typing kit (Yanhuang 34).
[0022] A third aspect of this application discloses a human microhaplotype genetic marker detection kit, which includes a primer combination for detecting the genetic marker composition of this application.
[0023] In one implementation of this application, the primer combination of the kit is the 358 micro-haplotype amplification primers shown in Table 2.
[0024] In one implementation of this application, the kit further includes PCR amplification reagents and enzymes, as well as PCR amplification product purification reagents. The PCR amplification product purification reagents may include, for example, magnetic beads used for DNA purification.
[0025] In one implementation of this application, the kit further includes reagents for detecting STR genetic markers.
[0026] It should be noted that the kit of this application can be supplemented with reagents for detecting STR genetic markers as needed, or reagents for detecting STR genetic markers can be purchased separately, such as the BGI Genomics IDentifier DNA Genotyping Kit (Yanhuang 34), and used in conjunction with the kit of this application to improve the ability to infer complex kinship and the accuracy of the inference results.
[0027] The fourth aspect of this application discloses a method for detecting microhaplotypes based on high-throughput sequencing technology, comprising constructing a sequencing library using the genetic marker composition of this application. For example, the sample to be tested is amplified by PCR using amplification primers for the 358 microhaplotypes shown in Table 2, and then a sequencing library is constructed based on this, followed by detection of microhaplotypes using high-throughput sequencing technology.
[0028] The beneficial effects of this application are as follows:
[0029] The human microhaplotype genetic marker composition of this application can simultaneously analyze 353 microhaplotype genetic markers, exhibiting advantages such as high population-individual identification rate, rich polymorphism, and good stability; furthermore, it can also be used to assist in sex determination. This human microhaplotype genetic marker composition can be used in forensic examinations such as forensic identification, individual identification, and kinship determination, offering advantages such as a larger number of loci, higher sensitivity, more accurate detection results, and wider applicability. Moreover, because it can detect more loci at once, it facilitates individual identification of human biological samples, better enabling the analysis and detection of complex mixed DNA and degradation or trace samples, better solving difficult and complex problems in human biological sample testing, providing clues for solving difficult cases, and offering a new solution and approach for forensic testing. Attached Figure Description
[0030] Figure 1 These are the statistical results of micro-haplotypes in the embodiments of this application;
[0031] Figure 2 This is a statistical chart showing the uniformity of standard sample 9948 in the embodiments of this application;
[0032] Figure 3 This is a statistical chart showing the differences in uniformity among different samples in the embodiments of this application;
[0033] Figure 4This is a sensitivity statistics chart of the 9948 standard sample in the embodiments of this application;
[0034] Figure 5 This is a graph showing the results of the anti-inhibition ability test in the embodiments of this application;
[0035] Figure 6 These are pedigree charts of 12 family samples in the embodiments of this application;
[0036] Figure 7 This is a diagram showing the results of inferring kinship through combinations of micro-haploid genetic markers in the embodiments of this application;
[0037] Figure 8 This is a diagram showing the ancestry inference results in an embodiment of this application. Detailed Implementation
[0038] Microhaplotypes (MH) combine the advantages of both STR and SNP genetic markers, effectively avoiding the influence of shadow peaks and demonstrating great potential in mixed DNA typing. They can better assist in individual identification, kinship determination, and the detection of complex and mixed cases. However, existing microhaplotype-based detection technologies have the following drawbacks:
[0039] 1. It contains a small number of sites, resulting in insufficient effectiveness in forensic applications.
[0040] 2. Existing loci do not show good polymorphism in the Chinese population.
[0041] 3. Similar genetic markers are only in the scientific and technological research stage, and have not been commercialized or made applicable, resulting in a lack of practical applications.
[0042] To address the above issues, this application presents a novel genetic marker composition for human microhaplotypes. This composition comprises 358 nucleic acid fragments with lengths between 60 and 400 bp. These 358 fragments consist of 353 microhaplotypes from 21 autosomes, the Amel_X and Amel_Y microhaplotypes at the Amelogenin locus, and 3 microhaplotypes at the Yindel locus. For example, Table 1 shows the genetic marker information for these microhaplotypes.
[0043] The human microhaplotype genetic marker composition of this application can simultaneously analyze 353 microhaplotype genetic markers, one Amel sex determination locus (including Amel_X microhaplotype and Amel_Y microhaplotype), and three Yindel loci for auxiliary sex determination. The human microhaplotype genetic marker composition of this application has advantages such as high population-individual recognition rate, rich polymorphism, good stability, and high individual recognition efficiency, which can better assist in the inference of complex kinship relationships and can also be used for auxiliary sex determination. The genetic marker combination of this application can effectively assist in solving related forensic problems and can be used for forensic testing such as individual identification and kinship determination. Compared with existing micro-haplotype genetic markers and their detection kits, this application has advantages such as more loci, higher sensitivity, stronger anti-inhibition ability, better species specificity, more accurate detection results, and wider applicability. It can detect more loci at once, which is more helpful for individual identification of human biological samples. At the same time, it demonstrates application value in difficult and complex problems such as complex mixed DNA analysis, degradation or trace sample detection, provides clues for solving difficult cases, and helps forensic cases.
[0044] This application is the first to propose a genetic marker combination containing 353 microhaplotype loci. This combination exhibits high individual recognition rate, rich polymorphism, and good stability, resulting in higher efficacy for individual identification and kinship determination. This application is also the first to establish a multiplex amplification system containing 358 loci and the first to commercialize microhaplotype genetic markers, creating a microhaplotype reagent kit with the largest number of loci, achieving a significant technological breakthrough in system establishment. The genetic marker combination in this application includes 5 sex identification loci, and compared to other microhaplotype products and systems, the corresponding system in this application can perform sex identification more accurately. In one implementation of this application, microhaplotype genetic markers are used in conjunction with STR genetic markers for the first time to determine complex kinship, further improving the accuracy of complex kinship inference. The genetic marker combination in this application has strong forensic efficacy, exhibiting higher individual identification and kinship determination capabilities compared to similar amplification systems and products. The multiplex amplification system based on the genetic marker composition of this application exhibits good stability and high sensitivity. Even with sample amounts as low as 62.5 pg, it can accurately detect over 99.5% of loci, better assisting in the detection of low-concentration samples in cases. It also demonstrates good species specificity and resistance to inhibition, effectively detecting more effective typing in special case samples. The detection kit based on the genetic marker composition of this application has no amplification bias or stutter peak influence, better assisting in the differentiation of mixed samples. The genetic marker composition of this application can better assist in kinship inference and ancestry tracing, combining ancestry inference with micro-haplotypes, which is more beneficial for forensic applications. The detection method based on the genetic marker composition of this application uses a two-step PCR amplification library construction method, greatly reducing the library construction operation steps and enhancing operability and applicability.
[0045] It is understood that, based on the human microhaplotype genetic marker composition of this application, a panel-like system can be formed by adding or removing microhaplotype loci (including SNPs), thereby achieving similar effects. The primer sequences for amplifying microhaplotypes in this application are all self-designed, and the resulting primer system is stable and accurate, as shown in Table 2. If the selected locus regions are the same, the primers and preparation ratios of this application can be used directly.
[0046] The present application will be further described in detail below through specific embodiments. The following embodiments are only for further illustration of the present application and should not be construed as limiting the present application.
[0047] Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.
[0048] Example 1 Site Screening
[0049] This example utilizes Vcftools to extract SNP genotyping data from the Asian population in the 1000 Genomes Study. Stable genetic regions of microhaplotypes containing two or more variant sites (SNPs) within 400 bp with a theoretical effective allele count (Ae) ≥ 2.5 are extracted. Sites are selected based on the following criteria: ① Allele frequency (MAF) of all microhaplotype SNPs in the 1000 Genomes Study > 0.1; ② Theoretical Ae value ≥ 2.5; ③ Each microhaplotype site must contain two or more SNP sites; ④ The physical location between microhaplotype sites on the same chromosome must be ≥ 0.1 Mb to avoid linkage disequilibrium between sites.
[0050] Statistical results are as follows Figure 1 As shown, this example ultimately screened out 353 microhaplotype sites that met the above conditions and had an Ae value ≥ 2.69.
[0051] Example 2: Establishment of a multiplex amplification system
[0052] For the sites screened in Example 1, mhAmel_X and mhAmel_Y of the Amelogenin locus and three SNP sites of the Yindel locus (rs771783753, rs2032678, and rs759551978) were added. Primers were designed for each site based on the hg38 reference sequence, with 2-3 pairs of specific primers designed for each site. The amplicon length was between 60-400 bp, the average annealing temperature was 59℃, and the SNP site coverage reached 100%. Primer evaluation software was used to evaluate primer dimers and non-specificity. Primers that produced a large number of dimers were replaced. For primers that passed the evaluation, single-pair primer tests were performed first to screen primers with good amplification effect for single-pair tests and enzymes with good multiplex amplification effect, and a multiplex amplification system was established.
[0053] Taking into account amplification efficiency, detection rate, and stability, the multiplex amplification system and reaction conditions were determined. The single primers were mixed to form a primer mix with a final concentration of 0.1-0.5 μM, and the sequencing library was constructed by a two-step PCR method.
[0054] This example ultimately yielded 358 microhaplotypes, consisting of 353 microhaplotypes from 21 autosomes, the Amel_X and Amel_Y microhaplotypes at the Amelogenin locus, and 3 microhaplotypes at the Yindel locus. The genetic marker information for these microhaplotypes is shown in Table 1. The forward and reverse primers for these 358 microhaplotypes are shown in Table 2. The Amel_X and Amel_Y microhaplotypes share a single primer pair.
[0055] Table 1. Genetic marker information for 358 microhaplotypes
[0056]
[0057]
[0058]
[0059]
[0060]
[0061]
[0062]
[0063]
[0064]
[0065] Table 2. Amplification primers for 358 micro-haplotypes
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075] In Tables 1 and 2, the microhaplotype numbers are as follows: in the number mh01FGI-001, mh represents the microhaplotype genetic marker, 01 represents chromosome 1, and FGI-001 represents the microhaplotype number. In Table 2, following the numbering order, the forward primers are sequenced from Seq ID No. 1 to Seq ID No. 357, and the reverse primers are sequenced from Seq ID No. 358 to Seq ID No. 714. Seq ID No. 1 and Seq ID No. 358 form a primer pair for amplifying microhaplotype mh01FGI-001; Seq ID No. 2 and Seq ID No. 359 form a primer pair for amplifying microhaplotype mh01FGI-002; and so on. Seq ID No. 357 and Seq ID No. 714 form a primer pair for amplifying mhAmel_X and mhAmel_Y.
[0076] This example demonstrates a two-step PCR method for constructing sequencing libraries. Specifically, the first step, PCR, involves preparing a reaction premix in a 25 μL system, including 2 ng DNA, 10 μL MH PCR Buffer (purchased from Nanjing Novizan, catalog number PM201-01), 2 μL L HPrimer Mix, 1 μL L Taq Enzyme (purchased from Nanjing Novizan, catalog number PM201-01), and adding water to a final volume of 25 μL.
[0077] The reaction conditions were 95℃ for 5 min, followed by 18 cycles: 95℃ for 15 s, 59℃ for 3 min, 70℃ for 1 min, and after the cycle was completed, 72℃ for 10 min, and then standby at 4℃.
[0078] The PCR product obtained from amplification was purified using magnetic beads (purchased from Nanjing Novizan, catalog number N411-03) at a ratio of 1.5×. The product was then dissolved in 23 μL of water, and 21 μL was recovered.
[0079] Step 2 PCR: Prepare a reaction premix in a 50 μL system, consisting of 21 μL of the first step PCR product, 25 μL of LMHPCR Master Mix, and 4 μL of Barcode Primer (synthesized by BGI Genomics).
[0080] The reaction conditions were: 95℃ for 3 min, followed by 12 cycles: 98℃ for 10 s, 62℃ for 1 min, 70℃ for 1 min, and after the cycle was completed, 72℃ for 10 min, and then standby at 4℃.
[0081] The PCR products were purified using magnetic beads (purchased from Nanjing Novizan, catalog number N411-03) at a ratio of 1.5×. The purified product was then used as a DNA library, quantified using Qubit, and temporarily stored at -20℃ for later use.
[0082] Example 3: Detection of micro-haplotype sites based on high-throughput sequencing technology
[0083] The libraries obtained in Example 2 were homogenized and pooled according to the same mass ratio. N libraries were homogenized and pooled, with a single library sample size (ng) of 350 ng / N and a single library sample volume (μL) equal to the single library sample size (ng) / the single library concentration (ng / μL). If the sample concentration was too high, to reduce sampling error, the homogenization process could be scaled up by X times from 350 ng. Finally, 350 ng was used for subsequent circularization. After pooling, 350 ng was used for DNA single-circularization and enzyme digestion according to the BGI Genomics circularization kit (purchased from BGI Genomics, catalog number 1000020570). The resulting single-stranded ssDNA was sequenced using the BGI Genomics DNBSEQ-G99RS SE400 sequencer. The DNB reagent was further replicated in rolling circles to prepare DNB nanospheres, which were then loaded onto a sequencing chip and sequenced using DNBseq-G99RS for SE400 sequencing. The sequencing time was approximately 20 hours. After sequencing, an FQ file was generated, and the data was analyzed using an automated analysis workflow.
[0084] In this example, 358 sites with good site polymorphism, good primer specificity, and good detection stability were finally selected to form a panel system. The site information is shown in Table 1 and Table 2.
[0085] Example 4: Balance, Sensitivity, Accuracy
[0086] Following the experimental methods of Examples 2 and 3, 9948 standard DNA was first prepared into test samples with concentrations of 2 ng / μL, 1 ng / μL, 0.5 ng / μL, 0.25 ng / μL, 0.125 ng / μL, 0.1 ng / μL, and 0.0625 ng / μL. According to the reaction system and conditions in Example 2, sequencing libraries were prepared using a two-step PCR method. The prepared libraries were quantified using Qubit. Different barcode libraries were mixed in equal mass ratios, and 350 ng of the mixed product was used for circularization and DNB preparation. Both the circularization and DNB preparation reagents were purchased from BGI Genomics (sequencing reagents were purchased from BGI Genomics, catalog number 940-000417-00). Base sequence reading was performed using the BGI Genomics DNBSEQ-G99 sequencing platform, and the site detection uniformity and site detection at low concentrations were verified through three replicate experiments. The statistical results of the uniformity of 9948 standard at the same concentration are shown below. Figure 2 As shown, the statistical results of the uniformity differences among samples of different concentrations are as follows: Figure 3 As shown in the figure, the detection rate statistics of 9948 standard at different dilution gradients are as follows: Figure 4 As shown.
[0087] Figures 2 to 4 The results show that when 1M Reads are truncated, the fractal depth at each point ranges from 397× to 8171×, with an average depth of 2548×. Specific equilibrium statistics are as follows: Figure 2 and Figure 3 As shown; even with a DNA input as low as 100 pg, it can still accurately detect genotyping at 358 loci, and the genotyping accuracy is high, as shown in the figure. Figure 4 As shown.
[0088] Example 5: Analysis of Mixed Samples
[0089] Based on the reaction conditions and system of Example 2 and the sequencing strategy of Example 3, mixtures were prepared using common DNA standards 9948 and 9947A. The genotypes of all DNA standards were known. 9948 and 9947A were mixed in weight ratios of 1:19, 2:18, 4:16, 8:12, 10:10, 12:8, 16:4, 18:2, and 19:1, respectively, to prepare mixed samples for two individuals. 2 ng of each mixed sample was used for library construction, sequencing, and analysis.
[0090] The results showed that even when the proportion of minor contributors was as low as 5% (1:19), over 99% of the minor contributor genotypes could still be accurately detected. Compared to STR genetic markers, the microhaplotypes in this application are not affected by the Stutter peak, allowing for more accurate detection of minor contributor genotypes.
[0091] Example 6 Species Specificity
[0092] Based on the reaction conditions and system of Example 2 and the sequencing strategy of Example 3, DNA samples from 12 common animals and microorganisms, including fish, horses, dogs, pigs, sheep, chickens, cats, ducks, rats, cattle, rabbits, and Escherichia coli, were selected and tested. The DNA samples from the 12 animals and E. coli were provided and preserved by Shenzhen BGI Forensic Technology Co., Ltd.
[0093] The results showed that when 5 ng of species DNA sample was added, the sequencing depth of all 358 sites was below 50×, with more than 91% of the sites having a depth below 10× or no reads. In contrast, no effective genotyping was detected in human DNA samples within the 100× threshold range. These results indicate that the genetic marker composition and its amplification primers of this application have good specificity for human DNA.
[0094] Example 7 Anti-inhibition ability
[0095] Based on the reaction conditions and system of Example 2 and the sequencing strategy of Example 3, 2 ng of 9948 standard DNA samples were added to test the detection rate in the presence of each inhibitor. Specifically, in the reaction system of Example 2, 2 ng of standard DNA sample was added, and heme was added to the reaction system at final concentrations of 20 μM, 40 μM, 80 μM, 160 μM, 320 μM, and 640 μM, respectively. An experiment without heme was set up as a control to test the inhibitory effect of heme on the detection system. In the reaction system of Example 2, 2 ng of standard DNA sample was added, and humic acid was added to the reaction system at final concentrations of 10 ng / μL, 20 ng / μL, 40 ng / μL, 80 ng / μL, and 100 ng / μL, respectively. An experiment without humic acid was set up as a control to test the inhibitory effect of humic acid on the detection system. In the reaction system of Example 2, 2 ng of standard DNA sample was added. Tannic acid was also added to the reaction system at final concentrations of 50 ng / μL, 100 ng / μL, 200 ng / μL, 300 ng / μL, and 400 ng / μL, respectively. An experiment without added tannic acid was set up as a control to test the inhibitory effect of tannic acid on the detection system. The results are as follows: Figure 5 As shown.
[0096] The results showed that all sites could still be detected with 100% accuracy even in the presence of 640 μM heme, 100 ng / μL humic acid, and 400 ng / μL tannic acid. This indicates that the system has strong resistance to inhibition of the samples.
[0097] Example 8: Kinship Inference
[0098] This embodiment, based on the genetic marker combinations selected in Examples 1, 2, and 3, further amplifies, constructs libraries, sequences, and genotypes DNA samples from 12 families with known kinship (open recruitment, ethics number BGI-IRB 23092) to test the ability of these genetic marker combinations to determine kinship. The pedigrees of the 12 family samples are shown below. Figure 6 As shown.
[0099] This example analyzes the results of kinship testing using a combination of human microhaplotype genetic markers, such as... Figure 7 As shown, the analysis results indicate that the human microhaplotype genetic marker combination of this application can accurately analyze and determine first- and second-degree kinship, including parent-child, identical twins, full siblings, half siblings, grandparents and grandchildren, uncles and nephews, with an accuracy rate of 100% and a sensitivity of 100%; for third-degree kinship, the accuracy rate is 19.04% and the sensitivity is 100%.
[0100] Example 9: Ancestry Deduction
[0101] This embodiment, based on the genetic marker combinations selected in Examples 1, 2, and 3, further performs cluster analysis on samples from different populations in the publicly available 1000 Genomes Database. The analysis method involves obtaining genotyping information of the microhaplotypes in this application, containing all SNP loci, and using Structure and Admixture software to perform population structure analysis based on the genotyping information. The results are as follows: Figure 8 As shown.
[0102] Figure 8 The results show that the genetic marker combination of this application can distinguish the three major population groups in the world well: Africa, Europe and Asia, and can basically distinguish the five major population groups in the world: Africa, Europe, South Asia, East Asia and the Americas, and can roughly distinguish the five population groups in East Asia: Northern Han, Southern Han, Dai, Jing and Japanese.
[0103] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. Those skilled in the art to which this application pertains can make several simple deductions or substitutions without departing from the concept of this application.
Claims
1. A genetic marker composition for human microhaplotypes, characterized in that: The genetic marker composition consists of 358 nucleic acid fragments with a length between 60 and 400 bp. These 358 nucleic acid fragments are composed of 353 microhaplotypes on 21 autosomes, the Amel_X microhaplotype and the Amel_Y microhaplotype at the Amelogenin locus, and 3 microhaplotypes at the Yindel locus. The genetic marker information of the microhaplotypes is shown in Table 1.
2. The genetic marker composition according to claim 1, characterized in that: The 358 nucleic acid fragments in the genetic marker composition are the fragments identified by the forward and reverse primers shown in Table 2.
3. The use of the genetic marker composition according to claim 1 or 2 in the preparation of reagents for forensic identification, individual identification or kinship identification.
4. The application according to claim 3, characterized in that: The reagents used for forensic identification, individual identification, or kinship identification include primers and / or probes that detect genetic marker information of 358 microhaplotypes; Preferably, the reagents for forensic identification, individual identification, or kinship identification are the amplification primers for the 358 micro-haplotypes shown in Table 2.
5. The application according to claim 3 or 4, characterized in that: The kinship identification includes parentage testing, full sibling testing, and two-level kinship testing involving only two individuals; Preferably, the application also includes using it in conjunction with STR genetic markers for kinship identification.
6. The application according to claim 3 or 4, characterized in that: The forensic examination includes case investigation and identification of the source of the body.
7. A genetic marker detection kit for human microhaplotypes, characterized in that: Includes primer combinations for detecting the genetic marker composition of claim 1 or 2.
8. The reagent kit according to claim 7, characterized in that: The primer combination consists of 358 micro-haplotype amplification primers as shown in Table 2.
9. The kit according to claim 7 or 8, characterized in that: It also includes PCR amplification reagents and enzymes, and PCR amplification product purification reagents; Preferably, it also includes reagents for detecting STR genetic markers.
10. A method for detecting micro-haplotypes based on high-throughput sequencing technology, characterized in that: This includes constructing sequencing libraries using the genetic marker composition of claim 1 or 2.