A primer set, a kit and a method for detecting a plurality of disease-related gene mutation types
By employing multiplex PCR amplification technology and high-throughput sequencing, the limitations of existing technologies in detecting hereditary deafness, thalassemia, and spinal muscular atrophy have been overcome. This enables rapid, low-cost, and accurate detection of multiple genetic diseases, simplifying the testing process and improving testing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PINFENG (JIANGSU) MEDICAL TECHNOLOGY CO LTD
- Filing Date
- 2023-02-09
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies for detecting hereditary deafness, thalassemia, and spinal muscular atrophy suffer from limitations in detection range, high cost, long detection time, high requirements for sample DNA volume, and difficulty in primer design, especially in detecting mitochondrial DNA heterogeneity and large fragment deletions.
A single-tube primer set was designed using multiplex PCR amplification technology combined with high-throughput sequencing. This primer set can simultaneously detect gene mutation types in hereditary deafness, thalassemia, and spinal muscular atrophy. Through primer optimization and bioinformatics analysis, uniform amplification of mitochondrial and chromosomal DNA was achieved, and amplification efficiency was corrected by β-actin gene analysis, simplifying the detection process.
It enables rapid, low-cost, and convenient testing for various genetic diseases, accurately identifying gene mutation types, including point mutations, small fragment insertions and deletions, large fragment deletions, and copy number variations, reducing false positive and false negative rates, and improving testing efficiency and accuracy.
Smart Images

Figure CN116287198B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology and medical research applications, specifically relating to primer sets, kits, and methods for detecting gene mutation types associated with hereditary deafness, thalassemia, and spinal muscular atrophy. Background Technology
[0002] To effectively reduce the occurrence of birth defects caused by genetics, the National Medical Products Administration has approved a number of genetic disease testing products in recent years, among which hereditary deafness, thalassemia, and spinal muscular atrophy have received particular attention.
[0003] Hereditary deafness is a typical single-gene disease with strong genetic heterogeneity. Currently, the technical platforms used for genetic screening of hereditary deafness include quantitative real-time PCR, PCR + flow-through hybridization, microarray technology, and high-throughput sequencing, among others. Their specific detection principles differ, and kits that only include PCR amplification or subsequent probe hybridization for site detection are currently more common. These products cover different ranges and pathogenic sites, but most products include mitochondrial DNA (mtDNA) mutation detection. Hereditary deafness caused by mtDNA mutations includes syndromic and non-syndromic hereditary deafness. Among these, the m.1555A>G and m.1494C>T mutations on the mtDNA 12S rRNA are susceptibility sites for aminoglycoside-induced deafness. In addition, other mtDNA mutations can also cause syndromic hereditary deafness, such as maternally inherited diabetes mellitus accompanied by hereditary deafness. However, unlike other germline mutations, mitochondrial DNA (mtDNA) mutations are randomly distributed during cell replication, and the mother-to-offspring mutation ratio can vary significantly. mtDNA exhibits a state threshold, meaning that a corresponding phenotype only occurs when the mutant DNA reaches a certain load or mtDNA function is deficient to a certain degree. Therefore, in recent years, many scholars have suggested increasing the detection of mitochondrial heterogeneity ratios to more accurately assess the role of mitochondrial mutations in disease development. Among currently known detection methods, high-throughput sequencing can detect autosomal hereditary deafness mutations and simultaneously analyze mitochondrial heterogeneity ratios. However, using amplicon library construction-based high-throughput sequencing for mitochondrial mutation detection presents another challenge: mitochondrial heterogeneity varies from person to person, and the copy number of the mitochondrial genome differs significantly from that of chromosomes. This results in different initial DNA template amounts (i.e., different initial template copy numbers) when multiplex PCR amplifies mitochondrial DNA and genomic DNA, posing challenges to subsequent PCR condition optimization and bioinformatics analysis.
[0004] The main types of thalassemia in my country are alpha thalassemia and beta thalassemia. The causes of these two types differ slightly. Alpha thalassemia is caused by the deletion of the alpha hemoglobin gene, with a minority caused by point mutations; while beta thalassemia is mainly caused by point mutations in the beta hemoglobin gene, with a minority caused by gene deletion. This difference leads to varying degrees of difficulty in detection techniques. Currently, in clinical diagnosis, PCR-based methods, such as Gap-PCR, can effectively distinguish between different types of thalassemia, but due to methodological limitations, they can lead to varying degrees of false positives or false negatives. In recent years, some researchers have attempted to use NGS sequencing, but this is mostly done using liquid-phase capture technology. Amplicon library construction for thalassemia detection is mainly focused on Beta thalassemia (CN105331694B), which is predominantly caused by point mutations. Detecting alpha thalassemia, which is predominantly caused by large fragment deletions, using conventional amplicon library construction methods still presents technical challenges. The difficulty stems primarily from the fact that the coding regions of the two key genes in thalassemia, HBA1 and HBA2, are completely identical, and thalassemia breakpoints often fall within intron regions. Generally, the GC content of multiplex PCR is between 40% and 60%, with 50% being optimal. However, the GC content in intron regions is high, reaching over 80% in some areas, severely impacting primer amplification efficiency and leading to uneven primer amplification, thus affecting the identification of thalassemia deletion fragments. Therefore, using amplicon library construction and balancing the amplification efficiency among primers to ensure amplification uniformity is particularly important for identifying large fragment deletions in thalassemia.
[0005] Spinal muscular atrophy (SMA) is a neuromuscular disease characterized by progressive muscle weakness and atrophy, caused by degeneration of motor neurons in the anterior horn of the spinal cord. It is the second leading cause of death among autosomal recessive genetic disorders. The disease is caused by a defect in the survival motor neuron 1 (SMN1) gene. Normally, at least one copy of the SMN1 gene is present on both chromosomes 5, while SMA patients either have a deletion of the SMN1 gene (specifically, a homozygous deletion in exon 7) or carry a mutated SMN1 gene. Studies show that approximately 95% of SMA patients have homozygous deletions of the SMN1 biallelic gene, while the remaining 5% have mutations (i.e., one allele is deleted, and the other allele has a minor pathogenic variation). Although the carrier rate in the population is as high as 1 / 40, and the incidence rate is approximately 1 / 10,000-1 / 6,000, it remains the leading cause of death among genetic diseases in children under 2 years old. Therefore, to reduce the incidence of SMA, efforts have begun to establish a three-tiered prevention and control system for SMA. The NMPA has also approved three PCR-based detection kits, all of which are designed for detecting 95% deletion of SMN1 and can be used for screening or auxiliary diagnosis of patients, carriers, and healthy individuals. Although SMN1 is the pathogenic gene for SMA, determining the occurrence of the disease, SMN2 is a phenotypic modifying gene that affects the severity and progression of the disease. SMN1 and SMN2 are highly homologous, differing by only 5 bases. CN112048548A provides a method for detecting SMN gene copy number using SMNP as a control, which simultaneously amplifies SMNP and SMN1 using one primer pair and simultaneously amplifies SMNP and SMN2 using another primer pair, and distinguishes them by capillary electrophoresis. This invention, however, uses the same primer pair to simultaneously amplify SMN1 and SMN2, and distinguishes them using high-throughput sequencing, making it simpler and more accurate. The limitations of currently available products are: 1) they cannot detect the 5% of minute variants; 2) they cannot be used to detect SMN2 copy number, thus they cannot be used to assess disease severity and progression. However, due to the high homology between SMN1 and SMN2, using next-generation sequencing (NGS) technology for detection presents significant challenges for primer design and analysis. For this reason, Berry Genomics in China uses third-generation sequencing technology with longer read lengths for detection; however, due to the high cost of NGS technology, its widespread clinical application will still take time. Summary of the Invention
[0006] Currently, hereditary deafness, thalassemia, and spinal muscular atrophy are three of the most common genetic diseases whose incidence can be effectively reduced through genetic screening. The key requirements for successful genetic disease screening are: 1. meeting the needs of high-throughput testing of large numbers of samples; 2. being able to detect multiple diseases, genes, or loci in a single test; 3. low cost, meeting the need for low-cost procurement; and 4. short turnaround time. High-throughput sequencing technology is undoubtedly an ideal tool for genetic disease screening. Currently, the most commonly used detection techniques based on high-throughput sequencing for genetic disease screening are liquid-phase capture technology and amplicon library preparation technology. Liquid-phase capture technology, compared to amplicon library preparation, has the following characteristics: 1. Longer turnaround time: Even with rapid hybridization, it takes 1-2 days to complete the process from DNA to sequencing, while amplicon library preparation can be completed in as little as 3 hours; 2. Significantly higher cost: Liquid-phase capture is much more expensive than amplicon library preparation; 3. Higher requirements for sample DNA volume; 4. Capture probe design is more tolerant of mismatches than amplicon PCR primer design, meaning that amplicon PCR primer design is far more difficult than capture probe design. Based on the descriptions of these two next-generation sequencing library preparation methods, using amplicon library preparation combined with next-generation sequencing technology for the detection of these three diseases is more in line with the needs of genetic disease screening, but it also faces more technical challenges.
[0007] In view of this, after extensive experimental verification, this invention has developed a method and kit for simultaneously detecting common gene mutation types of three genetic diseases—hereditary deafness, thalassemia, and spinal muscular atrophy—using a single primer tube, based on multiplex PCR amplification technology combined with high-throughput sequencing technology.
[0008] Specifically, according to a first aspect of the present invention, a primer set for detecting multiple disease-related gene mutation types is provided, comprising:
[0009] 1) Primers for detecting gene mutation types in hereditary deafness are shown in SEQ ID NO:1-30;
[0010] 2) Primers for detecting gene mutation types in thalassemia are shown in SEQ ID NO:31-340;
[0011] 3) Primers for detecting gene mutation types in spinal muscular atrophy are shown in SEQ ID NO:341-350;
[0012] 4) Primers for detecting the β-actin housekeeping gene are shown in SEQ ID NO:351-352.
[0013] In one embodiment, the gene mutation type of the hereditary deafness is a point mutation in genes GJB2, GJB3, SLC26A4, and 12S rRNA.
[0014] In one embodiment, the gene mutation types of thalassemia are point mutations, short fragment insertions / deletions, and large fragment deletions in genes HBA1, HBA2, and HBB.
[0015] In one embodiment, the gene mutation type of spinal muscular atrophy is an SMN1 point mutation and copy number variation in exons 7 and 8 of the SMN1 / SMN2 gene.
[0016] According to a second aspect of the present invention, a kit is provided for simultaneously detecting gene mutation types associated with hereditary deafness, thalassemia, and spinal muscular atrophy, comprising the primer set described in the first aspect of the present invention.
[0017] In one embodiment, the kit further comprises an amplification enzyme mix, a primer digest solution, an amplification enhancer, index primers, and purified magnetic beads.
[0018] According to a third aspect of the present invention, a method for simultaneously detecting multiple disease-related gene mutation types is provided, characterized by comprising the following steps:
[0019] S1. Using the whole genome DNA of the sample as a template, targeted ultramultiplex PCR amplification was performed in the same reaction tube using the primer set described in the first aspect of the present invention, and the amplification products were purified using magnetic beads to remove primers and primer dimers.
[0020] S2. Construct a library based on the amplicon obtained in S1;
[0021] S3. Based on the library constructed in S2, high-throughput sequencing was used for sequencing. Bioinformatics methods were used to analyze the sequencing results to obtain genetic variation information of various gene mutation-related diseases, and then to determine whether the sample has the corresponding gene mutation type.
[0022] The diseases mentioned are hereditary deafness, thalassemia, and spinal muscular atrophy.
[0023] In one embodiment, the gene mutation type of the hereditary deafness is a point mutation in genes GJB2, GJB3, SLC26A4, and 12S rRNA.
[0024] In one embodiment, the gene mutation types of thalassemia are point mutations, short fragment insertions / deletions, and large fragment deletions in genes HBA1, HBA2, and HBB.
[0025] In one embodiment, the gene mutation type of spinal muscular atrophy is an SMN1 point mutation and copy number variation in exons 7 and 8 of the SMN1 / SMN2 gene.
[0026] In one implementation, S3 further includes the following steps:
[0027] S31: Use a high-throughput sequencer to sequence the library and obtain raw sequencing data;
[0028] S32: Perform basic quality control on the raw sequencing data obtained in step S31, including filtering sequencing adapters and low-quality bases, to obtain valid data. Use BWT algorithm software to align the valid data with the human reference genome to obtain alignment information files with base positions.
[0029] S33: Based on the alignment information file obtained in step S32, the alignment file is split into two parts: autosomal alignment information and mitochondrial genome alignment information. Based on the autosomal alignment information, a Bayesian algorithm is used to detect SNP and INDEL mutations. SNP mutation information is obtained by scoring, sorting, and filtering the mitochondrial genome alignment results.
[0030] In one implementation, the method for determining the mutation type of the genes associated with hereditary deafness, thalassemia, and spinal muscular atrophy is as follows:
[0031] 1) Methods for determining autosomal SNP and INDEL mutations in hereditary deafness, thalassemia, and spinal muscular atrophy: If the variant allele frequency (VAF) is >95%, it is homozygous; if the mutation frequency is 15% ≤ mutation frequency ≤ 95%, it is heterozygous.
[0032] 2) Method for determining mitochondrial SNP mutations in hereditary deafness: If the variant allele frequency (VAF) is >3%, then a mutation exists; otherwise, it is wild-type.
[0033] 3) Diagnostic methods for spinal muscular atrophy caused by SMN1 gene deletion:
[0034] a) Extract the reads that are aligned to the β-actin, SMN1 exon7, exon8, and SMN2 exon7, exon8 regions in the alignment information file, filter out reads with an alignment quality value <60 and >3 base mismatches, and obtain the filtered sequence alignment information file.
[0035] b) Based on the sequence alignment information file filtered in step a), the average coverage depth of the target region is calculated, and the copy number of S1-7, S1-8, S2-7, and S2-8 is obtained by dividing the average coverage depth of exon7 and exon8 of the SMN1 and SMN2 genes by the average coverage depth of β-actin.
[0036] c) Obtain the data sets of negative samples S1-7, S1-8, S2-7, and S2-8 through known and verified SMA negative samples; calculate the median of S1-7 and S1-8;
[0037] d) Calculate the multiples of S1-7 and S1-8 between the test sample and the SMA negative sample set, and determine the deletion type of the sample;
[0038] e) Determine whether the test sample is a patient or carrier of spinal muscular atrophy; the specific determination method is as follows:
[0039] S1-7 = average coverage depth of SMN1 exon7 / average coverage depth of β-actin;
[0040] S1-8 = average coverage depth of SMN1 exon8 / average coverage depth of β-actin;
[0041] S2-7 = average coverage depth of SMN2 exon7 / average coverage depth of β-actin;
[0042] S2-8 = average coverage depth of SMN2 exon8 / average coverage depth of β-actin;
[0043] Ratio(S1-7) = S1-7 / median of S1-7 in the negative control set;
[0044] Ratio(S1-8) = S1-8 / median of S1-8 in the negative control set;
[0045] If the judgment value of the SMN1 gene Ratio(S1-7) ≤ 0.2 and Ratio(S1-8) ≤ 0.2, the test sample is homozygous deletion, that is, the SMN1 copy number is 0;
[0046] If Ratio(S1-7) ≤ 0.2 and 0.2 < Ratio(S1-8) ≤ 0.7, the test sample is homozygous deletion of exon7 and heterozygous deletion of exon8;
[0047] If 0.2 < Ratio(S1-7) ≤ 0.7 and Ratio(S1-8) ≤ 0.2, the test sample is heterozygous deletion of exon7 and homozygous deletion of exon8;
[0048] If 0.2 < Ratio(S1-7) ≤ 0.7 and 0.2 < Ratio(S1-8) ≤ 0.7, the test sample is heterozygous deletion of SMN1;
[0049] If S2-7 or S2-8 ≤ 0.3, the SMN2 exon7 or exon8 has 0 copies;
[0050] If 0.3 < S2-7 or S2-8 ≤ 0.7, then the copy number of SMN2 exon7 or exon8 is 1;
[0051] If 0.7 < S2-7 or S2-8 ≤ 1.2, then the copy number of SMN2 exon7 or exon8 is 2;
[0052] If 1.2 < S2-7 or S2-8 ≤ 1.7, then the copy number of SMN2 exon7 or exon8 is 3;
[0053] If 1.7 < S2-7 or S2-8 ≤ 2.2, then the copy number of SMN2 exon7 or exon8 is 4;
[0054] If S2-7 or S2-8 > 2.2, then the copy number of SMN2 exon7 or exon8 is 5 or more;
[0055] 3) Method for judging thalassemia caused by deletion of HBA2, HBA1, and HBB genes:
[0056] a) Statistically analyze the read length coverage depth of the target region, and normalize the sequencing depth of the target region based on the average sequencing depth of the β-actin gene to obtain the Z value, where the Z value = target region sequencing depth / β-actin average sequencing depth * 2000;
[0057] b) Calculate the ratio of the Z value of the test sample to the median of the Z values of the negative control sample set, and determine the deletion breakpoint based on the HMM algorithm;
[0058] c) Judge the deletion copy number of the thalassemia-related gene in the test sample, and the judgment criteria are as follows:
[0059] Judgment value B = log2(test sample Z value / median of negative control sample set Z values)
[0060] If the judgment value B ≤ -1.4, the test sample is homozygous deletion, that is, the gene copy number is 0;
[0061] If -1.4 < judgment value B ≤ -0.6, the test sample is heterozygous deletion, that is, the gene copy number is 1;
[0062] If -0.6 < judgment value B ≤ 0.6, the copy number of the test sample is normal, that is, the gene copy number is 2.
[0063] According to the fourth aspect of the present invention, it provides a system for detecting mutation types related to multiple diseases, including:
[0064] 1) Primer amplification module: Using the whole genome DNA of the sample as a template, targeted ultramultiplex PCR amplification is performed in the same reaction tube using the primer set described in the first aspect of the present invention, and the amplification products are purified using magnetic beads to remove primers and primer dimers.
[0065] 2) Library construction module: Perform library construction on the amplicon fragments obtained from module 1);
[0066] 3) Data processing module: Based on the library constructed in module 2), high-throughput sequencing is used for sequencing. Bioinformatics methods are used to analyze the sequencing results to obtain genetic variation information of various gene mutation-related diseases, and then determine whether the sample has the corresponding gene mutation type.
[0067] The diseases mentioned are hereditary deafness, thalassemia, and spinal muscular atrophy.
[0068] In one embodiment, the gene mutation type of the hereditary deafness is a point mutation in genes GJB2, GJB3, SLC26A4, and 12S rRNA.
[0069] In one embodiment, the gene mutation types of thalassemia are point mutations, short fragment insertions / deletions, and large fragment deletions in genes HBA1, HBA2, and HBB.
[0070] In one embodiment, the gene mutation type of spinal muscular atrophy is an SMN1 point mutation and copy number variation in exons 7 and 8 of the SMN1 / SMN2 gene.
[0071] In one implementation, the data processing module 3) further includes the following:
[0072] 3.1: The library was sequenced using a high-throughput sequencer to obtain raw sequencing data;
[0073] 3.2: Perform basic quality control on the raw sequencing data obtained in step 3.1, including filtering out sequencing adapters and low-quality bases, to obtain valid data. Use BWT algorithm software to align the valid data with the human reference genome to obtain alignment information files with base positions.
[0074] 3.3: Based on the alignment information file obtained in step 3.2, the alignment file is split into two parts: an autosomal alignment information file and a mitochondrial genome alignment information file; according to the autosomal alignment information file, a Bayesian algorithm is used to detect autosomal SNPs and INDEL mutations; mitochondrial SNP mutation information is obtained by scoring, sorting, and filtering the results of the mitochondrial genome alignment information file.
[0075] In one implementation, the method for determining the mutation type of the genes associated with hereditary deafness, thalassemia, and spinal muscular atrophy is as follows:
[0076] 1) Methods for determining autosomal SNP and INDEL mutations in hereditary deafness, thalassemia, and spinal muscular atrophy: If the variant allele frequency (VAF) is >95%, it is homozygous; if the mutation frequency is 15% ≤ mutation frequency ≤ 95%, it is heterozygous.
[0077] 2) Mitochondrial SNP mutations in hereditary deafness: If the variant allele frequency (VAF) is >3%, the mutation is present; otherwise, it is wild-type.
[0078] 3) Diagnostic methods for spinal muscular atrophy caused by SMN1 gene deletion:
[0079] a) Extract the read lengths of the regions aligned to β-actin, SMN1 exon7, exon8, and SMN2 exon7, exon8 in the bam file, and filter out read lengths with an alignment quality value <60 and >3 base mismatches.
[0080] b) Based on the above documents, the average coverage depth of the target region is calculated, and the copy number of S1-7, S1-8, S2-7, and S2-8 is obtained by dividing the average coverage depth of exon7 and exon8 of the SMN1 and SMN2 genes by the average coverage depth of β-actin.
[0081] c) Obtain the datasets of negative samples S1-7, S1-8, S2-7, and S2-8 using known validated SMA negative samples; calculate the median of S1-7 and S1-8.
[0082] d) Calculate the multiples of S1-7 and S1-8 of the sample to be tested and the SMA negative sample set to determine the missing type of the sample;
[0083] e) Determine whether the sample to be tested is from a patient or carrier of spinal muscular atrophy; the specific determination method is as follows:
[0084] S1-7 = Average coverage depth of SMN1 exon7 / Average coverage depth of β-actin;
[0085] S1-8 = Average coverage depth of SMN1 exon8 / Average coverage depth of β-actin;
[0086] S2-7 = Average coverage depth of SMN2 exon7 / Average coverage depth of β-actin;
[0087] S2-8 = Average coverage depth of SMN2 exon8 / Average coverage depth of β-actin;
[0088] Ratio(S1-7) = S1-7 / Median of negative control set S1-7;
[0089] Ratio(S1-8) = S1-8 / Median of negative control set S1-8;
[0090] If the judgment value of SMN1 gene Ratio(S1-7) ≤ 0.2 and Ratio(S1-8) ≤ 0.2, then the test sample is homozygous deletion, that is, the SMN1 copy number is 0;
[0091] If Ratio(S1-7) ≤ 0.2 and 0.2 < Ratio(S1-8) ≤ 0.7, then the test sample is homozygous deletion of exon7 and heterozygous deletion of exon8;
[0092] If 0.2 < Ratio(S1-7) ≤ 0.7 and Ratio(S1-8) ≤ 0.2, then the test sample is heterozygous deletion of exon7 and homozygous deletion of exon8;
[0093] If 0.2 < Ratio(S1-7) ≤ 0.7 and 0.2 < Ratio(S1-8) ≤ 0.7, then the test sample is heterozygous deletion of SMN1;
[0094] If S2-7 or S2-8 ≤ 0.3, then the copy number of SMN2 exon7 or exon8 is 0;
[0095] If 0.3 < S2-7 or S2-8 ≤ 0.7, then the copy number of SMN2 exon7 or exon8 is 1;
[0096] If 0.7 < S2-7 or S2-8 ≤ 1.2, then the copy number of SMN2 exon7 or exon8 is 2;
[0097] If 1.2 < S2-7 or S2-8 ≤ 1.7, then the copy number of SMN2 exon7 or exon8 is 3;
[0098] If 1.7 < S2-7 or S2-8 ≤ 2.2, then the copy number of SMN2 exon7 or exon8 is 4;
[0099] If S2-7 or S2-8 > 2.2, then the copy number of SMN2 exon7 or exon8 is 5 or more;
[0100] 4) Judgment method for thalassemia caused by deletion of HBA2, HBA1, and HBB genes:
[0101] a) Calculate the read coverage depth of the target region and standardize the sequencing depth of the target region based on the average sequencing depth of the β-actin gene to obtain the Z value, where Z value = sequencing depth of the target region / average sequencing depth of β-actin * 2000.
[0102] b) Calculate the ratio of the Z-value of the test sample to the median Z-value of the negative control sample set, and determine the mutation breakpoint based on the HMM algorithm;
[0103] c) Determine the copy number of thalassemia-related genes deleted in the sample to be tested, using the following criteria:
[0104] Judgment value B = log2(test sample Z-score / median Z-score of negative control sample set)
[0105] If the judgment value B ≤ -1.4, then the sample to be tested is homozygous deletion, that is, the copy number of the gene is 0;
[0106] If -1.4 < judgment value B ≤ -0.6, the sample to be tested is a heterozygous deletion, that is, the copy number of the gene is 1;
[0107] If -0.6 < judgment value B ≤ 0.6, the copy number of the sample to be tested is normal, that is, the copy number of the gene is 2.
[0108] According to a fifth aspect of the present invention, the application of the primer set described in the first aspect of the present invention in the preparation of a kit for detecting multiple disease-related gene mutation types is provided.
[0109] In one embodiment, the disease is hereditary deafness, thalassemia, and spinal muscular atrophy.
[0110] The superior technical effects of the primer set, method, system, and kit described in this invention are mainly in the following aspects:
[0111] 1. This invention employs a single-tube, two-step PCR amplification library preparation technique to simultaneously detect three genetic diseases, including hereditary deafness gene mutation, thalassemia (α+β), and SMA. Currently, similar techniques mostly use liquid-phase capture for detection or multiplex PCR amplification methods for different single diseases. Compared to liquid-phase capture library preparation techniques, this invention offers advantages such as speed, ease of operation, high efficiency, and low cost, making it more suitable for widespread clinical application.
[0112] 2. This invention uses single-tube PCR to detect multiple diseases, placing higher demands on primer design. It requires consideration of the compatibility of different primers within the same reaction system and the uniformity of amplification efficiency. This places higher demands on primer design, system optimization, and subsequent bioinformatics analysis. Among the three genetic diseases described in this invention, the mutation types are complex, including point mutations, small fragment insertions or deletions, large fragment deletions, and copy number variations. Among these mutation types, large fragment deletions are particularly challenging for multiplex PCR library design, especially in α-thalassemia and SMA, which both contain large fragment deletions. Existing technologies typically employ two approaches: the first involves designing primers for specific breakpoint regions and analyzing the length of the amplified products to identify the deletion region. This method can only detect fixed-type deletions; if the deletion breakpoint is outside the designed region, it will miss the detection. The second approach involves covering the large fragment deletion region with a PCR fragment and using bioinformatics analysis to determine the deletion breakpoint by observing changes in the PCR signal depth within the covered region. This method is more flexible in detecting deletion types, and can detect all deletion breakpoints occurring within the PCR covered region. This invention employs the second method, designing PCR primer pairs at relatively fixed intervals (generally designed to be within 500 bp) to detect deletion mutations. Since the deletion region in alpha-thalassemia reaches 35 kb, following conventional design principles and considering other diseases, the entire product would require nearly 400 pairs of PCR amplification primers. However, a large number of primers in the same reaction system increases the possibility of primer interference and system instability, while also increasing costs. In this study, we used a characteristic SNP method, comprehensively considering common disease gene mutation types and primer interference, and optimized the method to reduce the number of PCR primer pairs to 176 pairs, ensuring both the stability of the experimental system and balancing cost and detection performance.
[0113] 3. This invention can simultaneously detect heterogeneity variations in mitochondrial DNA and germline mutations in chromosomes. Each mitochondrion contains approximately 2-10 copies of mitochondrial DNA, and the normal human somatic cell contains approximately 10 copies of mitochondrial DNA. 3 -10 4 The heterogeneity between individuals is very significant. Given the heterogeneity of mitochondrial DNA and the difference in copy number between it and chromosomal DNA, other studies typically amplify mitochondrial DNA and chromosomal DNA separately. This invention's kit, through optimized reaction conditions and adjustments using bioinformatics methods, achieves uniform and stable amplification of mitochondrial DNA and chromosomal DNA within the same tube.
[0114] 4. This invention is used to detect large deletion mutations in thalassemia. To meet the needs of disease detection and cover as many deletion mutation breakpoint regions as possible, multiple primer pairs need to be designed within a 35kb range on chromosome 16. However, this region is characterized by high sequence repetition and high GC content (some high GC regions (over 80%)). Conventional methods separate primers for different deletion types to avoid mutual interference between primers or systems. This method optimizes the amplification experimental system and bioinformatics analysis methods by adjusting the primer design regions and the ratio between primers, ensuring the detection of multiple types of thalassemia deletions. Furthermore, it can be performed in a single PCR reaction tube, reducing experimental complexity and improving detection efficiency.
[0115] 5. Conventional SMA detection methods primarily detect homozygous deletion mutations in exon 7 of SMN1, which contributes to the highest incidence rate. However, the copy number of SMN2 has a compensatory effect on the symptoms, so simultaneously detecting the copy number of SMN2 would greatly aid clinicians in diagnosing the disease. However, SMN1 and SMN2 are highly homologous, differing by only 5 bases, making the design of primers that can distinguish between them technically challenging. This method, through primer design adjustments, ensures that exon 7 and exon 8 of the SMN1 and SMN2 genes are amplified simultaneously, guaranteeing consistent amplification efficiency in this region. Furthermore, it incorporates the batch effect of the β-actin housekeeping gene to correct the amplification system, enabling specific detection of the copy numbers of SMN1 and SMN2.
[0116] 6. Conventional methods for detecting large deletion fragments involve data volume correction, GC content correction, and comparison of the test sample with a control set to obtain the copy number change in the test sample. However, this method uses multiplex PCR amplification technology, which, due to variations in amplification efficiency, increases batch-to-batch instability, leading to reduced detection accuracy. This method adds β-actin housekeeping gene primers to the amplification system, using the average sequencing depth of this gene to correct for the amplification efficiency of the sample, reducing the impact of amplification. Furthermore, it employs intra- and inter-sample controls to identify large deletion fragments in SMA and thalassemia, improving detection stability. Attached Figure Description
[0117] Figure 1 Primer design diagram;
[0118] Figure 2 QC spectra of amplified fragments from the library show that fragment sizes are concentrated between 200bp and 500bp, with the main peak around 380bp.
[0119] Figure 3 The technical route for multiplex PCR amplification;
[0120] Figure 4Heatmap of depth distribution before and after adjustment of mitochondrial primer ratio;
[0121] Figure 5 The validation results were obtained using the "Human Motor Neuron Survival Gene 1 (SMN1) Detection Kit (PCR-Melting Curve Method)". Figure 5 A represents the detection result of NA03813 SMN1. Figure 5 B indicates the detection result of NA03814 SMN1;
[0122] Figure 6 Distribution of mitochondrial heterogeneity mutations in PF02-L-14 samples. Detailed Implementation
[0123] The preferred embodiments of this application will be further described in detail below with reference to the accompanying drawings. The following description is exemplary and not intended to limit the application. Any other similar situations also fall within the protection scope of this application.
[0124] definition
[0125] To better understand this invention, the definitions and explanations of relevant terms are provided below.
[0126] As used in this article, “detection” refers to the qualitative analysis of the presence or absence of a substance.
[0127] As used herein, the words “comprising,” “including,” “having,” “containing,” or any other variations thereof are intended to cover a non-exclusive inclusion. For example, a composition, step, method, article, or apparatus that includes the listed elements is not necessarily limited to those elements, but may include other elements not expressly listed or elements inherent to such a composition, step, method, article, or apparatus.
[0128] As used herein, the conjunction "composed of..." excludes any unspecified element, step, or component. If used in a claim, this phrase makes the claim closed, excluding any material other than those described, except for conventional impurities associated with them. When the phrase "composed of..." appears in a clause of the body of a claim rather than immediately following it, it limits only the element described in that clause; other elements are not excluded from the claim as a whole.
[0129] As used herein, when a ratio or other value or parameter is expressed as a range, a preferred range, or a range defined by a series of upper and lower preferred values, this should be understood as specifically disclosing all ranges formed by any pair of any upper or preferred value with any lower or preferred value, regardless of whether the range is disclosed individually. For example, when the range “1–5” is disclosed, the described range should be interpreted as including the ranges “1–4”, “1–3”, “1–2”, “1–2 and 4–5”, “1–3 and 5”, etc. When numerical ranges are described herein, unless otherwise stated, the range is intended to include its endpoints and all integers and fractions within that range.
[0130] As used herein, “and / or” is used to indicate that one or both of the described situations may occur, for example, A and / or B includes (A and B) and (A or B).
[0131] As used in this article, "multiple" refers to two or more.
[0132] As used in this article, the term "clean data" or "effective data" refers to the raw sequencing data that has been processed to remove sequences containing sequencing adapters and low-quality bases, resulting in effective data that can be used for subsequent analysis.
[0133] As used in this article, the term "read" or "read length" refers to the sequence information of a DNA fragment.
[0134] As used herein, the term "amplification enzyme mix" refers to an amplification mixture of DNA polymerase, dNTPs, and buffer.
[0135] As used in this article, the term "primer digestion solution" refers to the reagent used to remove residual dNTPs and primers after PCR.
[0136] As used in this article, the term "amplification enhancer" refers to an additive containing a certain concentration of dimethyl sulfoxide (DMSO), betaine, and trehalose, which can improve the PCR amplification performance of high GC templates.
[0137] As used in this article, the term "Index primer" refers to a primer that carries a specific base sequence and is capable of amplifying a complete sequencing adapter, where "Index" refers to a specific base sequence used to distinguish samples during library mixing sequencing.
[0138] As used in this article, the term "purification magnetic beads" refers to magnetic bead reagents that can be used to separate and purify DNA.
[0139] As used in this article, the term "base mutation frequency" refers to the fact that a base mutation is a change in the gene structure caused by the substitution, addition, or deletion of base pairs in a DNA molecule; frequency refers to the proportion of base sequences that have undergone the above mutations, and is calculated as the proportion of the number of mutated reads to the total number of reads at that site.
[0140] As used in this article, the term "HMM algorithm" refers to Hidden Markov Model.
[0141] Sequence List Overview
[0142] This application is accompanied by a sequence listing containing numerous nucleic acids. Table A provides information on all included mutation sites, and Table B provides an overview of the included primer sequences.
[0143] Table A
[0144]
[0145]
[0146]
[0147]
[0148] Table B
[0149]
[0150]
[0151]
[0152]
[0153]
[0154]
[0155]
[0156]
[0157]
[0158]
[0159]
[0160]
[0161]
[0162]
[0163]
[0164]
[0165]
[0166]
[0167]
[0168]
[0169]
[0170] Example 1 Preparation of a kit for detecting gene mutations in three genetic diseases
[0171] 1.1 Primer design
[0172] 1.1.1 First round of primers
[0173] In this invention, we employ a signature SNP method, comprehensively considering common disease gene mutation types and primer interference. Through optimization, the number of PCR primers is reduced, ensuring the stability of the experimental system while balancing cost and detection performance. A schematic diagram of the primer design is shown below. Figure 1 As shown.
[0174] The detection range of genes related to three genetic diseases—hereditary deafness, thalassemia, and spinal muscular atrophy—was determined, and primers were designed targeting mutations within this detection range. For large deletion regions, marker SNPs were selected, with a minimum allele frequency (MAF) close to 50% and an interval of approximately 500 bp between each marker SNP. During primer design, factors such as primer dimerization, high GC content, and annealing temperature were considered. All designed primers were cross-aligned using BLAST and compared with the human reference genome to ensure the absence of complementary sequences and high specificity. A total of 176 primer pairs were obtained for subsequent experimental evaluation. During primer synthesis, the universal primer sequences for sequencing were added to the specific upstream and downstream primers to obtain the primer sequences for the first round of PCR: first-round primer F (universal sequence F + specific primer F) and first-round primer R (universal sequence R + specific primer R).
[0175] 1.1.2 The primers for the second round of primers are complete adapter sequence primers for the sequencing platform, which include adapter sequences, universal sequences, and indexes for differentiating different samples. The structure of the complete adapter sequence with indexes is shown in Table 1.
[0176] Table 1 contains the complete connector sequence with index.
[0177]
[0178] Note: XXXXXXXX is an 8bp index, and different samples contain different index combinations.
[0179] 1.2 Preparation of a kit for detecting gene mutations in three genetic diseases
[0180] The specific components of the kit for detecting gene mutations in three genetic diseases are shown in Table 2:
[0181] Table 2. Three Genetic Disease Gene Mutation Detection Kits
[0182] Reagent Name Storage temperature Multiple primers -20℃±5℃ amplification enzyme mix -20℃±5℃ Amplification Enhancer 1 -20℃±5℃ Amplification Enhancer 2 -20℃±5℃ Index primer (I5 / I7) -20℃±5℃ Primer digestion solution 4℃ Purified magnetic beads 4℃
[0183] Example 2 Optimization of multiplex PCR reaction system
[0184] During the research process, the inventors screened several pairs of PCR primers targeting the target region, ultimately selecting primers with good specificity, amplification efficiency, and multiplex amplification that met the experimental requirements. The following describes the systematic optimization of high-GC template primer design and system optimization, the balancing of SMN1 and SMN2 amplification efficiencies, the simultaneous detection of mitochondrial DNA and chromosomes in the same tube, and the ratio of 176 primer pairs.
[0185] 2.1 Primer design and amplification system optimization for high GC content regions of chromosome 16
[0186] Large deletion mutations in thalassemia are located within a 35kb range on chromosome 16. The breakpoints of these deletions are distributed in various ways. To cover as many possible breakpoint regions as possible, multiple primer pairs are designed to cover more areas, which is more helpful for detection and diagnosis. However, this region has a complex base composition, with repetitive sequences and high GC content (some high GC regions (over 80%)). Conventional amplification methods separate primers for detection based on different deletion types to avoid mutual interference between primers or systems.
[0187] To address the aforementioned issues, this invention designs multiple primer pairs to cover the corresponding regions. To ensure compatibility with amplification efficiency in high-GC regions, the experimental reaction system is optimized. By adjusting the final concentration of the amplification enhancer (previously 2%, now 6%), the amplification efficiency in this region is significantly improved, with the average sequencing depth increasing from 100-300X to 2000-3000X, thus enabling stable detection and ensuring subsequent analysis. The specific differences before and after the adjustment are shown in Table 3.
[0188] Table 3. Optimization test of high-GC fragments in large fragment deletion regions of thalassemia.
[0189]
[0190] 2.2 Balancing the amplification efficiency of SMN1 and SMN2
[0191] The pathogenesis of spinal muscular atrophy (SMA) is well understood: deletion of exon 7 or exon 7-8 of SMN1 leads to more than 95% of patients developing the disease. SMN1 and SMN2 genes are homologous genes, with SMN2 being a pseudogene of SMN1. Their copy numbers vary across the population, and increased copy numbers can have a mitigating effect on SMA patients. Understanding the copy numbers of SMN1 and SMN2 during testing helps clinicians assess the disease and influence treatment decisions. To ensure this kit can stably detect the copy numbers of SMN1 and SMN2 and reduce the impact of amplification efficiency, we adjusted the primer positions so that exon 7 and exon 8 of SMN1 and SMN2 can be amplified simultaneously with a single primer pair, achieving the required coverage depth for analysis, as shown in Table 4.
[0192] Table 4 shows the simultaneous amplification of SMN1 and SMN2 after optimization.
[0193]
[0194]
[0195] 2.3 Simultaneous Detection of Mitochondrial DNA and Chromosomal DNA Mitochondrial genome copy number varies from person to person; however, the normal human nuclear genome contains two copies of chromosomes. Therefore, when mitochondrial DNA and chromosomal DNA are amplified and detected in a single tube, the ratio of mitochondrial primers needs to be optimized to ensure the uniformity of the amplification products. Before adjustment, the ratio of mitochondrial to autosomal primers was 1:1, resulting in a much higher amplification depth in the mitochondrial region than in the autosomal region. After adjustment, the ratio was 1:10, resulting in nearly identical amplification depths in both the mitochondrial and autosomal regions. The difference in primer ratios before and after adjustment is illustrated in the image. Figure 4 As shown.
[0196] 2.4 Adjustment of overall primer concentration ratio
[0197] Based on the adjustments made to the aforementioned special regions, and combined with the primer amplification efficiency of other regions, the amplification efficiency of primers was improved by balancing the input ratio of each primer and optimizing the primers, thus ensuring uniform sequencing depth in all regions, especially for CNV detection regions. Table 5 shows the sequencing depth before and after some primer adjustments.
[0198] Table 5. Sequencing alignment before and after primer 118 adjustment:
[0199]
[0200] Example 3 National reference materials and simulated clinical sample testing
[0201] Using the kit obtained in Example 1, a library was constructed using whole-genome DNA as a template and multiplex PCR was used. The library was sequenced using high-throughput sequencing technology, and the sequencing data was analyzed by bioinformatics to detect and analyze gene mutations.
[0202] 3.1 Sample Sources: The samples used for hereditary deafness and thalassemia were national reference materials; SMA was genomic DNA purchased from Coriell and was verified using the "Human Motor Neuron Survival Gene 1 (SMN1) Detection Kit (PCR-Melting Curve Method)" (National Medical Device Registration Certificate No. 20213400937), as shown in Table 6.
[0203] Table 6. Sources of samples used in the study
[0204]
[0205] 3.2 Detection Method
[0206] 3.2.1 Whole genome DNA extraction: Use the specified nucleic acid extraction kit to extract whole genome DNA from blood or blood card.
[0207] 3.2.2 First round of PCR: Target sequence amplification
[0208] Prepare the reaction system for each sample in the PCR tube according to Table 7, and place the prepared system on the PCR instrument to react according to the reaction conditions in Table 8.
[0209] Table 7 Reaction System
[0210]
[0211]
[0212] [1]Based on the DNA Qubit quantification results, determine the input volume, requiring an initial amount of 5-100 ng / reaction tube.
[0213] Table 8 Reaction conditions
[0214]
[0215] 3.3 Magnetic Bead Purification
[0216] 3.3.1 After amplification, remove the PCR tube, add 0.9 times the volume of magnetic beads (27 μL) to each reaction system (30 μL), mix well, let stand for 5 min, centrifuge briefly, place the PCR tube on a magnetic rack for 3 min, and wait for the solution to become clear.
[0217] 3.3.2 Thoroughly remove the supernatant, remove the PCR tube from the magnetic rack, add 50 μL of primer digestion solution to the tube, mix well, let stand at room temperature for 5 min, centrifuge briefly, place the PCR tube on the magnetic rack for 3 min, and wait for the solution to become clear.
[0218] 3.3.3 Discard the supernatant, add 180 μL of 80% ethanol solution, and let stand for 30 seconds; discard the supernatant again, add 180 μL of 80% ethanol solution, and let stand for 30 seconds.
[0219] 3.3.4 Discard the supernatant, briefly centrifuge, and remove any residual ethanol at the bottom.
[0220] 3.3.5 Dry the magnetic beads, add 20 μL of ddH2O, remove the PCR tube from the magnetic rack, mix well by pipetting, and let stand at room temperature for 2 min. Centrifuge briefly, and then let stand on the magnetic rack for 2 min until the solution becomes clear.
[0221] 3.3.6 Take the supernatant and place it in a new PCR tube for subsequent library construction.
[0222] 3.4 Second round of PCR: Adding adapter sequences to the PCR amplification products obtained in step 3.3. The reaction system is shown in Table 9, and the reaction conditions are shown in Table 10.
[0223] Table 9. Reaction System with Added Connector Sequence
[0224] reagents Volume (μL) <![CDATA[ddH2O]]> 2 Amplification Enhancer 2 2.5 Index primer (I5 / I7) (5μM) 2 Previous PCR products 13.5 amplase Mix 10 total 30
[0225] Table 10 Reaction Conditions
[0226]
[0227] 3.5 Second round of magnetic bead purification
[0228] Same as step 3.3.
[0229] 3.6 Library Quantitative Analysis and Quality Control
[0230] Take 1 μL of the library and use Qubit to determine the library concentration, and record the library concentration; take 1 μL of the library and use Agilent DNA 1000kit or equivalent instruments and reagents to determine the library fragment length.
[0231] 3.7 High-throughput sequencing and data analysis
[0232] 3.7.1 The library obtained in step 3.5 was subjected to high-throughput sequencing using second-generation sequencing platforms such as Illumina Novaseq and Nextseq CN500 to obtain the raw sequencing sequence.
[0233] 3.7.2 The raw sequencing data was processed using the FASTP software (https: / / github.com / OpenGene / fastp) to remove sequencing adapters, low-quality bases (base quality <10), and filter reads shorter than 40 bp to obtain clean data. The clean data was then aligned with the hg38 reference genome using the MEM algorithm in the BWA software to obtain the corresponding read positions.
[0234] 3.7.3 After completing step 3.7.2, use the GATK software (https: / / github.com / broadinstitute / gatk / releases) Haplotype Caller to detect single nucleotide variants (SNVs) and insertion / deletion mutations (INDELs) in the nuclear genome, and use Mutect2 to detect single nucleotide variants in the mitochondrial genome to obtain information such as read length coverage depth and variant frequency at mutation sites.
[0235] 3.7.4 The information obtained in step C was annotated using ANNOVAR software (https: / / annovar.openbioinformatics.org / en / latest / ). Based on the mutation frequencies of SNVs and INDELs in the nuclear genome, the homozygous heterozygous information of nuclear gene variants was determined. If the mutation frequency was ≥0.95, the site was homozygous; if the mutation frequency was 0.15 ≤ mutation frequency ≤0.95, it was heterozygous. Sites with a mitochondrial genome mutation frequency <5% were filtered out and classified as wild-type.
[0236] 3.7.5 Obtain the gene mutation sites of the sample to be tested for the three single-gene genetic diseases, including only SNP and Indel mutations, and output an Excel file.
[0237] 3.7.6 Extract the read lengths mapped to the regions of β-actin, SMN1 exon7, exon8, SMN2 exon7, and exon8 from the bam file, and filter out the read lengths with mapping quality value <60 and >3 base mismatches;
[0238] 3.7.7 Based on the above file, calculate the average coverage depth of the target region. Divide the average coverage depths of exon7 and exon8 of the SMN1 and SMN2 genes by the average coverage depth of β-actin to obtain the copy numbers of S1-7, S1-8, S2-7, and S2-8;
[0239] 3.7.8 Through the known verified SMA negative samples, obtain the data sets of S1-7, S1-8, S2-7, and S2-8 of the negative samples; calculate the median of S1-7 and S1-8;
[0240] 3.7.9 Calculate the multiples of S1-7 and S1-8 between the test sample and the SMA negative sample set, and judge the deletion type of the sample;
[0241] 3.7.10 Judge whether the test sample is a patient or carrier of spinal muscular atrophy; the specific judgment method is as follows:
[0242] S1-7 = average coverage depth of SMN1 exon7 / average coverage depth of β-actin;
[0243] S1-8 = average coverage depth of SMN1 exon8 / average coverage depth of β-actin;
[0244] S2-7 = average coverage depth of SMN2 exon7 / average coverage depth of β-actin;
[0245] S2-8 = average coverage depth of SMN2 exon8 / average coverage depth of β-actin;
[0246]
[0250] If 0.2 < Ratio(S1-7) ≤ 0.7 and Ratio(S1-8) ≤ 0.2, the test sample is homozygous deletion of exon7 and homozygous deletion of exon8;
[0251] If 0.2 < Ratio(S1-7) ≤ 0.7 and 0.2 < Ratio(S1-8) ≤ 0.7, the test sample is heterozygous deletion of SMN1;
[0252] If S2-7 or S2-8 ≤ 0.3, the copy number of SMN2 exon7 or exon of exon8 is 0;
[0253] If 0.3 < S2-7 or S2-8 ≤ 0.7, the copy number of SMN2 exon7 or exon of exon8 is 1;
[0254] If 0.7 < S2-7 or S2-8 ≤ 1.2, the copy number of SMN2 exon7 or exon of exon8 is 2;
[0255] If 1.2 < S2-7 or S2-8 ≤ 1.7, the copy number of SMN2 exon7 or exon of exon8 is 3;
[0256] If 1.7 < S2-7 or S2-8 ≤ 2.2, the copy number of SMN2 exon7 or exon of exon8 is 4;
[0257] If S2-7 or S2-8 > 2.2, the copy number of SMN2 exon7 or exon of exon8 is 5 or more;
[0258] 3.7.11 Statistically analyze the read coverage depth of the target region, and normalize the sequencing depth of the target region of HBB, HBA1, and HBA2 genes based on the average sequencing depth of the β-actin gene (Z value = target region sequencing depth / B-actin average sequencing depth * 2000)) to obtain the Z value.
[0259] 3.7.12 Calculate the ratio of the Z value of the test sample to the median of the Z values of the negative control sample set, and determine the mutation breakpoint based on the HMM algorithm.
[0260] 3.7.13 Judge the deletion copy number of the thalassemia-related genes in the test sample. The judgment criteria are as follows:
[0261] Judgment value B = log2(test sample Z value / median of negative control sample set Z values)
[0262] If the judgment value B ≤ -1.4, the test sample is homozygous deletion, that is, the copy number of this gene is 0;
[0263] If -1.4 < judgment value B ≤ -0.6, the sample to be tested is a heterozygous deletion, that is, the copy number of the gene is 1;
[0264] If -0.6 < judgment value B ≤ 0.6, the copy number of the sample to be tested is normal, that is, the copy number of the gene is 2.
[0265] 3.8 Test results of national reference materials and simulated clinical samples
[0266] Forty-one national reference samples and simulated clinical samples were tested using the kit of this invention. The overall performance met expectations, with the target region coverage of negative samples approaching 100% at 100X. Basic quality control information of the samples is shown in Table 11.
[0267] Table 11 Sample Basic Data Quality Control
[0268]
[0269]
[0270]
[0271] Of the 39 national reference samples for hereditary deafness and thalassemia, 30 were positive for SNP and Indel mutations, 3 were negative for hereditary deafness, and 2 were negative for thalassemia. In this kit experiment, all SNP and Indel mutation sites were detected. The detailed detection results are shown in Table 12.
[0272] Table 12 Detection of SNP and INDEL mutation sites
[0273]
[0274]
[0275]
[0276]
[0277] The kit used in this invention was used to test 11 national reference samples of thalassemia containing large fragment deletions and 2 clinical dummy samples of SMA containing exon 7 and 8 deletions. All mutations were accurately detected, as shown in Table 13.
[0278] Table 13 List of Copy Number Variations Detected
[0279]
[0280]
[0281] The PF02-L-14 sample contained approximately 2.5% mitochondrial heterogeneity mutations. Using the detection reagent of this invention, with a sequencing data volume of 0.15G, the coverage depth of this site was 517X, and 21 mutation reads were detected, with a mutation frequency of 4%, which was successfully detected. Figure 6 As shown, this is a feature that currently available products do not possess.
[0282] It should be understood that although the present invention has been described by way of example according to its preferred embodiments, it should not be limited to the above embodiments. Various modifications and variations can be made to the present invention by those skilled in the art. The reaction reagents, reaction conditions, etc., involved in the preparation of the kit for detecting multiple disease-related gene mutation types and the method for detecting multiple disease-related gene mutation types can be adjusted and changed according to specific needs. Therefore, those skilled in the art can make several simple substitutions without departing from the concept and principles of the present invention, and these should all be included within the scope of protection of the present invention.
Claims
1. A primer set for detecting multiple disease-related gene mutation types, comprising: 1) Primers for detecting gene mutation types in hereditary deafness are shown in SEQ ID NO: 1-30, wherein, The gene mutation type of the hereditary deafness is a gene. GJB2 , GJB3 , SLC26A4 and 12S rRNA Point mutations; 2) Primers for detecting gene mutation types in thalassemia are shown in SEQ ID NO: 31~340, wherein the gene mutation type of thalassemia is a gene... HBA1 , HBA2 , HBB Point mutations, short-segment insertions and deletions, and large-segment deletions; 3) Primers for detecting gene mutation types in spinal muscular atrophy are shown in SEQ ID NO: 341-350, wherein the gene mutation type of spinal muscular atrophy is... SMN1 Point mutation and SMN1 / SMN2 Copy number variation in exons 7 and 8 of the gene; 4) Detection β-actin The primers for the housekeeping gene are shown in SEQ ID NO: 351-352.
2. A kit for simultaneously detecting gene mutation types associated with hereditary deafness, thalassemia, and spinal muscular atrophy, comprising the primer set as described in claim 1.
3. The kit according to claim 2, wherein, The kit also includes an amplification enzyme mix, primer digestion solution, amplification enhancer, index primers, and purified magnetic beads.
4. A system for detecting multiple disease-related gene mutation types, comprising: 1) Primer amplification module: Using the whole genome DNA of the sample as a template, targeted ultramultiplex PCR amplification is performed in the same reaction tube using the primer set described in claim 1, and the amplified fragments are purified using magnetic beads to remove primers and primer dimers; 2) Library Construction Module: Perform library construction on the amplified fragments obtained from module 1); 3) Data processing module: Based on the library constructed in module 2), high-throughput sequencing is used for sequencing. Bioinformatics methods are used to analyze the sequencing results to obtain genetic variation information of various gene mutation-related diseases, and then determine whether the sample has the corresponding gene mutation type. The diseases mentioned are hereditary deafness, thalassemia, and spinal muscular atrophy; Among them, the gene mutation type of the hereditary deafness is a gene GJB2 , GJB3 , SLC26A4 and 12S rRNA Point mutations; The gene mutation type of thalassemia mentioned above is a gene. HBA1 , HBA2 , HBB Point mutations, short-segment insertions and deletions, and large-segment deletions; The gene mutation type of spinal muscular atrophy is as follows: SMN1 Point mutation and SMN1 / SMN2 Copy number variation in exons 7 and 8 of the gene.
5. The system according to claim 4, wherein, The data processing module described in 3) further includes the following: 3.1: The library was sequenced using a high-throughput sequencer to obtain raw sequencing data; 3.2: Perform basic quality control on the raw sequencing data obtained in step 3.1, including filtering out sequencing adapters and low-quality bases, to obtain valid data. Use BWT algorithm software to align the valid data with the human reference genome to obtain alignment information files with base positions. 3.3: Based on the alignment information file obtained in step 3.2, the alignment information file is split into two parts: an autosomal alignment information file and a mitochondrial genome alignment information file; according to the autosomal alignment information file, a Bayesian algorithm is used to detect autosomal SNPs and INDEL mutations; mitochondrial SNP mutation information is obtained by scoring, sorting, and filtering the results of the mitochondrial genome alignment information file.
6. The system according to claim 4, wherein, The methods for determining the gene mutation types associated with hereditary deafness, thalassemia, and spinal muscular atrophy are as follows: 1) Methods for determining autosomal SNP and INDEL mutations in hereditary deafness, thalassemia, and spinal muscular atrophy: If the base mutation frequency (VAF) is >95%, it is homozygous; if the base mutation frequency is 15% ≤ 95%, it is heterozygous. 2) Method for determining mitochondrial SNP mutations in hereditary deafness: If the base mutation frequency (VAF) is >3%, then a mutation exists; otherwise, no mutation exists. 3) By SMN1 Methods for diagnosing spinal muscular atrophy caused by gene deletion: a) Extract the compared information from the comparison file β-actin , SMN1 exon7, exon8 SMN2 The read lengths of the exon7 and exon8 regions are filtered out, and read lengths with an alignment quality value <60 and >3 base mismatches are removed to obtain the filtered alignment information file. b) Based on the filtered alignment information file from step a), calculate the average coverage depth of the target region using the following methods. SMN1 , SMN2 The average coverage depth of Exon7 and Exon8 divided by β-actin The average coverage depth was used to obtain the copy number of S1-7, S1-8, S2-7, and S2-8; c) Obtain the data sets of negative samples S1-7, S1-8, S2-7, and S2-8 through known and verified SMA negative samples; calculate the median of S1-7 and S1-8; d) Calculate the multiples of S1-7 and S1-8 of the test sample and the SMA negative sample set, and determine the deletion type of the sample; e) Determine whether the test sample is a patient or carrier of spinal muscular atrophy; the specific determination method is as follows: S1-7= SMN1 The average coverage depth of the Exon7 / β-actin The average coverage depth; S1-8= SMN1 average coverage depth of exon8 / β-actin The average coverage depth; S2-7= SMN2 The average coverage depth of the Exon7 / β-actin The average coverage depth; S2-8= SMN2 average coverage depth of exon8 / β-actin The average coverage depth; Ratio(S1-7) = S1-7 / median of S1-7 in the negative control set; Ratio(S1-8) = S1-8 / median of S1-8 in the negative control set; like SMN1 If the gene determination values Ratio(S1-7) ≤ 0.2 and Ratio(S1-8) ≤ 0.2, then the sample to be tested is a homozygous deletion, i.e. SMN1 The copy number is 0; If Ratio(S1-7) ≤ 0.2 and 0.2 < Ratio(S1-8) ≤ 0.7, the test sample has a homozygous deletion of exon7 and a heterozygous deletion of exon8; If 0.2 < Ratio(S1-7) ≤ 0.7 and Ratio(S1-8) ≤ 0.2, the test sample has a heterozygous deletion of exon7 and a homozygous deletion of exon8; If 0.2 < Ratio(S1-7) ≤ 0.7 and 0.2 < Ratio(S1-8) ≤ 0.7, then the test sample is SMN1 heterozygous deletion; If S2-7 or S2-8 ≤ 0.3, then SMN2 exon7 or exon8 is 0 copies; If 0.3 < S2-7 or S2-8 ≤ 0.7, then SMN2 exon7 or exon8 is in one copy; If 0.7 < S2-7 or S2-8 ≤ 1.2, then SMN2 exon7 or exon8 is in two copies; If 1.2 < S2-7 or S2-8 ≤ 1.7, then SMN2 exon7 or exon8 is in three copies; If 1.7 < S2-7 or S2-8 ≤ 2.2, then SMN2 exon7 or exon8 is 4 copies; If S2-7 or S2-8 > 2.2, then SMN2 Exon7 or Exon8 requires 5 or more copies; 4) By HBA2 , HBA1 , HBB Methods for diagnosing thalassemia caused by gene deletion: a) Statistically determine the read length and coverage depth of the target area, and based on... β-actin The average sequencing depth of the gene is normalized to the sequencing depth of the target region to obtain the Z-value, where Z-value = sequencing depth of target region / β-actin Average sequencing depth ; b) Calculate the ratio of the Z value of the test sample to the median of the Z values of the negative control sample set, and determine the mutation breakpoint based on the HMM algorithm; c) Determine the deletion copy number of the thalassemia-related gene in the test sample, and the judgment criteria are as follows: Judgment value B = log2(test sample Z value / median of Z values of the negative control sample set) If the judgment value B ≤ -1.4, the test sample has a homozygous deletion, that is, the gene copy number is 0; If -1.4 < judgment value B ≤ -0.6, the test sample has a heterozygous deletion, that is, the gene copy number is 1; If -0.6 < judgment value B ≤ 0.6, the test sample has a normal copy number, that is, the gene copy number is 2.
7. The use of the primer set according to claim 1 in the preparation of a kit for detecting multiple disease-related gene mutation types, wherein, The diseases are hereditary deafness, thalassemia, and spinal muscular atrophy.