HBA and HBB mutation detection system and methods for non-disease diagnostic purposes
The HBA and HBB mutation detection system employs high molecular weight DNA extraction, targeted enrichment, and single-molecule real-time sequencing technologies to solve the challenges of detecting complex variant types and linkage detection of HBA and HBB genes in existing technologies. This enables high-precision gene mutation detection, particularly for the molecular diagnosis of thalassemia.
Patent Information
- Application Number
- CN202411980667.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing technologies are insufficient to efficiently and accurately detect complex variant types of HBA and HBB genes and determine whether they are linked, especially in terms of their ability to detect complex structural variants and unknown mutations. Furthermore, traditional methods cannot detect multiple mutation types simultaneously.
The HBA and HBB mutation detection system includes a collection module, a Y-shaped pre-library preparation module, a targeted enrichment module, a single-molecule real-time sequencing library construction module, and a single-molecule real-time sequencing analysis module. Through high molecular weight DNA extraction, end repair, long fragment PCR amplification, targeted enrichment, and single-molecule real-time sequencing, it achieves accurate detection of HBA and HBB genes.
It enables precise detection of HBA and HBB genes, distinguishes complex structural variations beyond routine testing, accurately determines gene mutation types, including large fragment deletions and unknown mutations, and has high detection accuracy with a base accuracy greater than 99%, making it suitable for the molecular diagnosis of thalassemia.
Smart Images

Figure CN119753115B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene detection, and in particular to an HBA and HBB mutation detection system and a method for non-disease diagnosis purposes. Background Art
[0002] Thalassemia is a common inherited anemia with high morbidity and mortality worldwide, particularly in Southeast Asia, the Middle East, and the Mediterranean region. Alpha- and beta-thalassemia are the two main types of thalassemia, caused by mutations in the HBA1 / 2 and HBB genes, resulting in abnormal alpha and beta globin synthesis and structural defects in hemoglobin, respectively.
[0003] Because of the complexity of the HBA gene locus and the need to analyze both HBA and HBB genes, it is very challenging to perform a joint analysis for the molecular diagnosis of different stages of thalassemia (silent, mild and moderate) and disease states. The HBA locus contains two essentially identical genes, HBA1 and HBA2, which encode the α-hemoglobin chain. α-thalassemia is the most common genetic mutation caused by the partial or complete loss of function of the HBA1 and HBA2 genes. It includes large deletions -THAI, -FIL, -MED and -SEA and small deletions -a 3.7 and -a 4.2 According to the Ithanet, HbVar, LOVD, and LOVD-China thalassemia databases, approximately 50 deletions of the α-globin gene cluster region and more than 900 non-deletion mutations of the HBA1 and HBA2 genes have been found worldwide, which can lead to reduced or absent α-globin expression. In addition, there are some rare structural variations, such as ααα anti3.7 , ααα anti4.2 , HKαα, and anti-HKαα. β-thalassemia is relatively simpler, consisting of a single HBB gene encoding the β-hemoglobin chain. The primary genetic lesions found in the HBB gene are single nucleotide variants (SNVs), insertions, and deletions. To date, over 200 pathogenic variants have been identified worldwide.
[0004] Molecular diagnostic methods can effectively detect mutations in the HBA1 / 2 and HBB genes. There are currently a variety of molecular diagnostic methods for HBA1 / 2 and HBB genes, such as gap-crossing PCR (Gap-PCR), PCR reverse dot hybridization (PCR-RDB), PCR oligonucleotide probe method (PCR-ASO), and next-generation sequencing (NGS) technology. However, the current traditional detection methods are not only cumbersome to operate, but also cannot accurately detect complex structural variations, such as the inability to distinguish between HKαα and -α 3.7, and cannot distinguish between antiHKαα and -α 4.2 At the same time, it is difficult to simultaneously detect all mutation types of HBA and HBB, especially uncommon mutation types. In addition, some patients may have two HBA and HBB gene mutations at the same time. These traditional detection methods are unable to detect whether the two mutations are linked. Although some detection methods have been derived based on third-generation sequencing, such as using different primers to perform long-PCR on the target fragment to detect the target mutation, these methods have the limitation of only being able to detect known mutations, and PCR has bias. Summary of the Invention
[0005] In view of the above-mentioned deficiencies in the prior art, the present invention aims to provide an HBA and HBB mutation detection system and a method for non-disease diagnosis purposes, so as to detect the complex variation types of the HBA gene and whether the two mutations of the HBA and HBB genes are linked.
[0006] In order to solve the above problems, the present invention adopts the following technical solutions:
[0007] In one aspect, the present invention provides an HBA and HBB mutation detection system, comprising an acquisition module, a Y-shaped pre-library preparation module, a targeted enrichment module, a single-molecule real-time sequencing library construction module, and a single-molecule real-time sequencing analysis module;
[0008] The collection module is used to collect peripheral blood from the subject to be tested and to extract and shear high molecular weight DNA;
[0009] The Y-shaped pre-library preparation module is used to perform end repair and 3'A-tailing reaction on the extracted and sheared high molecular weight DNA, and perform Y-shaped adapter ligation, and then use a long-fragment amplification enzyme to perform long-fragment PCR amplification to obtain a Y-shaped pre-library;
[0010] The targeted enrichment module is used to hybridize and target-enrich the target gene using the target region capture probe based on the Y-shaped pre-library to obtain a capture library;
[0011] The single-molecule real-time sequencing library construction module is used to perform sequencing adapter ligation based on the capture library to obtain a single-molecule real-time sequencing library;
[0012] The single-molecule real-time sequencing analysis module is used to perform single-molecule real-time sequencing using a single-molecule real-time sequencing library to obtain sequencing data of a target region, decompose the sequencing data of the target region, and analyze the variation types of HBA and HBB.
[0013] Furthermore, the target region capture range of the target region capture probe labeled with biotin is: HBA1: chr16: 203679-207521, HBA2: chr16: 202875-203709 and HBB: chr11: 5226694-5228301;
[0014] The nucleotide sequences of the target region capture probes labeled with biotin are shown in SEQ ID No: 1 to SEQ ID No: 232.
[0015] In another aspect, the present invention provides a method for detecting HBA and HBB mutations for non-disease diagnosis purposes, comprising:
[0016] Collect peripheral blood from the subjects and perform high molecular weight DNA extraction and shearing;
[0017] The extracted and sheared high molecular weight DNA is subjected to end repair and 3'A-tailing reaction, and Y-shaped adapter ligation is performed. Then, long-fragment PCR amplification is performed using a long-fragment amplification enzyme to obtain a Y-shaped pre-library.
[0018] Based on the Y-shaped prelibrary, target region capture probes are used to hybridize and target enrich the target gene to obtain a capture library;
[0019] Based on the captured library, sequencing adapters are connected to obtain a single-molecule real-time sequencing library;
[0020] Single-molecule real-time sequencing was performed using a single-molecule real-time sequencing library to obtain sequencing data of the target region. The sequencing data of the target region was decomposed and analyzed to obtain the variation types of HBA and HBB.
[0021] The present invention offers the following benefits: precise detection results. Beyond conventional testing, it can accurately distinguish HBA1 / HBA2 mutations and precisely detect complex structural variations such as large deletions and triplets. It can also determine whether HBA and HBB gene mutations are linked, as well as detect previously unreported, unknown mutations. This high level of accuracy is evident in the PacBio dumbbell-shaped library, which can undergo multiple rounds of sequencing. After correction, the sequencing results achieve a base accuracy exceeding 99%. Furthermore, PacBio sequencing errors are random, and deep sequencing correction results in base accuracy exceeding 99.9%, enabling precise interpretation of gene mutations within the primer detection range. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a quality inspection graph of the fragment length distribution of the captured library of the present invention; wherein, LM20 represents a band with a size of 20 bp of DNA in the quality inspection marker, and LM1000 represents a band with a size of 1000 bp of DNA in the quality inspection marker.
[0023] Figure 2 HBA of the present invention, -α 3.7 PacBio sequencing results of / αα gene mutation samples.
[0024] Figure 3 This is a diagram showing the PacBio test results for the HBB gene mutation sample of the present invention. DETAILED DESCRIPTION
[0025] The present invention will be further described in detail below with reference to specific embodiments.
[0026] It should be noted that these embodiments are only used to illustrate the present invention, rather than to limit the present invention. Simple improvements to the method based on the concept of the present invention fall within the scope of protection claimed by the present invention.
[0027] Example 1 Construction of HBA / HBB mutation detection process
[0028] 1. Collect 10 mL of peripheral blood from the patient using an EDTA anticoagulant tube.
[0029] 2. Genomic DNA extraction.
[0030] The peripheral blood of the subjects to be tested was extracted using a high molecular weight DNA extraction kit to obtain high molecular weight gDNA, and the concentration was measured and the integrity of the genomic DNA was detected by agarose gel electrophoresis.
[0031] 3. Fragmentation of high molecular weight gDNA
[0032] 3 μg of genomic DNA was aspirated and sheared using a Covaris g-tube centrifuge at 6000 rpm for 2 min.
[0033] 4. Screening of fragmented DNA
[0034] (1) AmpureXP beads (Beckman, A63881) were diluted with 1 mM Tris-HCl pH 8 at a ratio of 1×, 2×, 3×, or 4×.
[0035] (2) Mix the fragmented DNA to be screened with the diluted fragment screening magnetic beads in a ratio of 1:1, 1:2, 1:3, or 1:4.
[0036] (3) Briefly centrifuge the tube to concentrate the magnetic bead / DNA mixture at the bottom of the tube. Place the tube on a magnetic rack for 5-10 minutes or until the magnetic beads gather on the side of the tube and the supernatant becomes clear. Slowly remove the clear supernatant to avoid disturbing the magnetic beads.
[0037] (4) Wash the magnetic beads twice with fresh 80% ethanol. After 30 seconds, remove the 80% ethanol with a pipette.
[0038] (5) The cleaned magnetic beads are eluted to obtain fragmented DNA after fragment screening.
[0039] 5. Perform Qsep100 quality control on the fragmented DNA after fragment screening using the S3 cartridge (Bioptic, c105106) for large fragment detection. The proportion of fragments below 3 kbp should be less than 10%, otherwise it will be considered as unqualified fragmented DNA.
[0040] 6. Preparation of pre-library
[0041] (1) Use a long fragment dedicated library construction kit to construct a long fragment Y-shaped pre-library. Dilute the purified fragmented DNA that has passed the quality inspection to 50 μL, and prepare the end repair and 3'A-tailing reactions according to Table 1.
[0042] Table 1 Preparation of end repair and 3'A-tailing reaction reagents
[0043]
[0044] Set up the PCR program according to Table 2 and perform end repair and 3'A-tailing reactions.
[0045] Table 2 PCR program for end repair and 3'A-tailing reaction
[0046]
[0047] (2) Adapter ligation: Add the following components in order according to Table 3 to the end repair and 3'A-tailing reaction products.
[0048] Table 3 Connector ligation reagents
[0049]
[0050] Vortex to mix, centrifuge briefly to ensure that all the solution is at the bottom of the tube, and immediately incubate in a PCR instrument at 20°C for 10 min, 20 min, 30 min, or 1 h with the heated lid set to 85°C.
[0051] Adapter sequence
[0052] A.ACACTCTTTCCCTACACGACGCTCTTCCGATC*T
[0053] BP-GATCGGAAGAGCACACGTCT
[0054] P: Phosphorylation*Phosphorylation
[0055] (3) Fragment screening of adapter-ligated products
[0056] ① Dilute AmpureXP beads (Beckman, A63881) with 1 mM Tris-HCl pH 8 at 1×, 2×, 3×, or 4× dilution ratios.
[0057] ② Mix the adapter-ligated products to be screened with the diluted fragment screening magnetic beads in a ratio of 1:1, 1:2, 1:3 or 1:4.
[0058] ③ Briefly centrifuge the tube to concentrate the magnetic bead / DNA mixture at the bottom of the tube. Place the tube on a magnetic rack for 5-10 minutes, or until the magnetic beads collect on the side of the tube and the supernatant becomes clear. Slowly remove the clear supernatant to avoid disturbing the magnetic beads.
[0059] ④ Wash the magnetic beads twice with fresh 80% ethanol. After 30 seconds, aspirate the 80% ethanol with a pipette.
[0060] ⑤ Elute the cleaned magnetic beads to obtain the adapter-ligated products after fragment screening.
[0061] (3) The ligation product was subjected to long-fragment PCR amplification using a dedicated enzyme for long-fragment amplification. The reaction system was configured according to Table 4; the dedicated enzyme for long-fragment amplification was Takara LA Taq DNA Polymerase Hot-Start Version (Takara, RR042A).
[0062] Table 4 Long fragment PCR amplification reagents
[0063]
[0064] Different tags can be single nucleotide sequences of lengths such as 8nt, 10nt, or 12nt.
[0065] Index primer sequence
[0066] Sequence 1:
[0067] CAAGCAGAAGACGGCATACGAGATNNNNNNNNNGTGACTGGAGTTCAGACGTGTGCTCTTCCGATC
[0068] Sequence 2:
[0069] AATGATACGGCGACCACCGAGATCTACACNNNNNNNNNAACACTCTTTCCCTACACGACGCTCTCTCCGATCT.
[0070] N is any base among A, T, C and G.
[0071] Perform PCR reaction according to Table 5.
[0072] Table 5 Long fragment PCR amplification program
[0073]
[0074] (4) Fragment screening was performed on the Y-shaped pre-library PCR products, aiming to control the short fragments below 3 kb to within 10%.
[0075] ① Dilute AmpureXP beads (Beckman, A63881) with 1 mM Tris-HCl pH 8 at 1×, 2×, 3×, or 4× dilution ratios.
[0076] ② Mix the Y-shaped pre-library PCR product to be screened with the diluted fragment screening magnetic beads in a ratio of 1:1, 1:2, 1:3 or 1:4.
[0077] ③ Centrifuge briefly to concentrate the magnetic bead / DNA mixture at the bottom of the tube. Place the tube on a magnetic rack for 5-10 minutes, or until the magnetic beads collect on the side of the tube and the supernatant becomes clear. Slowly remove the clear supernatant to avoid disturbing the magnetic beads.
[0078] ④ Wash the magnetic beads twice with fresh 80% ethanol. After 30 seconds, aspirate the 80% ethanol with a pipette.
[0079] ⑤ Elute the cleaned magnetic beads to obtain the Y-shaped pre-library PCR product after fragment screening.
[0080] ⑥ Use the S3 cartridge for large fragment detection (Houze, c105106) to perform Qsep100 quality inspection on the Y-shaped pre-library PCR products after fragment screening. The proportion of fragments below 3 kbp should be less than 10%, otherwise it will be regarded as unqualified Y-shaped pre-library PCR products.
[0081] 7. Hybridization of the Y-shaped prelibrary with the biotin-labeled double-stranded DNA probe
[0082] The components of the capture kit used in this example refer to patent ZL202011284834.1
[0083] (1) In this embodiment, the target gene region is shown in Table 6. The GRCh37 / hg19 gene sequence is used as the reference genome sequence from databases such as UCSC or NCBI. The target gene region and the 500 bp-1.5 kb upstream and downstream extension sequence are extracted. After removing the repeated regions, the probe is designed. 120 bp from the first base is cut as the first probe sequence. At the same time, the probe density is adjusted according to the GC content. The HBA probe region is GRCh37 / hg19, chr16: 201565-208687; the HBB probe region is GRCh37 / hg19, chr11: 5225768-5229935. The capture probe used to enrich the target gene and flanking regions is a biotin-labeled double-stranded DNA probe. The probe sequence table is shown in Table 7.
[0084] Table 6 Target area capture range table
[0085] Gene chromosome Start Finish HBA1 chr16 203679 207521 HBA2 chr16 202875 203709 HBB chr11 5226694 5228301
[0086] Table 7 Probe sequence list
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101] (2) Hybridization reaction can be performed as single hybrid or multi-hybrid. For multi-hybrid, multiple Y-shaped prelibraries need to be mixed with equal molecular weight according to the following formula: fmol = (concentration * 10 6 ) / [650*length(bp)].
[0102] (3) Vacuum concentrate a single library or a mixed library of equal molecular weight to a dry powder state, and resuspend the library according to the composition in Table 8.
[0103] Table 8 Library resuspension reagents
[0104]
[0105] (4) Denature the resuspended library and probe according to the PCR program in Table 9.
[0106] Table 9 Denaturation PCR program
[0107]
[0108]
[0109] (5) Add 20 μl of hybridization buffer preheated at 65°C for at least 5 min to the denatured prelibrary and probe mixture. Incubate at 65°C with a heated lid at 85°C for 16-18 hours.
[0110] 8. Targeted Enrichment
[0111] (1) Preheat the streptavidin magnetic beads at 65°C for at least 10 minutes, then transfer the hybridization product directly from the 65°C condition to the 65°C streptavidin magnetic beads. Be sure not to cool the hybridization product temperature to room temperature, otherwise it will cause serious off-target effects.
[0112] (2) After the hybridization product is bound to the streptavidin magnetic beads, the hybridization product and the streptavidin magnetic beads are rinsed and eluted.
[0113] (3) The eluted product was captured and then PCR was performed using a long fragment specific amplification enzyme. The amplification system is shown in Table 10.
[0114] Table 10 Capture library amplification system
[0115]
[0116] Universal primer sequence:
[0117] AATGATACGGCGACCACCGAG
[0118] CAAGCAGAAGACGGCATACGA.
[0119] Amplify according to the amplification program in Table 11.
[0120] Table 11 Capture library amplification program
[0121]
[0122]
[0123] (4) Purify the PCR product using 0.8* purification magnetic beads to obtain the capture library.
[0124] (5) The capture library is subjected to concentration and fragment length distribution quality inspection by Qubit and Qsep100. The capture library length should be 3k-10k, see Figure 1 .
[0125] 9. PacBio third-generation library construction (kit from PacBio)
[0126] (1) DNA damage repair
[0127] Configure the DNA damage repair system according to Table 12.
[0128] Table 12 Damage repair system
[0129]
[0130] The prepared reaction solution was shaken and mixed. The reaction conditions were: 37°C, 20 min; maintain at 4°C.
[0131] (2) End repair and purification
[0132] Configure the end repair and purification reaction system according to Table 13.
[0133] Table 13 End repair and purification reaction system
[0134]
[0135] The prepared reaction solution was shaken and mixed. The reaction conditions were: 25°C, 5 min; maintain at 4°C.
[0136] (3) Connector connection
[0137] Configure the adapter ligation reaction system according to Table 14.
[0138] Table 14 Linker ligation reaction system
[0139]
[0140] The prepared reaction solution was shaken and mixed. The reaction conditions were: 25°C, 24h; 65°C, 10min; and maintained at 4°C.
[0141] (4) Digestion of DNA and adapter sequences where ligation failed.
[0142] The reaction system was configured according to Table 15.
[0143] Table 15 Digestion reaction system
[0144]
[0145] The prepared reaction solution was shaken and mixed. The reaction conditions were: 37°C, 1h; maintain at 4°C.
[0146] 7. Sequencing and result analysis
[0147] The prepared library was sequenced using a third-generation single-molecule sequencer (PacBio Revio), so that the average sequencing depth of the target region reached more than 50×.
[0148] The obtained data is split and analyzed. The specific analysis method is as follows:
[0149] 1) Sample data splitting:
[0150] Lima software was used to split the PacBio sequencing data based on the sample index and remove the index sequence to obtain clean data. This step ensures the purity and accuracy of the sample data and lays the foundation for subsequent analysis.
[0151] 2) Alignment to the reference genome:
[0152] Sequencing data were aligned to the human reference genome (hg19) using Minimap2. During the alignment process, all non-human sequences that did not map to the reference genome were removed to ensure data accuracy and specificity.
[0153] 3) Data quality control:
[0154] NanoPlot was used to perform quality control statistical analysis on the sequencing data of the target region, covering the following indicators:
[0155] Data volume: Count the total amount of sequencing data to ensure sufficient data;
[0156] Mapping rate: the ratio of sequencing data of the target region mapped to the reference genome (hg19);
[0157] Fragment length: Evaluate the fragment length distribution of the data to determine the quality of the sequencing data;
[0158] Capture efficiency: Evaluate the capture efficiency of the target area to ensure that the target area is adequately covered.
[0159] Coverage statistics: Ensure that the coverage depth of the target region is greater than 50X. Coverage is a key indicator for determining whether sequencing data is sufficient. High coverage helps improve the accuracy of variant detection.
[0160] 4) SNV analysis:
[0161] Single nucleotide variant (SNV) analysis was performed on the BAM files of the samples using DeepVariant. DeepVariant is a deep learning-based variant detection tool with high accuracy and sensitivity, capable of identifying various SNVs in the HBA and HBB genes.
[0162] 5)SV analysis:
[0163] Use cuteSV to perform structural variation (SV) analysis on the sample BAM files. SV analysis can identify a wide range of genomic rearrangements, such as deletions, insertions, and inversions. Compare the SV analysis results with known common variant types (such as 3.7, 4.2, 2.4, SEA, Hong Kong, and Thai types) to accurately determine the mutation type in the target region.
[0164] 6) Variant annotation and disease association analysis:
[0165] Annovar and VEP were used to annotate SNVs in detail. The annotation contents included:
[0166] Mutation information: such as nucleotide changes, amino acid changes, HGVS nomenclature, genotype, etc.;
[0167] Population frequency: for example, the frequency of variants in the 1000g and gnomAD databases;
[0168] Disease and phenotype information: including ACMG pathogenicity judgment of the variant, inheritance mode (such as dominant, recessive), disease phenotypic characteristics, etc.
[0169] Prediction tools and scoring: such as spliceAI, SIFT, PolyPhen-2, REVEL_score, GERP++, etc., are used to assess the potential impact of variants on gene function.
[0170] According to the above construction process, an HBA and HBB mutation detection system was constructed, including an acquisition module, a Y-shaped pre-library preparation module, a targeted enrichment module, a single-molecule real-time sequencing library construction module, and a single-molecule real-time sequencing analysis module;
[0171] The collection module is used to collect peripheral blood from the patient and extract and shear high molecular weight DNA;
[0172] The Y-shaped pre-library preparation module is used to perform end repair and 3'A-tailing reaction on the extracted and sheared high-molecular-weight DNA, and perform Y-shaped adapter ligation. Then, long-fragment PCR amplification is performed using a long-fragment amplification enzyme to obtain a Y-shaped pre-library.
[0173] The targeted enrichment module is used to hybridize and target-enrich the target gene using the target region capture probe based on the Y-shaped pre-library to obtain the capture library;
[0174] The single-molecule real-time sequencing library construction module is used to perform sequencing adapter ligation based on the captured library to obtain a single-molecule real-time sequencing library;
[0175] The single-molecule real-time sequencing analysis module is used to perform single-molecule real-time sequencing using a single-molecule real-time sequencing library to obtain sequencing data of the target region, decompose the sequencing data of the target region, and analyze the variation types of HBA and HBB.
[0176] Example 2
[0177] Ten positive test samples were tested using the HBA and HBB mutation detection systems described in Example 1. The results demonstrate that the present invention can intuitively and accurately detect the mutation types of all positive samples (Table 16). Furthermore, for samples 3, 5, 8, and 9 where the traditional method had ambiguous judgments, the mutation types can be clearly determined, and the breakpoint locations can be accurately identified. Samples with inconsistent results were also validated using PCR or PCR-Sanger sequencing, and the results were consistent with those of the present invention.
[0178] Table 16 Test sample test results
[0179]
[0180]
[0181] illustrate:
[0182] Sample 1: The detection results of the two methods are consistent, and the present invention can clearly see the breakpoint position.
[0183] Sample 2: The detection results of the two methods are consistent, and the present invention can clearly see the breakpoint position.
[0184] Sample 3: See Figure 2 、 Figure 3 The present invention has a wider detection range. The second-generation sequencing data fluctuated greatly and it was impossible to determine whether HBA was a 3.7kb deletion. After repeated verification by PCR-Sanger, it was indeed HBA.
[0185] -α3.7 / αα, the method used in the present invention directly indicates HBA, -α3.7 / αα, and the breakpoint position can be clearly seen; the second generation of HBB point mutation is consistent with the method used in the present invention.
[0186] Sample 4: The detection results of the two methods are consistent, and the present invention can clearly see the breakpoint position.
[0187] Sample 5: The present invention can be used to distinguish -α 3.7 The second generation judgment results are relatively vague, and the method used in the present invention can directly judge it as ααα anti4.2 / -α 3.7 Compound mutations were detected, and the breakpoint positions were clearly visible. PCR-Sanger analysis confirmed the results, which were consistent with those of the present invention.
[0188] Sample 6: The detection results of the two methods are consistent, and the present invention can clearly see the breakpoint position.
[0189] Sample 7: The detection results of the two methods are consistent, and the present invention can clearly see the breakpoint position.
[0190] Sample 8: The present invention has a wider detection range. The missing 29kb fragment region includes 3.7kb and 2.4kb regions. The second generation sequencing can only roughly determine it as HBA. 3.7 / -α 2.4 The method used in the present invention can directly determine the unknown 29kb deletion mutation and clearly see the breakpoint position.
[0191] Sample 9: The present invention has a wider range of structural variation detection. The method of the present invention can directly determine the variation type as HBA, αααα anti3.7 / ααα anti3.7 , and the breakpoint position can be clearly seen. This is consistent with the PCR-Sanger repeated verification results.
[0192] Sample 10: The detection results of the two methods are consistent, and the present invention can clearly see the breakpoint position.
[0193] In summary, the present invention can detect complex mutation types such as HBA / HBB point mutations, large deletions, multi-copy mutations, Hong Kong type and trans-Hong Kong type, and can accurately identify cis- and trans-mutations without the need for family verification and negative sample controls.
[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described with reference to the preferred embodiments of the present invention, it should be understood by those skilled in the art that various changes can be made in form and details without departing from the spirit and scope of the present invention as defined in the appended claims.
Claims
1. A HBA and HBB mutation detection system, characterized in that: It includes acquisition module, Y-shaped pre-library preparation module, targeted enrichment module, single-molecule real-time sequencing library construction module and single-molecule real-time sequencing analysis module; The collection module is used to collect peripheral blood from the subject to be tested and to extract and shear high molecular weight DNA; The Y-shaped pre-library preparation module is used to perform end repair and 3'A-tailing reaction on the extracted and sheared high molecular weight DNA, and perform Y-shaped adapter ligation, and then use a long-fragment amplification enzyme to perform long-fragment PCR amplification to obtain a Y-shaped pre-library; The targeted enrichment module is used to hybridize and target-enrich the target gene using the target region capture probe based on the Y-shaped pre-library to obtain a capture library; The single-molecule real-time sequencing library construction module is used to perform sequencing adapter ligation based on the capture library to obtain a single-molecule real-time sequencing library; The single-molecule real-time sequencing analysis module is used to perform single-molecule real-time sequencing using a single-molecule real-time sequencing library to obtain sequencing data of a target region, decompose the sequencing data of the target region, and analyze the variation types of HBA and HBB; The target region capture ranges of the biotin-labeled target region capture probes are: HBA1: chr16: 203679-207521, HBA2: chr16: 202875-203709, and HBB: chr11: 5226694-5228301; The nucleotide sequences of the target region capture probes labeled with biotin are shown in SEQ ID No: 1 to SEQ ID No:
232.
2. The HBA and HBB mutation detection system according to claim 1, characterized in that: The short fragments of less than 3 kb in the Y-shaped pre-library are less than 10%.
3. The HBA and HBB mutation detection system according to claim 1, characterized in that: The PCR program for end repair and 3'A-tailing reaction of the extracted and sheared high molecular weight DNA is: 20°C, 30 minutes, 65°C, 30 minutes, and 4°C hold; the PCR program for long-fragment PCR amplification using a long-fragment amplification enzyme is: 94°C, 2 minutes; 98°C, 10 seconds, 58.8°C, 30 seconds, 68°C, 15 minutes, repeated 6 times; 68°C, 10 minutes; and 4°C hold.
4. The HBA and HBB mutation detection system according to claim 1, characterized in that: The method comprises hybridizing the Y-shaped prelibrary with a target region capture probe and performing targeted enrichment on the target gene to obtain a capture library, comprising: hybridizing the Y-shaped prelibrary with a target region capture probe labeled with biotin, binding the hybridized target DNA fragments with magnetic beads labeled with streptavidin, eluting the target DNA fragments, and performing long-fragment PCR amplification on the eluted target DNA fragments to obtain a capture library.
5. The HBA and HBB mutation detection system according to claim 4, characterized in that: The program of the long-fragment PCR amplification is 94° C., 2 min; 98° C., 10 s, 58.8° C., 30 s, 68° C., 15 min, repeated 12 times; 68° C., 10 min; and hold at 4° C.
6. The HBA and HBB mutation detection system according to claim 1, characterized in that: The method of performing sequencing adapter ligation based on the captured library to obtain a single-molecule real-time sequencing library comprises: performing DNA damage repair on the captured library; Then perform end repair and purification; The purified products were subjected to third-generation sequencing adapter ligation; Digest unligated DNA fragments and next-generation sequencing adapters.
7. The HBA and HBB mutation detection system according to claim 1, characterized in that: The single molecule real-time sequencing analysis module is used for: Lima software was used to split the sequencing data of the target region into samples; Minimap2 was used to align the sequencing data to the hg19 genome and remove unaligned non-human sequences; NanoPlot was used to perform quality control statistical analysis on the sequencing data of the target region, including data volume, mapping rate, fragment length, capture efficiency, and coverage; Use DeepVariant to perform SNV analysis of HBA and HBB genes on the sample BAM files; Use cutSV to perform SV analysis on the sample BAM files; Annovar and VEP were used to annotate SNVs in detail, including variant information, population frequency, disease and phenotype information, and prediction tools and scores.
Citation Information
Patent Citations
Safe sequencing system
CN103748236A
Method and kit for simultaneously detecting multiple mutations of HBA1 / 2 and HBB gene loci
CN112708674A