High-throughput DNA (deoxyribonucleic acid) detection method, platform, probe composition and kit for hereditary hemoglobinopathy

By designing probe compositions and high-throughput DNA detection methods, combined with the Bioinformatics analysis platform, the problem of insufficient coverage and misdiagnosis of hereditary hemoglobin disease detection in traditional methods is solved, and efficient and accurate detection of a variety of hemoglobin mutations is achieved, supporting clinical diagnosis and treatment.

CN120290702AActive Publication Date: 2025-07-11RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510325231.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-11
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and comprehensively detect genetic mutations in hereditary hemoglobinopathy, especially inadequate coverage of common and rare mutations, resulting in misdiagnosis and misdiagnosis. The traditional methods are costly and time-consuming, and it is impossible to detect multiple target sites at the same time.

Method used

A probe composition is designed, including multiple types of probes covering the coding region and regulatory region of the hemoglobin gene, which can detect mutations in single nucleotides and mutations in highly homologous sequence regions, and combine high-throughput DNA detection methods and a bioinformatic analysis platform to conduct full coverage screening and differential diagnosis through NGS sequencing and data processing software.

Benefits of technology

It has achieved efficient detection of various hemoglobin mutation types, reduced the missed detection rate, accurately identified the genetic mutations carried by patients, supported clinical diagnosis and intervention treatment, and improved detection efficiency and coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120290702A_ABST
    Figure CN120290702A_ABST
Patent Text Reader

Abstract

The invention provides a high-throughput DNA (deoxyribonucleic acid) detection method, platform, probe composition and kit for hereditary hemoglobinopathy. The high-throughput DNA detection method comprises the following steps: S1, designing probes for detecting known and unknown mutations in abnormal hemoglobinopathy related genes according to a specific rule; s2, synthesizing a probe; s3, extracting genome DNA from peripheral blood of a patient, constructing a library, hybridizing and capturing a target sequence, recovering a target region library, and performing library quality detection; s4, loading the library into a sequencing platform for high-throughput sequencing to obtain sequencing original data; and S5, performing data analysis on the sequencing original data. According to the application, a probe with a special sequence is designed, and sequencing data is subjected to information analysis, so that multiple hemoglobin mutation types can be detected at the same time, the detection efficiency and coverage are greatly improved, missing detection is avoided, and differential diagnosis can be performed on diseases with similar clinical manifestations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of detection technologies, and in particular to a high-throughput DNA detection method and a bioinformatics analysis platform for hemoglobin and genes related to its regulation. Background Art

[0002] Although certain acquired factors can cause abnormal hemoglobin function or anemia, most anemias or hemolysis caused by hemoglobin defects are hereditary and are usually caused by mutations in globin genes and their regulatory genes. Hereditary hemoglobinopathies can be roughly divided into structural variations and synthesis disorders of hemoglobin. Essentially, they are all caused by changes in the bases of the primary structure. If the base change only slightly affects the spatial structure of hemoglobin, it is a structural variation; if it causes a disorder in the production of hemoglobin subunits, it is thalassemia, also known as Mediterranean anemia (thalassemia). Hemoglobin structural variations include HBC, HBD, HBE, HBM, HBS, HBJ, etc., and more than 1300 hemoglobin variants have been reported; hemoglobin synthesis disorders include various types of thalassemia such as α / β / δβ / γδβ. The pathogenesis and clinical manifestations of abnormal hemoglobin diseases are also different. Generally speaking, most hemoglobin structural variations with clinical significance will cause hemolysis; while thalassemia is ineffective erythropoiesis and the destruction of erythroid progenitor cells in the bone marrow; sickle cell disease (SCD) can also cause vaso-occlusive phenomena. At the same time, some red blood cell membrane diseases and enzyme diseases, etc., will also have similar clinical manifestations and need to be differentiated.

[0003] Most hemoglobin structural variations are autosomal dominant inheritances. Although the symptoms are mild, as long as the gene is carried, there will be symptoms. Thalassemia has been considered a heterogeneous hereditary disease, which can range from the mildest asymptomatic carriers to transfusion-dependent anemia. However, in the case of infection or stress, some asymptomatic carriers on weekdays will show mild anemia symptoms. For a long time, preventing the birth of severely thalassemic children has been the focus of laboratory tests. However, with the large-scale migration of the population and people's higher demand for improving the quality of life, accurately diagnosing carriers of hemoglobin gene mutations has increasingly shown its necessity.

[0004] Traditional diagnostic methods have various limitations. Currently, the methods for detecting gene mutations in molecular biology mainly include first-generation sequencing (Sanger), next-generation sequencing (NGS), real-time fluorescence PCR (RT-PCR), digital PCR (dPCR), etc. Among them, most of Sanger, RT-PCR, and dPCR are used to detect known mutation sites in genes, and the costs of primers and probes are relatively high during the experimental process, making it less feasible for multi-gene mutation screening. Traditional methods for detecting hemoglobinopathies include hematological examinations, hemoglobin electrophoresis, PCR detection, and Sanger sequencing. Hematological examinations mainly rely on peripheral blood red blood cell parameters (such as MCV, MCH), combined with hemoglobin electrophoresis to analyze the proportions of different types of hemoglobin, and initially screen for suspected carriers. However, these methods cannot distinguish specific gene mutation types, and are less sensitive to deletion mutations of α-thalassemia, especially in mild or silent carriers, who are easily misdiagnosed. For the detection methods of hemoglobin, currently commonly used techniques include HPLC, IEF, capillary electrophoresis, etc. Although these methods have relatively low costs, due to the characteristics of the technology, many abnormal hemoglobin peaks overlap in similar positions and cannot be distinguished, resulting in misdiagnosis or missed diagnosis.

[0005] Although the role of first-line DNA testing is evolving, DNA testing is rarely used as a screening test for abnormal hemoglobinopathies. Routinely using DNA testing as a second-line test method is becoming the medical standard. For patients suspected of having hemoglobinopathies, molecular testing is the gold standard for diagnosing the disease. Currently, the commonly used methods in China are GAP-PCR, reverse dot blot, etc. designed based on the known high-frequency mutations in the Chinese population. Although they have high cost performance and can diagnose most thalassemia patients, they cannot clearly diagnose uncommon mutations, especially hemoglobinopathies other than HBE, and basically do not cover symptomatically very mild carriers. Traditional research-oriented PCR methods design primers for specific known point mutations or deletions to detect common mutations, but are ineffective for rare or unknown mutations and cannot cover multiple target sites simultaneously. Sanger sequencing is widely used for gene mutation analysis, but this method is time-consuming, has low throughput, is not suitable for complex gene detection and large-scale screening, and has limited ability to detect gene copy number variations (CNVs) and tandem repeat sequences (such as HBA1 and HBA2 genes). Although the MLPA method can detect some repeat and rearrangement sequences of globin, it still has the disadvantages of a relatively small detectable region and high cost. Although several kits for detecting thalassemia using NGS technology have been patented and some are already used in third-party testing centers, we found that some use the amplicon method, some have imperfect coverage, need to be continuously updated, and still have a certain missed detection rate. Summary of the Invention

[0006] To solve the above technical problems in the prior art, the present application provides four aspects of technical solutions, which are specifically as follows:

[0007] The first aspect of the present application provides a probe composition, the nucleic acid sequence of which includes, but is not limited to, the probes shown in SEQ ID NO. 1 to 64.

[0008] The second aspect of the present application provides a kit for detecting genes related to hereditary hemoglobinopathy, which includes the following two types of probes:

[0009] The first type of probe, which is used to detect all coding regions of hemoglobin genes, all non-coding regions including the regulatory regions of the hemoglobin gene cluster, transcription factors, the gene coding region and regulatory region genes of the gene HBS1L-MYB that regulates hemoglobin switching;

[0010] The second type of probe, which is used to detect single nucleotide mutations in the highly homologous sequence regions of hereditary hemoglobinopathy.

[0011] Furthermore, the kit for detecting genes related to hereditary hemoglobinopathy further includes a third type of probe, which is used to detect related genes for excluding red blood cell membrane defects (hereditary spherocytosis, hereditary elliptocytosis, hereditary acanthocytosis, hereditary abetalipoproteinemia), enzyme defects (glucose-6-phosphate dehydrogenase deficiency, pyruvate kinase deficiency, pyrimidine-5'-nucleotidase deficiency) and hereditary erythropoiesis abnormalities that have similar phenotypes to hereditary hemoglobinopathy.

[0012] Furthermore, the kit for detecting genes related to hereditary hemoglobinopathy further includes a fourth type of probe, which is used to detect HLA-A, HLA-B and HLA-C loci.

[0013] The technical concept of the present application is as follows: The first type of probe and the second type of probe are used to screen all patients with mutated hemoglobin genes. When encountering patients with no detected gene abnormalities but with similar clinical symptoms, then the third type of probe plays a role in differential diagnosis. At the same time, the fourth type of probe can assist in family diagnosis. Exemplarily, the nucleic acid sequence of the first type of probe includes, but is not limited to, those shown in SEQ ID NO. 1 to 13; the nucleic acid sequence of the second type of probe includes, but is not limited to, those shown in SEQ ID NO. 14 to 35; the nucleic acid sequence of the third type of probe includes, but is not limited to, those shown in SEQ ID NO. 36 to 61; the nucleic acid sequence of the fourth type of probe includes, but is not limited to, those shown in SEQ ID NO. 62 to 64.

[0014] The third aspect of the present application provides a high-throughput DNA detection method for hereditary hemoglobinopathy, which includes:

[0015] Step S1: Design probes for detecting known and unknown mutations in genes related to abnormal hemoglobinopathy according to specific rules;

[0016] Step S2: Synthesize the probes;

[0017] Step S3: Extract genomic DNA from the peripheral blood of the patient, construct a library, hybridize and capture the target sequence, recover the target region library and perform library quality detection to meet the requirements of NGS sequencing;

[0018] Step S4: Load the library onto the sequencing platform for high-throughput sequencing to obtain raw sequencing data;

[0019] Step S5: Perform data analysis on the raw sequencing data: First, perform data quality control, genome alignment, repetitive sequence marking and clipping quality correction on the raw sequencing data, re-align near indels, evaluate the coverage of the target region and the overall repetition rate, perform SNV and INDEL mutation analysis, and copy number variation CNV analysis; Next, perform gene mutation annotation, filter out pathogenic sites according to the established SOP, and finally, confirm the authenticity of the mutation through IGV review, perform visual analysis on the CNV analysis, and perform image visualization correspondence between the genomic interval and the clinical report protein fragment.

[0020] Further, in step S1, the specific rules are as follows: Based on the human genome database, for all coding regions of hemoglobin genes and all non-coding regions including the regulatory regions of the hemoglobin gene cluster; transcription factors; the gene coding regions and regulatory regions of the gene HBS1L-MYB that regulates hemoglobin switching; HLA-A, B, C loci; design probes for genes related to diseases with similar phenotypes to exclude:

[0021] (1) Design probe length and density: The length of the designed probes is 120bp; for the target capture region, for exons / intervals <120bp that do not overlap with the genomic repetitive region, the probe length will be truncated to 80bp or 60bp. Specifically, for exons ≤ 60bp, 60bp probes are designed, and for >60 ≤120bp, 80bp probes are designed to cover the entire region;

[0022] (2) Design the probe density to be 1× coverage: For the highly homologous sequences of HBA2 exon2 and downstream 2kb, design 2× probe coverage; design heterozygous SNP probes with a density of 5kb in the possible CNV region that are frequent in the Chinese population;

[0023] (3) Blast the designed probes against the NCBI database to ensure the specificity of the probes;

[0024] (4) Evaluate the synthesized probes by clinical sample sequencing: Apply the synthesized probes to clinical actual sample capture sequencing. The evaluation indicators are as follows: 1) Probe capture efficiency > 0.7; 2) Evaluation of capture uniformity; 3) Average coverage depth of the core gene region: Average sequencing depth is 500×, and the region with more than 100× coverage should be higher than 95%, and the region with more than 200× coverage should exceed 90%; 4) The results of next-generation sequencing are 100% consistent with the results of first-generation sequencing; 5) The gene mutation frequency is 100% consistent with the results of first-generation sequencing;

[0025] (5) Optimize the probes according to the evaluation results: 1) Perform repetitive sequence analysis and motif analysis on the off-target regions. For repetitive regions and high-proportion motifs, adjust or delete the corresponding probe positions; 2) For low-coverage regions, without affecting the overall capture efficiency, increase the probe coverage of low-capture regions and add specific probes. The positions of specific probes are preferably partially overlapped with the original probes; 3) Increase the probe density or mutant probes for the core regions of key hot-spot mutations with low coverage depth; 4) For gene mutation sites with large differences in gene frequencies, increase the probe density, design mutant probes, or design CNV probes with a probe density of 10K within 100K around them.

[0026] Further, the step S5 includes S5.1 to S5.7:

[0027] S5.1 Sequencing data quality control: Use fastp to perform preliminary quality control on the raw DNA sequencing data, remove low-quality reads and adapter sequences to ensure data quality; After downloading the raw data and performing md5 verification, use the fastp software to filter the raw data reads with the following parameters: --trim_poly_g --trim_poly_x --length_required 60 --low_complexity_filter --complexity_threshold 20 –correction, filter the reads of poly G, polyN, reads with a length < 60bp, and reads with a sequence complexity lower than 20 respectively, and correct the overlapping parts of paired reads at the same time;

[0028] S5.2 Data comparison and pre - processing: Use bwa - 0.7.18 and samtools - 1.9.1 for alignment (hg19), bam conversion and sorting, and then use 4.6.1.0 respectively for ReorderSam, MarkDuplicates, BaseRecalibrator, ApplyBQSR according to GATK best practices. Finally, obtain the bam file within the captured region through PrintReads for subsequent gene mutation analysis;

[0029] S5.3 Gene mutation analysis: Use Mutect2 and VarDict - 1.8.3 software to analyze the sequencing data of this panel. Specifically, for Mutect2, first, use the default parameters for mutation analysis. Second, use the default parameters for Mutect2 mutation filter to obtain the vcf file with filtered tags. Finally, perform the first - round filtering of Mutect2 results. Use SnpSift software to filter the Mutect2 filtered tags with the following specific parameters: filter - n "(GEN[1].AF < 0.05) |(FILTER has 'panel_of_normals'| FILTER has 'normal_artifact' | FILTER has 'contamination' | FILTER has 'germline' | FILTER has 'weak_evidence'| FILTERhas 'low_allele_frac' | (FILTER has'multiallelic' & (FILTER has'slippage'|FILTER has 'panel_of_normals')))" to obtain the results of the first - round filtering of Mutect2, filtering out obvious systematic errors for subsequent integrated analysis; For VarDict, first, add the following parameters in mutation analysis: - r 7 - f0.05 –nosv; Then, process the non - standard part of the original vcf file, and at the same time, forcefully replace the icons with Q0 filtering parameters with PASS; After that, standardize the vcf to obtain the results of the first - round filtering of VarDict for subsequent integrated analysis;

[0030] S5.4 Integrated analysis: First, use the CombineVariants tool of GATK3 to integrate the first-round filtering results of Mutect2 and VarDict obtained previously. Through mutual verification of the results of the two software, reliable gene mutation information is obtained, and a vcf file of the second-round filtering results is generated.

[0031] S5.5 Gene mutation annotation: Annotate the vcf file of the second-round filtering results obtained finally. The annotation content includes: functional annotation and filtering annotation; for functional annotation, use SnpEff and Annovar to annotate the protein-coding changes of mutations. The principle is to annotate the splice sites 5bp upstream and downstream of CDS and exon; filtering annotation includes genetic diversity annotation of the normal population, disease database annotation, protein change function prediction filtering, and self-built negative and positive database filtering; specifically, use the SnpSift annotate tool to mark the processed database vcf. The genetic diversity databases include: dbsnp151, 1000Gp3, ESP6500, ExAC, gnomAD exomes, UK10K, ICGCV27; the disease databases include: cosmic85, clinvar, ICGCV27 and HGMD, and the protein change function prediction database uses dbNSFPv3.5a; the self-built negative database is the PON files trained separately in normal control samples using three software, Mutect1, Mutect2 and VarDict, on two different sequencing platforms, Illumina PE-150 and BGI PE-90, and three different capture kits, AgilentV5, IDTv2 and OncoWESuper; through the above annotation, classify the pathogenicity of the variation and provide an explanation for important mutations.

[0032] S5.6 Gene mutation filtering and authenticity confirmation: For the above gene mutation annotation results, screen the gene mutations according to the clinically formulated SOP, and reconfirm the authenticity of the gene mutations in the IGV software through the bam file. Finally, generate a genotype analysis report, including all detected pathogenic mutations and their clinical relevance.

[0033] S5.7 Copy number variation analysis: The genomes of several normal individuals were captured and sequenced. For the bam file with duplicate sequences marked after bwa alignment, a normal reference genome database for this panel was trained using the cnvkit software. Then, the cbs algorithm of the cnvkit software was used to perform paired analysis on the bam file of each clinical sample under the same alignment conditions with the normal control to obtain normalized coverage information. Then, the call tool of cnvkit was used to analyze the CNV intervals, and the mapping tool of this software was used to visualize the intervals or genes of interest. In the HBA and HBB regions, the drawing coordinates were normalized to obtain a genomic image relative to clinical protein electrophoresis.

[0034] Furthermore, in step S2, the actual synthesis of the probe was completed by solid-phase phosphoramidite chemistry: (1) Single nucleotide extension: Phosphoramidite monomers of C, G, A, and T were gradually added to the solid-phase support; (2) Chemical reaction cycle: a. Deprotection: Removal of the protecting group to expose the active group; b. Extension: Addition of the next nucleotide monomer, which binds to the 5'-hydroxyl group of the previous nucleotide; c. Oxidation stabilization: Treatment with an oxidant to stabilize the phosphodiester bond; (3) Synthesis completion: The above steps were repeated until a 120 bp probe sequence was completed.

[0035] During the synthesis process, biotin modification was added to the 5' end of the probe for subsequent binding to streptavidin magnetic beads; terminal modification was added to improve the probe stability.

[0036] Purification treatment was carried out during the solid-phase synthesis process: (1) PAGE purification: Separation by polyacrylamide gel electrophoresis to distinguish the complete 120 bp probe according to the molecular weight; (2) HPLC purification: Further improvement of the probe purity by high-performance liquid chromatography to remove impurities and incorrect products; (3) Desalting treatment: Removal of residual salts and buffer components to improve the downstream experimental performance of the probe.

[0037] Furthermore, step S3 includes:

[0038] S3.1 Nucleic acid extraction: In a biosafety cabinet, 200 μL of blood cell samples were taken into 1.5 mL centrifuge tubes; a commercial kit was used to extract genomic DNA according to the standard operating procedure.

[0039] S3.2 Nucleic acid quality control:

[0040] S3.2.1 The concentration of the extracted DNA was measured using a Qubit 4.0 fluorometer.

[0041] S3.2.2 Detect the quality of the extracted DNA using the Agilent High Sensitivity DNA Kit;

[0042] S3.2.3 Nucleic acid quality control standard: The total amount of DNA ≥ 300 ng, and the quality is qualified;

[0043] S3.3. DNA fragmentation:

[0044] S3.3.1 According to the starting sample amount of 300 ng, the Qubit concentration of the positive and negative quality control products needs to be verified. For a starting amount of 300 ng, add 1×TE buffer to 50 μL. The buffer is pH 8.0, 10 mM Tris-HCl, 0.1 mM EDTA-2Na, and calculate the input volume of each sample and the positive and negative quality control products; S3.3.2 Prepare labeled 1.5 mL centrifuge tubes according to the number of samples, the number of DNA + positive and negative quality control products; S3.3.3 Add the corresponding DNA sample and 1×TE buffer to each centrifuge tube, mix well and centrifuge briefly; S3.3.4 Use the Covaris M220 instrument for DNA fragmentation;

[0045] S3.4. Library construction:

[0046] S3.4.1 The end repair and A addition reagent is melted and shaken well in an ice box and centrifuged briefly; Prepare the "end repair and A addition" mixed solution, shake well and place it in an ice box for later use; Transfer the reaction system to a PCR instrument for the end repair and A addition reaction;

[0047] S3.4.2 Adapter ligation: Prepare the adapter ligation Mix, shake well and centrifuge briefly, and place it in an ice box for later use; According to the experimental record form, take 5 μL of the corresponding numbered diluted adapter and add it to the end repair and A addition product, vortex mix for 20 seconds; Add the adapter ligation Mix to the end repair product and adapter mixture, shake gently to mix well and centrifuge briefly; Set the PCR program, and transfer the reaction system to the PCR instrument after the hot lid and module temperature are stable;

[0048] S3.4.3 Purification after joint ligation: Add 88 μL (0.8×) purification magnetic beads to the ligation product, pipette and mix well. After brief centrifugation, let it stand at room temperature for 10 minutes; Place the 8-strip tube on the magnetic rack and adsorb for 8 minutes. After the liquid becomes clear, aspirate and discard the supernatant; Add 200 μL of 80% ethanol to the tube, let it stand at room temperature for 30 seconds, discard the supernatant, and repeat once; Briefly centrifuge to collect the liquid droplets on the tube wall, and use a pipette again to completely aspirate the residual ethanol at the bottom. Dry at room temperature for 3 - 5 minutes; Add 22 μL of ultrapure water, vortex and mix well for 10 seconds. After brief centrifugation, let it stand for 2 minutes; Place the tube on the magnetic rack and adsorb for 3 minutes. After the liquid becomes clear, take 2 μL of the sample for Qubit quantification, and transfer 20 μL of the supernatant to a new 8-strip tube for standby;

[0049] S3.4.4 Pre-PCR reaction: Take out the pre-PCR reagent, thaw it on an ice box, mix well before use, and briefly centrifuge; Prepare the pre-PCR Mix according to the number of libraries; Add 30 μL of the pre-PCR Mix to a 0.2 mL 8-strip tube containing the purified product, vortex and mix well, and set aside after brief centrifugation;

[0050] S3.4.5 Sorting and purification after Pre-PCR reaction: After Pre-PCR, briefly centrifuge, add 25 μL of purification magnetic beads to the tube, vortex and mix well, then let it stand for 10 minutes; Place the 8-strip tube on the magnetic rack, adsorb for 5 minutes. After the liquid becomes clear, transfer 75 μL of the supernatant to a new tube, and add 25 μL of purification magnetic beads; Repeat the steps to ensure the removal of impurities; Add 200 μL of 80% ethanol, let it stand for 30 seconds, aspirate and discard the supernatant, and repeat once; Briefly centrifuge, place the 8-strip tube on the magnetic rack, aspirate and discard the residual ethanol, and dry at room temperature for 5 minutes; Add 42 μL of ultrapure water, vortex and mix well, then let it stand for 5 minutes; Place the tube on the magnetic rack and adsorb for 1 minute, take the supernatant for Qubit quantification, and transfer 40 μL of the supernatant to a labeled centrifuge tube for standby;

[0051] S3.5. Library hybridization and capture:

[0052] S3.5.1 Calculate the input volume according to the library concentration measured by Qubit; Ensure that the total amount of each captured library is between 1000 ng - 4000 ng; S3.5.2 Mix the blocking sequence and Cot DNA, vortex and mix well, and briefly centrifuge; S3.5.3 Put the sample into a vacuum centrifuge and dry it at 60 °C for 20 minutes to 1 hour; S3.5.4 Prepare the hybridization mixture, add 17 μL to the dried sample, mix well and briefly centrifuge, let it stand at room temperature for 5 minutes, and repeat mixing twice; S3.5.5 Set the PCR program to ensure that the hot lid temperature is 100 °C; Transfer the hybridization system to a PCR tube and run the program;

[0053] S3.6. Library washing after capture:

[0054] S3.6.1 Prepare relevant buffers and preheat the magnetic stand to 65 °C in advance;

[0055] S3.6.2 Operating steps: Add 50 μL of magnetic beads into a PCR tube, vortex and mix well, adsorb, and discard the supernatant; Repeat the washing twice and add 100 μL of 1× magnetic bead elution buffer; After adding Bead Resuspension Mix, vortex and mix well, incubate at 65 °C for 15 minutes; Mix the magnetic beads with the hybridization solution, incubate at 65 °C for 45 minutes, and pipette once in a while; Add 100 μL of elution buffer I at room temperature into the tube, adsorb and discard the supernatant after adsorption; Repeat the washing twice; Add 100 μL of elution buffer I at room temperature into the tube and transfer it to a new tube; Add 150 μL of elution buffers II and III at room temperature into the PCR tube, incubate at room temperature for 2 minutes, centrifuge briefly, and aspirate and discard the supernatant; Add 20 μL of ultrapure water, transfer it to a new tube for standby;

[0056] S3.7. Library enrichment:

[0057] S3.7.1 Post-PCR reaction: Mix 20 μL of the captured product with 30 μL of post-PCR Mix, vortex and mix well, and gently shake to keep the magnetic beads suspended; Set the PCR program, with the hot lid at 105 °C, and run the Post-PCR reaction;

[0058] S3.7.2 Purification after Post-PCR reaction: Take out the purification magnetic beads, vortex and mix well, and let stand for 30 minutes; After the Post-PCR reaction, place the PCR tube on the magnetic stand, adsorb and transfer 50 μL of the supernatant to a new tube; Add 75 μL of purification magnetic beads to the new tube and let stand at room temperature for 5 minutes; After brief centrifugation, aspirate and discard the supernatant, add 200 μL of 80% ethanol for washing, and repeat once; Aspirate all the residual liquid and dry at room temperature until the surface of the magnetic beads does not reflect light; Add 42 μL of ultrapure water, vortex and mix well, after brief centrifugation, transfer 40 μL of the supernatant to a new centrifuge tube for standby;

[0059] S3.8. Post-PCR library quality control:

[0060] S3.8.1 Use a Qubit fluorometer to quantify the library concentration to ensure that the library concentration meets the requirements;

[0061] S3.8.2 Use an Agilent or Bioanalyzer to detect the library quality to ensure that it meets the requirements of NGS sequencing.

[0062] Furthermore, step S4 includes:

[0063] S4.1 Library dilution: Dilute the library to an appropriate concentration according to the requirements of the sequencing platform;

[0064] S4.2 Library mixing: Mix different libraries in proportion to ensure that the final concentration of each library is consistent;

[0065] S4.3 Library loading:

[0066] S4.3.1 Loading the library: The loading amount is between 10 - 15 pM;

[0067] S4.3.2 Adding PhiX control: Add a certain amount of PhiX control library, and the proportion of PhiX is 5 - 10%;

[0068] S4.4. Sequencing platform setup: Select appropriate sequencing mode, sequencing depth, and number of cycles.

[0069] S4.5. Start sequencing;

[0070] S4.6. Monitoring the sequencing process;

[0071] S4.7. Post - sequencing processing:

[0072] S4.7.1 Data acquisition: After sequencing is completed, extract the sequencing data from the instrument and output it in FASTQ format;

[0073] S4.7.2 Library quality inspection: Use FastQC to perform quality control on the sequencing data, and check the quality of the library, sequencing depth, GC content, and sequence distribution;

[0074] S4.7.3 Data backup: Back up all the original data.

[0075] The fourth aspect of this application provides a high - throughput DNA sequencing data analysis platform for hereditary hemoglobinopathy, which includes: a sequencing data quality control component, a data alignment and pre - processing component, a gene mutation analysis component, an integration analysis component, a gene mutation annotation component, a gene mutation filtering and authenticity confirmation component, and a copy number variation analysis component;

[0076] The sequencing data quality control component is used for: performing preliminary quality control on the raw DNA sequencing data using fastp to remove low-quality reads and adapter sequences to ensure data quality; after downloading the raw data and performing md5 verification, using the fastp software to filter the raw data reads with the following parameters: --trim_poly_g --trim_poly_x --length_required 60 --low_complexity_filter --complexity_threshold 20 –correction, filtering reads with poly G, polyN, reads with a length < 60bp, and reads with a sequence complexity lower than 20 respectively, and at the same time performing base correction on the overlapping parts of paired reads;

[0077] The data alignment and preprocessing component is used for: performing alignment (hg19), bam conversion and sorting using bwa-0.7.18 and samtools-1.9.1, and then respectively performing ReorderSam, MarkDuplicates, BaseRecalibrator, ApplyBQSR according to GATK best practices using 4.6.1.0. Finally, obtaining a bam file within the capture region through PrintReads for subsequent gene mutation analysis;

[0078] The gene mutation analysis component is used for: analyzing paired WES samples using Mutect2 and VarDict-1.8.3 software; specifically, for Mutect2, first, the mutation analysis uses default parameters, and second, Mutect2 mutationfilter uses default parameters to obtain a vcf file with filtered tags; finally, the first-round filtering of Mutect2 results is performed. The Mutect2 filtered tags are filtered through SnpSift software, and the specific parameters are as follows: filter -n "(GEN[1].AF <0.05) | (FILTER has 'panel_of_normals'| FILTER has 'normal_artifact' | FILTER has 'contamination' | FILTER has 'germline' | FILTER has 'weak_evidence'|FILTER has 'low_allele_frac' | (FILTER has'multiallelic' & (FILTER has'slippage'|FILTER has 'panel_of_normals')))" to obtain the results of the first-round filtering of Mutect2, filtering out obvious systematic errors for subsequent integrated analysis; for VarDict, first, the following parameters are added in the mutation analysis: -r 7 -f 0.05 –nosv; then, the non-standard part of the original vcf file is processed, and at the same time, the icons with Q0 filtering parameters are forced to be replaced with PASS; after that, the vcf is standardized to obtain the results of the first-round filtering of VarDict for subsequent integrated analysis;

[0079] The integrated analysis component is used for: first, using the CombineVariants tool of GATK3 to integrate the results of the first-round filtering of Mutect2 and VarDict obtained previously, and verifying each other through the results of the two software to obtain reliable gene mutation information and generate a vcf file of the second-round filtering results;

[0080] The gene mutation annotation component is used for: annotating the vcf file of the second-round filtering result finally filtered, and the annotation content includes: functional annotation and filtering annotation; for functional annotation, SnpEff and Annovar are used to annotate the protein-coding changes of mutations, and the principle is to annotate the splice sites 5bp upstream and downstream of CDS and exon; the filtering annotation includes normal population genetic diversity annotation, disease database annotation and protein change function prediction filtering, as well as self-built negative and positive database filtering; specifically, the processed database vcf is marked by using the SnpSift annotate tool, and the genetic diversity databases include: dbsnp151, 1000Gp3, ESP6500, ExAC, gnomAD exomes, UK10K, ICGCV27; the disease databases include: cosmic85, clinvar, ICGCV27 and HGMD, and the protein change function prediction database uses dbNSFPv3.5a; the self-built negative database is the PON files respectively trained in normal control samples by using three software, namely Mutect1, Mutect2 and VarDict, on two different sequencing platforms, namely Illumina PE-150 and BGI PE-90, and three different capture kits, namely AgilentV5, IDTv2 and OncoWESuper; through the above annotation, the pathogenicity of the variation is classified and an explanation is provided for important mutations;

[0081] The gene mutation filtering and authenticity confirmation component is used for: screening the gene mutations according to the clinically formulated SOP for the above gene mutation annotation results, reconfirming the authenticity of the gene mutations in the IGV software through the bam file, and finally generating a genotype analysis report, including all detected pathogenic mutations and their clinical relevance;

[0082] The copy number variation analysis component is used for: capturing and sequencing the genomes of several normal people, using the cnvkit software to train a normal reference genome database for this panel for the bam file marked with duplicate sequences after bwa alignment; then using the cbs algorithm of the cnvkit software to perform paired analysis on the bam file of each clinical sample under the same alignment conditions with the normal control to obtain normalized coverage information; then using the call tool of cnvkit to analyze the CNV interval, and then using the mapping tool of this software to visualize the interval or gene of interest; in the HBA and HBB regions, the drawing coordinates are normalized to obtain a genomic image relative to clinical protein electrophoresis.

[0083] After the present application adopts the above technical solutions, it has the following excellent technical effects:

[0084] In this application, a large-scale overlapping probe is designed to cover the hemoglobin gene, and probes with special sequences are used to distinguish the similar regions of HBA1 and HBA2. Probes are also designed for genes with common hemoglobin membrane and enzyme defects. Further bioinformatics analysis is performed on the data obtained after large-scale sequencing. The present invention provides a method that can simultaneously detect multiple types of hemoglobin mutations, including point mutations, insertions / deletions, gene rearrangements, etc., greatly improving the detection efficiency and coverage, avoiding missed detections, and making differential diagnoses for diseases with similar clinical manifestations. It can comprehensively and accurately identify the gene variations carried by patients, so as to timely intervene and treat, and improve the prognosis of patients. The present invention can detect all variations in the genomes of hemoglobin α and β subunits and regulatory genes. The variations described in the present invention include point mutations, small fragment deletions and insertions, large fragment deletions and duplications, inversions, translocations and fusion variations. The bioinformatics analysis platform described in the present invention can directly give the mutation types of common mutations. In this application, a targeted DNA sequencing data analysis platform is developed for the characteristics of the customized panel and target mutation types, taking into account clinical practicability, and mainly has the following functions: 1) automatic alignment, analysis, filtering and annotation; 2) multi-algorithm analysis process, especially for mutations in homologous sequence regions such as HBA1 and HBA2, to perform reassembly analysis, greatly reducing the missed detection rate of mutations; 3) parallel analysis of multiple samples, with 16-20 samples analyzed in parallel at a time, and the average analysis time for each sample is 10 minutes; 4) support for large fragment indel analysis, and indels within 300 bp can be accurately detected; 5) support for CNV detection at the single exon level within the captured region. Description of the Drawings

[0085] Figure 1 It is a flowchart of a high-throughput DNA detection method for hereditary hemoglobinopathy in an embodiment of this application. Detailed Embodiments

[0086] The advantages of the present invention are further elaborated below in conjunction with the drawings and specific embodiments. Those skilled in the art should understand that the content specifically described below is illustrative rather than restrictive, and should not be used to limit the protection scope of the present invention.

[0087] As Figure 1 shown, this embodiment provides a high-throughput DNA detection method for hereditary hemoglobinopathy. High-throughput sequencing technology can cover a wider range and has higher sensitivity and specificity. The high-throughput DNA detection method for hereditary hemoglobinopathy includes the following steps S1-step S5:

[0088] Step S1: Design probes for detecting known and unknown mutations in genes related to abnormal hemoglobinopathy according to specific rules;

[0089] In this step, a series of probes with high specificity, high sensitivity, and high coverage are designed to detect known and unknown mutations in genes related to abnormal hemoglobinopathy. Probe design is a key step in this embodiment, with the goal of ensuring that the probes can specifically hybridize with the target sequence while avoiding non-specific hybridization. The experimental performance of the probes in this embodiment is shown in Table 1.

[0090] Table 1 Preliminary experimental performance of the probes

[0091]

[0092] Probe design: Based on the human genome database (GRCh38), probes are designed for all coding regions of hemoglobin genes (mainly including HBA1, HBA2, HBB, HBBP1, HBD, HBE1, HBG1, HBG2, HBM, HBQ, HBZ, etc.) and all non-coding regions including the regulatory regions of the hemoglobin gene cluster (HS-40, HS1, HS2, HS3, HS4, HS5); transcription factors (GATA1, NFE2, KLF1, SP1, BLC11A); the gene coding regions and regulatory regions of the gene HBS1L-MYB that regulates hemoglobin switching; HLA-A, B, C loci (due to the characteristic of verifying genetic relationship required in some application scenarios of this detection); probes are designed to exclude genes related to diseases with similar phenotypes (G6PD, ANK1, SLC4A1, SPTB, SPTA1, EPB42, EPB41, GYPC, GYPA, ADD2, ANK1, XK, GSR, SEC23B, CDAN1).

[0093] 1.1 Probe length and density: The length of the designed probes is 120bp. For the target capture region, for exons / intervals <120bp that do not overlap with the genomic repeat region, the probe length will be truncated to 80bp or 60bp (specifically, for exons ≤ 60bp, 60bp probes are designed, and for >60 ≤120bp, 80bp probes are designed to cover the entire region).

[0094] 1.2 The probe density is 1× coverage. For the highly homologous sequences of HBA2 exon2 and downstream 2kb, 2× probe coverage is designed. For possible CNV regions, heterozygous SNP probes with a density of 5kb and high frequency in the Chinese population are designed.

[0095] 1.3 The designed probes are aligned with databases such as NCBI by Blast, and it is necessary to ensure that the Tm value of the probes is moderate, there are no special structures such as palindromic sequences, and the sequence homology is between 50% and 80%. These values are used to ensure the specificity of the probes.

[0096] 1.4 Evaluate the synthesized probes by clinical sample sequencing: Apply the synthesized probes to the capture sequencing of clinical actual samples. The main evaluation indicators are: 1) Probe capture efficiency (>0.7); 2) Evaluation of capture uniformity; 3) Average coverage depth of the core gene region (average sequencing depth is 500×, the region exceeding 100× should be higher than 95%, and the region exceeding 200× should exceed 90%); 4) Consistency between the next-generation sequencing results and the first-generation sequencing results (100%); 5) Consistency between the gene mutation frequency and the first-generation sequencing results (100%).

[0097] 1.5 Optimize the probes according to the evaluation results: The main contents of probe optimization include: (1) Conduct repetitive sequence analysis and motif analysis on the off-target regions. For repetitive regions and motifs with a proportion higher than 10%, adjust or delete the corresponding probe positions; (2) For low-coverage regions (the average coverage of the probes is lower than 0.7X), without affecting the overall capture efficiency, increase the probe coverage of the low-capture regions to above 0.9X, and add specific probes. The positions of the specific probes are preferably partially overlapped with the original probes; (3) For the key hot-spot mutation core region (the entire exon of the mutation region with clear clinical significance) where the coverage depth is low (<0.8X), it is necessary to increase the probe density (to >0.95X) or mutant probes; (4) For gene mutation sites with a large difference in gene frequency (the difference in VAF of mutations in the previous R & D repeatability test is greater than 5% or more), it is necessary to increase the probe density, design mutant probes, or design CNV probes with a probe density of 10K within 100K around them.

[0098] Exemplarily, the nucleic acid sequences of the four types of probes in this embodiment are shown in Table 2:

[0099] Table 2 Nucleic acid sequences corresponding to four types of probes

[0100]

[0101]

[0102]

[0103] When the sequences in Table 2 are inconsistent with the sequence listing, Table 2 shall prevail.

[0104] Step S2: Synthesize probes;

[0105] The actual synthesis of the probe is completed by solid-phase phosphoramidite chemistry synthesis: (1) Single nucleotide extension: Phosphoramidite monomers of C, G, A, and T are gradually added to a solid-phase support (usually polystyrene microbeads). (2) Chemical reaction cycle: a. Deprotection: Remove the protecting groups to expose the reactive groups. b. Extension: Add the next nucleotide monomer, which binds to the 5'-hydroxyl group of the previous nucleotide. c. Oxidation stabilization: Treat with an oxidizing agent to stabilize the phosphodiester bond. (3) Synthesis completion: Repeat the above steps until the 120-bp probe sequence is completed.

[0106] During the synthesis process, biotin modification is added to the 5' end of the probe for subsequent binding to streptavidin magnetic beads. Terminal modifications are added to improve probe stability (such as phosphate groups).

[0107] Since short fragments with incomplete extension may be generated during the solid-phase synthesis process, purification is required: (1) PAGE purification: Separate by polyacrylamide gel electrophoresis to distinguish the complete 120-bp probe according to molecular weight. (2) HPLC purification: Further improve the probe purity by high-performance liquid chromatography to remove impurities and incorrect products. (3) Desalting treatment: Remove residual salts and buffer components to improve the downstream experimental performance of the probe.

[0108] Step S3: Extract genomic DNA from the patient's peripheral blood, construct a library, hybridize and capture the target sequence, recover the target region library, and perform library quality detection to meet the requirements of NGS sequencing;

[0109] Specimen processing: Extract genomic DNA from the patient's peripheral blood to obtain high-purity DNA (A260 / 280 ratio of 1.8 - 2.0) and sufficient amount of DNA (≥50 ng). Construct a library, hybridize and capture the target sequence. Recover the target region library and perform quality inspection to meet the requirements of NGS sequencing. NGS (Next Generation Sequencing) result analysis is a complex and multi-step process, involving multiple links such as raw data quality control, sequence alignment, variant detection, gene expression analysis, and functional annotation. When performing target region sequencing using NGS technology, probe capture technology is often used to enrich specific DNA fragments to improve sequencing efficiency and reduce costs. The specific operations are as follows:

[0110] 3.1 Nucleic acid extraction: In a biosafety cabinet, transfer 200 μL of the blood cell sample to a 1.5-mL centrifuge tube. Use a commercial kit (such as QIAamp DNA Blood Mini Kit) to extract genomic DNA according to the standard operating procedure.

[0111] 3.2 Nucleic acid quality control

[0112] 3.2.1 Measure the concentration of the extracted DNA using a Qubit 4.0 fluorometer.

[0113] 3.2.2 Detect the quality of the extracted DNA using an Agilent High Sensitivity DNA Kit.

[0114] 3.2.3 Nucleic acid quality control standard: The total amount of DNA ≥ 300 ng and the quality is qualified.

[0115] 3.3 DNA fragmentation

[0116] 3.3.1 According to the starting sample amount of 300 ng (the Qubit concentration of positive and negative quality control products needs to be verified, starting amount 300 ng), add 1×TE buffer (pH 8.0, 10 mM Tris-HCl, 0.1 mM EDTA-2Na) to 50 μL, and calculate the input volume of each sample and positive and negative quality control products. 3.3.2 Prepare labeled 1.5 mL centrifuge tubes according to the number of samples (number of DNA + positive and negative quality control products). 3.3.3 Add the corresponding DNA sample and 1×TE buffer to each centrifuge tube, mix well and centrifuge briefly. 3.3.4 Use a Covaris M220 instrument for DNA fragmentation.

[0117] 3.4 Library construction

[0118] 3.4.1 The end repair and A-addition reagent is melted and shaken well in an ice box and centrifuged briefly. Prepare an "end repair and A-addition" mixed solution, shake well and place it in an ice box for later use. Transfer the reaction system to a PCR instrument for end repair and A-addition reaction.

[0119] 3.4.2 Adapter ligation Prepare an adapter ligation Mix, shake well and centrifuge briefly, and place it in an ice box for later use. According to the experimental record form, take 5 μL of the corresponding numbered diluted adapter and add it to the end repair and A-addition product, vortex mix for 20 seconds. Add the adapter ligation Mix to the end repair product and adapter mixture, shake gently to mix well, and centrifuge briefly. Set the PCR program, and transfer the reaction system to a PCR instrument after the hot lid and module temperature are stable.

[0120] 3.4.3 Purification after adapter ligation Add 88 μL (0.8×) purification magnetic beads to the ligation product, pipette and mix well. After brief centrifugation, let it stand at room temperature for 10 minutes. Place the 8-strip tube on the magnetic rack and adsorb for 8 minutes. After the liquid becomes clear, aspirate and discard the supernatant. Add 200 μL of 80% ethanol to the tube, let it stand at room temperature for 30 seconds, discard the supernatant, and repeat once. Centrifuge briefly to collect the liquid droplets on the tube wall, and then use a pipette to aspirate the residual ethanol at the bottom as completely as possible. Dry at room temperature for 3 - 5 minutes. Add 22 μL of ultrapure water, vortex and mix well for 10 seconds. After brief centrifugation, let it stand for 2 minutes. Place the tube on the magnetic rack and adsorb for 3 minutes. After the liquid becomes clear, take 2 μL of the sample for Qubit quantification, and transfer 20 μL of the supernatant to a new 8-strip tube for standby.

[0121] 3.4.4 Pre-PCR reaction Take out the pre-PCR reagents, thaw them on an ice box, mix well before use, and centrifuge briefly. Prepare the pre-PCR Mix according to the number of libraries. Add 30 μL of the pre-PCR Mix to a 0.2 mL 8-strip tube containing the purified product, vortex and mix well, and centrifuge briefly for standby.

[0122] 3.4.5 Sorting and purification after Pre-PCR reaction After the Pre-PCR is completed, centrifuge briefly, add 25 μL of purification magnetic beads to the tube, vortex and mix well, and then let it stand for 10 minutes. Place the 8-strip tube on the magnetic rack, adsorb for 5 minutes. After the liquid becomes clear, transfer 75 μL of the supernatant to a new tube, and add 25 μL of purification magnetic beads. Repeat the steps to ensure the removal of impurities. Add 200 μL of 80% ethanol, let it stand for 30 seconds, aspirate and discard the supernatant, and repeat once. Centrifuge briefly, place the 8-strip tube on the magnetic rack, aspirate and discard the residual ethanol, and dry at room temperature for 5 minutes. Add 42 μL of ultrapure water, vortex and mix well, and then let it stand for 5 minutes. Place the tube on the magnetic rack and adsorb for 1 minute. Take the supernatant for Qubit quantification, and transfer 40 μL of the supernatant to a labeled centrifuge tube for standby.

[0123] 3.5. Library hybridization and capture

[0124] 3.5.1 Calculate the input volume according to the library concentration measured by Qubit. Ensure that the total amount of each captured library is between 1000 ng and 4000 ng. 3.5.2 Mix the blocking sequence and Cot DNA, vortex and mix well, and centrifuge briefly. 3.5.3 Place the sample in a vacuum centrifuge and dry it at 60 °C for 20 minutes to 1 hour. 3.5.4 Prepare the hybridization mixture, add 17 μL to the dried sample, mix well and centrifuge briefly, let it stand at room temperature for 5 minutes, and repeat the mixing twice. 3.5.5 Set the PCR program to ensure that the hot lid temperature is 100 °C. Transfer the hybridization system to a PCR tube and run the program.

[0125] 3.6. Washing of the library after capture

[0126] 3.6.1 Prepare relevant buffer solutions and preheat the magnetic stand to 65 °C in advance.

[0127] 3.6.2 Operating steps: Add 50 μL of magnetic beads into a PCR tube, vortex to mix evenly, adsorb, and discard the supernatant. Repeat the washing twice, and add 100 μL of 1× magnetic bead elution buffer. After adding Bead Resuspension Mix, vortex to mix evenly and incubate at 65 °C for 15 minutes. Mix the magnetic beads with the hybridization solution, incubate at 65 °C for 45 minutes, and pipette gently every once in a while. Add 100 μL of elution buffer I at room temperature into the tube, adsorb and discard the supernatant after adsorption. Repeat the washing twice. Add 100 μL of elution buffer I at room temperature into the tube and transfer it to a new tube. Add 150 μL of elution buffer II and III at room temperature into the PCR tube, incubate at room temperature for 2 minutes, centrifuge briefly, and aspirate and discard the supernatant. Add 20 μL of ultrapure water and transfer it to a new tube for standby.

[0128] 3.7. Library enrichment

[0129] 3.7.1 Post-PCR reaction Mix 20 μL of the captured product with 30 μL of post-PCR Mix, vortex to mix evenly, and gently shake to keep the magnetic beads suspended. Set the PCR program with the hot lid at 105 °C and run the Post-PCR reaction.

[0130] 3.7.2 Purification after Post-PCR reaction Take out the purification magnetic beads, vortex to mix evenly, and let it stand for 30 minutes. After the Post-PCR reaction, place the PCR tube on the magnetic stand, adsorb and transfer 50 μL of the supernatant to a new tube. Add 75 μL of purification magnetic beads into the new tube and let it stand at room temperature for 5 minutes. After centrifuging briefly, aspirate and discard the supernatant, add 200 μL of 80% ethanol for washing, and repeat once. Aspirate all the residual liquid and dry it at room temperature until the surface of the magnetic beads does not reflect light. Add 42 μL of ultrapure water, vortex to mix evenly, after centrifuging briefly, transfer 40 μL of the supernatant to a new centrifuge tube for standby.

[0131] 3.8. Post-PCR library quality control

[0132] 3.8.1 Use a Qubit fluorometer to quantify the library concentration and ensure that the library concentration meets the requirements.

[0133] 3.8.2 Use an Agilent or Bioanalyzer to detect the library quality and ensure that it meets the requirements for NGS sequencing.

[0134] Step S4: Load the library into the sequencing platform for high-throughput sequencing to obtain the raw sequencing data;

[0135] Load the library onto the sequencing platform (Illumina Nextseq 500) for high-throughput sequencing. The sequencer generates raw data (FASTQ files), and use a tool (FastQC) to check the data quality, including Q-value distribution, adapter contamination, GC content, etc. The specific steps are as follows:

[0136] 4.1 Library dilution: Dilute the library to an appropriate concentration according to the requirements of the sequencing platform. Usually, for the NextSeq platform, the recommended concentration is 12 - 15 nM.

[0137] 4.2 Library mixing: Different libraries can be mixed in proportion according to the experimental design to ensure that the final concentration of each library is the same. Ensure that the quantity ratio between libraries is reasonable during mixing.

[0138] 4.3 Library loading

[0139] 4.3.1 Loading the library: The loading amount is generally adjusted according to the specific sequencing depth and data volume requirements, usually between 10 - 15 pM. For high-throughput platforms, the library concentration can be appropriately lowered to ensure uniform distribution of each library.

[0140] 4.3.2 Adding PhiX control: To improve sequencing quality and correct sequence biases, a certain amount of PhiX control library is usually added. The proportion of PhiX is generally 5 - 10%, which helps to increase the sequencing depth in regions rich in GC content and correct sequencing errors.

[0141] 4.4. Sequencing platform settings: Select the appropriate sequencing mode, sequencing depth, and number of cycles.

[0142] 4.5. Start sequencing

[0143] 4.6. Monitoring the sequencing process

[0144] 4.7. Post-sequencing processing

[0145] 4.7.1 Data acquisition: After sequencing is completed, extract the sequencing data from the instrument, usually output in FASTQ format. Ensure that all raw data is correctly saved to avoid loss.

[0146] 4.7.2 Library quality inspection: Use appropriate tools (such as FastQC) to perform quality control on the sequencing data, and check the library quality, sequencing depth, GC content, sequence distribution, etc.

[0147] 4.7.3 Data backup: Back up all raw data to ensure data security. The data can be uploaded to the cloud platform or saved to an external storage device.

[0148] Step S5: Perform data analysis on the raw sequencing data

[0149] Perform genome alignment, repetitive sequence marking, and clipping quality correction on the raw sequencing data, re-align near possible indels, and then evaluate the coverage of the target region and the overall repetition rate, SNV and INDEL mutation analysis, copy number variation CNV analysis; Next, perform gene mutation annotation, filter to obtain pathogenic sites according to the established SOP, and finally confirm the authenticity of the mutation through IGV review, perform visual analysis on the CNV analysis, and visually correspond the genomic interval to the clinical report protein fragment.

[0150] 5.1 Sequencing data quality control

[0151] Use fastp to perform preliminary quality control on the sequencing data, remove low-quality reads and adapter sequences to ensure data quality. After downloading the raw data and performing md5 verification, use the fastp software to filter the raw data reads with the following parameters: --trim_poly_g --trim_poly_x --length_required 60 --low_complexity_filter --complexity_threshold 20 –correction, filter reads with poly G, polyN, reads with length < 60bp, and reads with sequence complexity lower than 20 respectively, and perform base correction on the overlapping parts of paired reads.

[0152] 5.2 Data alignment and preprocessing Use bwa-0.7.18 and samtools-1.9.1 for alignment (hg19), bam conversion and sorting, and then use 4.6.1.0 respectively according to the recommended process of GATK best practices for ReorderSam, MarkDuplicates, BaseRecalibrator, ApplyBQSR. Finally, obtain the bam file within the captured region through PrintReads for subsequent gene mutation analysis.

[0153] 5.3 Gene mutation analysis

[0154] This application uses two common software, Mutect2 (consistent with the GATK version, GATK-4.6.1.0) and VarDict-1.8.3, to analyze paired WES samples.

[0155] Specifically, for Mutect2, first, the mutation analysis uses default parameters. Second, Mutect2 mutationfilter uses default parameters to obtain a vcf file with filtered tags. Finally, the Mutect2 results are filtered in the first round. The filtered tags of Mutect2 are filtered by the SnpSift software. The specific parameters are as follows: filter -n "(GEN[1].AF <0.05) | (FILTER has 'panel_of_normals'| FILTER has 'normal_artifact' | FILTER has 'contamination' | FILTER has 'germline' | FILTER has 'weak_evidence'|FILTER has 'low_allele_frac' | (FILTER has'multiallelic' & (FILTER has'slippage'|FILTER has 'panel_of_normals')))" to obtain the results of the first-round filtering of Mutect2, filtering out obvious systematic errors for subsequent integrated analysis.

[0156] For VarDict, first, the following parameters are added in the mutation analysis: -r 7 -f 0.05 –nosv. Then, the non-standard part of the original vcf file is processed, and at the same time, the icons with Q0 filtering parameters are forced to be replaced with PASS. After that, the vcf is standardized to obtain the results of the first-round filtering of VarDict for subsequent integrated analysis.

[0157] 5.4 Multi-software integrated analysis

[0158] First, use the CombineVariants tool of GATK3 to integrate the results of the first-round filtering of Mutect2 and VarDict obtained previously. Through the mutual verification of the results of the two software, reliable gene mutation information is obtained, and a vcf file of the second-round filtering results is generated. Among them, the mutual verification of the results includes any one of the following three situations: 1. Exactly the same; 2. Partially overlapping; 3. Independently detected, but the respective results have passed the first-round filtering parameters.

[0159] 5.5 Gene mutation annotation

[0160] Annotate the integrated vcf file obtained after the final filtering. The main contents of the annotation include: functional annotation and filtering annotation. For functional annotation, SnpEff and Annovar are mainly used to annotate the protein-coding changes of mutations. The main principle is to annotate the splice sites 5bp upstream and downstream of CDS and exon.

[0161] Filtering annotation includes genetic diversity annotation of the normal population, disease database annotation, protein change function prediction filtering, and self-built negative and positive database filtering. Specifically, use the SnpSift annotate tool to mark the processed database vcf. The genetic diversity databases include: dbsnp151, 1000Gp3, ESP6500, ExAC, gnomADexomes, UK10K, ICGCV27

[0162] The disease databases include: cosmic85, clinvar, ICGCV27 and HGMD. The protein change function prediction database uses dbNSFPv3.5a; the self-built negative database is the PON files trained respectively in normal control samples by using three software, Mutect1, Mutect2 and VarDict, on two different sequencing platforms, Illumina PE-150 and BGI PE-90, and three different capture kits, AgilentV5, IDTv2 and OncoWESuper. Through the above annotations, classify the pathogenicity of the variations and provide explanations for important mutations.

[0163] 5.6 Gene Mutation Filtering and Authenticity Confirmation

[0164] For the above gene mutation annotation results, screen the gene mutations according to the clinically formulated SOP, and reconfirm the authenticity of the gene mutations in the IGV software through the bam file. Finally, generate a genotype analysis report, including all detected pathogenic mutations and their clinical relevance. The report can be used by clinicians for the diagnosis of hemoglobinopathies (such as thalassemia, sickle cell anemia) and other hereditary hemoglobinopathies.

[0165] 5.7 Copy Number Variation Analysis

[0166] First, at the initial stage of probe design for the panel, we considered the capture interval distance of the target region or genes of interest and the coverage of heterozygous SNV probes. This design is beneficial for the subsequent analysis algorithm to accurately judge CNV. Specifically, we captured and sequenced the genomes of 40 normal individuals, and for the bam file with duplicate sequences marked after bwa alignment, we used the cnvkit software to train a normal reference genome database for this panel. Then, we used the cbs algorithm of the cnvkit software to perform paired analysis on the bam file of each clinical sample under the same alignment conditions with the normal control to obtain normalized coverage information. Then, we used the call tool of cnvkit to analyze the CNV interval, and then used the mapping tool of this software to visualize the interval or gene of interest. In the HBA and HBB regions, by normalizing the drawing coordinates, a genomic image relative to clinical protein electrophoresis can be obtained, which is convenient for clinical report issuance.

[0167] In one of the above embodiments, it has the following beneficial effects:

[0168] Verification of probe effectiveness: High capture efficiency, good specificity, and can cover all target regions, including known mutation sites.

[0169] In samples of thalassemia patients, multiple known mutations were successfully detected, including point mutations, small fragment insertions / deletions, large fragment deletions / duplications, fragment recombinations, etc. In some samples, new mutation types were also detected.

[0170] Deletions of α-thalassemia were detected, including -- SEA Deletion, -α 3.7 Deletion, -α 4.2 Deletion, -- THAI Deletion, HKαα deletion;

[0171] Point mutations of α-thalassemia were detected, including common mutations cs, ws, qs, and rare α2 CD41 point mutation;

[0172] Common point mutations of β-thalassemia were detected, including CD41-42, IVS-II-654, -28, -29, CD71-72, CD27 / 28, CD43, IntM, IVS-I-1, IVS-1-5, CD17, CD14-15; β-thalassemia Chinese type was detected;

[0173] Rare mutations of β-thalassemia were detected, including IVS-II-1 (G→A), CD37, CD72 / 73(-TG), (S.E.A-HPFH), c.316_578delinsAAGTAGA, etc.;

[0174] Hemoglobin variants such as HbC, HbE, HbS, HbTW, HbG Siriraj, Hb G-Taipei, Hb Q-Thailand, HbGuangzhou-Hangzhou, Hb Youngstown were detected;

[0175] Different mutations in genes such as SPTB, ANK1, SPTA1, SLC4A1, G6PD, KIF23, SEC23B, CDAN1, EPB41, EPB42, GSR, ADD2, ALDOA, GYB5A, GPI, NT5C3A were detected.

[0176] LOH phenomena of βCD22 and CD26 were detected.

[0177] It should be noted that the embodiments of the present invention have better implementability and do not impose any form of limitation on the present invention. Any person skilled in the art may use the disclosed technical content to modify or transform it into equivalent effective embodiments. However, as long as it does not depart from the technical solution of the present invention, any modification, equivalent change or modification made to the above embodiments based on the technical essence of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A probe composition, characterized in that, Its nucleic acid sequence is as shown in SEQ ID NO. 1 to 64.

2. A kit for genetic detection of hereditary hemoglobinopathy, characterized in that, It includes the following two types of probes: The first type of probe, which is used to detect all coding regions of the hemoglobin gene, all non-coding regions including the regulatory region of the hemoglobin gene cluster, transcription factors, the gene coding region and regulatory region genes of the HBS1L-MYB gene that regulates hemoglobin switching; The second type of probe, which is used to detect single nucleotide mutations in the highly homologous sequence region of hereditary hemoglobinopathy.

3. The kit for detecting genes of hereditary hemoglobinopathy according to claim 2, characterized in that, It further includes a third type of probe, which is used to detect related genes that exclude red blood cell membrane defects, enzyme defects, and hereditary erythropoiesis abnormalities with phenotypes similar to hereditary hemoglobinopathy.

4. The kit for detecting genes of hereditary hemoglobinopathy according to claim 3, characterized in that, It further includes a fourth type of probe, which is used to detect HLA-A, HLA-B, and HLA-C loci.

5. A high-throughput DNA detection method for a hereditary hemoglobinopathy, characterized in that, It includes: Step S1: Design probes for detecting known and unknown mutations in genes related to abnormal hemoglobinopathy according to specific rules; Step S2: Synthesize the probes; Step S3: Extract genomic DNA from the patient's peripheral blood, construct a library, hybridize and capture the target sequence, recover the target region library and perform library quality detection to meet the requirements of NGS sequencing; Step S4: Load the library into the sequencing platform for high-throughput sequencing to obtain raw sequencing data; Step S5: Analyze the raw sequencing data: First, perform data quality control, genome alignment, repetitive sequence marking and clip quality correction on the raw sequencing data, re-align near indels, evaluate the coverage of the target region and the overall repetition rate, perform SNV and INDEL mutation analysis, and copy number variation CNV analysis; Next, perform gene mutation annotation, filter to obtain pathogenic sites according to the established SOP, and finally, confirm the authenticity of the mutation through IGV review, perform visual analysis on the CNV analysis, and perform image visualization correspondence between the genomic interval and the clinical report protein fragment.

6. The high-throughput DNA detection method for hereditary hemoglobinopathy according to claim 5, characterized in that In step S1, the specific rules are: based on the human genome database, for all coding regions of the hemoglobin gene and all non-coding regions including the regulatory region of the hemoglobin gene cluster; transcription factors; The gene coding region and regulatory region of the HBS1L-MYB gene that regulates hemoglobin switching; HLA-A, B, C loci; Design probes for related genes that exclude diseases with similar phenotypes: (1) Design probe length and density: The length of the designed probes is all 120bp; for the target capture region, for exons / intervals <120bp that do not overlap with the genomic repetitive region, the probe length will be truncated to 80bp or 60bp. Specifically, for exons ≤ 60bp, 60bp probes are designed, and for >60 ≤120bp, 80bp probes are designed to cover the entire region; (2) Design the probe density to be 1× coverage: For the highly homologous sequence of HBA2 exon2 and downstream 2kb, design 2× probe coverage; design 5kb density heterozygous SNP probes that are frequent in the Chinese population for possible CNV regions; (3) Blast the designed probes against the NCBI database to ensure the specificity of the probes; (4) Evaluate the synthesized probes through clinical sample sequencing: Apply the synthesized probes to the capture sequencing of clinical actual samples. The evaluation indicators are as follows: 1) Probe capture efficiency > 0.7; 2) Evaluation of capture uniformity; 3) Average coverage depth of the core gene region: Average sequencing depth is 500×, the region with a depth exceeding 100× should be higher than 95%, and the region with a depth exceeding 200× should exceed 90%; 4) The results of next-generation sequencing are 100% consistent with those of first-generation sequencing; 5) The gene mutation frequency is 100% consistent with the results of first-generation sequencing; (5) Optimize the probes according to the evaluation results: 1) Perform repetitive sequence analysis and motif analysis on the off-target regions. For repetitive regions and high-proportion motifs, adjust or delete the corresponding probe positions; 2) For low-coverage regions, without affecting the overall capture efficiency, increase the probe coverage of the low-capture regions and add specific probes. The positions of the specific probes are preferably partially overlapped with the original probes; 3) Increase the probe density or mutant probes for the core regions with low coverage depth of key hot-spot mutations; 4) For gene mutation sites with a large difference in gene frequency, increase the probe density, design mutant probes, or design CNV probes with a probe density of 10K within 100K around them.

7. The high-throughput DNA detection method for hereditary hemoglobinopathy according to claim 5, characterized in that, The step S5 includes S5.1 to S5.7: S5.1 Sequencing data quality control: Use FastQC to perform preliminary quality control on the raw DNA sequencing data, remove low-quality reads and adapter sequences to ensure data quality; After downloading the raw data and performing md5 verification, use the fastp software to filter the raw data reads with the following parameters: --trim_poly_g --trim_poly_x --length_required 60 --low_complexity_filter --complexity_threshold 20 –correction, filter the reads of polyG, polyN, reads with a length < 60bp, and reads with a sequence complexity lower than 20 respectively, and correct the overlapping parts of the paired reads; S5.2 Data alignment and preprocessing: Use bwa-0.7.18 and samtools-1.9.1 for alignment (hg19), bam conversion and sorting, and then perform ReorderSam, MarkDuplicates, BaseRecalibrator, ApplyBQSR respectively according to GATK best practices using 4.6.1.

0. Finally, obtain the bam file within the capture region through PrintReads for subsequent gene mutation analysis; S5.3 Gene Mutation Analysis: The paired WES samples were analyzed using Mutect2 and VarDict-1.8.3 software; specifically, for Mutect2, first, the mutation analysis used default parameters, and second, the Mutect2 mutation filter used default parameters to obtain a vcf file with filtered tags; finally, the Mutect2 results were filtered in the first round. The Mutect2 filtered tags were filtered by the SnpSift software with the following specific parameters: filter -n "(GEN[1].AF < 0.05) |(FILTER has 'panel_of_normals'| FILTER has 'normal_artifact' | FILTER has 'contamination' | FILTER has 'germline' | FILTER has 'weak_evidence'| FILTERhas 'low_allele_frac' | (FILTER has'multiallelic' & (FILTER has'slippage'|FILTER has 'panel_of_normals')))" to obtain the results of the first-round filtering of Mutect2, filtering out obvious systematic errors for subsequent integrated analysis; for VarDict, first, the following parameters were added in the mutation analysis: -r 7 -f0.05 –nosv; then, the non-standard part of the original vcf file was processed, and at the same time, the icons with the Q0 filtering parameter were forced to be replaced with PASS; after that, the vcf was standardized to obtain the results of the first-round filtering of VarDict for subsequent integrated analysis; S5.4 Integrated Analysis: First, the CombineVariants tool of GATK3 was used to integrate the results of the first-round filtering of Mutect2 and VarDict obtained previously. Through the mutual verification of the results of the two software, reliable gene mutation information was obtained, and a vcf file of the second-round filtering results was generated; Annotation of S5.5 gene mutations: Annotate the vcf file of the second-round filtering results obtained after final filtering. The annotation content includes: functional annotation and filtering annotation; for functional annotation, use SnpEff and Annovar to annotate the protein-coding changes of mutations. The principle is to annotate the splice sites 5bp upstream and downstream of CDS and exon; filtering annotation includes genetic diversity annotation of the normal population, disease database annotation, protein change function prediction filtering, and self-built negative and positive database filtering; specifically, use the SnpSift annotate tool to mark the processed database vcf. The genetic diversity databases include: dbsnp151, 1000Gp3, ESP6500, ExAC, gnomAD exomes, UK10K, ICGCV27; the disease databases include: cosmic85, clinvar, ICGCV27 and HGMD, and the protein change function prediction database uses dbNSFPv3.5a; the self-built negative database is the PON files trained respectively by using three software, Mutect1, Mutect2 and VarDict, on two different sequencing platforms, Illumina PE-150 and BGI PE-90, and three different capture kits, AgilentV5, IDTv2 and OncoWESuper, in normal control samples; through the above annotation, classify the pathogenicity of the variations and provide explanations for important mutations; Filtering and authenticity confirmation of S5.6 gene mutations: For the above gene mutation annotation results, screen the gene mutations according to the SOP formulated clinically, and reconfirm the authenticity of the gene mutations in the IGV software through the bam file. Finally, generate a genotype analysis report, including all detected pathogenic mutations and their clinical relevance; Copy number variation analysis: Capture and sequence the genomes of several normal people. For the bam file marked with duplicate sequences after bwa alignment, use the cnvkit software to train a normal reference genome database for this panel; then use the cbs algorithm of the cnvkit software to perform paired analysis on the bam file of each clinical sample under the same alignment conditions with the normal control to obtain standardized coverage information; then use the call tool of cnvkit to analyze the CNV interval, and then use the mapping tool of this software to visualize the interval or gene of interest; in the HBA and HBB regions, standardize the drawing coordinates to obtain a genomic image relative to clinical protein electrophoresis.

8. The high-throughput DNA detection method for hereditary hemoglobinopathy according to claim 5, characterized in that, In step S2, the actual synthesis of the probe is completed by solid-phase phosphoramidite chemical synthesis: (1) Single nucleotide extension: Phosphoramidite monomers of C, G, A, and T are gradually added to the solid-phase support; (2) Chemical reaction cycle: a. Deprotection: Remove the protecting group to expose the active group; b. Extension: Add the next nucleotide monomer and bind it to the 5'-hydroxyl group of the previous nucleotide; c. Oxidation stabilization: Treat with an oxidant to stabilize the phosphodiester bond; (3) Synthesis completion: Repeat the above steps until the 120-bp probe sequence is completed; During the synthesis process, biotin modification is added to the 5' end of the probe for subsequent binding to streptavidin magnetic beads; terminal modification is added to improve the probe stability; Purification treatment is carried out during the solid-phase synthesis process: (1) PAGE purification: Separate by polyacrylamide gel electrophoresis and distinguish the complete 120-bp probe according to the molecular weight; (2) HPLC purification: Further improve the probe purity by high-performance liquid chromatography to remove impurities and incorrect products; (3) Desalting treatment: Remove the residual salt and buffer components to improve the downstream experimental performance of the probe.

9. The high-throughput DNA detection method for hereditary hemoglobinopathy according to claim 5, wherein Step S3 includes: S3.1 Nucleic acid extraction: In a biosafety cabinet, take 200 μL of blood cell sample into a 1.5-mL centrifuge tube; use a commercial kit to extract genomic DNA according to the standard operating procedure; S3.2 Nucleic acid quality control: S3.2.1 Measure the concentration of the extracted DNA using a Qubit 4.0 fluorometer; S3.2.2 Detect the quality of the extracted DNA using an Agilent High Sensitivity DNA Kit; S3.2.3 Nucleic acid quality control standard: The total amount of DNA ≥ 300 ng and the quality is qualified; S3.

3. DNA fragmentation: S3.3.1 According to the starting sample amount of 300 ng, the Qubit concentration of the positive and negative quality control products needs to be verified. For the starting amount of 300 ng, add 1×TE buffer to 50 μL. The buffer is pH 8.0, 10 mM Tris-HCl, 0.1 mM EDTA-2Na, and calculate the input volume of each sample and the positive and negative quality control products; S3.3.2 Prepare labeled 1.5-mL centrifuge tubes according to the number of samples, the number of DNAs + positive and negative quality control products; S3.3.3 Add the corresponding DNA sample and 1×TE buffer to each centrifuge tube, mix well and perform a short centrifugation; S3.3.4 Use a Covaris M220 instrument for DNA fragmentation; S3.

4. Library construction: S3.4.1 The end repair plus A reagent is melted and shaken well in an ice box and centrifuged briefly; Prepare an "end repair plus A" mixed solution, shake well and place it in an ice box for later use; Transfer the reaction system to a PCR instrument for end repair plus A reaction; S3.4.2 Adapter Ligation: Prepare the Adapter Ligation Mix, vortex to mix well and centrifuge briefly, then place it on an ice box for standby; according to the experimental record sheet, take 5 μL of the diluted adapter with the corresponding number and add it to the end-repaired and A-tailed product, vortex for 20 seconds; add the Adapter Ligation Mix to the end-repaired product and adapter mixture, gently vortex to mix well and centrifuge briefly; set the PCR program, and after the hot lid and module temperature are stable, transfer the reaction system to the PCR instrument; S3.4.3 Purification after Adapter Ligation: Add 88 μL (0.8×) of purification magnetic beads to the ligation product, pipette to mix well, centrifuge briefly, and then let it stand at room temperature for 10 minutes; place the 8-strip tube on the magnetic rack to adsorb for 8 minutes, and aspirate the supernatant after the liquid becomes clear; add 200 μL of 80% ethanol to the tube, let it stand at room temperature for 30 seconds, discard the supernatant, and repeat once; centrifuge briefly to collect the liquid droplets on the tube wall, and use the pipette again to aspirate all the residual ethanol at the bottom, and dry at room temperature for 3 - 5 minutes; add 22 μL of ultrapure water, vortex for 10 seconds, centrifuge briefly, and then let it stand for 2 minutes; place the tube on the magnetic rack to adsorb for 3 minutes, and after the liquid becomes clear, take 2 μL of the sample for Qubit quantification, and transfer 20 μL of the supernatant to a new 8-strip tube for standby; S3.4.4 Pre-PCR Reaction: Take out the pre-PCR reagent, thaw it on an ice box, mix well before use and centrifuge briefly; prepare the pre-PCR Mix according to the number of libraries; add 30 μL of the pre-PCR Mix to a 0.2 mL 8-strip tube containing the purified product, vortex to mix well, centrifuge briefly and then set aside for standby; S3.4.5 Sorting and Purification after Pre-PCR Reaction: After the Pre-PCR is completed, centrifuge briefly, add 25 μL of purification magnetic beads to the tube, vortex to mix well and then let it stand for 10 minutes; place the 8-strip tube on the magnetic rack, adsorb for 5 minutes, and after the liquid becomes clear, transfer 75 μL of the supernatant to a new tube, and add 25 μL of purification magnetic beads; repeat the steps to ensure impurity removal; add 200 μL of 80% ethanol, let it stand for 30 seconds, aspirate the supernatant, and repeat once; centrifuge briefly, place the 8-strip tube on the magnetic rack, aspirate the residual ethanol, and dry at room temperature for 5 minutes; add 42 μL of ultrapure water, vortex to mix well, and then let it stand for 5 minutes; place the tube on the magnetic rack to adsorb for 1 minute, take the supernatant for Qubit quantification, and transfer 40 μL of the supernatant to a labeled centrifuge tube for standby; S3.

5. Library Hybridization and Capture: S3.5.1 Calculate the input volume according to the library concentration measured by Qubit; ensure that the total amount of each captured library is between 1000 ng and 4000 ng; S3.5.2 Mix the blocking sequence and Cot DNA, vortex to mix well and centrifuge briefly; S3.5.3 Place the sample in a vacuum centrifuge and dry it at 60°C for 20 minutes to 1 hour; S3.5.4 Prepare the hybridization mixture, add 17 μL to the dried sample, mix well and centrifuge briefly, let it stand at room temperature for 5 minutes, and repeat the mixing twice; S3.5.5 Set the PCR program to ensure that the hot lid temperature is 100°C; transfer the hybridization system to a PCR tube and run the program; S3.

6. Post-capture library cleaning: S3.6.1 Prepare relevant buffers and preheat the magnetic stand to 65 °C in advance; S3.6.2 Operating steps: Add 50 μL of magnetic beads to a PCR tube, vortex and mix well, adsorb, and discard the supernatant; Repeat the washing twice and add 100 μL of 1× magnetic bead elution buffer; After adding Bead Resuspension Mix, vortex and mix well, incubate at 65 °C for 15 minutes; Mix the magnetic beads with the hybridization solution, incubate at 65 °C for 45 minutes, and pipette gently every once in a while; Add 100 μL of elution buffer I at room temperature to the tube, adsorb and discard the supernatant after adsorption; Repeat the washing twice; Add 100 μL of elution buffer I at room temperature to the tube and transfer it to a new tube; Add 150 μL of elution buffers II and III at room temperature to the PCR tube, incubate at room temperature for 2 minutes, centrifuge briefly, and aspirate and discard the supernatant; Add 20 μL of ultrapure water and transfer it to a new tube for standby; S3.

7. Library enrichment: S3.7.1 Post-PCR reaction: Mix 20 μL of the post-capture product with 30 μL of post-PCR Mix, vortex and mix well, and gently shake to keep the magnetic beads suspended; Set the PCR program, with the hot lid at 105 °C, and run the Post-PCR reaction; S3.7.2 Purification after Post-PCR reaction: Take out the purification magnetic beads, vortex and mix well, and let stand for 30 minutes; After the Post-PCR is completed, place the PCR tube on the magnetic stand, adsorb and transfer 50 μL of the supernatant to a new tube; Add 75 μL of purification magnetic beads to the new tube and let stand at room temperature for 5 minutes; After brief centrifugation, aspirate and discard the supernatant, add 200 μL of 80% ethanol for washing and repeat once; Aspirate all the residual liquid and dry at room temperature until the surface of the magnetic beads is not reflective; Add 42 μL of ultrapure water, vortex and mix well, after brief centrifugation, transfer 40 μL of the supernatant to a new centrifuge tube for standby; S3.

8. Post-PCR library quality control: S3.8.1 Use a Qubit fluorometer to quantify the library concentration to ensure that the library concentration meets the requirements; S3.8.2 Use an Agilent or Bioanalyzer to detect the library quality to ensure compliance with NGS sequencing requirements.

10. The high-throughput DNA detection method for hereditary hemoglobinopathy according to claim 5, wherein Step S4 includes: S4.1 Library dilution: Dilute the library to an appropriate concentration according to the requirements of the sequencing platform; S4.2 Library mixing: Mix different libraries in proportion to ensure that the final concentration of each library is the same; S4.3 Library loading: S4.3.1 Load the library: The loading amount is between 10 - 15 pM; S4.3.2 Add PhiX control: Add a certain amount of PhiX control library, and the proportion of PhiX is 5 - 10%; S4.

4. Sequencing platform setting: Select an appropriate sequencing mode, sequencing depth, and number of cycles; S4.

5. Start sequencing; S4.

6. Monitor the sequencing process; S4.

7. Post-sequencing processing: S4.7.1 Data acquisition: After the sequencing is completed, extract the sequencing data from the instrument and output it in FASTQ format; S4.7.2 Library quality inspection: Use FastQC to perform quality control on the sequencing data, and check the library quality, sequencing depth, GC content, and sequence distribution; S4.7.3 Data backup: Back up all the original data.

11. A high-throughput DNA sequencing data analysis platform for a hereditary hemoglobinopathy, characterized in that Including: Sequencing data quality control component, data alignment and preprocessing component, gene mutation analysis component, integrated analysis component, gene mutation annotation component, gene mutation filtering and authenticity confirmation component, and copy number variation analysis component; The sequencing data quality control component is used for: performing preliminary quality control on the sequencing data using fastp, removing low-quality reads and adapter sequences to ensure data quality; after downloading the original data and performing md5 verification, using the fastp software to filter the original data reads with the following parameters: --trim_poly_g --trim_poly_x --length_required 60 --low_complexity_filter --complexity_threshold 20 –correction, filtering the reads of polyG, polyN, reads with length < 60bp, and reads with sequence complexity lower than 20 respectively, and at the same time performing base correction on the overlapping parts of paired reads; The data alignment and preprocessing component is used for: performing alignment (hg19), bam conversion and sorting using bwa-0.7.18 and samtools-1.9.1, and then performing ReorderSam, MarkDuplicates, BaseRecalibrator, ApplyBQSR respectively using 4.6.1.0 according to GATK best practices. Finally, obtain the bam file within the capture region through PrintReads for subsequent gene mutation analysis; The gene mutation analysis component is used for: analyzing the sequencing data captured by this panel using Mutect2 and VarDict-1.8.3 software; specifically, for Mutect2, first, the mutation analysis uses default parameters, and second, the Mutect2 mutation filter uses default parameters to obtain a vcf file with filtered tags; finally, the Mutect2 results are filtered in the first round. The Mutect2 filtered tags are filtered through the SnpSift software. The specific parameters are as follows: filter -n "(GEN[1].AF < 0.05) | (FILTER has 'panel_of_normals'| FILTER has 'normal_artifact'| FILTER has 'contamination' | FILTER has 'germline' | FILTER has 'weak_evidence'| FILTER has 'low_allele_frac' | (FILTER has'multiallelic' &(FILTER has'slippage'|FILTER has 'panel_of_normals')))" to obtain the results of the first-round filtering of Mutect2, filtering out obvious systematic errors for subsequent integrated analysis; for VarDict, first, the following parameters are added in the mutation analysis: -r 7 -f 0.05 –nosv; then, the non-standard part of the original vcf file is processed, and at the same time, the icons with Q0 filtering parameters are forced to be replaced with PASS; after that, the vcf is standardized to obtain the results of the first-round filtering of VarDict for subsequent integrated analysis; The integration analysis component is used for: first, using the CombineVariants tool of GATK3 to integrate the results of the first-round filtering of Mutect2 and VarDict obtained previously, and through the mutual verification of the results of the two software, obtaining reliable gene mutation information and generating a vcf file of the second-round filtering results; The gene mutation annotation component is used for: annotating the vcf file of the second-round filtering result finally filtered, and the annotation content includes: functional annotation and filtering annotation; for functional annotation, SnpEff and Annovar are used to annotate the protein-coding changes of mutations, and the principle is to annotate the splice sites 5bp upstream and downstream of CDS and exon; filtering annotation includes genetic diversity annotation of the normal population, disease database annotation, protein change function prediction filtering, and self-built negative and positive database filtering; specifically, the SnpSift annotate tool is used to mark the processed database vcf, and the genetic diversity databases include: dbsnp151, 1000Gp3, ESP6500, ExAC, gnomAD exomes, UK10K, ICGCV27; the disease databases include: cosmic85, clinvar, ICGCV27 and HGMD, and the protein change function prediction database uses dbNSFPv3.5a; the self-built negative database is the PON file trained respectively by using three software, namely Mutect1, Mutect2 and VarDict, on two different sequencing platforms, namely Illumina PE-150 and BGI PE-90, and three different capture kits, namely AgilentV5, IDTv2 and OncoWESuper, in normal control samples; through the above annotations, the pathogenicity of the variation is classified and an explanation is provided for important mutations; The gene mutation filtering and authenticity confirmation component is used for: screening the gene mutations according to the clinically formulated SOP for the above gene mutation annotation results, reconfirming the authenticity of the gene mutations in the IGV software through the bam file, and finally generating a genotype analysis report, including all detected pathogenic mutations and their clinical relevance; The copy number variation analysis component is used for: capturing and sequencing the genomes of several normal people, using the cnvkit software to train a normal reference genome database for this panel for the bam file marked with duplicate sequences after bwa alignment; then using the cbs algorithm of the cnvkit software to perform paired analysis on the bam file of each clinical sample under the same alignment conditions with the normal control to obtain standardized coverage information; then using the call tool of cnvkit to analyze the CNV interval, and then using the mapping tool of this software to visualize the interval or gene of interest; in the HBA and HBB regions, standardize the drawing coordinates to obtain a genome image relative to the clinical protein electrophoresis.

Citation Information

Patent Citations

  • Gene detection probe group for newborn inherited metabolic diseases and hemoglobinopathy and application thereof

    CN110938685A

  • Probe set and kit for detecting alpha thalassemia and beta thalassemia related pathogenic genes

    CN112359109A

  • Assay for hemoglobin A (hba) detection and genotyping

    CN115605615A

  • Device for noninvasive prenatal detection of thalassemia and application thereof

    CN118127143A

  • DNA chip for diagnosing hereditary anaemia related gene mutation

    CN1335406A