High-throughput DNA detection methods, platforms, probe compositions, and kits for hereditary hemoglobinopathies
By designing highly specific probe compositions and high-throughput sequencing platforms, combined with bioinformatics analysis, the problems of accuracy and coverage in the detection of hereditary hemoglobinopathies have been solved, enabling comprehensive detection of multiple mutations, reducing the false negative rate, and improving detection efficiency and accuracy.
Patent Information
- Application Number
- CN202510325231.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-03-19
AI Technical Summary
Existing technologies are insufficient to efficiently and accurately detect multiple mutations in hereditary hemoglobinopathies, especially common and rare mutations, leading to high rates of missed and misdiagnosed cases. Traditional methods are costly and time-consuming, and cannot cover multiple target sites. NGS technology has imperfect coverage and a rate of missed detection.
We design highly specific and sensitive probe compositions that cover all coding and regulatory regions of the hemoglobin gene. By combining high-throughput sequencing and bioinformatics analysis platforms, we can screen for hemoglobin gene mutations in a comprehensive manner. We use multiple probe types to distinguish different types of gene mutations and perform data quality control, analysis, and annotation to ensure the accuracy and coverage of the detection.
It enables comprehensive and accurate detection of hereditary hemoglobinopathies, reduces the false negative rate, can detect multiple mutation types simultaneously, including point mutations, insertions/deletions, etc., supports large-scale screening, improves detection efficiency, and provides reliable clinical diagnostic results.
Smart Images

Figure CN120290702B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of detection technology, and in particular to a high-throughput DNA detection method and bioinformatics analysis platform for hemoglobin and its regulatory genes. Background Technology
[0002] While some acquired factors can cause hemoglobin dysfunction or anemia, most anemia or hemolysis caused by hemoglobin defects are hereditary, usually caused by mutations in globin genes and their regulatory genes. Hereditary hemoglobinopathies can be broadly classified into structural variations and synthetic disorders of hemoglobin. Essentially, both are caused by changes in the bases of the primary structure. If the base changes only slightly affect the spatial structure of hemoglobin, it is a structural variation; if it causes a disorder in the formation of hemoglobin subunits, it is thalassemia, also known as thalassemia. Hemoglobin structural variations include HBC, HBD, HBE, HBM, HBS, HBJ, etc., with over 1300 hemoglobin variants reported to date; hemoglobin synthetic disorders include various types of thalassemia such as α / β / δβ / γδβ. The pathogenesis and clinical manifestations of abnormal hemoglobinopathies vary. Generally, most clinically significant hemoglobin structural variations lead to hemolysis; thalassemia involves ineffective erythropoiesis and destruction of erythroid precursor cells in the bone marrow; sickle cell disease (SCD) can also cause vascular occlusion. Meanwhile, some erythrocyte membrane diseases and enzyme disorders can also present with similar clinical manifestations, requiring differential diagnosis.
[0003] Most structural variations in hemoglobin are inherited in an autosomal dominant manner. Although the symptoms are mild, anyone carrying the gene will experience symptoms. Thalassemia, however, is considered a heterogeneous genetic disease, ranging from asymptomatic carriers to transfusion-dependent anemia. In cases of infection or stress, some normally asymptomatic carriers may exhibit mild anemia symptoms. For a long time, preventing the birth of children with severe thalassemia was the focus of laboratory testing. However, with large-scale population migration and increased demands for a higher quality of life, accurate diagnosis of hemoglobin gene mutation carriers has become increasingly necessary.
[0004] Traditional diagnostic methods have various limitations. Currently, methods for detecting gene mutations in molecular biology mainly include first-generation sequencing (Sanger), next-generation sequencing (NGS), real-time fluorescence PCR (RT-PCR), and digital PCR (dPCR). Sanger, RT-PCR, and dPCR primarily target known mutation sites in genes, and the cost of primers and probes is high, making multi-gene mutation screening less feasible. Traditional hemoglobinopathies detection methods include hematological examination, hemoglobin electrophoresis, PCR detection, and Sanger sequencing. Hematological examination mainly relies on peripheral blood erythrocyte parameters (such as MCV and MCH), combined with hemoglobin electrophoresis analysis of the proportion of different hemoglobin types for preliminary screening of suspected carriers. However, these methods cannot distinguish specific gene mutation types and have low sensitivity to deletion mutations in α-thalassemia, especially in mild or silent carriers, leading to missed diagnoses. Commonly used techniques for detecting hemoglobin include HPLC, IEF, and capillary electrophoresis. Although these methods are relatively inexpensive, due to the nature of the technology, many abnormal hemoglobin peaks overlap in close proximity and cannot be distinguished, leading to missed diagnoses or misdiagnoses.
[0005] Although the role of first-line DNA testing is evolving, it is rarely used as a screening test for abnormal hemoglobinopathies. Routine use of DNA testing as a second-line diagnostic method is becoming the medical standard. For patients suspected of having hemoglobinopathies, molecular testing is the gold standard for diagnosis. Currently, commonly used methods in my country include GAP-PCR and reverse dot blot hybridization, designed based on known high-frequency mutations in the Chinese population. While these methods offer high cost-effectiveness and can diagnose most thalassemia patients, they cannot definitively diagnose hemoglobinopathies with uncommon mutations, especially those other than HBE. They are largely unused for carriers with very mild symptoms. Traditional research-oriented PCR methods design primers for specific known point mutations or deletions to detect common mutations, but they are ineffective against rare or unknown mutations and cannot simultaneously cover multiple target sites. Sanger sequencing is widely used for gene mutation analysis, but this method is time-consuming and has low throughput, making it unsuitable for complex gene testing and large-scale screening, especially for detecting copy number variations (CNVs) and tandem repeat sequences (such as HBA1 and HBA2 genes). While MLPA can detect some repeats and rearrangements of globin sequences, it still suffers from drawbacks such as a limited detectable region and high cost. Although several kits for detecting thalassemia using NGS technology have been patented and some are already in use at third-party testing centers, we have found that some use amplicon methods, some have imperfect coverage and require continuous updates, and still have a certain rate of false negatives. Summary of the Invention
[0006] To address the aforementioned technical problems in the prior art, this application provides four technical solutions, as follows:
[0007] The first aspect of this application provides a probe composition having a nucleic acid sequence including, but not limited to, probes as shown in SEQ ID NO. 1 to 64.
[0008] A second aspect of this application provides a kit for gene detection of hereditary hemoglobinopathies, comprising the following two types of probes:
[0009] Type 1 probes are used to detect all coding regions of the hemoglobin gene, all non-coding regions including the regulatory region of the hemoglobin gene cluster, transcription factors, the coding region and regulatory region of the gene HBS1L-MYB that regulates hemoglobin switching.
[0010] The second type of probe is used to detect single nucleotide mutations in highly homologous sequence regions of hereditary hemoglobinopathies.
[0011] Furthermore, the kit for genetic testing of hereditary hemoglobinopathies further includes a third type of probe, which is used to detect genes associated with erythrocyte membrane defects (hereditary spherocytosis, hereditary elliptic erythrocytes, hereditary acanthosis, hereditary abeloinemia), enzyme defects (glucose-6-phosphate dehydrogenase deficiency, pyruvate kinase deficiency, pyrimidine-5'-nucleotidase deficiency), and hereditary abnormal erythropoiesis that have similar phenotypes to hereditary hemoglobinopathies.
[0012] Furthermore, the kit for genetic testing of hereditary hemoglobinopathies further includes a fourth type of probe for detecting HLA-A, HLA-B, and HLA-C loci.
[0013] The technical concept of this application is to use a first-type probe and a second-type probe to comprehensively screen patients with mutations in the hemoglobin gene. When a patient with similar clinical symptoms but no detected gene abnormality is encountered, a third-type probe plays a role in differential diagnosis. Simultaneously, a fourth-type probe can assist in family diagnosis. For example, the nucleic acid sequences of the first-type probe include, but are not limited to, those shown in SEQ ID NO. 1–13; the nucleic acid sequences of the second-type probe include, but are not limited to, those shown in SEQ ID NO. 14–35; the nucleic acid sequences of the third-type probe include, but are not limited to, those shown in SEQ ID NO. 36–61; and the nucleic acid sequences of the fourth-type probe include, but are not limited to, those shown in SEQ ID NO. 62–64.
[0014] The third aspect of this application proposes a high-throughput DNA detection method for hereditary hemoglobinopathies, comprising:
[0015] Step S1: Design probes according to specific rules to detect known and unknown mutations in genes related to abnormal hemoglobinopathies;
[0016] Step S2: Synthesize the probe;
[0017] Step S3: Extract genomic DNA from the patient's peripheral blood, construct a library, hybridize and capture the target sequence, recover the target region library and perform library quality testing to meet NGS sequencing requirements;
[0018] Step S4: Load the library into the sequencing platform for high-throughput sequencing to obtain raw sequencing data;
[0019] Step S5: Data analysis of raw sequencing data: First, perform data quality control, genome alignment, repetitive sequence labeling and editing quality correction on the raw sequencing data, re-align near indels, evaluate the coverage of the target region and the overall repetition rate, perform SNV and INDEL mutation analysis, and copy number variation (CNV) analysis; next, perform gene mutation annotation, filter to obtain pathogenic sites according to the established SOP, and finally, confirm the authenticity of mutations through IGV review, perform visualization analysis of CNV analysis, and visualize the correspondence between genomic regions and clinical report protein fragments.
[0020] Further, in step S1, the specific rules are: based on the human genome database, targeting all coding regions of the hemoglobin gene and all non-coding regions, including the regulatory region of the hemoglobin gene cluster; transcription factors; the coding and regulatory regions of the gene HBS1L-MYB, which regulates hemoglobin switching; HLA-A, B, and C loci; and probes designed to exclude diseases with similar phenotypes.
[0021] (1) Design of probe length and density: The length of the designed probes is 120bp; for the target capture region, the probe length will be shortened to 80bp or 60bp for exons / intervals that do not overlap with the genomic repetitive region and are <120bp. Specifically, for exons ≤ 60bp, design 60bp probes, and for exons > 60 ≤ 120bp, design 80bp probes to fully cover the entire region;
[0022] (2) Design probe density of 1× coverage: For HBA2 exon2 and downstream 2kb highly homologous sequences, design 2× probe coverage; for possible CNV regions, design 5kb density of high-frequency heterozygous SNP probes in the Chinese population.
[0023] (3) Perform a BLAST comparison between the designed probe and the NCBI database to ensure the specificity of the probe;
[0024] (4) Evaluation of the synthesized probes by clinical sample sequencing: The synthesized probes were applied to the actual clinical sample capture sequencing, and the evaluation indicators were: 1) Probe capture efficiency > 0.7; 2) Capture uniformity evaluation; 3) Average coverage depth of the core gene region: the average measurement depth was 500×, the region with more than 100× should be higher than 95%, and the region with more than 200× should be higher than 90%; 4) The second-generation sequencing results were 100% consistent with the first-generation sequencing results; 5) The gene mutation frequency was 100% consistent with the first-generation sequencing results.
[0025] (5) Optimize the probes based on the evaluation results: 1) Perform repetitive sequence analysis and motif analysis on the off-target region, and adjust or delete the corresponding probe positions for repetitive regions and high proportion of motifs; 2) For low coverage regions, increase the probe coverage of low capture regions and increase specific probes without affecting the overall low capture efficiency. The position of specific probes should ideally overlap with the original probes; 3) For key hot spot mutation core regions with low coverage depth, increase probe density or mutant probes; 4) For gene mutation sites with large differences in gene frequency, increase probe density, design mutant probes, or design CNV probes with a probe density of 10K within a 100K radius.
[0026] Further, step S5 includes S5.1 to S5.7:
[0027] S5.1 Sequencing Data Quality Control: FastP is used to perform preliminary quality control on the raw DNA sequencing data, removing low-quality reads and adapter sequences to ensure data quality. After downloading the raw data and performing MD5 verification, the following parameters are used in FastP software to filter the raw data reads: --trim_poly_g --trim_poly_x --length_required 60 --low_complexity_filter --complexity_threshold 20 –correction. This filters poly G and poly N reads, reads with a length <60bp, and reads with a sequence complexity lower than 20, respectively. At the same time, base correction is performed on the overlapping parts of paired reads.
[0028] S5.2 Data alignment and preprocessing: bwa-0.7.18 and samtools-1.9.1 were used for alignment (hg19), BAM conversion and sorting. Then, according to GATK best practices, ReorderSam, MarkDuplicates, BaseRecalibrator, and ApplyBQSR were performed using version 4.6.1.0. Finally, BAM files within the captured region were obtained through PrintReads for subsequent gene mutation analysis.
[0029] S5.3 Gene Mutation Analysis: The sequencing data of this panel was analyzed using Mutect2 and VarDict-1.8.3 software. Specifically, for Mutect2, firstly, the mutation analysis used the default parameters, and secondly, the Mutect2 mutation filter used the default parameters to obtain a vcf file with filter tags. Finally, the first round of filtering of the Mutect2 results was performed using SnpSift software to filter the Mutect2 filter tags. The specific parameters are as follows: filter -n "(GEN[1].AF < 0.05) |(FILTER has 'panel_of_normals'| FILTER has 'normal_artifact' | FILTER has 'contamination' | FILTER has 'germline' | FILTER has 'weak_evidence'| FILTER has 'low_allele_frac' | (FILTER has 'multiallelic' & (FILTER has 'slippage'|FILTER has 'panel_of_normals')))" The first round of filtering for Mutect2 is obtained, filtering out obvious systematic errors for subsequent integration analysis. For VarDict, the following parameters are first added to the mutation analysis: -r 7 -f0.05 –nosv. Then, the non-standard parts of the original vcf file are processed, and icons with the Q0 filtering parameter are forcibly replaced with PASS. After that, the vcf is standardized to obtain the first round of VarDict filtering results for subsequent integration analysis.
[0030] S5.4 Integration Analysis: First, the CombineVariants tool of GATK3 is used to integrate the first-round filtering results of Mutect2 and VarDict obtained earlier. The results of the two software are cross-validated to obtain reliable gene mutation information and generate the second-round filtering result vcf file.
[0031] S5.5 Gene Mutation Annotation: The VCF files from the second round of filtering are annotated, including functional and filtering annotations. Functional annotations use SnpEff and Annovar to annotate protein-coding changes caused by mutations, focusing on splice sites 5 bp upstream and downstream of the CDS and exon. Filtering annotations include annotations of genetic diversity in the normal population, disease database annotations, protein change function prediction filtering, and filtering of self-built negative and positive databases. Specifically, the processed database VCFs are labeled using the SnpSift annotate tool. Genetic diversity databases include: dbsnp151, 1000Gp3, ESP6500, ExAC, gnomAD exomes, UK10K, and ICGCV27; disease databases include: cosmic85, clinvar, ICGCV27, and HGMD; the protein change function prediction database uses dbNSFPv3.5a; and the self-built negative database uses Mutect1, PON files were trained on normal control samples using three software programs: Mutect2 and VarDict, on two different sequencing platforms (Illumina PE-150 and BGI PE-90) and three different capture kits (Agilent V5, IDTv2, and OncoWESuper). The pathogenicity of the variants was classified and explanations were provided for important mutations through the above annotations.
[0032] S5.6 Gene mutation filtering and authenticity verification: Based on the above gene mutation annotation results, gene mutations are screened according to the clinically established SOP. The authenticity of gene mutations is verified again in IGV software through BAM files. Finally, a genotype analysis report is generated, including all detected pathogenic mutations and their clinical relevance.
[0033] S5.7 Copy Number Variation Analysis: The genomes of several normal individuals were captured and sequenced. BAM files with repetitive sequences were labeled after BWA alignment. A normal reference genome database for this panel was trained using cnvkit software. Then, the cbs algorithm of cnvkit software was used to perform paired analysis of BAM files from each clinical sample under the same alignment conditions with normal controls to obtain standardized coverage information. Next, the call tool of cnvkit was used to analyze CNV regions, and the software's plotting tools were used to visualize regions or genes of interest. In the HBA and HBB regions, the plotting coordinates were standardized to obtain genomic images relative to clinical protein electrophoresis.
[0034] Further, in step S2, the actual synthesis of the probe is completed through solid-phase phosphoramide chemical synthesis: (1) Single nucleotide extension: phosphoramide monomers of C, G, A, and T are gradually added to the solid-phase support; (2) Chemical reaction cycle: a. Deprotection: removing the protecting group to expose the active group; b. Extension: adding the next nucleotide monomer and binding it to the 5'-hydroxyl group of the previous nucleotide; c. Oxidative stabilization: stabilizing the phosphodiester bond by treatment with an oxidant; (3) Synthesis completion: repeating the above steps until a 120 bp probe sequence is completed;
[0035] During the synthesis process, biotin was added to the 5' end of the probe for subsequent binding with streptavidin magnetic beads; end modification was added to improve probe stability.
[0036] Purification was carried out during solid-phase synthesis: (1) PAGE purification: the complete 120 bp probe was separated by polyacrylamide gel electrophoresis according to the molecular weight; (2) HPLC purification: the purity of the probe was further improved by high performance liquid chromatography to remove impurities and erroneous products; (3) Desalting treatment: residual salt and buffer components were removed to improve the downstream experimental performance of the probe.
[0037] Further, step S3 includes:
[0038] S3.1 Nucleic acid extraction: In a biosafety cabinet, transfer 200 μL of blood cell sample into a 1.5 mL centrifuge tube; use a commercial kit to extract genomic DNA according to standard operating procedures.
[0039] S3.2 Nucleic Acid Quality Control:
[0040] S3.2.1 The concentration of the extracted DNA was determined using a Qubit 4.0 fluorescence spectrometer;
[0041] S3.2.2 The quality of the extracted DNA was tested using the Agilent High Sensitivity DNA Kit;
[0042] S3.2.3 Nucleic acid quality control standard: Total DNA amount ≥300ng, quality qualified;
[0043] S3.3. DNA Fragmentation:
[0044] S3.3.1 Based on an initial sample volume of 300 ng, and considering the need to verify the Qubit concentration for positive and negative control samples, add 1×TE buffer to a final volume of 50 μL. The buffer solution is pH 8.0, 10 mM Tris-HCl, and 0.1 mM EDTA-2Na. Calculate the input volume for each sample and the positive and negative control samples. S3.3.2 Based on the sample quantity, the number of DNA samples, and the positive and negative control samples, prepare labeled 1.5 mL centrifuge tubes. S3.3.3 Add the corresponding DNA sample and 1×TE buffer to each centrifuge tube, mix well, and perform a brief centrifugation. S3.3.4 Perform DNA fragmentation using a Covaris M220 instrument.
[0045] S3.4. Library Construction:
[0046] S3.4.1 Melt the end-repair and A reagent in an ice box and vortex to mix well, then briefly centrifuge; prepare the "end-repair and A" mixed solution, vortex to mix well, and place in an ice box for later use; transfer the reaction system to a PCR instrument to perform the end-repair and A reaction;
[0047] S3.4.2 Adapter ligation: Prepare the adapter ligation mix, vortex to mix and briefly centrifuge, then place in an ice box for later use; according to the experimental record sheet, take 5 μL of the corresponding number of diluted adapter and add it to the end repair product A, vortex to mix for 20 seconds; add the adapter ligation mix to the end repair product and adapter mixture, vortex to mix and briefly centrifuge; set the PCR program, and after the temperature of the hot lid and module has stabilized, transfer the reaction system to the PCR instrument;
[0048] S3.4.3 Purification after ligation: Add 88 μL (0.8×) purification magnetic beads to the ligation product, mix by pipetting, briefly centrifuge, and let stand at room temperature for 10 minutes; place the eight-tube strip on a magnetic rack for adsorption for 8 minutes, and discard the supernatant after the liquid becomes clear; add 200 μL of 80% ethanol to the tube, let stand at room temperature for 30 seconds, discard the supernatant, and repeat once; collect the droplets on the tube wall by brief centrifugation, and again use a pipette to remove the residual ethanol at the bottom, and dry at room temperature for 3-5 minutes; add 22 μL of ultrapure water, vortex to mix for 10 seconds, briefly centrifuge, and let stand for 2 minutes; place the tube on a magnetic rack for adsorption for 3 minutes, and after the liquid becomes clear, take 2 μL of sample for Qubit quantification, and transfer 20 μL of supernatant to a new eight-tube strip for later use;
[0049] S3.4.4 Pre-PCR reaction: Take out the pre-PCR reagent, place it on an ice box to thaw, mix well before use, and centrifuge briefly; prepare the pre-PCR Mix according to the number of libraries; add 30 μL of pre-PCR Mix to a 0.2 mL eight-tube containing the purified product, vortex to mix, and centrifuge briefly before use.
[0050] S3.4.5 Sorting and Purification after Pre-PCR: After Pre-PCR, briefly centrifuge, add 25 μL of purification magnetic beads to the tube, vortex to mix, and let stand for 10 minutes; place the eight-tube strip on a magnetic rack and let it adsorb for 5 minutes. After the liquid becomes clear, transfer 75 μL of supernatant to a new tube and add 25 μL of purification magnetic beads; repeat the steps to ensure impurity removal; add 200 μL of 80% ethanol, let stand for 30 seconds, discard the supernatant, and repeat once; briefly centrifuge, place the eight-tube strip on a magnetic rack, discard the residual ethanol, and dry at room temperature for 5 minutes; add 42 μL of ultrapure water, vortex to mix, and let stand for 5 minutes; place the tube on a magnetic rack and let it adsorb for 1 minute, take the supernatant for Qubit quantification, and transfer 40 μL of supernatant to a labeled centrifuge tube for later use;
[0051] S3.5. Library hybridization and capture:
[0052] S3.5.1 Calculate the input volume based on the library concentration measured by Qubit; ensure the total amount of each captured library is between 1000ng and 4000ng; S3.5.2 Mix the blocking sequence and Cot DNA, vortex to mix, and briefly centrifuge; S3.5.3 Place the sample in a vacuum centrifuge and evaporate to dryness at 60℃ for 20 minutes to 1 hour; S3.5.4 Prepare the hybridization mixture by adding 17μL to the evaporated sample, mixing well, briefly centrifuging, and incubating at room temperature for 5 minutes. Repeat the mixing process twice; S3.5.5 Set the PCR program, ensuring the hot cap temperature is 100℃; transfer the hybridization system to PCR tubes and run the program;
[0053] S3.6. Post-capture library cleaning:
[0054] S3.6.1 Prepare the relevant buffer solution and preheat the magnetic rack to 65°C;
[0055] S3.6.2 Procedure: Add 50 μL of magnetic beads to a PCR tube, vortex to mix, adsorb, and discard the supernatant; repeat the washing process twice, and add 100 μL of 1× magnetic bead elution buffer; add Bead Resuspension Mix, vortex to mix, and incubate at 65°C for 15 minutes; mix the magnetic beads with the hybridization solution, incubate at 65°C for 45 minutes, and pipette the mixture periodically; add 100 μL of elution buffer I to the tube, adsorb, and discard the supernatant; repeat the washing process twice; add 100 μL of room temperature elution buffer I to the tube and transfer to a new tube; add 150 μL of room temperature elution buffers II and III to the PCR tube, incubate at room temperature for 2 minutes, briefly centrifuge, and discard the supernatant; add 20 μL of ultrapure water and transfer to a new tube for later use;
[0056] S3.7. Document Enrichment:
[0057] S3.7.1 Post-PCR reaction: Mix 20 μL of the captured product with 30 μL of post-PCR Mix, vortex to mix, and gently shake to keep the magnetic beads suspended; set the PCR program, heat the lid to 105℃, and run the Post-PCR reaction;
[0058] S3.7.2 Post-PCR purification: Remove the purification magnetic beads, vortex to mix, and let stand for 30 minutes; after the post-PCR is completed, place the PCR tube on a magnetic rack, adsorb, and transfer 50 μL of supernatant to a new tube; add 75 μL of purification magnetic beads to the new tube and let stand at room temperature for 5 minutes; after a short centrifugation, discard the supernatant, add 200 μL of 80% ethanol to wash, and repeat once; aspirate the residual liquid and dry at room temperature until the surface of the magnetic beads is no longer reflective; add 42 μL of ultrapure water, vortex to mix, and after a short centrifugation, transfer 40 μL of supernatant to a new centrifuge tube for later use;
[0059] S3.8. Post-PCR library quality control:
[0060] S3.8.1 Use a Qubit fluorescence spectrometer to quantify the library concentration to ensure that the library concentration meets the requirements;
[0061] S3.8.2 Use Agilent or Bioanalyzer to perform library quality checks to ensure compliance with NGS sequencing requirements.
[0062] Further, step S4 includes:
[0063] S4.1 Library dilution: Dilute the library to the appropriate concentration according to the requirements of the sequencing platform;
[0064] S4.2 Library Mixing: Mix different libraries in a certain proportion to ensure that the final concentration of each library is consistent;
[0065] S4.3 Library Loading:
[0066] S4.3.1 Loading libraries: The loading size is between 10-15 pM;
[0067] S4.3.2 Add PhiX control: Add a certain amount of PhiX control library, with the proportion of PhiX being 5-10%;
[0068] S4.4. Sequencing platform settings: Select the appropriate sequencing mode, sequencing depth, and cycle number.
[0069] S4.5. Start sequencing;
[0070] S4.6. Monitoring the sequencing process;
[0071] S4.7. Post-sequencing processing:
[0072] S4.7.1 Data Acquisition: After sequencing is completed, extract the sequencing data from the instrument and output it in FASTQ format;
[0073] S4.7.2 Library quality check: Use FastQC to perform quality control on sequencing data, checking library quality, sequencing depth, GC content, and sequence distribution;
[0074] S4.7.3 Data Backup: Back up all original data.
[0075] The fourth aspect of this application provides a high-throughput DNA sequencing data analysis platform for hereditary hemoglobinopathies, which includes: a sequencing data quality control component, a data alignment and preprocessing component, a gene mutation analysis component, an integrated analysis component, a gene mutation annotation component, a gene mutation filtering and authenticity verification component, and a copy number variation analysis component.
[0076] The sequencing data quality control component is used to: perform preliminary quality control on the raw DNA sequencing data using FastP, remove low-quality reads and adapter sequences, and ensure data quality; after downloading the raw data and performing MD5 verification, use FastP software with the following parameters to filter the raw data reads: --trim_poly_g --trim_poly_x --length_required 60 --low_complexity_filter --complexity_threshold 20 –correction, to filter poly G and poly N reads, reads with a length <60bp, and reads with a sequence complexity lower than 20, respectively, and to perform base correction on the overlapping parts of paired reads;
[0077] The data alignment and preprocessing components are used to: perform alignment (hg19), BAM conversion and sorting using bwa-0.7.18 and samtools-1.9.1, then perform ReorderSam, MarkDuplicates, BaseRecalibrator, and ApplyBQSR according to GATK best practices using version 4.6.1.0, and finally obtain the BAM file within the captured region using PrintReads for subsequent gene mutation analysis.
[0078] The gene mutation analysis component is used to analyze paired WES samples using Mutect2 and VarDict-1.8.3 software. Specifically, for Mutect2, firstly, the mutation analysis uses default parameters, and secondly, the Mutect2 mutationfilter uses default parameters to obtain a vcf file with filter tags. Finally, the first round of filtering of the Mutect2 results is performed by filtering the Mutect2 filter tags using SnpSift software. The specific parameters are as follows: filter -n "(GEN[1].AF <0.05) | (FILTER has 'panel_of_normals'| FILTER has 'normal_artifact' | FILTER has 'contamination' | FILTER has 'germline' | FILTER has 'weak_evidence'|FILTER has 'low_allele_frac' | (FILTER has 'multiallelic' & (FILTER has 'slippage'|FILTER has 'panel_of_normals')))" The first round of filtering for Mutect2 is obtained, filtering out obvious systematic errors for subsequent integration analysis. For VarDict, firstly, the following parameters are added to the mutation analysis: -r 7 -f 0.05 –nosv; then, the non-standard parts of the original vcf file are processed, and icons with the Q0 filtering parameter are forcibly replaced with PASS; then, the vcf is standardized to obtain the first round of VarDict filtering results for subsequent integration analysis.
[0079] The integrated analysis component is used to: firstly, integrate the first-round filtering results of Mutect2 and VarDict obtained earlier using the CombineVariants tool of GATK3, verify the results of the two software programs to obtain reliable gene mutation information, and generate a vcf file of the second-round filtering results;
[0080] The gene mutation annotation component is used to annotate the VCF files of the second-round filtered results obtained from the final filtering process. The annotations include functional annotations and filtering annotations. Functional annotations use SnpEff and Annovar to annotate the protein-coding changes of the mutations, primarily annotating the splice sites 5 bp upstream and downstream of the CDS and exon. Filtering annotations include annotations of genetic diversity in the normal population, disease database annotations, and protein change function prediction filtering, as well as filtering of self-built negative and positive databases. Specifically, the SnpSift annotate tool is used to label the processed database VCFs. Genetic diversity databases include: dbsnp151, 1000Gp3, ESP6500, ExAC, gnomAD exomes, UK10K, and ICGCV27; disease databases include: cosmic85, clinvar, ICGCV27, and HGMD; the protein change function prediction database uses dbNSFPv3.5a; and the self-built negative database is generated using Mutect1, Mutect2, and VarDict software in Illumina. PON files trained on normal control samples using two different sequencing platforms, PE-150 and BGI PE-90, and three different capture kits, Agilent V5, IDTv2, and OncoWESuper; the pathogenicity of the variants was classified and explanations were provided for important mutations through the above annotations.
[0081] The gene mutation filtering and authenticity verification component is used to: screen gene mutations according to clinically established SOPs based on the above gene mutation annotation results, reconfirm the authenticity of gene mutations in IGV software through BAM files, and finally generate a genotype analysis report, including all detected pathogenic mutations and their clinical relevance.
[0082] The copy number variation analysis component is used for: capturing and sequencing the genomes of several normal individuals; marking repetitive sequences in BAM files after BWA alignment; training a normal reference genome database for this panel using cnvkit software; then using the cbs algorithm of cnvkit software to perform paired analysis on BAM files of each clinical sample under the same alignment conditions with normal controls to obtain standardized coverage information; then using the call tool of cnvkit to analyze CNV intervals; and then using the software's plotting tools to visualize the intervals or genes of interest; in the HBA and HBB regions, standardizing the plotting coordinates to obtain a genome image relative to clinical protein electrophoresis.
[0083] After adopting the above technical solution, this application has the following excellent technical effects:
[0084] This application utilizes a large-scale shingled probe design to cover the hemoglobin gene, along with probes with specific sequences to differentiate similar regions of HBA1 and HBA2, and designs probes targeting common hemoglobin membrane and enzyme defects. Further bioinformatics analysis is performed on data obtained after large-scale sequencing. This invention provides a method for simultaneously detecting multiple hemoglobin mutation types, including point mutations, insertions / deletions, and gene rearrangements, significantly improving detection efficiency and coverage, avoiding missed detections, and enabling differential diagnosis of diseases with similar clinical manifestations. It achieves comprehensive and accurate identification of gene variations carried by patients, thereby facilitating timely intervention and treatment, and improving patient prognosis. This invention can detect all genomic variations of the hemoglobin α-subunit, β-subunit, and regulatory genes. The variations described in this invention include point mutations, small deletions and insertions, large deletions and duplications, inversions, translocations, and fusion variations. The bioinformatics analysis platform described in this invention can directly provide the mutation types of common mutations. This application develops a targeted DNA sequencing data analysis platform for customized panel features and target variant types, while also considering clinical applicability. The platform primarily features the following functions: 1) Automated alignment, analysis, filtering, and annotation; 2) Multi-algorithm analysis workflow, particularly for reassembly analysis of mutations in homologous sequence regions such as HBA1 and HBA2, significantly reducing the false negative rate; 3) Parallel analysis of multiple samples, with 16-20 samples analyzed in parallel at a time, averaging 10 minutes per sample; 4) Support for large indel analysis, accurately detecting indels up to 300 bp; 5) Support for CNV detection at the exon level within the capture region. Attached Figure Description
[0085] Figure 1 This is a flowchart of a high-throughput DNA detection method for hereditary hemoglobinopathies according to one embodiment of this application. Detailed Implementation
[0086] The advantages of the present invention are further illustrated below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that the following detailed description is illustrative rather than restrictive and should not be construed as limiting the scope of protection of the present invention.
[0087] like Figure 1 As shown, this embodiment provides a high-throughput DNA detection method for hereditary hemoglobinopathies. High-throughput sequencing technology can cover a wider range and has higher sensitivity and specificity. The high-throughput DNA detection method for hereditary hemoglobinopathies includes the following steps S1-S5:
[0088] Step S1: Design probes according to specific rules to detect known and unknown mutations in genes related to abnormal hemoglobinopathies;
[0089] In this step, a series of highly specific, highly sensitive, and high-coverage probes are designed to detect known and unknown mutations in genes related to abnormal hemoglobinopathies. Probe design is a key step in this embodiment, aiming to ensure that the probes specifically hybridize with the target sequence while avoiding non-specific hybridization. The experimental performance of the probes in this embodiment is shown in Table 1.
[0090] Table 1. Preliminary experimental performance of the probe
[0091]
[0092] Probe design: Based on the Human Genome Database (GRCh38), the probes target all coding regions of the hemoglobin gene (mainly including HBA1, HBA2, HBB, HBBP1, HBD, HBE1, HBG1, HBG2, HBM, HBQ, HBZ, etc.) and all non-coding regions, including the hemoglobin gene cluster regulatory regions (HS-40, HS1, HS2, HS3, HS4, HS5); transcription factors (GATA1, NFE2, KLF1, SP1, BLC11A); the coding and regulatory regions of the gene HBS1L-MYB, which regulates hemoglobin switching; HLA-A, B, and C loci (due to the need to verify kinship in some applications); and genes related to diseases with similar phenotypes (G6PD, ANK1, SLC4A1, SPTB, SPTA1, EPB42, EPB41, GYPC, GYPA, ADD2, ANK1, XK, GSR). SEC23B, CDAN1) probe design.
[0093] 1.1 Probe Length and Density: All designed probes are 120bp in length. For the target capture region, the probe length will be truncated to 80bp or 60bp for exons / intervals that do not overlap with the genomic repetitive region and are <120bp (specifically, 60bp probes are designed for exons ≤ 60bp, and 80bp probes are designed for exons > 60 ≤ 120bp to fully cover the entire region).
[0094] 1.2 The probe density is 1× coverage. For the highly homologous sequence HBA2 exon2 and downstream 2kb, a 2× probe coverage is designed. Hybrid SNP probes with a density of 5kb are designed for potential CNV regions, representing high frequencies found in the Chinese population.
[0095] 1.3 The designed probe is compared with databases such as NCBI using BLAST, and it is necessary to ensure that the probe has a moderate Tm value, no special structures such as palindromic sequences, and sequence homology between 50% and 80%. These values are used to ensure the specificity of the probe.
[0096] 1.4 Evaluation of synthesized probes through clinical sample sequencing: The synthesized probes were applied to capture sequencing of actual clinical samples. The main evaluation indicators were: 1) probe capture efficiency (>0.7); 2) capture uniformity evaluation; 3) average coverage depth of core gene regions (average measurement of 500×, regions exceeding 100× should be higher than 95%, and regions exceeding 200× should be higher than 90%); 4) consistency between second-generation sequencing results and first-generation sequencing results (100%); 5) consistency between gene mutation frequency and first-generation sequencing results (100%).
[0097] 1.5 Optimize probes based on evaluation results: The main contents of probe optimization include: (1) Perform repetitive sequence analysis and motif analysis on off-target regions. For repetitive regions and motifs with a proportion higher than 10%, adjust or delete the corresponding probe positions; (2) For low-coverage regions (the average probe coverage is less than 0.7X), increase the probe coverage of low-capture regions to above 0.9X without affecting the overall capture efficiency, and increase specific probes. The position of specific probes should ideally overlap with the original probes; (3) For key hotspot mutation core regions (the entire exon of mutation regions with clear clinical significance) with low domain coverage depth (< 0.8X), increase probe density (to > 0.95X) or mutant probes; (4) For gene mutation sites with large differences in gene frequency (the difference in VAF of mutations in the previous reproducibility test is greater than 5%), increase probe density, design mutant probes, or design CNV probes with a probe density of 10K within a 100K radius.
[0098] For example, the four types of probe sequences in this embodiment are shown in Table 2:
[0099] Table 2. Nucleic acid sequences corresponding to the four types of probes
[0100]
[0101]
[0102]
[0103] When the sequence in Table 2 is inconsistent with the sequence list, Table 2 shall prevail.
[0104] Step S2: Synthesize the probe;
[0105] The actual synthesis of the probe is accomplished through solid-phase phosphoramide chemical synthesis: (1) Single nucleotide extension: phosphoramide monomers of C, G, A, and T are gradually added to a solid support (usually polystyrene microbeads). (2) Chemical reaction cycle: a. Deprotection: removing the protecting group to expose the active group. b. Extension: adding the next nucleotide monomer, which binds to the 5'-hydroxyl group of the previous nucleotide. c. Oxidative stabilization: stabilizing the phosphodiester bond by treatment with an oxidant. (3) Synthesis completion: repeating the above steps until a 120 bp probe sequence is completed.
[0106] During synthesis, biotin was added to the 5' end of the probe for subsequent binding with streptavidin magnetic beads. Terminal modifications (such as phosphate groups) were added to improve probe stability.
[0107] Since incompletely extended short fragments may be generated during solid-phase synthesis, purification is required: (1) PAGE purification: separate by polyacrylamide gel electrophoresis to distinguish complete 120 bp probes according to molecular weight. (2) HPLC purification: further improve probe purity by high-performance liquid chromatography to remove impurities and erroneous products. (3) Desalting treatment: remove residual salt and buffer components to improve the downstream experimental performance of the probe.
[0108] Step S3: Extract genomic DNA from the patient's peripheral blood, construct a library, hybridize and capture the target sequence, recover the target region library and perform library quality testing to meet NGS sequencing requirements;
[0109] Sample processing: Genomic DNA is extracted from the patient's peripheral blood to obtain high-purity (A260 / 280 ratio of 1.8–2.0) and sufficient DNA quantity (≥50 ng). A library is constructed, hybridized, and the target sequence is captured. The target region library is recovered and quality-checked to meet NGS sequencing requirements. NGS (Next Generation Sequencing) result analysis is a complex and multi-step process involving raw data quality control, sequence alignment, variant detection, gene expression analysis, and functional annotation. When performing target region sequencing, NGS technology often uses probe capture technology to enrich specific DNA fragments to improve sequencing efficiency and reduce costs. The specific operation is as follows:
[0110] 3.1 Nucleic acid extraction: In a biosafety cabinet, transfer 200 μL of blood cell sample into a 1.5 mL centrifuge tube. Extract genomic DNA using a commercial kit (such as the QIAamp DNA Blood Mini Kit) following standard operating procedures.
[0111] 3.2 Nucleic Acid Quality Control
[0112] 3.2.1 The concentration of the extracted DNA was determined using a Qubit 4.0 fluorescence spectrometer.
[0113] 3.2.2 The quality of the extracted DNA was tested using the Agilent High Sensitivity DNA Kit.
[0114] 3.2.3 Nucleic acid quality control standard: Total DNA amount ≥300ng, quality qualified.
[0115] 3.3 DNA Fragmentation
[0116] 3.3.1 Based on an initial sample volume of 300 ng (for positive and negative control samples, the Qubit concentration needs to be verified, with an initial volume of 300 ng), add 1×TE buffer (pH 8.0, 10 mM Tris-HCl, 0.1 mM EDTA-2Na) to a final volume of 50 μL, and calculate the input volume for each sample and positive / negative control sample. 3.3.2 Prepare labeled 1.5 mL centrifuge tubes according to the sample quantity (number of DNA samples + positive / negative control samples). 3.3.3 Add the corresponding DNA sample and 1×TE buffer to each centrifuge tube, mix well, and briefly centrifuge. 3.3.4 Perform DNA fragmentation using a Covaris M220 instrument.
[0117] 3.4. Library Construction
[0118] 3.4.1 End-repair with A reagent: Melt the reagent in an ice box and vortex to mix, then briefly centrifuge. Prepare the "end-repair with A" mixture, vortex to mix, and store in an ice box for later use. Transfer the reaction system to a PCR instrument to perform the end-repair with A reaction.
[0119] 3.4.2 Adapter Ligation Preparation: Prepare the adapter ligation mix, vortex to mix, and briefly centrifuge. Place in an ice box for later use. According to the experimental record sheet, take 5 μL of the corresponding numbered diluted adapter and add it to the end-repair product A. Vortex to mix for 20 seconds. Add the adapter ligation mix to the end-repair product and adapter mixture, vortex gently to mix, and briefly centrifuge. Set the PCR program, and after the hot lid and module temperatures stabilize, transfer the reaction system to the PCR instrument.
[0120] 3.4.3 Post-Connection Purification: Add 88 μL (0.8×) purification magnetic beads to the ligation product, mix by pipetting, briefly centrifuge, and let stand at room temperature for 10 minutes. Place the eight-tube strip on a magnetic rack for adsorption for 8 minutes. After the liquid becomes clear, discard the supernatant. Add 200 μL of 80% ethanol to the tube, let stand at room temperature for 30 seconds, discard the supernatant, and repeat once. Briefly centrifuge to collect droplets on the tube wall, and again use a pipette to remove any residual ethanol at the bottom. Dry at room temperature for 3-5 minutes. Add 22 μL of ultrapure water, vortex for 10 seconds, briefly centrifuge, and let stand for 2 minutes. Place the tube on a magnetic rack for adsorption for 3 minutes. After the liquid becomes clear, take 2 μL of sample for Qubit quantification, and transfer 20 μL of supernatant to a new eight-tube strip for later use.
[0121] 3.4.4 Pre-PCR Reaction: Remove the pre-PCR reagent and thaw it on an ice pack. Mix well and centrifuge briefly before use. Prepare the pre-PCR Mix according to the number of libraries. Add 30 μL of the pre-PCR Mix to an eight-tube 0.2 mL tube containing the purified product, vortex to mix, and centrifuge briefly before use.
[0122] 3.4.5 Sorting and Purification after Pre-PCR Reaction After Pre-PCR, briefly centrifuge, add 25 μL of purification magnetic beads to the tube, vortex to mix, and let stand for 10 minutes. Place the eight-tube strip on a magnetic rack and allow it to adsorb for 5 minutes. After the liquid becomes clear, transfer 75 μL of supernatant to a new tube and add 25 μL of purification magnetic beads. Repeat the steps to ensure impurity removal. Add 200 μL of 80% ethanol, let stand for 30 seconds, discard the supernatant, and repeat once. Briefly centrifuge, place the eight-tube strip on a magnetic rack, discard residual ethanol, and dry at room temperature for 5 minutes. Add 42 μL of ultrapure water, vortex to mix, and let stand for 5 minutes. Place the tube on a magnetic rack and allow it to adsorb for 1 minute. Take the supernatant for Qubit quantification, and transfer 40 μL of the supernatant to a labeled centrifuge tube for later use.
[0123] 3.5. Library hybridization and capture
[0124] 3.5.1 Calculate the input volume based on the library concentration measured by Qubit. Ensure the total amount of each captured library is between 1000 ng and 4000 ng. 3.5.2 Mix the blocking sequence and Cot DNA, vortex to mix, and briefly centrifuge. 3.5.3 Place the sample in a vacuum centrifuge and evaporate to dryness at 60°C for 20 minutes to 1 hour. 3.5.4 Prepare the hybridization mixture by adding 17 μL to the evaporated sample, mixing well, briefly centrifuging, and incubating at room temperature for 5 minutes. Repeat the mixing process twice. 3.5.5 Set the PCR program, ensuring the heated lid temperature is 100°C. Transfer the hybridization system to PCR tubes and run the program.
[0125] 3.6. Post-capture library cleaning
[0126] 3.6.1 Prepare the relevant buffer solution and preheat the magnetic rack to 65°C.
[0127] 3.6.2 Procedure: Add 50 μL of magnetic beads to a PCR tube, vortex to mix, adsorb, and discard the supernatant. Repeat the washing process twice, adding 100 μL of 1× magnetic bead elution buffer. After adding Bead Resuspension Mix, vortex to mix and incubate at 65°C for 15 minutes. Mix the magnetic beads with the hybridization solution and incubate at 65°C for 45 minutes, repositioning the beads periodically. Add 100 μL of elution buffer I to the tube, adsorb, and discard the supernatant. Repeat the washing process twice. Add 100 μL of room temperature elution buffer I to the tube and transfer to a new tube. Add 150 μL of room temperature elution buffers II and III to the PCR tube, incubate at room temperature for 2 minutes, briefly centrifuge, and discard the supernatant. Add 20 μL of ultrapure water and transfer to a new tube for later use.
[0128] 3.7. Library Enrichment
[0129] 3.7.1 Post-PCR Reaction: Mix 20 μL of the captured product with 30 μL of post-PCR Mix, vortex to mix, and gently shake to keep the magnetic beads suspended. Set the PCR program, heat to 105°C, and run the Post-PCR reaction.
[0130] 3.7.2 Post-PCR Purification: Remove the purified magnetic beads, vortex to mix, and let stand for 30 minutes. After post-PCR, place the PCR tube on a magnetic rack, adsorb, and transfer 50 μL of supernatant to a new tube. Add 75 μL of purified magnetic beads to the new tube and let stand at room temperature for 5 minutes. After brief centrifugation, discard the supernatant, add 200 μL of 80% ethanol to wash, and repeat once. Aspirate all residual liquid and dry at room temperature until the surface of the magnetic beads is no longer reflective. Add 42 μL of ultrapure water, vortex to mix, and after brief centrifugation, transfer 40 μL of supernatant to a new centrifuge tube for later use.
[0131] 3.8. Post-PCR Library Quality Control
[0132] 3.8.1 Use a Qubit fluorometer to quantify the library concentration to ensure that the library concentration meets the requirements.
[0133] 3.8.2 Use Agilent or Bioanalyzer to perform library quality checks to ensure compliance with NGS sequencing requirements.
[0134] Step S4: Load the library into the sequencing platform for high-throughput sequencing to obtain raw sequencing data;
[0135] The library was loaded into the sequencing platform (Illumina Nextseq 500) for high-throughput sequencing. The sequencer generated raw data (FASTQ files), and the data quality was checked using tools (FastQC), including Q-value distribution, adapter contamination, and GC content. The specific steps are as follows:
[0136] 4.1 Library dilution: Dilute the library to a suitable concentration according to the requirements of the sequencing platform. For the NextSeq platform, the recommended concentration is 12-15 nM.
[0137] 4.2 Library Mixing: Different libraries can be mixed in proportion according to the experimental design to ensure that the final concentration of each library is consistent. During mixing, it is essential to ensure that the proportions between the libraries are appropriate.
[0138] 4.3 Library Loading
[0139] 4.3.1 Library Loading: The loading volume is generally adjusted according to the specific sequencing depth and data volume requirements, typically between 10-15 pM. For high-throughput platforms, the library concentration can be appropriately reduced to ensure uniform distribution of each library.
[0140] 4.3.2 Adding PhiX control: To improve sequencing quality and correct sequence bias, a certain amount of PhiX control library is usually added. The proportion of PhiX is generally 5-10%, which helps to increase the sequencing depth of GC-rich regions and correct sequencing errors.
[0141] 4.4. Sequencing platform settings: Select the appropriate sequencing mode, sequencing depth, and cycle number.
[0142] 4.5. Start sequencing
[0143] 4.6. Sequencing process monitoring
[0144] 4.7. Post-sequencing processing
[0145] 4.7.1 Data Acquisition: After sequencing is complete, extract the sequencing data from the instrument, typically output in FASTQ format. Ensure all raw data is correctly saved to avoid loss.
[0146] 4.7.2 Library quality control: Use appropriate tools (such as FastQC) to perform quality control on sequencing data, and check library quality, sequencing depth, GC content, sequence distribution, etc.
[0147] 4.7.3 Data Backup: Back up all original data to ensure data security. Data can be uploaded to a cloud platform or saved to external storage devices.
[0148] Step S5: Perform data analysis on the raw sequencing data.
[0149] The sequencing raw data undergoes genome alignment, repetitive sequence labeling and editing quality correction, re-alignment near possible indels, evaluation of target region coverage and overall repetition rate, SNV and INDEL mutation analysis, and copy number variation (CNV) analysis. Next, gene mutation annotation is performed, filtering pathogenic sites according to established SOPs. Finally, IGV review confirms the authenticity of mutations, and CNV analysis is visualized, with genomic regions visually corresponding to clinically reported protein fragments.
[0150] 5.1 Sequencing Data Quality Control
[0151] FastP was used for initial quality control of sequencing data to remove low-quality reads and adapter sequences, ensuring data quality. After downloading the raw data and performing MD5 verification, the following parameters were used in FastP software to filter the raw data reads: `--trim_poly_g --trim_poly_x --length_required 60 --low_complexity_filter --complexity_threshold 20 --correction`. This filtered reads of poly G and poly N, reads with a length <60bp, and reads with a sequence complexity lower than 20, respectively. Base correction was also performed on overlapping portions of paired reads.
[0152] 5.2 Data alignment and preprocessing: bwa-0.7.18 and samtools-1.9.1 were used for alignment (hg19), BAM conversion, and sorting. Then, according to the recommended workflow of GATK best practices, ReorderSam, MarkDuplicates, BaseRecalibrator, and ApplyBQSR were performed using 4.6.1.0. Finally, BAM files within the captured region were obtained through PrintReads for subsequent gene mutation analysis.
[0153] 5.3 Gene Mutation Analysis
[0154] This application uses two commonly used software programs, Mutect2 (same version as GATK, GATK-4.6.1.0) and VarDict-1.8.3, to analyze paired WES samples.
[0155] Specifically, for Mutect2, firstly, the mutation analysis uses the default parameters, and secondly, the Mutect2 mutationfilter uses the default parameters to obtain a vcf file with filter tags. Finally, the first round of filtering of the Mutect2 results is performed by filtering the Mutect2 filter tags using SnpSift software. The specific parameters are as follows: filter -n "(GEN[1].AF <0.05) | (FILTER has 'panel_of_normals'| FILTER has 'normal_artifact' | FILTER has 'contamination' | FILTER has 'germline' | FILTER has 'weak_evidence'|FILTER has 'low_allele_frac' | (FILTER has 'multiallelic' & (FILTER has 'slippage'|FILTER has 'panel_of_normals')))" This yields the results of the first round of filtering of Mutect2, filtering out obvious systematic errors for subsequent integration analysis.
[0156] For VarDict, first, add the following parameters to the mutation analysis: -r 7 -f 0.05 –nosv; then, process the non-standard parts of the original vcf file, and force the icons with the Q0 filtering parameter to be replaced with PASS; then, standardize the vcf file to obtain the first round of VarDict filtered results, which are used for subsequent integration analysis.
[0157] 5.4 Multi-software integration analysis
[0158] First, the CombineVariants tool in GATK3 is used to integrate the first-round filtering results from Mutect2 and VarDict. The results from both software programs are cross-validated to obtain reliable gene mutation information, generating a VCF file containing the second-round filtering results. Cross-validation includes any one of the following three scenarios: 1. Completely identical; 2. Partial overlap; 3. Detected independently, but each result passed the first-round filtering parameters.
[0159] 5.5 Gene Mutation Annotation
[0160] The final filtered and integrated VCF file is annotated, with the main content including functional annotation and filtering annotation. Functional annotation mainly uses SnpEff and Annovar to annotate the protein coding changes of the mutation, with the main principle being to annotate the splice sites 5 bp upstream and downstream of the CDS and exon.
[0161] Filtering annotations include genetic diversity annotations for normal populations, disease database annotations, protein alteration function prediction filtering, and filtering of self-built negative and positive databases. Specifically, the SnpSift annotate tool is used to label the processed database vcf, including the following genetic diversity databases: dbsnp151, 1000Gp3, ESP6500, ExAC, gnomADexomes, UK10K, and ICGCV27.
[0162] The disease databases include cosmic85, clinvar, ICGCV27, and HGMD; the protein alteration function prediction database uses dbNSFPv3.5a; the self-built negative database consists of PON files trained on normal control samples using Mutect1, Mutect2, and VarDict software on two different sequencing platforms (Illumina PE-150 and BGI PE-90) and three different capture kits (Agilent V5, IDTv2, and OncoWESuper). These annotations are used to classify the pathogenicity of variants and provide explanations for important mutations.
[0163] 5.6 Gene mutation filtering and authenticity verification
[0164] Based on the above gene mutation annotation results, gene mutations were screened according to clinically established SOPs. The authenticity of gene mutations was then confirmed again in IGV software using BAM files. Finally, a genotype analysis report was generated, including all detected pathogenic mutations and their clinical relevance. The report can be used by clinicians to diagnose hemoglobinopathies (such as thalassemia and sickle cell anemia) and other hereditary hemoglobinopathies.
[0165] 5.7 Copy Number Variation Analysis
[0166] Firstly, we considered the capture interval distance of the target region or gene of interest and the coverage of heterozygous SNV probes from the initial probe design of the panel. This design facilitates accurate CNV identification by subsequent analysis algorithms. Specifically, we captured and sequenced the genomes of 40 normal individuals, labeled the BAM files with repetitive sequences after BWA alignment, and trained a normal reference genome database for this panel using cnvkit software. Then, we used the cbs algorithm of cnvkit software to perform paired analysis on the BAM files of each clinical sample under the same alignment conditions with the normal controls to obtain standardized coverage information. Next, we used the call tool of cnvkit to analyze the CNV intervals, and then used the software's plotting tools to visualize the intervals or genes of interest. In the HBA and HBB regions, we standardized the plotting coordinates to obtain a genomic image relative to clinical protein electrophoresis, which is convenient for clinical reporting.
[0167] In one example of the above embodiments, the following beneficial effects are achieved:
[0168] Validation of probe effectiveness: high capture efficiency, good specificity, and coverage of all target regions, including known mutation sites.
[0169] In samples from patients with thalassemia, a variety of known mutations were successfully detected, including point mutations, small fragment insertions / deletions, large fragment deletions / duplications, and fragment recombination. Novel mutation types were also detected in some samples.
[0170] Deletion of alpha-thalassemia was detected, including -- SEA Missing, -α 3.7 Missing, -α 4.2 Missing, -- THAI Missing, HKαα missing;
[0171] Point mutations in alpha thalassemia were detected, including common mutations cs, ws, and qs, and rare alpha2 mutations. CD41 point mutation;
[0172] Common point mutations in β-thalassemia were detected, including CD41-42, IVS-II-654, -28, -29, CD71-72, CD27 / 28, CD43, IntM, IVS-I-1, IVS-1-5, CD17, and CD14-15; Chinese β-thalassemia was also detected.
[0173] Rare mutations associated with β-thalassemia were detected, including IVS-II-1 (G→A), CD37, CD72 / 73 (-TG), (SEA-HPFH), c.316_578delinsAAGTAGA, etc.
[0174] Hemoglobin variants such as HbC, HbE, HbS, HbTW, HbG Siriraj, Hb G-Taipei, Hb Q-Thailand, HbGuangzhou-Hangzhou, and Hb Youngstown were detected.
[0175] Different mutations were detected in genes such as SPTB, ANK1, SPTA1, SLC4A1, G6PD, KIF23, SEC23B, CDAN1, EPB41, EPB42, GSR, ADD2, ALDOA, GYB5A, GPI, and NT5C3A.
[0176] LOH phenomenon was detected in βCD22 and CD26.
[0177] It should be noted that the embodiments of the present invention have better practicability and do not impose any form of limitation on the present invention. Any technician familiar with the field may use the technical content disclosed above to change or modify it into an equivalent effective embodiment. However, any modification or equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A probe composition, characterized in that, The nucleic acid sequences of the probe composition are shown in SEQ ID NO. 1 to 64.
2. A kit for detecting genes in hereditary hemoglobinopathies, characterized in that, Includes the probe composition of claim 1.
Citation Information
Patent Citations
Gene detection probe group for newborn inherited metabolic diseases and hemoglobinopathy and application thereof
CN110938685A
Probe set and kit for detecting alpha thalassemia and beta thalassemia related pathogenic genes
CN112359109A