A nucleic acid probe set for identifying mycobacterium abscessus complex mlst typing and use thereof
Patent Information
- Application Number
- CN202610535066.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-22
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-04-22
AI Technical Summary
[0004]然而,现有技术在脓肿分枝杆菌MLST分型的可及性和时效性方面存在严重不足:
本申请首次实现了对MLST分型相关靶标的高效富集与同步检测。其技术效果主要体现在:1)高靶向效率与完整性:靶标探针设计靶向区域仅13.3kb,却实现了对脓肿分枝杆菌复合群亚种鉴定相关关键位点的100%覆盖,避免了全基因组测序的数据冗余。2)高灵敏度与快速检测:检测灵敏度达到10个基因组拷贝,突破了传统方法对长时间纯培养的依赖,将检测周期从数周缩短至2-3天。3)临床与防控价值:能够直接输出亚种、MLST型别及完整的耐药基因型信息,既可精准指导个体化抗感染治疗方案的制定,又能为院内感染溯源调查提供高分辨率的分子分型依据。
Smart Images

Figure CN122071747B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Mycobacterium abscessus gene detection, and more specifically, to a nucleic acid probe set for MLST typing identification of Mycobacterium abscessus complex and its uses. Background Technology
[0002] Mycobacterium abscessus complex (MABC) is a significant pathogen causing nontuberculous mycobacterial lung diseases and other infections. Its infection rate in patients with structural lung diseases such as cystic fibrosis and bronchiectasis is increasing annually, and it can cause clonal transmission within healthcare facilities, creating a persistent cycle of nosocomial infection. Therefore, in addition to accurate subspecies identification and drug resistance testing, molecular typing of strains to clarify their phylogenetic relationships has become a crucial means of infection tracing, outbreak investigation, and transmission interruption.
[0003] Multilocus sequence typing (MLST) is currently recognized as the gold standard for bacterial molecular typing. By analyzing the internal sequence variations of multiple housekeeping genes, strains are classified into different sequence types (STs). It has significant advantages such as data reproducibility and the ability to compare results across laboratories. For the Mycobacterium abscessus complex, a mature MLST typing scheme has been established internationally (typically covering 7 housekeeping genes). This scheme plays an irreplaceable role in distinguishing sporadic from epidemic strains, identifying nosocomial cross-infections, and elucidating community transmission networks.
[0004] However, existing technologies have serious shortcomings in terms of accessibility and timeliness of MLST typing of Mycobacterium abscesses: 1) Conventional detection techniques lack typing capabilities: Traditional microbiological methods (such as culture, biochemical identification, and MALDI-TOF MS) can only complete species identification and cannot provide MLST typing information at all. Although single-gene sequencing (such as rpoB and hsp65) can be used for preliminary intraspecific identification, its resolution is far from sufficient to support the parallel analysis of multiple housekeeping genes required for MLST typing.
[0005] 2) Low clinical accessibility of whole genome sequencing (WGS): WGS can theoretically derive complete MLST typing results, but its implementation heavily relies on 2-4 weeks of pure culture to obtain sufficient bacterial cells. The sequencing and data analysis cycle is as long as 1-2 months, and more than 95% of the massive amount of data generated is typing-irrelevant information. It is costly and has a high analysis threshold, making it difficult to popularize as a routine typing tool in clinical laboratories and disease control networks.
[0006] 3) Lack of genotyping coverage in existing molecular diagnostic products: Targeted capture sequencing (tNGS) products, which have emerged in recent years, primarily focus on pathogen identification and drug resistance gene detection. Their probe designs typically ignore or only sporadically cover MLST genotyping sites. Commercially available respiratory pathogen detection panels (such as Illumina RPIP and BGI PMseq™) list Mycobacterium abscessis as a species-level target, resulting in limited capture areas and a lack of comprehensive MLST genotyping capabilities. While genome-wide enrichment schemes for mycobacteria exist in the literature, their coverage is too broad, cost-inefficient, and not specifically optimized for MLST genotyping, making them insufficient for routine genotyping monitoring needs.
[0007] In summary, there has long been a lack of a technical tool in this field that can quickly, cost-effectively, and specifically perform MLST typing of Mycobacterium abscesses. Summary of the Invention
[0008] In view of this, this application provides a nucleic acid probe set for the identification of Mycobacterium abscessus complex MLST typing and its use.
[0009] In the first aspect, this application proposes a nucleic acid probe set for MLST typing identification of Mycobacterium abscessus complex, the nucleic acid probe set including target probes and high GC balanced probes; The target probe comprises a nucleotide sequence containing 0 to 20 (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20) nucleotides relative to the nucleotide sequence shown in SEQ ID NO. 1 to 188, or any combination thereof, including substitutions, deletions, or insertions. The high GC balanced probe comprises a nucleotide sequence as shown in SEQ ID NO.189 containing 0 to 20 (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20) nucleotide substitutions, deletions, or insertions, or any combination thereof.
[0010] In some specific embodiments, the target probe comprises nucleotide sequences as shown in SEQ ID NO.1 to 188.
[0011] In some specific embodiments, the 5' end of each nucleotide sequence in the target probe is labeled with biotin.
[0012] In some specific embodiments, the high GC balance probe comprises a nucleotide sequence as shown in SEQ ID NO.189.
[0013] In some more specific embodiments, the 5'-end of the nucleotide sequence of the high GC balanced probe is labeled with biotin.
[0014] In some alternative implementations, the nucleic acid probe set may also include an internal control probe.
[0015] In some alternative implementations, the internal control probe targets the RPPH1 gene. The internal control probe is used to detect the human RPPH1 gene to assess sample quality and nucleic acid extraction efficiency.
[0016] In some alternative embodiments, the internal control probe comprises a nucleotide sequence as shown in SEQ ID NO. 190 containing 0 to 20 (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20) nucleotide substitutions, deletions, or insertions, or any combination thereof.
[0017] In some specific embodiments, the internal reference probe comprises a nucleotide sequence as shown in SEQ ID NO.190.
[0018] In some more specific embodiments, the 5' end of the nucleotide sequence of the internal reference probe is labeled with biotin.
[0019] In some specific embodiments, the nucleic acid probe set includes: (I) A target probe comprising nucleotide sequences as shown in SEQ ID NO. 1 to 188; (II) A high GC balanced probe, wherein the high GC balanced probe comprises a nucleotide sequence as shown in SEQ ID NO.189; (III) An internal control probe comprising a nucleotide sequence as shown in SEQ ID NO.190.
[0020] In some more specific embodiments, each nucleotide sequence in the nucleic acid probe set is labeled with biotin at its 5' end.
[0021] Secondly, this application provides the use of the nucleic acid probe set described in the first aspect in the MLST typing identification of Mycobacterium abscessus complex.
[0022] Thirdly, this application provides the use of the nucleic acid probe set described in the first aspect in the preparation of a product for the identification of Mycobacterium abscessus complex MLST typing.
[0023] In some alternative embodiments, the product is selected from diagnostic reagents or diagnostic kits. In some specific embodiments, the product is selected from diagnostic reagents. In some specific embodiments, the product is selected from diagnostic kits.
[0024] Fourthly, this application provides a detection reagent for identifying Mycobacterium abscessus complex MLST typing, the detection reagent comprising the nucleic acid probe set described in the first aspect.
[0025] Fifthly, this application provides a detection kit for identifying Mycobacterium abscessus complex MLST typing, the detection kit comprising the nucleic acid probe set described in the first aspect.
[0026] Sixthly, this application provides a method for identifying the Mycobacterium abscessus complex MLST typing, the method comprising the following steps: S100. Extract genomic DNA from Mycobacterium abscessus in the sample to be tested; S200, Constructing sequencing libraries, including: Fragmentation of genomic DNA; End repair and tailing of DNA fragments; The adapter is ligated to the repaired and tailed DNA fragment; The ligation product was amplified and purified; S300, target capture of sequencing libraries, including: The sequencing library and the nucleic acid probe set described in the first aspect are hybridized to form a probe-target DNA complex; The probe-target DNA complex is bound to a solid support to obtain a solid support immobilized with the probe-target DNA complex. The solid-phase carrier is eluted to remove unbound DNA fragments, and the captured target DNA is amplified. S400, perform high-throughput sequencing on the captured products; S500 performs bioinformatics analysis on high-throughput sequencing data, including: Quality control of high-throughput sequencing data; The quality-controlled sequences were compared with the reference genome; Sequence information of the target region was extracted and MLST typing analysis was performed. Output the sequence analysis results of the target region.
[0027] In some alternative embodiments, in step S100, the method for extracting genomic DNA is selected from magnetic bead extraction or column-based genomic DNA extraction kits.
[0028] In some optional embodiments, in step S100, the quality requirements of the genomic DNA are: concentration ≥5ng / μL, total amount ≥100ng, A260 / A280 ratio between 1.8 and 2.0, and the size of the genomic DNA is mainly distributed in (10-20)kb.
[0029] In some alternative embodiments, in step S200, the size of the DNA fragment is concentrated in (300-500) bp.
[0030] In some optional embodiments, the hybridization reaction in step S300 is a liquid-phase hybridization reaction. In other optional embodiments, the conditions for the liquid-phase hybridization reaction are: denaturation at (90-100)°C (e.g., 90°C, 91°C, 92°C, 93°C, 94°C, 95°C, 96°C, 97°C, 98°C, 99°C, 100°C, or any range of two values) for (3-7) minutes (e.g., 3 minutes, 3.5 minutes, 4 minutes, 4.5 minutes, 5 minutes, 5.5 minutes, 6 minutes, 6.5 minutes, 7 minutes, or any range of two values), followed by rapid cooling to (60-70)°C (e.g., 60°C, 61°C, 62°C, 93°C, 94°C, 95°C, 96°C, 97°C, 98°C, 99°C, 100°C, or any range of two values). Incubate at 2℃, 63℃, 64℃, 65℃, 66℃, 67℃, 68℃, 69℃, 70℃ or any two of these values for 16-24 hours (e.g., 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hours, 24 hours or any two of these values) ...
[0031] In some optional embodiments, in step S300, the solid support is selected from streptavidin-labeled solid supports. In some optional embodiments, the solid support is selected from streptavidin-labeled magnetic beads.
[0032] This application has the following beneficial effects: This application achieves, for the first time, highly efficient enrichment and simultaneous detection of MLST-related targets. Its technical advantages are mainly reflected in: 1) High targeting efficiency and completeness: The target probe design targets a region of only 13.3kb, yet achieves 100% coverage of key sites related to the identification of Mycobacterium abscessus complex subspecies, avoiding data redundancy from whole-genome sequencing. 2) High sensitivity and rapid detection: The detection sensitivity reaches 10 genome copies, breaking through the dependence of traditional methods on long-term pure culture and shortening the detection cycle from several weeks to 2-3 days. 3) Clinical and prevention value: It can directly output subspecies, MLST type, and complete drug resistance genotype information, which can accurately guide the formulation of individualized anti-infective treatment plans and provide high-resolution molecular typing evidence for nosocomial infection tracing investigations. Attached Figure Description
[0033] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application. In the drawings: Figure 1 The sequencing coverage depth repeatability analysis results of the proposed method were evaluated through intra-batch and inter-batch repeatability experiments.
[0034] Figure 2 The consistency analysis results of the key indicators of the method in this application were evaluated through intra-batch and inter-batch repeated experiments. Detailed Implementation
[0035] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0036] General definitions and terminology: In this application, the term "GC content" refers to the percentage of guanine (G) and cytosine (C) residues in a polynucleotide sequence relative to the total number of bases in the sequence. GC content directly affects the melting temperature (Tm) of nucleic acid molecules and their tendency to form secondary structures.
[0037] Unless otherwise stated, all experimental methods involved in this application are conventional methods.
[0038] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses. To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to examples.
[0039] Example 1: Design of target probe: One of the core innovations of this application lies in the completely independently developed "targeted capture probe library for MLST typing identification of Mycobacterium abscessus complex" (Appendix 1, a total of 188 target probes, with a total target region size of 13630bp, and each target probe is labeled with biotin at its 5' end). Its design concept and implementation path are fundamentally different from any existing commercial or publicly available targeted capture NGS panel. 1) 114 high-quality complete / chromosome-level genomes and 105 self-sequencing clinical isolates were selected from the public database (NCBI Genbank), totaling 219 isolates, covering all three subspecies (107 abscess subspecies, 68 Marseille subspecies, and 44 Boleyn subspecies). The complete coding regions of 7 housekeeping genes (argH, cya, gnd, murC, pta, purH, rpoB) were extracted.
[0040] Based on 219 strains of the Mycobacterium abscessus complex, ensuring sufficient coverage of the genetic diversity of the Mycobacterium abscessus complex, the following high-value target regions were extracted one by one using a self-built algorithm through multiple sequence alignment and feature extraction. This is illustrated based on the Mycobacterium abscessus reference genome ATCC 19977 (RefSeq: NC_010397.1): High-resolution MLST classification region: The PubMLST Mycobacterium abscessuscomplext database scheme includes the complete coding regions of seven housekeeping genes (argH, cya, gnd, murC, pta, purH, rpoB) and their upstream and downstream 200 bp regions.
[0041] 2) Dynamic optimization strategy for target probe length: This application addresses the technical challenge of low target DNA capture efficiency caused by the formation of stable secondary structures in high-GC-content regions (average GC% > 64%) of the Mycobacterium abscess complex genome, particularly in some housekeeping gene regions (such as some regions of the argH gene with GC% > 70%). This results in the target DNA sequence being "hidden," preventing the target probe from effectively approaching and binding to the target DNA. The application innovatively proposes a dynamic optimization strategy for target probe length.
[0042] Specifically, this application does not adopt the traditional uniform length probe design, but dynamically adjusts the target probe length according to the local GC content and secondary structure prediction results of each target region: when the GC content of the target region is ≥68%, the target probe length is shortened to (100~105) nt (reducing the self-pairing probability and improving the binding efficiency of high GC regions); when the GC content of the target region is ≤58%, the target probe length is extended to (125~130) nt (increasing binding stability and reducing off-target); when the GC content of the target region is 58% < GC content < 68%, the target probe length remains at the standard length of 120 nt.
[0043] 3) Tiled high-density overlapping design: To ensure that complete spliced sequences can be obtained even in cases where clinical sample DNA is severely fragmented (<150bp) or exhibits significant degradation, this application employs a "tile-like overlapping coverage" strategy for all seven MLST housekeeping genes: The center-to-center spacing between adjacent probes is fixed at (60–80) bp; The overlap area between any two adjacent probes shall be no less than 40 bp (i.e., overlap rate (33-50)%). This design enables an average effective coverage density of 1 probe / 52bp (1 probe / 15bp in critical regions), and still allows for the complete assembly of all seven MLST housekeeping genes in 10ng of severely degraded sputum DNA (average fragment <120bp).
[0044] This strategy ensures that even if some probes fail to capture data, the complete sequence of the target region can still be obtained by sequence splicing through the overlapping areas of adjacent probes, which greatly improves the robustness of detection and data integrity.
[0045] Table 1. Target probes used for MLST typing identification of Mycobacterium abscessus complex:
[0046] This application is the first to introduce the design concept of "minimalist yet complete information" into the field of Mycobacterium abscessus complex diagnosis: the target region is only about 0.3% of the whole genome, but it covers 100% of MLST allele information (based on CARD and NTM-profiler database verification), achieving the optimal solution of "smallest target, largest information content".
[0047] The fundamental difference between this application and existing capture panels is that existing commercial respiratory pathogen capture panels (such as the Illumina Respiratory Pathogen ID / AMR Panel and SureSelect Infectious Disease Panel) only include Mycobacterium abscessus as a species-level target, with a capture region of less than 15 kb, mainly used for "presence detection" and completely lacking MLST typing capabilities; the NTM capture protocols reported in the literature (such as Bryant et al., 2021) are all based on the whole genome enrichment approach, with a capture region >4.5Mb, resulting in serious data redundancy and cost waste; Reference: Bryant JM, Brown KP, Burbaud S, et al. Stepwise pathogenicevolution of Mycobacterium abscessus[J]. Science, 2021, 372(6541): eabb8699. Example 2: Design and optimization of high GC balancing probes: This application addresses the high GC content of the Mycobacterium abscessis genome by designing a high-GC balancing probe. This probe, with a GC content as high as 74%, acts as a high-GC balancing element in the hybridization system, ensuring that the Tm values of target probes with different GC contents tend to be consistent, thus guaranteeing optimal capture efficiency for all target regions under single hybridization conditions.
[0048] Table 2. High GC Balance Probes:
[0049] Example 3: Design and optimization of internal reference probes: In response to the extremely high background (usually >90%) of human host DNA in clinical respiratory samples (such as sputum and bronchoalveolar lavage fluid), this application did not simply remove human DNA, but creatively introduced a "limited capture of human internal standard" strategy: selecting the single-copy conserved housekeeping gene RPPH1 (RNase p RNA component H1) in the human genome (GRCh38.p14, GenBank: GCA_000001405.29) as an internal control probe.
[0050] Table 3. Internal control probes:
[0051] The design of this internal control probe ensures that a small number of stable internal control reads are obtained only when sample DNA extraction is successful and library construction quality is up to standard. This allows the internal control probe to serve as the fundamental criterion for distinguishing between negative samples (no pathogen detected) and experimental failures (extraction failures), without wasting sequencing costs by over-capturing human sequences. For example, when the number of RPPH1 internal control reads is below a preset threshold (e.g., 10), the system will alert that "sample DNA extraction or library construction may have failed"; when the number of RPPH1 internal control reads is normal but the number of reads related to Mycobacterium abscessis complex is zero, it can be reported as "not detected".
[0052] Example 4: Nucleic acid probe set for subspecies identification, MLST typing, and drug resistance analysis of Mycobacterium abscessus complex: This nucleic acid probe set includes: (I) A target probe comprising nucleotide sequences as shown in SEQ ID NO. 1 to 188; (II) A high GC balanced probe, wherein the high GC balanced probe comprises a nucleotide sequence as shown in SEQ ID NO.189; (III) An internal control probe comprising a nucleotide sequence as shown in SEQ ID NO.190; Furthermore, each nucleotide sequence in this nucleic acid probe set is labeled with biotin at its 5' end.
[0053] Example 5: Detection reagents or kits for MLST typing identification of Mycobacterium abscessus complex: In this embodiment, the detection reagents or detection kits include the nucleic acid probe set of Example 4, the genomic DNA extraction kit, the end repair enzyme mix (such as the Illumina kit, NEB kit, KAPA kit, TransGen kit, etc.), the hybridization buffer (such as commercially available ultrahybridization buffers such as Rapid-hyb Buffer (GE), UltraHyb™ (Ambion) or QuickHyb (Stratagene), streptavidin-labeled magnetic beads, washing buffer (1% SDS, 100mM NaHCO3), and low-salt elution buffer (1% TritonX-100, 0.1% SDS, 2mM EDTA (pH 8.0), 150mM NaCl, 20mM Tris-HCl (pH 8.0)).
[0054] Example 6: Method for identifying Mycobacterium abscessus complex using the MLST typing method This method is applicable to a variety of sample types, including but not limited to sputum, bronchoalveolar lavage fluid, pus, and pure cultures of abscess mycobacteria.
[0055] In this embodiment, the method for identifying the Mycobacterium abscessus complex MLST typing specifically includes the following steps: S100. Extract genomic DNA from Mycobacterium abscessus in the sample to be tested: S101. Sample Pretreatment: a) For viscous samples such as sputum, add 1-2 volumes of 4% (wt) sodium hydroxide aqueous solution, vortex to mix, and let stand at room temperature for 15-20 minutes for digestion. For direct clinical samples (such as sputum), before DNA extraction, a step to remove host DNA or enrich pathogens (such as differential centrifugation or selective lysis) can be added as appropriate to increase the proportion of pathogen nucleic acids and further improve detection sensitivity. b) For pure cultures, DNA can be extracted directly from 1-2 loops of bacterial cells.
[0056] S102. DNA Extraction and Quality Control: Genomic DNA extraction should be performed using either a magnetic bead method or a column-based method kit. The extracted genomic DNA must meet the following quality requirements: concentration ≥ 5 ng / μL, total volume ≥ 100 ng, and A260 / A280 ratio between 1.8 and 2.0. DNA integrity should ensure that major fragments are distributed within 10-20 kb, which can be verified by agarose gel electrophoresis.
[0057] S200. Constructing the sequencing library, specifically including the following steps: S201. Use an ultrasonic fragmentation device or enzyme digestion method to fragment the genomic DNA, so that the DNA fragment size is concentrated in 300-500bp.
[0058] S202. End repair and tailing of DNA fragments: End repair enzymes (End Repair EnzymeMix) are used to repair the ends of DNA fragments. The purpose is to convert the sticky ends (overhangs) generated by sonication or enzyme digestion into blunt ends, so that the DNA fragments are in a state of 5'-phosphate and 3'-hydroxyl, and add an "A" base to the 3'-end to facilitate ligation with adapters.
[0059] S203. Ligate the Y-type adapter with molecular index to the repaired and tailed DNA fragment.
[0060] S204. Amplify and purify the ligation product; perform PCR amplification of the ligation product using universal primers with indexed sequences (usually 6-8 cycles), and then purify the amplification product using magnetic beads to remove residual primer dimers and byproducts.
[0061] S300. Target capture of the sequencing library, specifically including the following steps: S301. Mix the purified sequencing library, nucleic acid probe set, and hybridization buffer to form a hybridization system. Place the hybridization system in a PCR instrument or hybridization oven, denature at 95°C for 5 minutes, then rapidly reduce to 65°C and incubate at this temperature for 16-24 hours to allow the biotin-labeled probe to fully hybridize with the complementary target DNA in the sequencing library, thus obtaining the probe-target DNA complex.
[0062] S302. Add streptavidin-labeled magnetic beads to the hybridization system and incubate at room temperature for 30 minutes to allow the probe-target DNA complex to bind to the magnetic beads. Discard the supernatant under an external magnetic field. Then wash the magnetic beads immobilized with the probe-target DNA complex 2-3 times with washing buffer (1% SDS, 100mM NaHCO3) at 65°C to thoroughly remove unhybridized and non-specifically bound DNA fragments.
[0063] S303. Amplify the captured target DNA; use low-salt elution buffer (1% Triton X-100, 0.1% SDS, 2mM EDTA (pH 8.0), 150mM NaCl, 20mM Tris-HCl (pH 8.0)) or sterile water, and treat at 95°C for 10 minutes to elute the captured target DNA from the magnetic beads; then perform 12-15 cycles of PCR amplification on the elution product to enrich the target region DNA.
[0064] S400. Perform high-throughput sequencing on the captured products, specifically including the following steps: S401. Library quality control and quantification: The final captured and amplified library is accurately quantified using a fluorometer or Qubit, and the fragment size distribution of the library is detected using an Agilent Bioanalyzer 2100 or similar device.
[0065] S402. Sequencing: Mix the qualified libraries in an appropriate ratio and perform paired-end sequencing on an Illumina NovaSeq 6000, MGIDNBSEQ-T7, or equivalent next-generation sequencing platform. It is recommended that each sample yield at least 500 Mb of data to ensure an average sequencing depth of over 1000X in the target region.
[0066] S500, performing bioinformatics analysis on high-throughput sequencing data, specifically including the following steps: S501. Quality control of high-throughput sequencing data: Use FastQC software to assess the quality of the raw high-throughput sequencing data after sequencing, and use Fastp software to remove adapter sequences and low-quality bases.
[0067] S502. Align the quality-controlled sequences with the reference genome: Use BWA-MEM software to align the high-quality quality-controlled sequences with the Mycobacterium abscessus reference genome to generate a sequence alignment file.
[0068] S503. Extract sequence information from the target region and perform MLST typing analysis: The full-length sequences of seven housekeeping genes—argH, cya, gnd, murC, pta, purH, and rpoB—were extracted from the sequencing data and compared with the PubMLST database to determine the strain's sequence type. New alleles or sequence types were automatically labeled and recorded by the system.
[0069] S504. Output the sequence analysis results of the target region: Automatically generate a comprehensive clinical report containing MLST classification results. The report is output in PDF format for clinicians to review.
[0070] Through the complete and coherent technical process described above, this application provides a comprehensive, accurate, and rapid integrated analysis of Mycobacterium abscessus complex MLST typing, from sample analysis to clinical report analysis, which greatly improves the efficiency and level of clinical diagnosis and infection control.
[0071] Example 7, Performance Verification: To comprehensively and intuitively demonstrate the technical performance of this application, this application systematically showcases the superior performance of the nucleic acid probe kit in four key dimensions: clinical accuracy, direct sample detection capability, and method stability.
[0072] (1) Clinical validation of accuracy and specificity: Sample set: 178 clinical isolates (identified as Mycobacterium abscessus complex by traditional culture methods) were subjected to blinded testing.
[0073] Compared with the gold standard of whole-genome sequencing, the method in this application achieved 98.31% (175 / 178) consistency in MLST genotyping. Among them, 3 inconsistent samples were found to be new alleles after review (1 new argH allele and 2 new pta alleles). The system automatically labeled and recorded them, which can be directly used for subsequent submission to the PubMLST database.
[0074] Success rate of full-length sequence assembly of 7 housekeeping genes (argH, cya, gnd, murC, pta, purH, rpoB): 100%, average coverage depth: >1050×.
[0075] The accuracy of MLST typing exceeds the clinically acceptable threshold of 95%, demonstrating the reliability of the results of this application.
[0076] (2) Performance of direct detection of clinical samples: The traditional culture method and the method of this application were used to test 56 sputum samples with positive smears. The results are shown in Table 4.
[0077] Table 4. Comparison of detection results between the method of this application and the traditional culture method:
[0078] Clinical sensitivity: 96.9% (31 / 32); Clinical specificity: 91.7% (22 / 24); Overall compliance rate: 94.6% (53 / 56).
[0079] The only missed sample was found to have extremely low DNA concentration (<0.1 ng / μL) and high degradation, indicating that there are limitations to the detection of extremely small amounts of degraded samples using this method, but it is still significantly better than the detection capability of traditional culture.
[0080] The only two false positive samples were confirmed to be mixed infections upon further analysis. The low abundance target sequences detected by this method were masked by the dominant bacteria during culture, demonstrating the detection advantage of this method in samples with complex bacterial communities.
[0081] Confusion matrix analysis showed that the method in this application was in high agreement with the gold standard culture results (Kappa coefficient > 0.9), confirming its feasibility and accuracy in directly detecting clinical samples without culture.
[0082] (3) Repeatability and reproducibility: The stability of the method in this application was evaluated through repeated experiments within batches (same operator, same batch) and between batches (different operators, two-week interval).
[0083] a) Coverage depth stability Figure 1 The intra-batch coefficient of variation (CV) for the average coverage depth of the key target regions (7 housekeeping genes) was 8.4%, and the inter-batch CV was 18.3%.
[0084] b) Detection consistency ( Figure 2 ST type interpretation intra-batch / inter-batch consistency: 100% (all comparable samples).
[0085] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A nucleic acid probe set for MLST typing identification of Mycobacterium abscessus complex, characterized in that, The nucleic acid probe set includes target probes and high-GC balanced probes; The nucleotide sequences of the target probe are shown in SEQ ID NO.1 to 188; The nucleotide sequence of the high GC balanced probe is shown in SEQ ID NO.
189.
2. The nucleic acid probe kit according to claim 1, characterized in that, Each nucleotide sequence in the target probe is labeled with biotin at its 5' end.
3. The nucleic acid probe kit according to claim 1, characterized in that, The nucleotide sequence of the high GC balanced probe is labeled with biotin at the 5' end.
4. The nucleic acid probe kit according to claim 1, characterized in that, The nucleic acid probe set also includes an internal reference probe; the internal reference probe targets the RPPH1 gene.
5. The nucleic acid probe kit according to claim 4, characterized in that, The nucleotide sequence of the internal control probe is shown in SEQ ID NO.
190.
6. The nucleic acid probe kit according to claim 5, characterized in that, The nucleotide sequence of the internal control probe is labeled with biotin at the 5' end.
7. Use of the nucleic acid probe set according to any one of claims 1 to 6 in the preparation of products for identification of Mycobacterium abscessus complex MLST typing.
8. A detection reagent for identifying Mycobacterium abscessus complex using the MLST typing method, characterized in that, The detection reagent includes the nucleic acid probe set as described in any one of claims 1 to 6.
9. A detection kit for identifying Mycobacterium abscessus complex MLST typing, characterized in that, The test kit includes the nucleic acid probe set as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Digital PCR method and detection kit for detecting HER2 copy number variation in breast cancer
CN110616251A