A screening method for target microbial markers and application thereof
By screening for species-specific and intraspecific conserved sequences, optimizing target regions, and designing high-quality probe sets, the problem of low enrichment efficiency in targeted capture sequencing methods has been solved, enabling rapid and accurate detection of nontuberculous mycobacteria.
Patent Information
- Application Number
- CN202311866082.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2043-12-29
AI Technical Summary
Existing targeted capture sequencing methods are inefficient in the enrichment efficiency optimization process, making it difficult to predict the capture efficiency of target regions in the early stages, which leads to increased experimental time and costs. Furthermore, existing methods are not comprehensive enough in assessing the complexity and diversity of genome sequences.
By comparing the genome sequence of the target microorganism with the sequence of the whole species, interspecific and intraspecific conserved sequences are screened out. Combined with splitting and alignment techniques, the screening strategy for target regions is optimized, and high-quality probe sets are designed for capture.
It improves the efficiency and specificity of targeted capture, reduces sequencing costs, minimizes the impact of host background noise, and improves the signal-to-noise ratio of low-load nucleic acid detection, enabling rapid and accurate detection of nontuberculous mycobacteria.
Smart Images

Figure CN117737272B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biological analysis and detection technology, and to a method for screening markers for detecting target microorganisms and its application. Background Technology
[0002] In recent years, with the development of molecular diagnostic technology, next-generation sequencing (NGS) technology has been initially applied in clinical pathogen detection. For example, metagenomic next-generation sequencing (mNGS) can quickly generate test reports within 24-48 hours, but it is expensive, the experimental operation is complex, and the interpretation of the test results is also a major challenge, reducing its clinical practical value. Clinical diagnosis of molecular diagnostics requires a next-generation sequencing detection method that is more sensitive, more specific, has a shorter cycle time, is cheaper, and is easier to interpret than mNGS. Targeted capture sequencing (tNGS) has emerged to meet this need. Depending on the technical approach, tNGS can be divided into methods based on multiplex PCR or probe hybridization capture, which enrich and capture target regions, combine with high-throughput sequencing technology for sequencing, and finally use bioinformatics methods to detect the enriched pathogen sequences in the sample.
[0003] The complexity and diversity of genomic sequence features, including uneven GC content, repetitive sequence characteristics, and palindromic sequence structures, can all affect the enrichment efficiency of target regions. Current strategies typically involve initial screening of target regions, followed by optimization based on experimental results to achieve better product performance. While this strategy can ultimately yield relatively good performance, because it's results-driven optimization, researchers spend a significant amount of time on optimization experiments to screen for target regions with higher enrichment efficiency, leading to inefficiency. If it were possible to predict early on which target regions would result in low probe capture efficiency, researchers could adjust their target sequence screening strategy accordingly, reducing optimization time and significantly improving enrichment efficiency.
[0004] Existing research has shown that GC content is a crucial factor affecting enrichment efficiency. However, due to the complexity and diversity of gene sequences, other unknown factors may still influence this process. A more scientific and reasonable evaluation method is needed to comprehensively assess the impact of target sequence complexity and diversity on enrichment efficiency, thereby better guiding upstream target region screening strategies. This is a technical challenge that needs to be addressed in this field. Summary of the Invention
[0005] To address at least one of the aforementioned technical problems, this disclosure provides a method for screening markers for target microorganisms, wherein the target microorganisms include multiple subtypes, and the method includes the following steps:
[0006] (1) Align the genome sequence of a subtype of the target microorganism with the genome sequence of the entire species to obtain a sequence set with ≥80% sequence identity with the target microorganism and ≤30% sequence identity with other species, which shall be used as the species-specific sequences of the target microorganism; and
[0007] (2) The interspecific sequence is compared with the genome sequence of other subtypes of the target microorganism to obtain a first sequence set with ≥80% sequence identity with the other subtypes, which is used as the intraspecific conserved sequence of the target microorganism.
[0008] In some implementations, step 1) further includes:
[0009] Before comparison, repetitive sequences on the genome sequence of the target microorganism and sequences with ≥75%, ≥76%, ≥77%, ≥78%, ≥79%, ≥80%, ≥81%, ≥82%, ≥83%, ≥84%, ≥85%, ≥86%, ≥87%, ≥88%, ≥89%, ≥90%, ≥91%, ≥92%, ≥93%, ≥94%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99%, ≥99.1%, ≥99.2%, ≥99.3%, ≥99.4%, ≥99.5%, ≥99.6%, ≥99.7%, ≥99.8%, ≥99.9%, or 100% identity with the genome sequence of the host of the target microorganism are filtered out; and / or
[0010] Before alignment, the genome sequence of the target microorganism is split to obtain fragments of 10-200 bp. After alignment, the split fragments are spliced together to form the interspecific sequence.
[0011] Preferably, the splitting is performed with a step size of 1 to 100 bp.
[0012] In some embodiments, the splitting includes splitting the genome sequence of the target microorganism into sequences of lengths of 10bp, 11bp, 12bp, 13bp, 14bp, 15bp, 16bp, 17bp, 18bp, 19bp, 20bp, 21bp, 22bp, 23bp, 24bp, 25bp, 26bp, 27bp, 28bp, 29bp, 30bp, 31bp, 32bp, 33bp, 34bp, 35bp, 36bp, 37bp, 38bp, 39bp, 40bp, 41bp, 42bp, 43bp, 44bp, 45bp, 46bp, 47bp, 48bp, 49bp, 50bp, 51bp, and 52bp. 53bp, 54bp, 55bp, 56bp, 57bp, 58bp, 59bp, 60bp, 61bp, 62bp, 63bp, 64bp, 65bp, 66bp, 67bp, 68bp, 69bp, 70bp, 71bp, 72bp, 73bp, 74bp, 75bp, 76bp, 77bp, 78bp, 79bp, 80bp, 81bp, 82bp, 83bp, 84bp, 85bp, 86bp, 87bp, 88bp, 89bp, 90bp, 91bp, 92bp, 93bp, 94bp, 95bp, 96bp, 97bp, 98bp, 99bp, 100bp, 101bp, 102 bp, 103bp, 104bp, 105bp, 106bp, 107bp, 108bp, 109bp, 110bp, 111bp, 112bp, 113bp, 114bp, 115bp, 116bp, 117bp, 118bp, 119bp, 120bp, 121bp, 122bp, 12 3bp, 124bp, 125bp, 126bp, 127bp, 128bp, 129bp, 130bp, 131bp, 132bp, 133bp, 134bp, 135bp, 136bp, 137bp, 138bp, 139bp, 140bp, 141bp, 142bp, 143bp, 1 44bp, 145bp, 146bp, 147bp, 148bp, 149bp, 150bp, 151bp, 152bp, 153bp, 154bp, 155bp, 156bp, 157bp, 158bp, 159bp, 160bp, 161bp, 162bp, 163bp, 164bp, 165bp, 166bp, 167bp, 168bp, 169bp, 170bp, 171bp, 172bp, 173bp, 174bp, 175bp, 176bp, 177bp, 178bp, 179bp, 180bp, 181bp, 182bp, 183bp, 184bp, 185bp,Fragments of 186bp, 187bp, 188bp, 189bp, 190bp, 191bp, 192bp, 193bp, 194bp, 195bp, 196bp, 197bp, 198bp, 199bp, or 200bp.
[0013] In some implementations, the splitting is preferably performed according to the following values: 1bp, 2bp, 3bp, 4bp, 5bp, 6bp, 7bp, 8bp, 9bp, 10bp, 11bp, 12bp, 13bp, 14bp, 15bp, 16bp, 17bp, 18bp, 19bp, 20bp, 21bp, 22bp, 23bp, 24bp, 25bp, 26bp, 27bp, 28bp, 29bp, 30bp, 31bp, 32bp, 33bp, 34bp, 35bp, 36bp, 37bp, 38bp, 39bp, 40bp, 41bp, 42bp, 43bp, 44bp, 45bp, 46bp, 47bp, 48bp, 49bp, 50bp, 5 Sliding splits are performed with step sizes of 1bp, 52bp, 53bp, 54bp, 55bp, 56bp, 57bp, 58bp, 59bp, 60bp, 61bp, 62bp, 63bp, 64bp, 65bp, 66bp, 67bp, 68bp, 69bp, 70bp, 71bp, 72bp, 73bp, 74bp, 75bp, 76bp, 77bp, 78bp, 79bp, 80bp, 81bp, 82bp, 83bp, 84bp, 85bp, 86bp, 87bp, 88bp, 89bp, 90bp, 91bp, 92bp, 93bp, 94bp, 95bp, 96bp, 97bp, 98bp, 99bp, or 100bp.
[0014] In some implementations, the absolute start and absolute end positions of the split fragments on the reference genome can be marked during the splitting process, and the split fragments can be spliced into the interspecies-specific sequence according to the absolute start and absolute end positions after alignment.
[0015] In some embodiments, step 1) involves comparing the genome sequence of a subtype of the target microorganism with the genome sequence of the entire species to obtain sequences with sequence identity ≥80%, ≥81%, ≥82%, ≥83%, ≥84%, ≥85%, ≥86%, ≥87%, ≥88%, ≥89%, ≥90%, ≥91%, ≥92%, ≥93%, ≥94%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99%, ≥99.1%, ≥99.2%, ≥99.3%, ≥99.4%, ≥99.5%, or ≥99%. A set of sequences comprising 6%, ≥99.7%, ≥99.8%, ≥99.9%, or 100%, and having sequence identity with other species of ≤1%, ≤2%, ≤3%, ≤4%, ≤5%, ≤6%, ≤7%, ≤8%, ≤9%, ≤10%, ≤11%, ≤12%, ≤13%, ≤14%, ≤15%, ≤16%, ≤17%, ≤18%, ≤19%, ≤20%, ≤21%, ≤22%, ≤23%, ≤24%, ≤25%, ≤26%, ≤27%, ≤28%, ≤29%, or ≤30%, shall be designated as the species-specific sequences of the target microorganism.
[0016] In some embodiments, step 2) involves comparing the interspecific sequence with the genomic sequences of other subtypes of the target microorganism to obtain a first set of sequences with sequence identity ≥80%, ≥81%, ≥82%, ≥83%, ≥84%, ≥85%, ≥86%, ≥87%, ≥88%, ≥89%, ≥90%, ≥91%, ≥92%, ≥93%, ≥94%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99%, ≥99.1%, ≥99.2%, ≥99.3%, ≥99.4%, ≥99.5%, ≥99.6%, ≥99.7%, ≥99.8%, ≥99.9%, or 100% with the other subtypes. This first set is then used as the intraspecific conserved sequence of the target microorganism.
[0017] In some implementations, step 2) further includes:
[0018] A second set of sequences is selected, in which the interspecies-specific sequences of one subtype of the target microorganism have an alignment rate ≥80% with the genomic sequences of other subtypes of the target microorganism. The intersection of the first and second sequence sets is taken as the intraspecies-conserved sequences of the target microorganism.
[0019] The alignment rate is the number of sequences in the first sequence set divided by the total number of sequences.
[0020] In some implementations, step 2) further includes:
[0021] A second set of sequences is selected based on the alignment rate of the interspecies-specific sequences of one subtype of the target microorganism with the genomic sequences of other subtypes of the target microorganism, which is ≥80%, ≥81%, ≥82%, ≥83%, ≥84%, ≥85%, ≥86%, ≥87%, ≥88%, ≥89%, ≥90%, ≥91%, ≥92%, ≥93%, ≥94%, ≥95%, ≥96%, ≥97%, ≥98%, ≥99%, ≥99.1%, ≥99.2%, ≥99.3%, ≥99.4%, ≥99.5%, ≥99.6%, ≥99.7%, ≥99.8%, ≥99.9%, or 100%. The intersection of the first and second sequence sets is taken as the intraspecies conserved sequence of the target microorganism.
[0022] The alignment rate is the number of sequences in the first sequence set divided by the total number of sequences.
[0023] In some embodiments, one subtype of the target microorganism, together with other subtypes of the target microorganism, constitutes all subtypes of the target microorganism.
[0024] In some embodiments, the method further includes:
[0025] (3) The interspecific sequence of one subtype of the target microorganism is compared with the genome sequence of other subtypes of the target microorganism to obtain a third sequence set with ≤30% sequence identity with the other subtypes, which is used as the subspecies-specific sequence of the target microorganism.
[0026] In some embodiments, step 3) involves comparing the obtained interspecies-specific sequence of one subtype of the target microorganism with the genomic sequences of other subtypes of the target microorganism to obtain a third set of sequences with sequence identity ≤1%, ≤2%, ≤3%, ≤4%, ≤5%, ≤6%, ≤7%, ≤8%, ≤9%, ≤10%, ≤11%, ≤12%, ≤13%, ≤14%, ≤15%, ≤16%, ≤17%, ≤18%, ≤19%, ≤20%, ≤21%, ≤22%, ≤23%, ≤24%, ≤25%, ≤26%, ≤27%, ≤28%, ≤29%, or ≤30% with the other subtypes. This third set is then used as the subspecies-specific sequence of the target microorganism.
[0027] In some implementations, the number of mismatches in the comparison is ≤5, the gap number is 0, and the expected value (e-value) is ≤10. -5 .
[0028] In some embodiments, the target microorganism is a nontuberculous mycobacterium;
[0029] Preferably, the target microorganism includes one or more of the following: Mycobacterium intracellulare, Mycobacterium avium, Mycobacterium abscessus, Mycobacterium canettii, Mycobacterium interjectum, Mycobacterium kansasii, Mycobacterium marinum, Mycobacterium chelonae, Mycolicibacterium fortuitum, Mycobacterium gordonae, Mycobacterium malmoense, Mycobacterium scrofulaceum, Mycobacterium simiae, and Mycobacterium bufo. xenopi and Mycobacterium ulcerans.
[0030] In some embodiments, the target microorganism is a subtype of one of the three subspecies of Mycobacterium abscessus: Mycobacterium abscessus subsp. Abscessus, Mycobacterium abscessus subsp. Bolletii, and Mycobacterium abscessus subsp. Massiliense.
[0031] In some implementations, the host includes mammals, birds, fish, etc.
[0032] In some implementations, the host includes humans, chickens, ducks, geese, pigs, cattle, rats, horses, dogs, cats, fish, etc.
[0033] In some embodiments, in step 1), the species-specific sequence of the target microorganism is preferably located in a coding region, a multicopy region, or a core gene region, wherein the core gene region is a region of the core gene set (marker gene set) provided by the MetaPhlAn3 database that can be used for species detection.
[0034] In some implementations, the genome sequences of other subtypes mentioned in step 2) or step 3) are derived from strains clinically isolated from patients, humans, or mammals. Relevant information is searched and downloaded from various databases such as RefSeq, Genbank, and PATRIC by species Latin name, and strain genomes for verification are obtained after screening. Taking the search for Mycobacterium abscessus strain genomes in the PATRIC database as an example, the screening principles are: priority is given to domestic infection sources and those related to human pathogenicity, i.e., "Isolation Country" is "China" and "Host Common Name" is "Human"; secondly, genome assembly integrity is considered, i.e., "Genome Status" is preferably "Complete", and "Genome Quality" is "Good". If no usable subtype genomes meeting the conditions are found, those from other infections or other geographical locations are considered. For example, based on the above steps, seven strains can be obtained for verification of Mycobacterium abscessus, namely Mycobacterium abscessus strain G153, Mycobacterium abscessus strain G141, Mycobacterium abscessus strain 199, Mycobacterium abscessus strain G122, Mycobacterium abscessus strain G220, Mycobacteroides abscessus strain GZ002(R)strain GZ002(R)strain GZ01 and Mycobacteroides abscessus strain GZ002-S.
[0035] According to a second aspect of the invention, a marker obtained by screening using the method described in the first aspect is provided.
[0036] In some implementations, the marker is used to detect nontuberculous mycobacteria.
[0037] According to a third aspect of this disclosure, a probe set is provided for capturing the markers described in the second aspect.
[0038] In some embodiments, the probe set is used to detect nontuberculous mycobacteria.
[0039] In some embodiments, the probe set contains any one or more of the nucleotide sequences shown in SEQ ID NO:1 to 180 or any one or more of the nucleotide sequences that have at least 60% sequence identity with them.
[0040] In some embodiments, the probe group comprises one or more of the following:
[0041] A probe set for detecting Mycobacteroides abscessus, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:1-10, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it.
[0042] A probe set for detecting Mycobacterium abscessus subsp. abscessus, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:11-20, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it.
[0043] A probe set for detecting Mycobacterium abscessus subsp. bolletii, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:21-30, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it.
[0044] A probe set for detecting Mycobacterium abscessus subsp. massiliense, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:31-40, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it.
[0045] A probe set for detecting Mycobacterium intracellulare, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:41-50, or any one or more of the nucleotide sequences that have at least 60% sequence identity with them.
[0046] A probe set for detecting Mycobacterium avium, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:51-60, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it;
[0047] A probe set for detecting Mycobacterium canettii, comprising any one or more of the nucleotide sequences shown in SEQ ID NO: 61-70, or any one or more of the nucleotide sequences that have at least 60% sequence identity with them.
[0048] A probe set for detecting Mycobacterium interjectum, comprising any one or more of the nucleotide sequences shown in SEQ ID NO: 71-80, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it.
[0049] A probe set for detecting Mycobacterium kansasii, comprising any one or more of the nucleotide sequences shown in SEQ ID NO: 81-90, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it;
[0050] A probe set for detecting Mycobacterium marinum, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:91-100, or any one or more of the nucleotide sequences having at least 60% sequence identity with it.
[0051] A probe set for detecting Mycobacteroides chelonae, comprising any one or more of the nucleotide sequences shown in SEQ ID NO: 101-110, or any one or more of the nucleotide sequences that have at least 60% sequence identity with them.
[0052] A probe set for detecting Mycolicibacterium fortuitum, comprising any one or more of the nucleotide sequences shown in SEQ ID NO: 111-120, or any one or more of the nucleotide sequences that have at least 60% sequence identity with them.
[0053] A probe set for detecting Mycobacterium gordonae, comprising any one or more of the nucleotide sequences shown in SEQ ID NO: 121-130, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it;
[0054] A probe set for detecting Mycobacterium malmoense, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:131-140, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it.
[0055] A probe set for detecting Mycobacterium scrofulaceum, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:141-150, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it.
[0056] A probe set for detecting Mycobacterium simiae, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:151-160, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it.
[0057] A probe set for detecting Mycobacterium xenopi, comprising any one or more of the nucleotide sequences shown in SEQ ID NO:161-170, or any one or more of the nucleotide sequences that have at least 60% sequence identity with it.
[0058] A probe set for detecting Mycobacterium ulcerans, comprising any one or more of the nucleotide sequences shown in SEQ ID NO: 171-180, or any one or more of the nucleotide sequences having at least 60% sequence identity with them.
[0059] In some embodiments, the probes in the probe set are single-stranded DNA or RNA.
[0060] According to a fourth aspect of this disclosure, a kit is provided that includes the probe set described in the third aspect.
[0061] According to a fifth aspect of this disclosure, a method for detecting nontuberculous mycobacteria using the probe set described in the third aspect or the kit described in the fourth aspect is provided, the method comprising the following steps:
[0062] 1) Extract nucleic acid from the sample.
[0063] 2) Construct a library from the extracted nucleic acids.
[0064] 3) Use the probe to hybridize and capture the target sequence.
[0065] 4) Sequencing and data analysis of the captured products.
[0066] In some implementations, step 1) includes removing the host nucleic acid.
[0067] In some embodiments, step 3) further includes amplifying the captured product.
[0068] In some implementations, the host includes mammals, birds, fish, etc.
[0069] In some implementations, the host includes humans, chickens, ducks, geese, pigs, cattle, rats, horses, dogs, cats, fish, etc.
[0070] In some implementations, the sample includes a natural sample or a standard.
[0071] In some implementations, the sample includes the detection of tissue samples or body fluid samples.
[0072] In some embodiments, the bodily fluid samples include: saliva, whole blood, serum, plasma, milk, urine, lumbar or ventricular CSF, lymph, prostatic fluid, semen, sputum, feces, tears, tumor cells, bronchoalveolar lavage fluid, sputum, pus, nasopharyngeal swabs, oral swabs, cerebrospinal fluid, pleural effusion, peritoneal fluid, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, vaginal discharge, and their processed forms.
[0073] In some embodiments, the tissue sample includes tissue, paraffin sections, and their processed forms.
[0074] In some embodiments, the method is capable of detecting nontuberculous mycobacteria and their subspecies.
[0075] In some embodiments, the method is capable of identifying different nontuberculous mycobacteria and their different subspecies.
[0076] In some embodiments, the nontuberculous mycobacteria include at least one of the following: *Mycobacterium intracellulare*, *Mycobacterium avium*, *Mycobacterium canettii*, *Mycobacterium interjectum*, *Mycobacterium kansasii*, *Mycobacterium marinum*, *Mycobacterium chelonae*, *Mycolicibacterium fortuitum*, *Mycobacterium gordonae*, *Mycobacterium malmoense*, *Mycobacterium scrofulaceum*, *Mycobacterium simiae*, and *Mycobacterium bufo*. Mycobacterium xenopi, Mycobacterium ulcerans, Mycobacterium abscessus and its three subspecies: Mycobacterium abscessus subsp. Abscessus, Mycobacterium abscessus subsp. Bolletii, or Mycobacterium abscessus subsp. Massiliense.
[0077] According to a sixth aspect of this disclosure, a system for detecting nontuberculous mycobacteria is provided, the system comprising:
[0078] (1) Detection module; and
[0079] (2) Judgment module.
[0080] In some embodiments, the system is used to perform the method described in the fifth aspect.
[0081] According to the seventh aspect of this disclosure, the use of the probe set described in the third aspect in the preparation of a kit is provided.
[0082] In some embodiments, the application is used to detect at least one of the following nontuberculous mycobacteria: *Mycobacterium intracellulare*, *Mycobacterium avium*, *Mycobacterium canettii*, *Mycobacterium interjectum*, *Mycobacterium kansasii*, *Mycobacterium marinum*, *Mycobacterium chelonae*, *Mycolicibacterium fortuitum*, *Mycobacterium gordonae*, *Mycobacterium malmoense*, *Mycobacterium scrofulaceum*, *Mycobacterium simiae*, and *Mycobacterium bufo*. Mycobacterium xenopi, Mycobacterium ulcerans, Mycobacterium abscessus and its three subspecies: Mycobacterium abscessus subsp. Abscessus, Mycobacterium abscessus subsp. Bolletii, or Mycobacterium abscessus subsp. Massiliense.
[0083] According to an eighth aspect of this disclosure, a storage medium is provided that records a program for running the screening method of the first aspect, the detection method of the fifth aspect, or the system of the sixth aspect.
[0084] This disclosure provides a method for screening markers for target microorganisms. The screened markers and their detection probe sets can rapidly, accurately, with high sensitivity and high specificity, target and detect non-tuberculous mycobacteria, especially non-tuberculous mycobacteria in blood samples. Compared with existing technologies, this method has stronger pathogen detection targeting, is less affected by human nucleic acid, can effectively reduce host background noise, and has a better signal-to-noise ratio for low-load nucleic acid detection, which is beneficial to improving the detection positivity rate. Moreover, it effectively reduces the amount of data required for sequencing, greatly reducing sequencing costs. Attached Figure Description
[0085] Figure 1 A flowchart illustrating a method for screening markers for target microorganisms is shown. Detailed Implementation
[0086] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention in any way. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of this disclosure. Such structures and techniques have also been described in many publications.
[0087] To obtain a probe set that efficiently enriches pathogenic microorganisms, it is necessary to first screen for high-quality markers (target regions). To address the need for high-quality marker screening, this disclosure presents a unique screening strategy, the specific process of which is as follows: Figure 1 As shown, it includes the following steps:
[0088] 1. Determine the reference sequence
[0089] Select a representative genome or a high-quality, fully assembled genome of the target microorganism from the NCBI RefSeq database as a reference sequence for probe design. High-quality refers to a genome of strain with a quality rating of "Good" or supported by literature, and fully assembled refers to a genome with an assembly completion rate of "Complete Genome".
[0090] 2. Preliminary screening yielded species-specific region sequences.
[0091] After masking repetitive sequences and homology with the human genome, the selected reference genome sequence was segmented using a sliding window to obtain a set of k-mer sequences of length n (n=50) and step size k (k=5), and their absolute start and end positions on the reference genome were marked. These k-mer sequences were then aligned with the NT library using the BLAST method, retaining only the high-quality k-mer sequences that uniquely aligned to the target microorganism. High-quality alignment was defined as sequence identity ≥90%, mismatch number <5, gap number 0, and expected value (e-value) <10. -5 Sequence identity is defined as the percentage of matching length to the total sequence length. High-quality aligned sequences that meet the criteria are spliced together according to the labeled positional information to obtain preliminary species-specific region sequences. These preliminary screening sequences are then analyzed and selected, prioritizing regions located in coding regions, multi-copy regions, or core gene sets available for species detection in the MetaPhlAn3 database.
[0092] 3. Comparison and filtering to obtain intraspecific conserved sequences.
[0093] Furthermore, genomes of 105 clinically isolated strains from patients, humans, or mammals, representing 18 species, were screened and downloaded from NCBI's RefSeq and Genbank databases, as well as the PATRIC database, for strain / subtype coverage validation. The strain / subtype coverage validation included the following steps: BLAST alignment of the initially screened specific region sequences with the genomes of clinically isolated strains was performed, followed by high-quality alignment filtering. The filtering conditions had to meet the following criteria:
[0094] ① High-quality alignment: Sequence identity ≥ 90%, mismatch number < 5, gap number 0, expected value (e-value) < 10 -5 ;
[0095] ② The alignment rate is ≥80%, and the alignment rate (%) = (number of sequences that meet the high-quality alignment requirements / total number of sequences) × 100;
[0096] ③ The subtype coverage is 100%, and the subtype coverage = (number of strain genomes with a comparison rate ≥ 80% / total number of target microbial strain genomes) × 100%.
[0097] 4. Comparison and filtering to obtain subspecies-specific sequences
[0098] The species-specific region sequences obtained from the initial screening are compared with the genome sequences of multiple subspecies of the target microorganism. The high-quality alignment sequence that uniquely aligns to the target subtype is retained as the subspecies-specific sequence of the target microorganism. This sequence is used as a marker for detecting the target subspecies of the target microorganism. A high-quality alignment is defined as sequence identity ≥90%, mismatch number <5, gap number 0, and expected value (e-value) <10. -5 .
[0099] The area obtained after the above filtration is used as a marker for detecting the target microorganism.
[0100] The markers for detecting the target microorganisms are used to design probes. Based on the markers for detecting the target microorganisms, a probe set is conventionally designed using a flip-flop method. The probes are approximately 120 bp in length and have nucleotide sequences as shown in SEQ ID NO:1–180. Biotin-labeled nucleotide probes are also synthesized.
[0101] As one example, the above method is used to screen for markers of nontuberculous mycobacteria, and then use the markers of the nontuberculous mycobacteria to generate a probe set for capturing nontuberculous mycobacteria using a conventional overlay design.
[0102] Non-tuberculosis-mycobacteria (NTM) refers to a large group of mycobacteria excluding Mycobacterium tuberculosis and Mycobacterium leprae. Currently, more than 190 types of NTM have been identified, including more than a dozen opportunistic pathogens such as Mycobacterium intracellulare, Mycobacterium abscessum, Mycobacterium Kansas, and Mycobacterium avium. These pathogens can cause disease in humans, primarily affecting the lungs, but can also impact various organs and systems throughout the body, potentially causing systemic disseminated diseases. NTM are ubiquitous in the environment, and their numbers in China are showing a significant upward trend year by year, making them a potentially significant public health problem threatening human health.
[0103] Nontuberculous mycobacteria (NTMs) are morphologically very similar to Mycobacterium tuberculosis under a microscope, making them difficult to distinguish with the naked eye. Conventional acid-fast staining or culture methods also cannot completely differentiate between tuberculous and nontuberculous mycobacteria. The clinical manifestations of NTM are very similar to those of tuberculosis, easily leading to misdiagnosis. In the absence of bacterial testing results, empirical treatment for suspected tuberculosis is usually initiated first. However, the treatment regimens for the two diseases differ significantly, and most NTMs exhibit natural resistance to anti-tuberculosis drugs, resulting in treatment failure or poor efficacy. Therefore, accurate identification of NTMs is a crucial and challenging issue, urgently requiring rapid and accurate testing methods for detecting nontuberculous mycobacterial infections.
[0104] Currently, detection methods for NTM include bacterial culture, mass spectrometry, polymerase chain reaction (PCR), and next-generation sequencing. Culture methods have long been considered the "gold standard" for detecting pathogenic microorganisms, but they are time-consuming and labor-intensive, generally requiring 3-10 days, and have a low positive rate. They are also susceptible to environmental and colonizing contamination, leading to false positives. Therefore, for infective body fluids (such as pulmonary lavage fluid and sputum), multiple cultures and isolations of non-tuberculous bacilli are necessary. A comprehensive assessment, combining clinical symptoms, signs, and imaging findings, is required for a diagnosis of non-tuberculous mycobacterial disease. Isolation and detection of non-tuberculous bacilli from sterile body fluid samples (such as blood) are of even greater significance for the diagnosis of non-tuberculous mycobacterial disease.
[0105] In specific embodiments, the markers obtained by the above method are used to detect at least one of the following nontuberculous mycobacteria: *Mycobacterium intracellulare*, *Mycobacterium avium*, *Mycobacterium canettii*, *Mycobacterium interjectum*, *Mycobacterium kansasii*, *Mycobacterium marinum*, *Mycobacterium chelonae*, *Mycolicibacterium fortuitum*, *Mycobacterium gordonae*, *Mycobacterium malmoense*, *Mycobacterium scrofulaceum*, *Mycobacterium simiae*, and *Mycobacterium bufo*. Mycobacterium xenopi, Mycobacterium ulcerans, Mycobacterium abscessus and its three subspecies: Mycobacterium abscessus subsp. Abscessus, Mycobacterium abscessus subsp. Bolletii, or Mycobacterium abscessus subsp. Massiliense.
[0106] The sequences of the probe set used to capture nontuberculous mycobacteria are shown in the nucleotide sequences of SEQ ID NO:1 to 180 in Table 1.
[0107] Table 1. Probe sequences for detecting nontuberculous mycobacteria
[0108]
[0109] Table 1 (continued)
[0110]
[0111] Table 1 (continued)
[0112]
[0113] Table 1 (continued)
[0114]
[0115] Table 1 (continued)
[0116]
[0117] Table 1 (continued)
[0118]
[0119] Table 1 (continued)
[0120]
[0121] Table 1 (continued)
[0122]
[0123] Table 1 (continued)
[0124]
[0125] Table 1 (continued)
[0126]
[0127] Table 1 (continued)
[0128]
[0129] Table 1 (continued)
[0130]
[0131] Table 1 (continued)
[0132]
[0133] Table 1 (continued)
[0134]
[0135] In a specific implementation, the method for detecting nontuberculous mycobacteria using the marker or the probe set includes the following steps:
[0136] A certain amount of biological liquid sample (bronchoalveolar lavage fluid, sputum, etc.) is taken, and after removing the host nucleic acid, it is lysed to extract the nucleic acid of the pathogenic microorganism.
[0137] The extracted nucleic acids were used for library preparation, including fragmentation, end repair, and the addition of "A". Adapters were then added to the ends of the nucleic acid fragments, and PCR was used to amplify them, resulting in a pre-library library. The pre-library library, after quantification and fragment quality control, was prepared for hybridization with probes: the probe set used to detect non-tuberculous mycobacteria was hybridized with the quality-controlled pre-library library to capture the target sequence. After eluting non-target sequences and impurities, the captured target sequence was amplified and purified to obtain the final library. After quantification and quality control, the final library was sequenced.
[0138] Bioinformatics tools are used to analyze sequencing data to detect nontuberculous mycobacteria. Bioinformatics analysis involves the following steps: quality control, host sequence removal, alignment with a reference genome, and species detection.
[0139] In some implementations, the species detection includes annotation and filtering.
[0140] This disclosure provides a method for screening markers for target microorganisms. The markers obtained by the method can be used to design probe sets based on these markers, enabling specific capture of various nonmicroorganisms (NTMs). Compared to mNGS technology, this method offers stronger pathogen targeting, is less affected by human nucleic acids, effectively reduces host background noise, and provides a better signal-to-noise ratio for low-load nucleic acid detection, thus improving the positive detection rate. Furthermore, after targeted capture, the sequencing throughput can be reduced to approximately 2M, while mNGS typically requires over 30M, a difference of 15 times or more, significantly reducing sequencing costs.
[0141] definition
[0142] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly used in the field to which this invention pertains. For the purposes of interpreting this specification, the following definitions will apply, and where appropriate, terms used in the singular will also include the plural forms, and vice versa.
[0143] Unless the context clearly indicates otherwise, the terms “a” and “an” as used herein include plural references.
[0144] The term "about" as used herein is as understood by one of ordinary skill in the art and varies within a certain range depending on the context in which it is used. If one of ordinary skill in the art is unfamiliar with the use of this term in the context in which it is used, "about" will mean a particular value plus or minus 10%.
[0145] In this disclosure, the term "non-tuberculous mycobacteria" belongs to the genus Mycobacterium, which is mainly divided into three categories: Mycobacterium tuberculosis complex (MTC), Mycobacterium leprae, and non-tuberculous mycobacteria (NTM). The clinical incidence of NTM (non-tuberculous mycobacteria) infection is constantly rising, with the most common being Mycobacterium avium, Mycobacterium kansas, and Mycobacterium abscessus, accounting for more than 90% of all clinical NTM isolates. The nontuberculous mycobacteria mentioned in this article are specifically understood to include the following strains: Mycobacterium intracellulare, Mycobacterium avium, Mycobacterium abscessus and its three subspecies (Mycobacterium abscessus subsp. Abscessus, Mycobacterium abscessus subsp. Bolletii, Mycobacterium abscessus subsp. massiliense), Mycobacterium canettii, Mycobacterium interferenceum, Mycobacterium kansasii, Mycobacterium marinum, Mycobacterium chelonae, Mycolicibacterium fortuitum, and Mycobacterium gordonii. Mycobacterium gordonae, Mycobacterium malmoense, Mycobacterium scrofulaceum, Mycobacterium simiae, Mycobacterium xenopi, and Mycobacterium ulcerans.
[0146] In this disclosure, the term "RNA," short for "ribonucleic acid," is one of the four biological macromolecules found in living cells: nucleic acids. RNA is a large polymer composed of nucleotides. Nucleotides consist of a base, a sugar, and a phosphate group. There are four types of bases: adenine (A), guanine (G), uracil (C), and cytosine (C).
[0147] In this disclosure, the term "DNA," short for "deoxyribonucleic acid," is one of the four major biological macromolecules found in living cells—a type of nucleic acid. DNA carries the genetic information necessary for the synthesis of RNA and proteins and is an essential biological macromolecule for the development and normal functioning of organisms. DNA is a large polymer composed of deoxynucleotides. Deoxynucleotides consist of a base, a deoxyribose sugar, and a phosphate group. There are four bases: adenine (A), guanine (G), thymine (T), and cytosine (C).
[0148] In this disclosure, the term "probe" refers to an oligonucleotide capable of binding to a target nucleic acid with a complementary sequence via one or more types of chemical bonds (typically through complementary base pairing, typically through hydrogen bonding). Depending on the stringency of the hybridization conditions, a probe may bind to a target sequence that lacks complete complementarity to the probe sequence. There may be any number of base pairs that interfere with hybridization between the target sequence described herein and a single-stranded sequence. However, if the number of mutations is so large that hybridization does not occur even under the least stringent hybridization conditions, the sequence is not a complementary target sequence. Probes may be single-stranded, or partially single-stranded and partially double-stranded. The strandedness of a probe is described by its structure, composition, and the properties of the target sequence. Probes may be directly or indirectly labeled, for example, with biotin subsequently bound to a streptavidin complex. In some embodiments, the term "probe" comprises an oligonucleotide chain, or an oligonucleotide chain complementary to it.
[0149] In this disclosure, the term "probe set" generally refers to a collection of one or more probes that enable the localization and / or quantification of a target nucleic acid by recognizing and binding to a target sequence (through hybridization). Each probe in a probe set is typically an oligonucleotide, such as a single-stranded DNA molecule or RNA. In some embodiments, the term "probe" comprises an oligonucleotide chain or an oligonucleotide chain complementary to said oligonucleotide chain.
[0150] In this disclosure, the term "sequence identity" refers to the "sequence identity percentage" or "sequence identity percentage" between two polynucleotides, which is the number of identical matching positions shared by sequences within a comparison window, taking into account additions or deletions (i.e., vacancies) that must be introduced for optimal alignment of the two sequences. A matching position is any location where the same nucleotide is present in both the target and reference sequences. Since vacancies are not nucleotides, vacancies present in the target sequence are not counted. Similarly, vacancies present in the reference sequence are not counted because nucleotides from the target sequence are counted but nucleotides from the reference sequence are not. At least 60% sequence identity includes at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the total length of the sequence having sequence identity.
[0151] The percentage of sequence identity can be calculated as follows: determine the number of positions in both sequences where the same amino acid residue or nucleic acid base appears (the number of matching positions), divide the number of matching positions by the total number of positions in the comparison window, and multiply the result by 100 to obtain the percentage of sequence identity. Sequence comparison and determination of the percentage of sequence identity between two sequences can be accomplished using software that is readily available online and downloadable. Suitable software programs are available from various sources for protein and nucleotide sequence alignment. A suitable program for determining the percentage of sequence identity is bl2seq, which is part of the BLAST program suite available from the National Center for Biotechnology Information (NCBI) website (blast.ncbi.nlm.nih.gov). Bl2seq uses either the BLASTN or BLASTP algorithm for comparing two sequences. BLASTN is used to compare nucleic acid sequences, while BLASTP is used to compare amino acid sequences. Other suitable programs are, for example, Needle, Stretcher, Water, or Matcher, which are part of the EMBOSS suite of bioinformatics programs and are also available from the European Institute of Bioinformatics (EBI) at www.ebi.ac.uk / Tools / psa.
[0152] In this disclosure, the term "alignment" refers to the process of comparing a read or tag with a reference sequence and thereby determining whether the reference sequence contains a read sequence. If the reference sequence contains a read, that read can be mapped to the reference sequence, or in some embodiments, to a specific position within the reference sequence.
[0153] In this disclosure, the term "hybridization" or "specific hybridization" refers to a molecule that can bind, double-strand, or hybridize only with a specific polynucleotide sequence under strict conditions when that sequence is present in a complex mixture of DNA or RNA (e.g., total cells).
[0154] In this disclosure, the term "complementary" refers to the concept of sequence complementarity between regions of two polynucleotide chains or between two regions of the same polynucleotide chain. It is known that an adenine base in a first region of a polynucleotide can form a specific hydrogen bond ("base pair") with a base in a second polynucleotide region antiparallel to the first region (if that base is thymine or uracil). Similarly, it is known that a cytosine base in the first polynucleotide chain can pair with a base in a second polynucleotide chain antiparallel to the first region (if that base is guanine). If, when the first region of a polynucleotide is arranged antiparallel to a second region of the same or different polynucleotide, at least one nucleotide in the first region can pair with a base in the second region, then the two regions are complementary. Therefore, two complementary polynucleotides do not need to pair at every nucleotide position. "Complementarity" means that the first polynucleotide and the second polynucleotide are 100% or "completely" complementary, and therefore form a base pair at every nucleotide site. "Complementary" also refers to a first polynucleotide that is not 100% complementary (e.g., 90%, 80%, 70%, 60%, or 50% complementary) containing mismatched nucleotides at one or more nucleotide positions. In one embodiment, two complementary polynucleotides are capable of hybridizing with each other under highly stringent hybridization conditions.
[0155] In this disclosure, the term "strict hybridization conditions" refers to conditions under which a probe hybridizes with its target subsequence, typically in a complex mixture of nucleic acids, but not with other sequences. Strict conditions are sequence-dependent and will vary under different conditions. Longer sequences hybridize specifically at higher temperatures. Typically, at a defined ionic strength pH, the chosen strict conditions are about 5–10 °C lower than the thermal melting point (Tm) of the specific sequence. Tm is a temperature (at a specified ionic strength, pH, and nucleic acid concentration) at which 50% of the probe complementary to the target hybridizes with the target sequence at equilibrium (because the target sequence is in excess, at Tm, 50% of the probe is occupied at equilibrium). Strict conditions can also be achieved by adding a destabilizing agent (e.g., formamide). For selective or specific hybridization, the positive signal is at least twice the background, preferably 10 times the background hybridization. Exemplary stringent hybridization conditions may be as follows: incubation at 42°C with 50% formamide, 5×SSC and 1% SDS, or incubation at 65°C with 5×SSC and 1% SDS, followed by washing at 65°C with 0.2×SSC and 0.1% SDS.
[0156] In this disclosure, the term "sequence identity" refers to the percentage of the sequence length in which a matching length is consistent.
[0157] In this disclosure, the term "mismatch number" refers to the number of nucleotides that are misaligned on each sequence.
[0158] In this disclosure, the term "gap number" refers to a missing region of a sequence relative to a reference sequence, that is, a sequence that is missing one or more nucleotides at a certain position.
[0159] In this disclosure, the term "expected value e-value" refers to the expected number of sequences whose scores are greater than or equal to the current alignment score under random conditions in a specific database. Alternatively, the expected value E is the number of sequences whose scores are expected to be greater than or equal to the current alignment score under random conditions in a single database search. A larger e-value indicates that the identity between the query sequence and the retrieved sequence is likely random, while a smaller e-value indicates that the sequence identity may be due to homology (or potential convergent evolution). The e-value is a way of reflecting alignment significance and is widely used to evaluate the reliability of homology between query and target sequences.
[0160] In this disclosure, the term “sequencing” refers to the process of determining the sequence (e.g., identity and order of monomeric units) of a biomolecule, such as a nucleic acid, like DNA or RNA. Exemplary sequencing methods include, but are not limited to, targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, double-strand sequencing, cyclic sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, low denaturing temperature co-amplification PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, synthetic sequencing, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa genome analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, DNA nanosphere sequencing (DNBSEQ), complex probe anchored polymerization sequencing (cPAS), and combinations thereof. In some implementations, sequencing can be performed using a gene analyzer, such as those commercially available from Illumina, Inc., Pacific Biosciences, Inc., Applied Biosystems / Thermo Fisher Scientific, or BGI Genomics Co., Ltd. Examples include BGI's DNBseq sequencing platforms such as BGISEQ-500, BGISEQ-50, MGISEQ-2000, MGISEQ-200, DNBSEQ-T7, DNBSEQ-G99, and DNBSEQ-T20X2, or Illumina's HiSeq2000, HiSeq2500, HiSeq4000, HiSeqX10, and NovaSeq6000.
[0161] In this disclosure, the term "targeted sequencing" refers to a technique that uses biotin-labeled DNA or RNA probes to capture and sequence target fragments in a DNA sample. The probes may be biotin-labeled. The individual nucleotides in the probes of this invention can be synthesized chemically using, for example, a universal DNA synthesizer (e.g., the Applied Biosystems Model 394). Oligonucleotides, such as probes, can also be synthesized using any other methods well known in the art.
[0162] In this disclosure, the terms "computer-readable medium" (e.g., data storage, data storage, etc.) or "computer-readable storage medium" refer to any medium that participates in providing instructions to a processor for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Examples of non-volatile media include, but are not limited to, optical discs, solid-state drives, and magnetic disks, such as storage devices. Examples of volatile media include, but are not limited to, dynamic memory, such as RAM.
[0163] Common forms of computer-readable media include, for example, floppy disks, floppy disks, hard disks, magnetic tapes or any other magnetic media, CD-ROMs, any other optical media, punched cards, paper tapes, any other physical media with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cassette tapes, or any other tangible media from which a computer can read.
[0164] In addition to computer-readable media, data may be provided as signals on a transmission medium included in a communication device or system to provide one or more sequences of instructions to a processor of a computer system for execution. For example, a communication device may include a transceiver having signals indicating instructions and data. The instructions and data are configured to cause one or more processors to perform the functions outlined in this disclosure. Representative examples of data communication transmission connections may include, for example, telephone modem connections, wide area networks (WANs), local area networks (LANs), infrared data connections, NFC connections, etc.
[0165] The following embodiments and accompanying drawings are provided to aid in understanding the present invention. However, it should be understood that these embodiments and drawings are for illustrative purposes only and do not constitute any limitation. The actual scope of protection of the present invention is set forth in the claims. It should be understood that any modifications and changes can be made without departing from the spirit of the present invention.
[0166] Example
[0167] Example 1. Targeted detection of pathogenic microorganisms
[0168] 1. Composition of the capture kit for targeted detection of pathogenic microorganisms
[0169] The probe sets shown in Table 1 (such as the nucleotide sequences shown in SEQ ID NO:1 to 180) are used for one or more detections of NTM.
[0170] 2. The NTM detection method specifically includes the following steps:
[0171] 2.1 Plasma isolation and extraction of cell-free nucleic acids
[0172] Take the test sample, centrifuge at 1600 g for 10 min at 4°C to separate plasma, and take 280 μL of plasma sample to extract nucleic acid using the QIAamp Viral RNA Mini Kit (Qiagen).
[0173] 2.2 Library construction
[0174] Use the Hieff C37P4 OnePot cDNA&gDNA Library Prep Kit (Yeasen) to construct the library, and the library fragments are 150 - 200 bp. After purification with Hieff NGS DNA Selection Beads (Yeasen), the library concentration is quantitatively detected by qubit.
[0175] 2.3 Liquid hybridization capture
[0176] Quantitatively take 25 ng of each library, mix it with 0.1 fmol of probe (containing nucleotide sequences shown in SEQ ID NO: 1 - 180), and use Hybrid Capture Reagents kit (Nanoand) to hybridize the target fragment with the biotin-labeled probe, and anchor the target fragment on the streptavidin magnetic beads through the biotin-streptavidin reaction. After washing, the target library fragment is obtained.
[0177] 2.4 Library amplification
[0178] Amplify the captured product by PCR amplification. The number of cycles is 16. See Table 2 for details:
[0179] Table 2. Library amplification experimental system and procedure
[0180]
[0181] 2.5 Sequencing
[0182] Use the Geneplus G100 sequencer, load the samples at 2M / sample, and perform SE100 sequencing.
[0183] 2.6 Data splitting and bioinformatics analysis
[0184] Use the bioinformatics analysis process to analyze the captured sequencing data, and the detection results of NTM can be obtained.
[0185] Example 2. Comparative experiment between mNGS and hybridization capture using the probe set of the present disclosure
[0186] 1. Method
[0187] 1.1 Preparation of simulated mixed samples
[0188] The simulated pooled samples contained human A549 cells, *Mycobacterium abscessus*, *Mycobacterium kansasii*, and *Mycobacterium avium*. Each simulated pooled sample contained 1.14 × 10⁶ A549 cells. 4 The pathogenic microorganism mixture was diluted three times with sterile buffer (PBS) at a concentration of 1 / ml and then added to A549 cells in equal volumes. The final concentration of pathogenic microorganisms is shown in Table 3.
[0189] Table 3. Pathogenic microorganisms and their concentrations (copy / mL) in simulated mixed samples
[0190]
[0191] 1.2 Pre-capture library preparation
[0192] 280 μL of each of the four gradient simulation samples were taken, and DNA was extracted and purified according to Example 1. Library construction was then performed, and one copy of the library before hybridization capture was retained, which is m. NGS Library.
[0193] 1.3 Hybridization capture and purification
[0194] Another library was quantified and 2000 ng was used for hybridization. The quality control was qualified if the hybridization library concentration was 5-25 ng / μL. The library fragment size was 150-200 bp. The probe used was a probe containing the nucleotide sequence shown in SEQ ID NO:1-180 to obtain the hybridization capture library, which is the tNGS library.
[0195] 1.4 Sequencing
[0196] The mNGS and tNGS libraries were run on the Geneplus G100 platform for NGS sequencing. The mNGS library was sequenced with a minimum of 30M data per sample and a read length of SE50. The tNGS library was sequenced with a minimum of 2M data per sample and a read length of SE100.
[0197] 2. Results
[0198] Bioinformatics analysis was used to analyze the captured sequencing data, and the sequencing data volume and detection status of the samples were statistically analyzed. For mNGS libraries, the number of sequences per million detected species was calculated, i.e., the number of detected microbial sequences normalized to the detection rate per million sequences (RPMCR). For tNGS libraries, the number of sequences per million of the target regions of the detected species was calculated (Target_RPMCR). The results are shown in Tables 4 and 5.
[0199] Table 4. Data volume during sequencing of simulated mixed samples
[0200] Sample number tNGS Library mNGS Library L1 2.754 48.49 L2 2.570 45.73 L3 2.402 33.05 L4 1.791 36.68
[0201] Table 5. Detection results of pathogenic microorganisms in simulated mixed samples
[0202]
[0203] Note: "-" indicates not detected, enrichment factor = tNGS-Target_RPMCR / mNGS-RPMCR.
[0204] The comparison of the results in Tables 4 and 5 shows that: (1) In the detection of simulated samples, the non-adsorbing Mycobacterium tuberculosis was negative, indicating that the capture probe set of the present invention has good specificity (100%) and will not non-specifically capture other pathogenic microorganisms. (2) Compared with the mNGS method, the average enrichment fold of the probe hybridization capture method of the present invention is about 733 times, indicating that the probes conventionally designed based on the screened markers can effectively enrich the target pathogen sequences. (3) From the detection limit results, the probe hybridization capture method of the present invention achieves a lower detection limit than the mNGS method with a lower sequencing data volume. The probe hybridization capture method of this invention achieves a sensitivity as low as 12.5 copies / mL for *Mycobacterium avium*, *Mycobacterium kansasii*, and *Mycobacterium abscessus* at an average data volume of 2M, while the mNGS method, at an average data volume of 40M, only achieves a sensitivity of 100 copies / mL for *Mycobacterium kansasii* and *Mycobacterium abscessus*. In this detection limit experiment, all tests for *Mycobacterium avium* were negative, indicating that the mNGS method has limited performance in detecting *Mycobacterium avium*.
[0205] In summary, the method of the present invention can achieve low-cost, high-sensitivity, and high-specificity detection of NTM.
[0206] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.
Claims
1. A method for screening markers for nontuberculous mycobacteria, characterized in that, Nontuberculous mycobacteria include multiple subtypes, and the method includes the following steps: (1) The genome sequence of a subtype of the nontuberculous mycobacterium was compared with the genome sequence of the whole species, and the sequence sequence was found to have ≥90% sequence identity with the nontuberculous mycobacterium, less than 5 mismatches, 0 gaps, and an expected value of less than 10. -5 The sequence set with ≤30% sequence identity with other species is considered as the species-specific sequence of the nontuberculous mycobacteria; and (2) The interspecific sequence is compared with the genome sequences of other subtypes of the nontuberculous mycobacteria to obtain a first sequence set with ≥90% sequence identity with the other subtypes; A second sequence set is selected based on the alignment rate between the sequence number in the first sequence set and the genome sequence number of other subtypes of the nontuberculous mycobacteria being ≥80%, and the subtype coverage being 100%. The intersection of the first sequence set and the second sequence set is taken as the intraspecific conserved sequence of the nontuberculous mycobacteria. The alignment rate is the number of sequences in the first sequence set divided by the total number of sequences. The total sequence number refers to the number of genome sequences of other subtypes of the nontuberculous mycobacteria; The subtype coverage is defined as (number of strain genomes with a comparison rate ≥ 80% / total number of nontuberculous mycobacterial strain genomes) × 100%.
2. The method according to claim 1, characterized in that, Step (1) further includes: Before comparison, repetitive sequences on the genome sequence of the nontuberculous mycobacteria and sequences with ≥75% genomic identity to the host of the nontuberculous mycobacteria are filtered out; and / or Before alignment, the genome sequence of the nontuberculous mycobacteria is split to obtain fragments of 10-200 bp. After alignment, the split fragments are spliced together to form the interspecies-specific sequence.
3. The method according to claim 2, characterized in that, The splitting is performed in a sliding splitting step size of 1 to 100 bp.
4. The method according to claim 1, characterized in that, The nontuberculous mycobacteria include one or more of the following: intracellular mycobacteria ( Mycobacterium intracellulare ), Mycobacterium avium ( Mycobacteriumavium Mycobacterium abscessus ( ) Mycobacteroides abscessus ), Mycobacterium carnetti ( Mycobacterium canettii ), Mycobacterium medulans ( Mycobacterium interjectum ), Mycobacterium Kansas Mycobacteriumkansasii ), Mycobacterium marinum ( Mycobacterium marinum ), Mycobacterium tectorum ( Mycobacteroideschelonae Occasional mycobacteria ( Mycolicibacterium fortuitum ), Mycobacterium Gordonii ( Mycobacteriumgordonae ), Mycobacterium marmosetum ( Mycobacterium malmoense ), Mycobacterium scrofula ( Mycobacterium scrofulaceum ), Mycobacterium simianum ( Mycobacterium simiae ), Mycobacterium bufossa ( Mycobacteriumxenopi ), Mycobacterium ulcerans ( Mycobacterium ulcerans ).
5. The method according to claim 1, characterized in that, The nontuberculous mycobacteria are subtypes of Mycobacterium abscessus, specifically the three subspecies: Mycobacterium abscessus subsp. abscessus ( Mycobacteroides abscessus subsp . Abscessus Mycobacterium boroughi subsp. abscessus Mycobacteroides abscessus subsp . Bolletii Mycobacterium abscessus subsp. Marseilles ( ) Mycobacteroides abscessus subsp . massiliense ).
Citation Information
Patent Citations
Identification of nucleotide sequences specific for mycobacteria and development of differential diagnosis strategies for mycobacterial species
CA2354197A1
Method and primer for quick detection and classification of mycobacteria
CN105907861A
Method and device for obtaining species-specific consensus sequences of microorganisms and application
CN111477276A
Method and system for screening specific sequences of pathogenic species
CN115719616A
Methods for Simultaneously Detecting Nontuberculous mycobacteria and Kits Using the Same
KR101443716B1