Reagent and kit for species identification and wide drug resistance spectrum detection of mycobacterium tuberculosis based on single molecule sequencing method and application of reagent and kit

Through the single-molecule sequencing method of multiple primer combinations, the problem of rapid and low-cost identification of Mycobacterium tuberculosis species and drug resistance detection in tuberculosis diagnosis has been solved, achieving high sensitivity and high specificity detection effects, and guiding clinical drug use.

CN120608162APending Publication Date: 2025-09-09BEIJING HUADA BIO & INFORMATION FUSION TECHNOLOGY RESEARCH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510524948.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies lack rapid, low-cost, and highly sensitive methods for Mycobacterium tuberculosis species identification and drug resistance detection in tuberculosis diagnosis, resulting in complex, costly, and low success rates in treatment.

Method used

The single-molecule sequencing method using multiple primer combinations is used to amplify the target region related to Mycobacterium tuberculosis through the primer set, and combined with single-molecule sequencing and data analysis to achieve rapid and accurate species identification and drug resistance detection.

Benefits of technology

It achieves high-sensitivity, high-specificity, and low-cost identification of Mycobacterium tuberculosis species and drug resistance detection, which can quickly and comprehensively guide clinical drug use. The detection limit is as low as 100 copies/μL, shortening detection time and reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120608162A_ABST
    Figure CN120608162A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of molecular biological detection, and particularly provides a reagent for species identification and drug resistance detection of mycobacterium tuberculosis, which comprises: a set of a plurality of primers, each primer comprises a first sequence, the first sequences of the plurality of primers are respectively shown as SEQ ID NO: 1-56, and the first sequences of the plurality of primers are respectively shown as SEQ ID NO: 1-SEQ ID NO: 2-SEQ ID NO: 3-SEQ ID NO: 4-SEQ ID NO: 5-SEQ ID NO: 6; wherein species identification and drug resistance detection of mycobacterium tuberculosis are both carried out in the same primer pool containing the set, the set comprises a first primer subset, the first primer subset comprises primers with first sequences as shown in SEQ ID NO: 1-26, and the second primer subset comprises primers with second sequences as shown in SEQ ID NO: 2-26; and a second subset of primers, the second subset of primers comprising the primers having the first sequences as shown in SEQ ID NO: 27-56, respectively. According to the primer combination, the detection method based on the primer combination and the like, low-cost, high-sensitivity, high-specificity and high-throughput detection is achieved, meanwhile, the detection period is short, the accuracy rate is high, drug-resistant gene detection is comprehensive, all domestic clinical drug-related drug-resistant genes are covered, and clinical medication can be effectively and comprehensively guided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of molecular biological detection technology, and specifically to a reagent, a kit and its application for species identification of Mycobacterium tuberculosis and detection of its extensive drug resistance spectrum based on single-molecule sequencing. Background Art

[0002] Tuberculosis (TB), a major global public health issue, is particularly severe in developing countries. The emergence of drug-resistant TB complicates treatment, with treatment cycles lasting 18 to 20 months, significant drug side effects, high costs, and relatively low success rates. Accurately identifying the TB pathogen and detecting its drug resistance is crucial for curing TB.

[0003] At present, the diagnosis, treatment, monitoring and prevention of tuberculosis all rely on the identification of pathogenic species and multiple drug sensitivity tests. The commonly used methods in clinical practice mainly include pathogenic culture of Mycobacterium tuberculosis, acid-fast staining smear microscopy, molecular biological methods, immunological techniques and imaging techniques. In summary, pathogenic culture is the gold standard for the diagnosis of tuberculosis, but its culture cycle is as long as 2-8 weeks, which is not suitable for providing patients with rapid diagnostic results and may even delay the optimal treatment time. General laboratories are unable to carry out Mycobacterium tuberculosis culture due to their low biosafety level. Therefore, the feasible scenarios and efficiency of pathogenic culture are very limited. In addition, although the detection of Mycobacterium tuberculosis in sputum or other body fluids under a microscope is a faster and easier method, the detection sensitivity of this method is poor and it cannot predict the drug resistance of the strain. Molecular biological diagnostic methods, such as nucleic acid amplification detection method ( MTB / RIF can directly detect Mycobacterium tuberculosis and its resistance to rifampicin in sputum within 2 hours. However, the number of targets detected is relatively limited, making it difficult to comprehensively analyze the drug resistance of Mycobacterium tuberculosis. Immunological techniques, such as the interferon-gamma release assay (IGRA), can measure the release of interferon-gamma specific to Mycobacterium tuberculosis. However, these techniques are expensive and can cross-react with non-tuberculous mycobacteria, affecting test accuracy.

[0004] Therefore, there is an urgent need to provide a highly sensitive and accurate Mycobacterium tuberculosis species identification and drug resistance detection reagent and method that is fast in detection, covers a wide range of drug-resistant mutations, is simple, easy to use, and low in cost. Summary of the Invention

[0005] The first embodiment of the present application provides a reagent for Mycobacterium tuberculosis species identification and drug resistance detection, comprising: a set of multiple primers, each primer comprising a first sequence, the first sequences of the multiple primers being respectively shown as SEQ ID NOs: 1-56, wherein the Mycobacterium tuberculosis species identification and drug resistance detection are both performed in the same primer pool comprising the set.

[0006] In some embodiments, the set includes: a first primer subset, the first primer subset comprising the primers having the first sequences shown as SEQ ID NOs: 1-26, respectively; and a second primer subset, the second primer subset comprising the primers having the first sequences shown as SEQ ID NOs: 27-56, respectively.

[0007] In some embodiments, each of the primers further comprises a second sequence comprising a tag sequence for distinguishing the source of the sample, wherein the tag sequence is optionally located 5' to the first sequence. In some embodiments, the second sequence further comprises one or more universal sequences.

[0008] In some embodiments, the reagent is a kit. In some embodiments, the reagent further comprises one or more of the following: DNA polymerase, dNTP, reaction buffer and Mg 2+ In some embodiments, the working concentrations of the first primer subset and the second primer subset are each 10 μM, wherein the first primer subset comprises equal concentrations of the primers having the first sequences shown in SEQ ID NOs: 1-26, and the second primer subset comprises equal concentrations of the primers having the first sequences shown in SEQ ID NOs: 27-56.

[0009] In some embodiments, the first subset of primers is used to amplify a first target region, which includes a region in the following genes: rplC, ddn, gyrA, rpsL, ahpC, pncA, RV0678, rpoB, rrs, ethA, katG and embB; and the second subset of primers is used to amplify a second target region, which includes a region in the following genes: eis, hsp65, fabG1, gyrB, inhA, rrl, gidB, tlyA, fbiA, fgd1, rpoB, rrs, ethA, katG and embB.

[0010] In some embodiments, the length of the first target region and the second target region amplified by the first primer subset and the second primer subset, respectively, is 500 to 1500 bp.

[0011] The present invention also provides a method for amplifying a target region related to the identification and drug resistance detection of Mycobacterium tuberculosis species, comprising: combining a set of primers as defined in any of the above embodiments with a nucleic acid sample to be tested, a DNA polymerase, dNTPs, and an optional reaction buffer and Mg. 2+ Mix to form a first reaction solution; or the first primer subset and the second primer subset as defined in any of the above embodiments are respectively mixed with the nucleic acid sample to be tested, DNA polymerase, dNTP and optional reaction buffer and Mg 2+ mixing to form a second reaction solution and a third reaction solution; and placing i. the second reaction solution and the third reaction solution or ii. the first reaction solution in a thermal cycle program to obtain an amplicon of the target region.

[0012] In some embodiments, the working concentrations of the primer set, the first primer subset, and the second primer subset are each 10 μM, wherein the primer set is obtained by mixing the primers having the first sequence as shown in SEQ ID NO: 1-56 in equal concentrations; the first primer subset is obtained by mixing the primers having the first sequence as shown in SEQ ID NO: 1-26 in equal concentrations, and the second primer subset is obtained by mixing the primers having the first sequence as shown in SEQ ID NO: 27-56 in equal concentrations.

[0013] In some embodiments, the amount of the nucleic acid sample to be tested is 0.1-300 ng, preferably 10-250 ng.

[0014] In some embodiments, the detection limit of Mycobacterium tuberculosis in the nucleic acid sample to be tested is 100 copies / μL.

[0015] In some embodiments, the thermal cycling program includes: a pre-denaturation stage: 98°C, 2-5 min; a cycling stage: 98°C, 15 s; 62°C, 30 s; 58°C, 30 s, 72°C, 1 min 30 s, 25-40 cycles, preferably 30-35 cycles; an extension stage: 72°C, 5-10 min; and an optional storage stage: 4°C-10°C.

[0016] The present application also provides a method for constructing a library for a nucleic acid sample, wherein the library is used for identifying the species of Mycobacterium tuberculosis and detecting drug resistance in the nucleic acid sample by single-molecule sequencing. The method comprises: amplifying the target region of the nucleic acid sample according to the method described in any of the above embodiments to obtain amplicons of the target region; and introducing a single-molecule sequencing adapter into one side of each amplicon to obtain the library.

[0017] In some embodiments, before introducing the single-molecule sequencing adapter, the method further comprises: performing end repair and optional tag sequence ligation and purification on the amplicon.

[0018] An embodiment of the present application also provides a method for constructing a library for multiple nucleic acid samples, wherein the library is used to perform batch species identification and drug resistance detection of Mycobacterium tuberculosis in the multiple nucleic acid samples through single-molecule sequencing. The method comprises: amplifying the target region of the nucleic acid sample according to the method defined in any of the above embodiments to obtain amplicons of the target region; and introducing a single-molecule sequencing adapter into one side of each amplicon to obtain the library, wherein the amplification is performed using a set of primers or a first primer subset and a second primer subset, wherein the set of primers comprises multiple primers, each primer comprising a first sequence, and the first sequences of the multiple primers are respectively shown in SEQ ID NOs: 1-56; the first primer subset comprises the primers having the first sequences respectively shown in SEQ ID NOs: 1-26; the second primer subset comprises the primers having the first sequences respectively shown in SEQ ID NOs: 27-56; wherein each primer further comprises a second sequence, and the second sequence comprises a tag sequence for distinguishing the source of the sample.

[0019] In some embodiments, before introducing the single-molecule sequencing adapter, the method further comprises: performing end repair and optional purification on the amplicon.

[0020] The embodiments of the present application also provide a library construction kit for performing species identification and drug resistance detection of Mycobacterium tuberculosis in a nucleic acid sample by single-molecule sequencing. The kit comprises: the reagents described in any of the above embodiments; and a single-molecule sequencing adapter, which optionally comprises a motor protein.

[0021] In some embodiments, the kit further comprises one or more of the following: end repair reagents, the single molecule sequencing adapter ligation reagents, purification reagents, and an optional tag sequence.

[0022] The present application also provides a method for species identification and drug resistance detection of Mycobacterium tuberculosis, comprising: constructing a library for single-molecule sequencing according to the method described in any of the above embodiments, wherein the library comprises amplicons of a target region of a nucleic acid sample; performing single-molecule sequencing on the library to obtain sequencing data of the amplicons; and parsing the sequencing data to perform species identification and drug resistance detection on Mycobacterium tuberculosis in the nucleic acid sample.

[0023] In some embodiments, parsing the sequencing data to perform species identification and drug resistance detection on Mycobacterium tuberculosis in the nucleic acid sample specifically includes: performing a first alignment of the sequencing data with a reference genome of Mycobacterium tuberculosis or with an hsp65 reference gene sequence of Mycobacterium tuberculosis, and reporting that the nucleic acid sample contains Mycobacterium tuberculosis based on the presence of a sequence matching the hsp65 reference gene in the sequencing data; performing a second alignment of the sequencing data with the reference genome of Mycobacterium tuberculosis or with reference target sequences corresponding to the remaining target regions to obtain a first alignment sequence aligning the sequencing data to each target region; based on the first alignment sequence, detecting variations in the first alignment sequence within the target region; and determining the drug resistance spectrum of Mycobacterium tuberculosis in the nucleic acid sample based on the number and variation of the first alignment sequences within each target region, wherein the Mycobacterium tuberculosis in the nucleic acid sample is confirmed to have drug resistance corresponding to the gene in the target region based on the number of the first alignment sequences within the target region being ≥100 and the variation frequency within the target region being ≥0.5.

[0024] In some embodiments, the method further comprises: confirming that Mycobacterium tuberculosis in the nucleic acid sample has drug resistance corresponding to the gene in the target region based on the number of first aligned sequences of the target site in the target region being ≥100 and the mutation frequency of the target site being ≥0.5, wherein the mutation frequency optionally includes a point mutation frequency.

[0025] In some embodiments, the method further includes: based on the number of the first aligned sequences of the target site in the target region being ≥100 and the mutation frequency of the target site being ≥0.5, using TB-Profiler to determine the drug resistance of Mycobacterium tuberculosis in the nucleic acid sample corresponding to the gene in the target region, wherein based on the TB-Profiler report that the Mycobacterium tuberculosis in the nucleic acid sample is highly correlated with the drug resistance corresponding to the gene in the target region, it is determined that the Mycobacterium tuberculosis in the nucleic acid sample has the drug resistance corresponding to the gene in the target region.

[0026] In some embodiments, before performing species identification and drug resistance testing on Mycobacterium tuberculosis in the nucleic acid sample, the method further includes: performing quality control filtering on the original sequencing data to obtain filtered sequencing data; and performing a third alignment of the filtered sequencing data with a reference genome of Mycobacterium tuberculosis or with reference gene sequences of all target regions of Mycobacterium tuberculosis. Based on the fact that the amount of alignment data in the filtered sequencing data that is aligned to the target region is greater than 7% of the total amount of the filtered sequencing data, the filtered sequencing data is determined to be a qualified sample, and downstream species identification and drug resistance testing are performed.

[0027] The present application also provides a method for species identification and drug resistance detection of Mycobacterium tuberculosis, wherein the method is performed by using the reagent or kit described in any of the above embodiments.

[0028] The present application also provides an example of using the reagent or kit described in any of the above examples in the identification of Mycobacterium tuberculosis species and the detection of drug resistance.

[0029] The present application also provides an integrated system for Mycobacterium tuberculosis species identification and drug resistance detection, the system comprising: i. an amplification-pre-built library module, for amplifying target regions related to Mycobacterium tuberculosis species identification and drug resistance detection in a nucleic acid sample to be tested, and obtaining amplicons of each target region, wherein the amplification is performed using a first primer set and a second primer set, wherein the amplification is performed using a set of primers or a first primer subset and a second primer subset, wherein the set of primers comprises a plurality of primers, each primer comprising a first sequence, and the first sequences of the plurality of primers are respectively as shown in SEQ ID NOs: 1-56; the first primer subset comprises the primers having the first sequences respectively as shown in SEQ ID NOs: 1-26; the second primer subset comprises the primers having the first sequences respectively as shown in SEQ ID NOs: NO: The primers of the first sequence shown in 27-56; wherein each of the primers further comprises a second sequence, and the second sequence comprises a tag sequence for distinguishing the source of the sample; ii. a single-molecule sequencing adapter introduction module, used to introduce a single-molecule sequencing adapter into one side of each amplicon to obtain a single-molecule sequencing library of the nucleic acid sample to be tested; iii. a single-molecule sequencing module, used to perform single-molecule sequencing on the single-molecule sequencing library to obtain sequencing data of each amplicon; and iv. a data analysis module, used to analyze the sequencing data to perform species identification and drug resistance detection on Mycobacterium tuberculosis in the nucleic acid sample.

[0030] The technical solution of this application achieves the following technical effects:

[0031] 1. The embodiment of the present application directly performs multi-target primer combination amplification on clinical samples to obtain long amplicons of target fragments that are adaptable to single-molecule long-read sequencers. The primer combination used can simultaneously detect one species identification gene and 21 drug resistance-related genes. For the first time, it covers all domestic clinical drug-related drug resistance genes, including drug resistance genes corresponding to new anti-tuberculosis drugs. Based on single-molecule sequencing, it can achieve rapid, accurate, and high-precision identification of Mycobacterium tuberculosis species and one-time detection of its comprehensive drug resistance spectrum. The detection limit is as low as 100 copies / μL, showing high sensitivity and high specificity (100%), which can effectively and comprehensively guide clinical drug use.

[0032] 2. The primer combinations proposed in the examples of this application have a streamlined number of primers that work in concert with each other. Compared to the primer designs of traditional second-generation sequencing platforms, they exhibit no cross-reactions and extremely low mismatch rates when conducting comprehensive drug resistance spectrum detection, demonstrating relatively balanced amplification efficiency and extremely high detection efficiency, especially in terms of sensitivity and specificity. At the same time, the special design of the multiple primer combinations also greatly reduces the demand for templates and detection reagents, and can save approximately 2 hours of tag sequence connection time in traditional library construction, thereby significantly reducing detection time while reducing detection costs.

[0033] 3. The multiplex PCR detection method proposed in the present application also enables simultaneous testing of multiple samples. By combining this detection method with drug resistance analysis software, an integrated "detection-analysis-reporting" system can be formed. This integrated system enables high-throughput, rapid "sample-to-report" testing within hours.

[0034] 4. The embodiments of the present application can also be equipped with technical analysis software designed specifically for single-molecule sequencing data of each amplicon. For clinical samples, generally only 300Mb of sequencing data is required to achieve rapid and accurate prediction of Mycobacterium tuberculosis resistance through such technical analysis software. For example, an ordinary laptop equipped with 2 CPU cores and 8GB of memory can complete the entire process analysis within one hour, thereby greatly reducing the data volume required in traditional drug resistance analysis and reducing the demand for computing power, making detection faster, more convenient and economical. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0036] Figure 1 The present invention provides an analytical process for species identification and drug resistance detection of Mycobacterium tuberculosis in a nucleic acid sample according to an embodiment of the present application.

[0037] Figure 2 Schematic diagram of the primer structure according to the embodiment of the present application;

[0038] Figure 3 This is the increase in the minimum coverage depth of the target in the single-molecule sequencing data within 30 minutes according to Example 1 of the present application. DETAILED DESCRIPTION

[0039] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.

[0040] This application creatively designs primer combinations for the identification of Mycobacterium tuberculosis species and the detection of drug-resistant gene mutations, multi-sample efficient library construction methods, and sequencing data analysis processes. By performing super-multiplex PCR amplification on the target gene, constructing a single-molecule length long sequencing library for the amplicon, and performing single-molecule sequencing and data analysis on the library, accurate, sensitive, and comprehensive identification of Mycobacterium tuberculosis and detection of its drug-resistant mutations are achieved. Through the special design of the primer combination and the detection process, the primer combination proposed in the embodiment of this application and the detection method based thereon achieve low-cost, high-sensitivity, high-specificity, and high-throughput detection. At the same time, the detection cycle is short, the accuracy rate is high, and the drug-resistant gene detection is comprehensive, covering all domestic clinical drug-related drug-resistant genes for the first time, which can effectively and comprehensively guide clinical drug use.

[0041] The first embodiment of the present application proposes a reagent for Mycobacterium tuberculosis species identification and drug resistance detection, comprising a set of multiple primers, each primer comprising a first sequence, the first sequences of the multiple primers being respectively shown as SEQ ID NOs: 1-56, wherein Mycobacterium tuberculosis species identification and drug resistance detection are both performed in the same primer pool comprising the set.

[0042] Table 1

[0043]

[0044]

[0045] In the examples of this application, the proposed primer combination detects drug-resistant genes covering a. Commonly used anti-tuberculosis drugs in clinical practice, including: first-line drugs such as rifampicin, isoniazid, ethambutol, pyrazinamide, and streptomycin; second-line drugs such as fluoroquinolones (such as levofloxacin, moxifloxacin, etc.), kanamycin, amikacin, capreomycin, linezolid, clofazimine, ethionamide, and prothionamide; and b. Common resistance sites for novel drugs such as delamanid. The 56 designed primers have an average length of 19 bases and an average GC content of 56.4%. The amplicon length is adapted to the single-molecule long-read sequencing platform, ranging from 578 to 1485 bases. It captures a total of 21 drug-resistant genes (involving resistance identification for 11 categories and more than 15 anti-tuberculosis drugs) and one species identification gene (hsp65), with a target length of 29,316 bases.

[0046] Compared with the primer design of the traditional second-generation sequencing platform, the primer combination proposed in the embodiment of the present application has a streamlined number of primers and cooperates with each other. It can realize the simultaneous detection of 1 species identification gene and 21 drug resistance-related genes (these drug resistance-related genes cover all domestic clinical drug-related drug resistance genes, including drug resistance genes corresponding to new anti-tuberculosis drugs), and there is no cross-reaction and extremely low mismatch rate when conducting comprehensive drug resistance spectrum detection, showing a relatively balanced amplification efficiency and extremely high detection efficiency, especially in terms of sensitivity and specificity. The primer combination proposed in the embodiment of the present application is designed for a single-molecule long-read sequencing platform, which can achieve rapid, accurate, and high-precision identification of Mycobacterium tuberculosis species and one-time detection of its comprehensive drug resistance spectrum, with a detection limit as low as 100 copies / μL, showing high sensitivity and high specificity (100%), and can effectively and comprehensively guide clinical drug use.

[0047] In some embodiments, the primer set includes: a first primer subset comprising primers having first sequences as shown in SEQ ID NOs: 1-26; and a second primer subset comprising primers having first sequences as shown in SEQ ID NOs: 27-56. In the embodiments of the present application, by grouping the primer combinations and performing group amplification during amplification, cross-reactions between primers are largely avoided, thereby improving the sensitivity, specificity, and accuracy of detection.

[0048] In some embodiments, a first subset of primers is used to amplify a first target region involving a region within the following genes: rplC, ddn, gyrA, rpsL, ahpC, pncA, RV0678, rpoB, rrs, ethA, katG, and embB; and a second subset of primers is used to amplify a second target region involving a region within the following genes: eis, hsp65, fabG1, gyrB, inhA, rrl, gidB, tlyA, fbiA, fgd1, rpoB, rrs, ethA, katG, and embB. The genes detected by the primer combination proposed in the examples of the present application include the hsp65 gene, the gold standard for Mycobacterium tuberculosis species identification, and cover common resistance sites of commonly used anti-tuberculosis drugs in clinical practice as well as newer anti-tuberculosis drugs that are not yet common. This achieves rapid, accurate, and highly precise identification of Mycobacterium tuberculosis species and one-time detection of its comprehensive drug resistance spectrum, and exhibits high sensitivity and high specificity (100%), which can effectively and comprehensively guide clinical drug use.

[0049] In some embodiments, the length of the first target region and the second target region amplified by the first primer subset and the second primer subset, respectively, is 500 to 1500 bp. Specifically, the amplicon length of the first primer subset ranges from 578 to 1485 bases, and the amplicon length of the second primer subset ranges from 996 to 1384 bases. The primer combination proposed in the embodiment of the present application can obtain long amplicons suitable for downstream single-molecule sequencing and can perform real-time reading of the entire single-molecule sequence, thereby conveniently and quickly achieving high-precision identification of Mycobacterium tuberculosis from sample to report and one-time detection of multiple drug-resistant mutations in hours.

[0050] In some embodiments, in addition to the first sequence shown in SEQ ID NO: 1-56, each primer may further comprise a second sequence, which may comprise a tag sequence (such as a barcode) for distinguishing the source of the sample. In some embodiments, the tag sequence may be located at the 5' side of the first sequence (such as Figure 2 As shown). It is understandable that the second sequence (label sequence) carried by the primers for amplifying the same sample is the same; the second sequence (label sequence) carried by the primers for amplifying different samples is different. In an embodiment of the present application, by introducing a label sequence into each primer respectively, the introduction of labels for multi-sample detection in the conventional single-molecule sequencing library construction process can be avoided, saving the operation time of the label introduction step (about 2h), and realizing the one-step "amplification-label labeling (i.e., library pre-preparation)", thereby greatly simplifying the experimental operation, shortening the detection time, and improving the detection throughput. It is understandable that, in addition to the label sequence, the second sequence may also include one or more other sequences, such as universal sequences for enrichment and purification, sequencing adapter identification / connection, etc., and the present application does not limit the types of other sequences contained in the second sequence.

[0051] In some embodiments, the reagents for Mycobacterium tuberculosis species identification and drug resistance detection can be in the form of a kit. In addition to the primer set described in the embodiments of the present application, the reagents can also include one or more of the following: DNA polymerase, dNTP, reaction buffer and Mg 2+ In some embodiments, based on the fact that each primer does not contain the second sequence containing the tag sequence, the reagent may further contain a tag sequence for distinguishing different samples. In some specific embodiments, the working concentration of the primer set, the first primer subset, and the second primer subset is 10 μM, respectively, wherein each primer contained in the primer set, the first primer subset, and the second primer subset is provided in the form of equal concentration.

[0052] In some embodiments, the reagents for Mycobacterium tuberculosis species identification and drug resistance detection may also introduce primers targeting other target genes or target mutation sites according to specific needs, for example, primers targeting other targets for identifying Mycobacterium tuberculosis species, such as 16S rRNA, IS6110, 16S-23S rRNA gene spacer (ITS), etc.; and primers targeting other drug resistance target genes, such as mutations in thyA that confer resistance to para-aminosalicylic acid (PAS), and atpE mutations that confer resistance to bedaquiline.

[0053] The second embodiment of the present application also proposes a method for amplifying a target region related to the identification and drug resistance detection of Mycobacterium tuberculosis species, comprising: combining a set of primers as defined in any embodiment of the present application with a nucleic acid sample to be tested, a DNA polymerase, dNTPs, and an optional reaction buffer and Mg. 2+ Mix to form a first reaction solution; or respectively mix the first primer subset and the second primer subset as defined in any embodiment of the present application with the nucleic acid sample to be tested, DNA polymerase, dNTP and optional reaction buffer and Mg 2+ mixing to form a second reaction solution and a third reaction solution; and subjecting i. the second reaction solution and the third reaction solution or ii. the first reaction solution to a thermal cycle program to obtain amplicons of the target region.

[0054] In some embodiments, the working concentrations of the primer set, the first primer subset, and the second primer subset are each 10 μM, wherein the primer set is obtained by mixing the primers having the first sequence as shown in SEQ ID NO: 1-56 in equal concentrations; the first primer subset is obtained by mixing the primers having the first sequence as shown in SEQ ID NO: 1-26 in equal concentrations, and the second primer subset is obtained by mixing equal concentrations of the primers having the first sequence as shown in SEQ ID NO: 27-56 in equal concentrations.

[0055] In some embodiments, the amount of nucleic acid sample to be tested is 0.1-300 ng, for example, 10-250 ng. In some embodiments, taking a 25 μL system as an example, the amplification system for amplifying target regions related to species identification and drug resistance detection of Mycobacterium tuberculosis may include: a. PCR mixture, including DNA polymerase, dNTP, Mg 2+ and buffer, a total of 12.5 μL; b. primer mixture (primer set or first and second primer subsets, preferably first and second primer subsets): 5 μL, 10 μM, where if the first and second primer subsets are used, take 2.5 μL of each at a working concentration of 10 μM; c. nucleic acid sample to be tested: 1 μL, with a concentration of 10-250 ng / μL; and d. sterile water: add to a total volume of 25 μL.

[0056] In some embodiments, the thermal cycling program includes: a preliminary denaturation stage: 98°C, 2-5 min; a cycling stage: 98°C, 15 s; 62°C, 30 s; 58°C, 30 s, 72°C, 1 min 30 s, 25-40 cycles, preferably 30-35 cycles; an extension stage: 72°C, 5-10 min; and an optional storage stage: 4°C-10°C.

[0057] In the examples of the present application, the ratio of each component in the amplification system, primer concentration, etc. are adaptively optimized based on the primer combination, which reduces the cross-reaction between primers and solves the compatibility problems in traditional multiplex PCR (such as primer dimers, amplification efficiency differences, etc.); at the same time, by adaptively adjusting the amplification program, optimized reaction parameters are provided to ensure that multiple targets can be amplified simultaneously in the same reaction system, thereby providing a basis for high-throughput, low-cost, high-sensitivity, and high-accuracy Mycobacterium tuberculosis species identification and drug resistance detection.

[0058] In a third aspect, an embodiment of the present application provides a method for constructing a library for a nucleic acid sample, the library being used for species identification and drug resistance detection of Mycobacterium tuberculosis in the nucleic acid sample by single-molecule sequencing, the method comprising: amplifying a target region of the nucleic acid sample according to a method as described in any embodiment of the present application to obtain amplicons of each target region; and introducing a single-molecule sequencing adapter to one side of each amplicon to obtain the library. In some embodiments, before introducing the single-molecule sequencing adapter, the method further comprises: performing end-repair and optional tag sequence ligation and purification on the amplicons.

[0059] In a fourth aspect, an embodiment of the present application proposes a method for constructing a library for multiple nucleic acid samples, the library being used for batch species identification and drug resistance detection of Mycobacterium tuberculosis in multiple nucleic acid samples by single-molecule sequencing, the method comprising: amplifying the target region of the nucleic acid sample according to the method defined in any embodiment of the present application to obtain an amplicon of the target region; and introducing a single-molecule sequencing connector to one side of each amplicon to obtain a library, wherein a set of primers or a first primer subset and a second primer subset are used for the amplification, wherein each primer in the set of primers or the first primer subset and the second primer subset contains a second sequence, which contains a label sequence for distinguishing the source of the sample. In the embodiment of the present application, the special design of the multiple primer combination greatly reduces the demand for templates and detection reagents, and can save about 2 hours of label sequence connection time in traditional library construction, thereby greatly shortening the detection time while reducing the detection cost, thereby realizing rapid detection of multiple samples.

[0060] In the examples of this application, prior to introducing the single-molecule sequencing adapter, the method further includes performing end-repair and optional purification on the amplicons. It will be appreciated that appropriate single-molecule sequencing adapters can be selected based on the compatibility requirements of a specific single-molecule sequencing platform, and other desired sequences or any other desired steps can be introduced during library construction based on specific needs. This application does not limit other specific reagents and operational steps involved in library construction based on multiple nucleic acid samples.

[0061] The fifth embodiment of the present application provides a method for species identification and drug resistance detection of Mycobacterium tuberculosis, comprising: constructing a library for single-molecule sequencing according to the method described in any embodiment of the present application, the library comprising each amplicon of the target region of the nucleic acid sample; performing single-molecule sequencing on the library to obtain sequencing data of each amplicon; and parsing the sequencing data to perform species identification and drug resistance detection on Mycobacterium tuberculosis in the nucleic acid sample.

[0062] In some embodiments, reference Figure 1 , parsing the sequencing data to identify the species and detect drug resistance of Mycobacterium tuberculosis in the nucleic acid sample specifically includes: performing a first alignment of the sequencing data with a reference genome of Mycobacterium tuberculosis or with a reference gene sequence of hsp65 of Mycobacterium tuberculosis, and reporting that the nucleic acid sample contains Mycobacterium tuberculosis based on the presence of a sequence matching the hsp65 reference gene in the sequencing data; performing a second alignment of the sequencing data with the reference genome of Mycobacterium tuberculosis or with reference target sequences corresponding to the remaining target regions to obtain a first alignment sequence that aligns the sequencing data to each target region; based on the first alignment sequence, detecting variations in the first alignment sequence within the target region; and determining the drug resistance spectrum of Mycobacterium tuberculosis in the nucleic acid sample based on the number and variation of the first alignment sequences in each target region. In some embodiments, Mycobacterium tuberculosis in the nucleic acid sample is confirmed to have drug resistance corresponding to the gene in the target region based on the number of first alignment sequences in the target region being ≥100 and the variation frequency in the target region being ≥0.5.

[0063] In some specific embodiments, alignment tools such as Minimap2 and TB-Profiler can be used to align the sequencing data to the reference sequence or reference genome of the 22 targets of Mycobacterium tuberculosis detected by the primer combination proposed in the embodiment of the present application, so as to obtain the coverage and / or number of alignments of each amplicon on the target gene in the sequencing data, and calculate the coverage depth. In some embodiments, tools such as Clair3 or Bcftools can be used to detect mutations and corresponding mutation frequencies in the sequencing data on each target gene. When the sequencing data is aligned to 21 drug-resistant gene targets and meets the number of aligned sequences on a certain target of ≥100 and the mutation frequency of ≥0.5, it is confirmed that the Mycobacterium tuberculosis carried in the sample has drug resistance corresponding to these target genes. The embodiment of the present application improves the reliability and accuracy of Mycobacterium tuberculosis drug resistance detection by setting a reporting threshold.

[0064] In some embodiments, Mycobacterium tuberculosis in the nucleic acid sample is confirmed to possess drug resistance corresponding to the gene in the target region based on the number of first aligned sequences of the target site in the target region being ≥100 and the mutation frequency of the target site being ≥0.5. In some embodiments, the mutation frequency may optionally include the frequency of point mutations.

[0065] In some specific embodiments, Figure 1 As shown, TB-Profiler can also be used to judge the drug resistance of Mycobacterium tuberculosis in the nucleic acid sample to the gene corresponding to the target region, wherein 21 drug-resistant gene targets are compared based on sequencing data and the number of comparison sequences on a certain target is ≥100 and the mutation frequency is ≥0.5, and TB-Profiler reports that Mycobacterium tuberculosis in the nucleic acid sample is highly correlated with the drug resistance corresponding to the gene in the target region (i.e., TB-Profiler displays "Assoc wR"), then it is determined that Mycobacterium tuberculosis in the nucleic acid sample has the drug resistance corresponding to the gene in the target region. In the embodiment of the present application, by using TB-Profiler for drug resistance judgment, the reliability and accuracy of Mycobacterium tuberculosis drug resistance detection are further improved.

[0066] In some embodiments, before performing species identification and drug resistance testing on Mycobacterium tuberculosis in a nucleic acid sample, the method further includes: performing quality control filtering on the original sequencing data, including removing linker sequences and performing quality control (such as Qscore-based quality filtering, such as Q7 quality filtering) to obtain filtered sequencing data; and performing a third alignment of the filtered sequencing data with a reference genome of Mycobacterium tuberculosis or with reference gene sequences of all target regions of Mycobacterium tuberculosis. Based on the fact that the amount of alignment data in the filtered sequencing data that is aligned to the target region is greater than 7% of the total amount of the filtered sequencing data, the filtered sequencing data is determined to be a qualified sample, and downstream species identification and drug resistance testing are performed.

[0067] In some embodiments, the sequencing data used for Mycobacterium tuberculosis species identification and drug resistance detection is less than 1Gb, preferably less than 500Mb, for example, 300Mb. Correspondingly, the sequencing used for Mycobacterium tuberculosis species identification and drug resistance detection only requires the time required to generate these data volumes. For example, taking the nanopore sequencing platform as an example, sequencing can be performed in less than 1 hour, less than half an hour, or even only 10 minutes. The Mycobacterium tuberculosis species identification and drug resistance detection method proposed in the embodiment of the present application can achieve rapid and accurate Mycobacterium tuberculosis species identification and drug resistance determination with extremely low data volume, thereby greatly reducing the data volume demand in traditional drug resistance analysis, and also reducing the demand for computing power, making detection faster, more convenient and economical.

[0068] In the examples of this application, any or all of the above-mentioned quality control filtering, first to third alignment, and drug resistance prediction steps can be combined to form a more efficient automated script. Furthermore, it is understood that other command lines, scripts, or software can also be used to analyze sequencing data, and this application does not limit the specific analysis script used.

[0069] The sixth aspect of the present application also proposes an integrated system for species identification and drug resistance detection of Mycobacterium tuberculosis, including: i. an amplification-pre-library module, used to amplify target regions related to species identification and drug resistance detection of Mycobacterium tuberculosis in the nucleic acid sample to be tested, and obtain amplicons of each target region, wherein the amplification is performed using the set of primers or the first primer subset and the second primer subset described in any embodiment of the present application, wherein each primer in the set of primers or the first primer subset and the second primer subset contains a second sequence, which contains a label sequence for distinguishing the source of the sample; ii. a single-molecule sequencing adapter introduction module, used to introduce a single-molecule sequencing adapter into one side of each amplicon to obtain a single-molecule sequencing library of the nucleic acid sample to be tested; iii. a single-molecule sequencing module, used to perform single-molecule sequencing on the single-molecule sequencing library to obtain sequencing data of each amplicon; and iv. a data analysis module, used to parse the sequencing data to perform species identification and drug resistance detection on Mycobacterium tuberculosis in the nucleic acid sample.

[0070] In the examples of the present application, by specially designing primer combinations for species identification and drug-resistant gene mutation detection of Mycobacterium tuberculosis, target enrichment based on super-multiplex PCR and efficient library construction of multiple samples are achieved, and combined with single-molecule sequencing technology and the analysis process of sequencing data, an integrated "detection-analysis-reporting" system can be formed. This integrated system can achieve high-throughput, hourly "from sample to report" rapid detection, while having high sensitivity and accuracy. The integrated detection method and integrated system proposed in the examples of the present application achieve low-cost, high-sensitivity, high-specificity, and high-throughput detection, while having a short detection cycle, high accuracy, and comprehensive drug-resistant gene detection, which can effectively and comprehensively guide clinical drug use.

[0071] The present application also provides a library construction kit for species identification and drug resistance testing of Mycobacterium tuberculosis in nucleic acid samples by single-molecule sequencing. The kit may include: reagents as described in any of the embodiments of the present application; and single-molecule sequencing adapters, which may optionally include a motor protein. In some embodiments, the kit further includes one or more of the following: end repair reagents, single-molecule sequencing adapter ligation reagents, purification reagents, and optional tag sequences. The formulations and / or sequences of these reagents are commonly available in the art.

[0072] The examples of the present application also provide a method for Mycobacterium tuberculosis species identification and drug resistance detection, wherein the method is performed by using the reagents or kits described in any of the examples of the present application.

[0073] The examples of the present application also provide the use of the reagent or kit described in any of the examples of the present application in the identification of Mycobacterium tuberculosis species and the detection of drug resistance.

[0074] It should be noted that the explanation of the embodiments of "reagents for Mycobacterium tuberculosis species identification and drug resistance detection" in this application is also applicable to methods for amplifying target regions related to Mycobacterium tuberculosis species identification and drug resistance detection, methods for constructing libraries for nucleic acid samples, methods for constructing libraries for multiple nucleic acid samples, library construction kits and their applications, methods for Mycobacterium tuberculosis species identification and drug resistance detection, and integrated systems for Mycobacterium tuberculosis species identification and drug resistance detection, which also have the same technical effects and advantages, and this application will not repeat them here.

[0075] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials and reagents used in the following examples, unless otherwise specified, were all commercially available.

[0076] Unless otherwise specified, the quantitative tests in the following examples were performed three times, and the results were averaged.

[0077] Except for the special parameters indicated, the data analysis in the following examples all refers to conventional procedures in the art, such as using conventional analysis methods and / or default parameters of relevant software.

[0078] Example 1

[0079] In this example, the primer combination shown in SEQ ID NO: 1-56 was used to PCR amplify the sensitive Mycobacterium tuberculosis strain H37R vDNA standard (S1, 10 4 copies / μL; 22 target genes derived from the Mycobacterium tuberculosis rifampicin resistance gene detection reagent were detected using the national reference material YJ230033 to verify the feasibility of the primer combination and detection method proposed in the examples of this application. The specific steps are as follows:

[0080] 1.1 Multiplex PCR

[0081] 1.1.1 The primer combination shown in SEQ ID NOs: 1-26 is used as primer set A (i.e., the first primer subset), and the primer combination shown in SEQ ID NOs: 26-56 is used as primer set B (i.e., the second primer subset). Equal volumes of primers from groups A and B are mixed to form a 10 μM primer set mixture of primer set A and primer set B. Reaction solutions from group A and group B are prepared according to the system shown in Table 2:

[0082] Table 2

[0083] Components volume rTaq Multiplex 2X Master Mix (Sinochem, Cat: LS-EZ-K-00003O) 12.5μL Primer set mixture (Group A or Group B) 5 μL, 10 μM S1 H37Rv DNA template 1 μL Sterile water Make up to 25 μL

[0084] 1.1.2 Place the reaction solution of group A and the reaction solution of group B in a PCR instrument respectively and perform amplification using the following thermal cycle program: (1) 98°C, 2 min; (2) 98°C, 15 s; 62°C, 30 s; 58°C, 30 s, 72°C, 1 min 30 s, 35 cycles; (3) 72°C, 5 min; (4) cool to 4°C and hold.

[0085] 1.1.3 Use magnetic beads to purify the reaction products and perform quality testing on the purified reaction products.

[0086] 1.2 Single-molecule long-read library construction for single samples

[0087] The purified reaction product was subjected to single-molecule library construction using the H940-000013 CycloneSEQ Universal Library Prep set (24RXN) in strict accordance with its instructions, i.e., single-molecule library construction of a single sample.

[0088] 1.3 Single-molecule long-read sequencing

[0089] Single-molecule long-read sequencing was performed on the single-molecule library in 1.2 using the H940-000016 CycloneSEQ WT Sequencing Kit (6T) in strict accordance with its instructions.

[0090] 1.4 Data Analysis

[0091] Refer to Figure 1 The analysis workflow shown here begins by removing adapters and quality filtering the raw sequencing data for each amplicon using Porechop. Using TB-Profiler, the filtered, high-quality sequencing data is aligned to the Mycobacterium tuberculosis reference genome H37Rv (NC_000962.3). A sample is considered qualified if: a. the data aligns to 22 targets; b. the data aligning to these 22 targets accounts for at least 7% of the total data; and c. the data aligns to the hsp65 gene of the reference genome H37Rv.

[0092] Based on the fact that the sample belongs to the Mycobacterium tuberculosis species and is qualified, the coverage and number of alignments for each target are counted, and the coverage depth is calculated. Clair3 and / or Bcftools are used to detect mutations and mutation frequencies for each target, and TB-Profiler is used to predict drug resistance. If a target has 1) ≥100 target alignments; 2) a mutation frequency ≥0.5; and 3) TB-Profiler indicates "Highly correlated with drug resistance phenotype (Assoc wR)", the sample is reported as having drug resistance for that target, and resistance results for all 21 targets are output.

[0093] 1.5 Results

[0094] In this example, the proposed primer combination was used to detect the sensitive strain H37Rv DNA standard of Mycobacterium tuberculosis (S1, 10 4 copies / μL) of amplification product concentration and single-molecule long-read sequencing data, with two repetitions in between. By comparing the corresponding data of the two repetitions, it was found that the uniformity was good. The results are shown in Table 3, indicating that the primer combination proposed in the embodiment of the present application can stably amplify the various targets of Mycobacterium tuberculosis. Among them, the data N50 length (referring to the sequence arranged from long to short, and then added, when it is added to a certain sequence to reach 50% of the total length, the length of the sequence) meets the design length of the amplicon, and the genome coverage of the full-length 29316 bases of the target amplified within 2 minutes of sequencing reaches 100% ( Figure 3), proving that the target is fully captured, the primers designed in this example are effective, and the sequencing data meet the requirements for full target detection. The sequence results of the rapid detection schemes of 10 minutes and 30 minutes are shown in Table 4. It can be seen that by applying the primer combination and detection method of this example, the sequencing data within 30 minutes can provide sufficient target coverage depth (>40) for mutation detection analysis; the detailed minimum target coverage depth within 30 minutes is shown in Table 4. Figure 3 The average gene coverage depth of the 22 target genes is shown in Table 5. TB-Profiler analysis did not detect any false positive results for drug-resistant mutations.

[0095] Table 3

[0096]

[0097] Table 4

[0098]

[0099] Table 5

[0100]

[0101]

[0102] The above experimental results demonstrate that the primer combination proposed in the examples of this application has excellent amplification performance for detecting a sensitive Mycobacterium tuberculosis strain H37RvDNA standard, and the primers are able to accurately and sensitively capture the designed Mycobacterium tuberculosis target. Sequencing results also demonstrate that the design of the present invention can be used on a single-molecule nanopore long-read platform, with good result uniformity. In addition, single-chip sequencing can rapidly detect Mycobacterium tuberculosis and its drug-resistant gene mutations within 30 minutes, with high sensitivity and strong specificity.

[0103] Example 2

[0104] In this example, the primer combination shown in SEQ ID NO: 1-56 was used to PCR amplify the sensitive Mycobacterium tuberculosis strain H37R vDNA standard (S1, 10 4 The 22 target genes of samples S2 (1000 copies / μL), S3 (100 copies / μL), and S4 (10 copies / μL) were detected using a ten-fold dilution of the Mycobacterium tuberculosis rifampicin resistance gene detection reagent (national reference material YJ230033) to determine the minimum detection limit of the primer combination proposed in the examples of this application and to establish a multi-sample library construction and sequencing process. The specific steps are as follows:

[0105] 2.1 Multiplex PCR

[0106] Referring to the method in step 1.1 of Example 1, samples S2-S4 were amplified using the primer combination shown in SEQ ID NOs: 1-26 as primer set A (i.e., the first primer subset) and the primer combination shown in SEQ ID NOs: 26-56 as primer set B (i.e., the second primer subset).

[0107] 2.2 Single-molecule long-read library construction for multiple samples

[0108] Using the H940-000018 CycloneSEQ 24 Barcode Library Prep Set, strictly following the instructions, the purified reaction products were subjected to multi-sample single-molecule library construction, including a barcode ligation step that took approximately 2 hours to distinguish different sample sources.

[0109] 2.3 Single-molecule long-read sequencing

[0110] Refer to step 1.3 of Example 1.

[0111] 2.4 Data Analysis

[0112] Refer to step 1.4 of Example 1.

[0113] 2.5 Results

[0114] In this example, the proposed primer combination was used to detect the concentration of amplified products and single-molecule long-read sequencing data from different concentrations of a sensitive Mycobacterium tuberculosis strain DNA standard, H37Rv (S2, 1000 copies / μL; S3, 100 copies / μL; S4, 10 copies / μL). The amplified product concentrations and sequence detection results are shown in Table 6. As can be seen from Table 6, the primer combination and detection method proposed in this application can achieve amplification with detectable product concentrations at sample concentrations as low as 10 copies / μL. The coverage depth of the 22 targets of S3 and S4 is shown in Table 7. Among them, the target gene detection coverage of the S3 (100 copies / μL) sample was normal, and the TB-Profiler analysis did not find false positive results of drug-resistant mutations. However, the S4 (10 copies / μL) sample did not detect the species identification gene hsp65 and there were targets with a gene coverage depth of <40. It was considered that the depth was insufficient to support accurate mutation detection. Therefore, the lower limit of detection concentration of Mycobacterium tuberculosis in the samples of this application was set to 100 copies / μL.

[0115] Table 6

[0116]

[0117] Table 7

[0118]

[0119] The above experimental results demonstrate that the primer combination proposed in the examples of this application can detect species and drug resistance of Mycobacterium tuberculosis at different concentrations. It can detect Mycobacterium tuberculosis nucleic acid as low as 100 copies / μL and has good amplification performance for Mycobacterium tuberculosis nucleic acid in samples. The primers can accurately and sensitively capture the designed Mycobacterium tuberculosis target. The sequencing results also show that the multi-sample mixed sample construction system can be well applied to the single-molecule nanopore long-read sequencing platform, suitable for clinical multi-sample Mycobacterium tuberculosis identification and drug-resistant gene mutation detection. The primers are highly sensitive and specific.

[0120] Example 3

[0121] This example uses the primer combination shown in SEQ ID NO: 1-56 to detect Mycobacterium tuberculosis species and drug resistance genes in clinical respiratory samples. The specific steps are as follows:

[0122] 3.1 Multiplex PCR

[0123] Referring to the method in step 1.1 of Example 1, the primer combination shown in SEQ ID NOs: 1-26 was used as primer group A (i.e., the first primer subset) and the primer combination shown in SEQ ID NOs: 26-56 was used as primer group B (i.e., the second primer subset) to perform bronchial alveolar lavage fluid analysis (provided by Shenzhen Third People's Hospital, MTB / RIF reported drug resistance mutation sites including rpsL88, embB306 and katG315 were amplified, where each primer was connected to a barcode sequence for distinguishing the source of the sample (such as Figure 2 In the reaction system, the input amount of the alveolar lavage fluid sample was 20 ng, 1 μL.

[0124] 3.2 Single-molecule long-read integrated library construction

[0125] Using the H940-000018 CycloneSEQ 24 Barcode Library Prep Set, strictly following the instructions, the purified reaction products were subjected to multi-sample single-molecule integrated library construction. The barcode ligation step (which takes approximately 2 hours) was omitted, and only the end-repair and sequencing adapter ligation steps were performed as follows:

[0126] 3.2.1 End Repair and Purification of End Repair Products

[0127] Mix 1 μg of the reaction product with sterile water to 45 μl, then add 12 μl of end-repair buffer and 3 μl of end-repair enzyme, incubate at 20°C for 10 minutes, then at 65°C for 10 minutes, and cool to 4°C to perform end-repair. Purify the end-repair product using 60 μl of magnetic beads, wash the beads with 200 μl of freshly prepared 80% ethanol, and elute the end-repaired amplicons with sterile water.

[0128] 3.2.2 Sequencing adapter ligation and ligation product purification

[0129] The purified end-repair products were mixed with 25 μl of ligation buffer, 10 μl of ligase, 2.5 μl of sterile water, and 2.5 μl of sequencing adapters. The mixture was reacted in a 25°C thermostatted metal bath for 30 minutes to ligate the sequencing adapters. The ligation products were then purified using 40 μl of magnetic beads and washed with 150 μl of fragment wash buffer (to prevent denaturation of the adapter protein after reaction with ethanol). The ligation products were then eluted using 17 μl of elution buffer. The purified ligation products were the sequencing libraries.

[0130] 3.3 Single-molecule long-read sequencing

[0131] Referring to step 1.3 of Example 1, >300 ng of library was mixed with 150 μl of sequencing reagent, 3 μl of anchoring reagent and sterile water to form a 250 μl reaction system, which was then loaded onto a sequencing chip for single-molecule long-read sequencing.

[0132] 3.4 Data Analysis

[0133] Refer to step 1.4 of Example 1.

[0134] 3.5 Results

[0135] The initial DNA concentration of the bronchoalveolar lavage fluid sample was 1.01 ng / μL. After amplification, the DNA concentration of the reaction product in Group A was 7.44 ng / μL, and the DNA concentration of the reaction product in Group B was 44.22 ng / μL (Table 8), indicating that the primer combination and detection method proposed in the embodiment of the present application can achieve effective amplification for each target. The results of single-molecule long-read sequencing after multi-sample integrated library construction are shown in Table 9, which includes the alignment sequence of hsp65 and the alignment sequences of the other 21 targets. The sample was detected to contain Mycobacterium tuberculosis and was a qualified sample. Furthermore, drug-resistant targets rpsL88, embB306 and katG315 were all detected, with the numbers of aligned sequences being 4088, 237 and 4665, respectively, all meeting the screening criteria of “number of target aligned sequences ≥ 100” set in the examples of this application; at the same time, the mutation frequencies of the three drug-resistant targets were 0.6, 0.6 and 0.5, respectively, all exceeding the screening criteria of “mutation frequency ≥ 0.5” set in the examples of this application; in addition, TB-Profiler reports of these three sites were all highly correlated with the drug-resistant phenotype (Assoc w R), and therefore a drug resistance report for the bronchoalveolar lavage fluid sample was output, including drug resistance associated with rpsL88, embB306 and katG315, which is consistent with the World Health Organization (WHO) recommendation as an important tool for tuberculosis diagnosis. The reported drug resistance patterns were consistent with those of MTB / RIF.

[0136] Table 8

[0137]

[0138] Table 9

[0139] target genes Number of aligned sequences (frequency of mutations detected at drug-resistant sites) hsp65 1 rplC 499 eis 14 ddn 10684 gyrA 266 fabG1 25 rpsL 4088(0.6) ahpC 993 pncA 405 RV0678 1549 gyrB 48 inhA 26 rrl 24 gidB 53 tlyA 146 fbiA 66 fgd1 92 rpoB 92 rrs 332 ethA 25164 katG 4665(0.5) embB 237(0.6)

[0140] The above experimental results show that the primer combination proposed in the embodiment of the present application has good performance in amplifying species identification targets and drug-resistant targets of Mycobacterium tuberculosis in clinical samples, and the primers can accurately, sensitively and specifically capture the designed Mycobacterium tuberculosis targets, thereby achieving targeted amplification of the target to be detected. The sequencing results also show that the primer combination and detection method designed in the present application can be used for the identification of Mycobacterium tuberculosis in clinical respiratory samples and the detection of drug-resistant gene mutations thereof, wherein by using specially designed primers, the label (barcode) introduction step that takes about two hours in the multi-sample library construction process can be omitted, thereby greatly shortening the detection time. In addition, the primer combination and detection method proposed in the embodiment of the present application can be adapted to the multi-sample integrated library construction system and the single-molecule nanopore long-read sequencing platform, and its detection speed is faster, the chip consumption is less, the sensitivity is high, the specificity is strong, and high-accuracy detection is achieved (compared to the World Health Organization (WHO) recommended as an important tool for tuberculosis diagnosis). resistance reported for MTB / RIF).

[0141] Industrial Applicability

[0142] This application creatively designs primer combinations for the identification of Mycobacterium tuberculosis species and the detection of drug-resistant gene mutations, multi-sample efficient library construction methods, and sequencing data analysis processes. By performing super-multiplex PCR amplification on the target gene, constructing a single-molecule length long sequencing library for the amplicon, and performing single-molecule sequencing and data analysis on the library, accurate, sensitive, and comprehensive identification of Mycobacterium tuberculosis and detection of its drug-resistant mutations are achieved. Through the special design of the primer combination and the detection process, the primer combination proposed in the embodiment of this application and the detection method based thereon achieve low-cost, high-sensitivity, high-specificity, and high-throughput detection. At the same time, the detection cycle is short, the accuracy rate is high, and the drug-resistant gene detection is comprehensive, covering all domestic clinical drug-related drug-resistant genes for the first time, which can effectively and comprehensively guide clinical drug use.

[0143] The applicant states that the present invention is intended to illustrate the detailed methods of the present invention through the above-described embodiments, but the present invention is not limited to the above-described detailed methods, that is, it does not mean that the present invention must rely on the above-described detailed methods in order to be implemented. Those skilled in the art should understand that any improvements to the present invention, equivalent substitutions for various raw materials in the products of the present invention, addition of auxiliary ingredients, and selection of specific methods, etc., are all within the scope of protection and disclosure of the present invention.

[0144] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0145] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A reagent for Mycobacterium tuberculosis species identification and drug resistance detection, comprising: A set of multiple primers, each primer comprising a first sequence, wherein the first sequences of the multiple primers are respectively shown as SEQ ID NOs: 1-56, The Mycobacterium tuberculosis species identification and drug resistance detection are both performed in the same primer pool comprising the set, wherein the set comprises: a first subset of primers, comprising the primers having the first sequences shown in SEQ ID NOs: 1-26, respectively; and A second primer subset comprises the primers having the first sequences shown in SEQ ID NOs: 27-56, respectively.

2. The reagent according to claim 1, wherein each of the primers further comprises a second sequence, wherein the second sequence comprises a tag sequence for distinguishing the source of the sample, and the tag sequence is optionally located at the 5' side of the first sequence. Optionally, the second sequence further comprises one or more universal sequences, Optionally, the first primer subset is used to amplify a first target region, the first target region comprising regions in the following genes: rplC, ddn, gyrA, rpsL, ahpC, pncA, RV0678, rpoB, rrs, ethA, katG, and embB; and The second primer subset is used to amplify a second target region, which includes regions in the following genes: eis, hsp65, fabG1, gyrB, inhA, rrl, gidB, tlyA, fbiA, fgd1, rpoB, rrs, ethA, katG, and embB, Optionally, the lengths of the first target region and the second target region amplified by the first primer subset and the second primer subset, respectively, are 500 to 1500 bp.

3. The reagent according to claim 1 or 2, wherein the reagent is a kit, Optionally, the reagents further include one or more of the following: DNA polymerase, dNTP, reaction buffer and Mg 2+ , Optionally, the working concentrations of the first primer subset and the second primer subset are respectively 10 μM, wherein the first primer subset comprises equal concentrations of the primers having the first sequences shown in SEQ ID NOs: 1-26, and the second primer subset comprises equal concentrations of the primers having the first sequences shown in SEQ ID NOs: 27-56, Optionally, the kit further comprises: a single molecule sequencing adapter, the single molecule sequencing adapter optionally comprising a motor protein, Optionally, the kit further comprises one or more of the following: an end repair reagent, the single molecule sequencing adapter ligation reagent, a purification reagent, and an optional tag sequence.

4. A method for amplifying a target region associated with species identification and drug resistance detection of Mycobacterium tuberculosis, comprising: The primer set as defined in claim 1 is mixed with a nucleic acid sample to be tested, a DNA polymerase, dNTPs and an optional reaction buffer and Mg. 2+ mixing to form a first reaction solution; or The first primer subset and the second primer subset as defined in any one of claims 1 to 3 are respectively mixed with the nucleic acid sample to be tested, DNA polymerase, dNTPs and optional reaction buffer and Mg 2+ mixing to form a second reaction liquid and a third reaction liquid; and i. the second reaction solution and the third reaction solution or ii. the first reaction solution are placed in a thermal cycle program to obtain an amplicon of the target region, Optionally, the thermal cycling program includes: Pre-denaturation stage: 98°C, 2-5 min; Cycling stage: 98°C, 15s; 62°C, 30s; 58°C, 30s, 72°C, 1min 30s, 25-40 cycles, preferably 30-35 cycles; Extension phase: 72°C, 5-10 min; and Optional storage stage: 4℃-10℃.

5. The method according to claim 4, wherein the working concentration of the primer set, the first primer subset and the second primer subset is 10 μM each, wherein The primer set is obtained by mixing primers having first sequences as shown in SEQ ID NOs: 1-56 in equal concentrations; The first primer subset is obtained by mixing primers having first sequences as shown in SEQ ID NOs: 1-26 in equal concentrations, and the second primer subset is obtained by mixing primers having first sequences as shown in SEQ ID NOs: 27-56 in equal concentrations, Optionally, the amount of the nucleic acid sample to be tested is 0.1-300 ng, preferably 10-250 ng. Optionally, the detection limit of Mycobacterium tuberculosis in the nucleic acid sample to be tested is 100 copies / μL.

6. A method for constructing a library for multiple nucleic acid samples, wherein the library is used to perform batch species identification and drug resistance detection of Mycobacterium tuberculosis in the multiple nucleic acid samples by single-molecule sequencing, the method comprising: amplifying the target region of the nucleic acid sample according to the method as defined in claim 4 or 5 to obtain an amplicon of the target region; and introducing a single-molecule sequencing adapter into one side of each amplicon to obtain the library, wherein said amplification is performed using a set of primers or a first subset of primers and a second subset of primers, wherein The primer set comprises a plurality of primers, each primer comprises a first sequence, and the first sequences of the plurality of primers are respectively shown as SEQ ID NOs: 1-56; The first primer subset comprises the primers having the first sequences shown in SEQ ID NOs: 1-26, respectively; The second primer subset comprises the primers having the first sequences shown in SEQ ID NOs: 27-56, respectively; Each of the primers further comprises a second sequence, wherein the second sequence comprises a tag sequence for distinguishing the source of the sample. Optionally, before introducing the single-molecule sequencing adapter, the method further comprises: The amplicons are end-repaired and optionally purified.

7. A method for species identification and drug resistance detection of Mycobacterium tuberculosis, comprising: Constructing a library for single-molecule sequencing according to the method of claim 6, wherein the library comprises each amplicon of the target region of the nucleic acid sample; performing single-molecule sequencing on the library to obtain sequencing data of each amplicon; and The sequencing data is analyzed to perform species identification and drug resistance detection on Mycobacterium tuberculosis in the nucleic acid sample.

8. The method according to claim 7, wherein parsing the sequencing data to perform species identification and drug resistance detection on Mycobacterium tuberculosis in the nucleic acid sample specifically comprises: performing a first alignment of the sequencing data with a reference genome of Mycobacterium tuberculosis or with a hsp65 reference gene sequence of Mycobacterium tuberculosis, and reporting that the nucleic acid sample contains Mycobacterium tuberculosis based on the presence of a sequence in the sequencing data that matches the hsp65 reference gene; Performing a second alignment of the sequencing data with the reference genome of Mycobacterium tuberculosis or the reference target sequences corresponding to the remaining target regions to obtain a first alignment sequence of the sequencing data aligned to each of the target regions; Based on the first aligned sequence, detecting variations in the first aligned sequence within the target region; Determine the drug resistance spectrum of Mycobacterium tuberculosis in the nucleic acid sample based on the number and variation of the first comparison sequences in each target region, Wherein, based on the fact that the number of the first aligned sequences in the target region is ≥100 and the mutation frequency in the target region is ≥0.5, it is confirmed that Mycobacterium tuberculosis in the nucleic acid sample has drug resistance corresponding to the gene in the target region.

9. The method according to claim 8, further comprising: Based on the number of first aligned sequences of the target site in the target region being ≥100 and the mutation frequency of the target site being ≥0.5, it is confirmed that the Mycobacterium tuberculosis in the nucleic acid sample has drug resistance corresponding to the gene in the target region, The variation frequency optionally includes point mutation frequency, Optionally, the method further includes: Based on the number of the first aligned sequences of the target site in the target region being ≥100 and the mutation frequency of the target site being ≥0.5, TB-Profiler is used to determine the drug resistance of Mycobacterium tuberculosis in the nucleic acid sample corresponding to the gene in the target region. Wherein, based on the TB-Profiler report that Mycobacterium tuberculosis in the nucleic acid sample is highly correlated with the drug resistance corresponding to the gene in the target region, it is determined that Mycobacterium tuberculosis in the nucleic acid sample has the drug resistance corresponding to the gene in the target region, Optionally, before performing species identification and drug resistance testing on the Mycobacterium tuberculosis in the nucleic acid sample, the method further comprises: Performing quality control filtering on the raw sequencing data to obtain filtered sequencing data; and The filtered sequencing data is compared with the reference genome of Mycobacterium tuberculosis or the reference gene sequence of all target regions of Mycobacterium tuberculosis for a third time. Based on the fact that the amount of alignment data in the filtered sequencing data that is aligned to the target region is greater than 7% of the total amount of the filtered sequencing data, the filtered sequencing data is determined to be a qualified sample, and downstream species identification and drug resistance testing are performed.

10. An integrated system for Mycobacterium tuberculosis species identification and drug resistance detection, comprising: i. Amplification-pre-built library module, used to amplify target regions related to Mycobacterium tuberculosis species identification and drug resistance detection in the nucleic acid sample to be tested, and obtain amplicons of each target region, wherein the amplification is performed using a first primer set and a second primer set, wherein The amplification is performed using a set of primers or a first subset of primers and a second subset of primers, wherein The primer set comprises a plurality of primers, each primer comprises a first sequence, and the first sequences of the plurality of primers are respectively shown as SEQ ID NOs: 1-56; The first primer subset comprises the primers having the first sequences shown in SEQ ID NOs: 1-26, respectively; The second primer subset comprises the primers having the first sequences shown in SEQ ID NOs: 27-56, respectively; Each of the primers further comprises a second sequence, wherein the second sequence comprises a tag sequence for distinguishing the source of the sample; ii. a single-molecule sequencing adapter introduction module for introducing a single-molecule sequencing adapter into one side of each amplicon to obtain a single-molecule sequencing library of the nucleic acid sample to be tested; iii. a single molecule sequencing module for performing single molecule sequencing on the single molecule sequencing library to obtain sequencing data of each amplicon; and iv. A data analysis module for analyzing the sequencing data to perform species identification and drug resistance detection on Mycobacterium tuberculosis in the nucleic acid sample.