Genital tract microorganism and drug-resistant gene analysis method based on targeted nanopore sequencing data
By using analytical methods based on targeted nanopore sequencing data, a database was constructed and targeted sequence identification and alignment were performed. This solved the problems of limited detection range and high cost of reproductive tract infections, and enabled rapid, low-cost, and highly sensitive detection of microorganisms and drug resistance genes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU DIAN BIOTECH CO LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies for detecting pathogens and drug resistance genes in reproductive tract infections suffer from limitations in detection range, high cost, long detection time, and insufficient sensitivity and specificity.
By employing an analysis method based on targeted nanopore sequencing data, a reference database of microorganisms and drug resistance genes is constructed to identify and compare targeted sequences. Combined with quality control and contamination filtration, this approach enables efficient and accurate detection of microorganisms and drug resistance genes.
It enables rapid, low-cost, highly sensitive, and specific detection of reproductive tract microorganisms and drug resistance genes, significantly shortening detection time, reducing false positives, and improving the reliability of test results.
Smart Images

Figure SMS_1 
Figure SMS_2
Abstract
Description
Technical Field
[0001] This invention belongs to the field of microbial detection technology, specifically relating to a method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data. Background Technology
[0002] Reproductive tract infections involve a variety of pathogenic microorganisms, including bacteria, fungi, and viruses. Rapid and accurate diagnosis is crucial for clinical treatment and prognosis. While traditional culture methods are the gold standard for etiological diagnosis, they are often limited by culture conditions and timelines, making it difficult to reflect the composition of pathogenic microorganisms in a timely and comprehensive manner.
[0003] Molecular detection methods such as PCR are widely used in the detection of specific pathogens and some drug resistance genes, but their detection range is limited and singular, making it difficult to handle mixed infections and rare infections. Furthermore, pathogen detection and drug resistance gene detection are performed separately, making it impossible to simultaneously detect the results of both microorganisms and drug resistance genes.
[0004] Next-generation sequencing (NGS)-based metagenomic sequencing of pathogens can detect thousands of pathogenic microorganisms in a single run, greatly expanding the detection range. However, clinical applications typically use single-end 50bp or 75bp sequencing, and these excessively short reads make it difficult to effectively distinguish homologous sequences, resulting in insufficient specificity in clinical samples, especially sensitive samples, as they struggle to differentiate closely related species. Third-generation nanopore metagenomic sequencing, with its long reads, significantly improves analytical specificity. However, both NGS and NGS are random and unbiased sequencing methods. Typically, over 95% of the reads obtained from a sample are host sequences and sequences from colonizing bacteria with no pathogenic significance, leading to substantial data waste and high costs.
[0005] Therefore, there is an urgent need to develop a data analysis method that is low-cost, has a short detection time, higher sensitivity, and better specificity. Summary of the Invention
[0006] To address at least one of the aforementioned problems, this invention provides a method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data.
[0007] To achieve the above objectives, the present invention employs the following technical means: The first aspect of this invention provides a method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data, comprising the following steps: S1. Construct a microbial reference database, including a microbial comparison database, a microbial annotation database, a drug resistance gene comparison database, and a drug resistance gene annotation database; S2. Obtain the raw sequencing data of the targeted nanopore sequencing of the sample; S3. Raw sequencing data preprocessing: First, the raw sequencing data is subjected to quality control to obtain a high-quality effective sequence set; then, the target sequence is identified and extracted after primer alignment. S4. Using the target sequence extracted in step S2 and the microbial reference database in step S1, perform species comparison and abundance estimation of microorganisms, calculate species unique comparison reads, and output microbial detection results after quality and batch contamination control. S5. After removing human reads from the target sequence extracted in step S2, the sequence of the drug resistance gene is compared with the drug resistance gene reference database in step S1, mutation detection and annotation are performed, the results are filtered, and then combined with the microbial detection results in step S4 to output the final result.
[0008] In some embodiments of the present invention, the microbial comparison database includes: (a) Bacterial 16S Database: Used for bacterial classification and species identification, mainly sourced from NCBI nt (ftp: / / ftp.ncbi.nlm.nih.gov / blast / db / FASTA / nt.gz) and RefSeq.
[0009] (b) Fungal ITS Database: Used for rapid identification of fungal species, data from UNITE (ftp: / / ftp.unite.ut.ee) and NCBI nt.
[0010] (c) Species-Specific Genome Bank: Contains high-quality strain sequences for reproductive tract-related species or genes, derived from published complete genomes or target gene sequences in GenBank (ftp: / / ftp.ncbi.nlm.nih.gov / genomes).
[0011] During the construction process, species names are mapped to NCBI Taxonomy, and unmatched species are labeled using the parent genus taxid. Redundant sequences are identified and removed at the species level, improving the coverage and comparison efficiency of microbial classification analysis.
[0012] In some embodiments of this invention, a microbial annotation database is used to annotate the identified microbial species. Classified by bacteria, fungi, mycoplasma, parasites, and viruses, the database integrates information such as family, genus, species, Chinese name, Latin name, colonization and infection sites, pathogenicity, associated diseases, and transmission methods. The database primarily covers common colonizing and pathogenic bacteria of the reproductive tract, containing approximately 600 microorganisms.
[0013] In some embodiments of this invention, a drug resistance gene alignment database is used for the alignment and identification of drug resistance genes in reproductive tract microbial samples obtained through targeted nanopore sequencing, integrating drug resistance gene sequences related to common reproductive tract pathogens. Reference sequences are primarily derived from authoritative databases CARD (https: / / card.mcmaster.ca) and ARG-ANNOT (https: / / github.com / katholt / srst2 / blob / master / data / ARGannot_r3.fasta), supplemented with information on common clinical drug resistance mechanisms. Redundant and low-confidence sequences are removed during the construction process to ensure the integrity and accuracy of the gene sequences, thereby improving the sensitivity and specificity of drug resistance gene detection.
[0014] In some embodiments of the present invention, the drug resistance gene annotation database is used to systematically annotate drug resistance genes identified by targeted nanopore sequencing, including gene name, mutation site, nucleotide / amino acid changes, drug resistance category, associated microorganism and its Chinese name, infection type, drug-resistant antibiotic, and other information.
[0015] In some embodiments of the present invention, the method for quality control of raw sequencing data is as follows: In some embodiments of the present invention, the targeted sequence identification method in step S3 is as follows: the fastplong (v0.2.2) tool is used to remove base fragments with an average quality value of less than 15 in a sliding window manner, while reads with an overall sequence length of less than 100 bp are also removed.
[0016] S3-1. Set a sequence similarity threshold of ≥75%, perform primer sequence alignment on the sequencing sequences, and retain only the best matching result for each sequence; S3-2, Set primer coverage ≥90%; primer mismatch number ≤3 bp; primer matching sites exist within 15 bp at both ends of the amplified fragment; amplified fragment length is in the range of 200–2000 bp, and only sequences that meet the directionality and site requirements are retained as target sequences.
[0017] In some embodiments of the present invention, the abundance threshold in the abundance estimation in step S4 is 0.1%; species below this threshold are determined as background signals and are not included in subsequent result statistics. In some embodiments of the present invention, the species unique alignment read calculation method in step S4 is as follows; S4-1. Extract candidate alignments for each read and calculate the alignment results; S4-2. Select the best and second-best alignments from the candidate alignments. If the ratio of the second-best to the second-best score is less than 0.9, then the reads are determined to belong uniquely to this species. S4-3. If there is only a single reliable alignment or no competing species, it is also recorded as a unique alignment. S4-4. Count the number of unique alignment reads for each species, which is used for species abundance correction and threshold determination. S4-5. Based on the number of uniquely matched reads, set a species positive determination threshold for different pathogenic microorganisms: bacteria ≥10, parasites ≥30, fungi ≥30, and viruses ≥3. Only when the test result meets the above threshold requirements is it determined to be a positive species; results below the threshold will not be reported.
[0018] In some embodiments of the present invention, in the quality and batch contamination control of step S4, the quality standard is set as follows: Q20 > 85%, Q30 > 75%, and the proportion of effective microbial sequences ≥ 10%; the batch standard is set as follows: if the proportion of reads of a certain species in a single sample to the total reads of that species in the batch is < 1%, it is determined to be contamination or background and is removed.
[0019] In some embodiments of the present invention, in step S4, the microbial detection results include species name, number of uniquely aligned reads, relative abundance, batch filtration status, and clinical annotations.
[0020] In some embodiments of the present invention, in step S5, the sequence alignment of the drug resistance gene must meet the following requirements: sequence consistency ≥ 90%, coverage ≥ 40%, and effective alignment length ≥ 450 bp.
[0021] In some embodiments of the present invention, the result filtering in step S5 adopts a two-layer filtering standard: gene presence / absence level: only positive results supporting read count ≥ 5 are retained; SNP level: the mutation site must simultaneously meet the requirements of coverage depth ≥ 10 and mutation frequency ≥ 25%.
[0022] In some embodiments of the present invention, in the final output of step S5, the relevant drug resistance genes and drug resistance SNP sites are reported only if the corresponding microorganism is detected in the sample; if the corresponding host microorganism is not detected, it is not included in the report.
[0023] In some embodiments of the present invention, in step S5, the drug resistance gene detection results include gene name, mutation site, number of homogenized reads, coverage depth, mutation frequency, and drug resistance determination.
[0024] This invention also discloses a system for detecting reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing, which includes three core modules: a database module, a pathogen detection module, and an analysis module.
[0025] In some embodiments of the present invention, the database module includes a microbial reference database module, a microbial annotation database module, a drug resistance gene database module, and a drug resistance gene annotation database module.
[0026] (1) The microbial reference database module includes bacterial 16S, fungal ITS and specific reproductive tract-related genome sequences (NCBI GenBank). All sequences are mapped to NCBI Taxonomy, and redundant sequences are identified and removed on a species-by-species basis to ensure accuracy and conciseness.
[0027] (2) The microbial annotation database module includes species classification, Chinese name, colonization and infection sites, pathogenicity and related diseases, etc., and contains about 600 common reproductive tract bacteria, providing annotation information for microbial detection results.
[0028] (3) The drug resistance gene database module is used for the identification of drug resistance genes of reproductive tract pathogens, referencing CARD and ARG-ANNOT, and supplemented with clinical drug resistance mechanisms. Redundant and low-confidence sequences are removed to ensure the integrity and accuracy of gene sequences.
[0029] (4) The drug resistance gene annotation database module contains gene name, mutation site, drug resistance category, associated bacteria and drug resistance information, providing annotation information for drug resistance gene results.
[0030] The present invention also discloses the following method for detecting pathogenic microorganisms: (1) Obtain targeted nanopore sequencing data of the sample; (2) Sequencing data preprocessing: Quality control of raw sequencing data is performed to remove low-quality and excessively short sequences; (3) Target sequence identification: Combining sequence feature screening and primer coverage determination, non-specific amplified sequences are filtered out to obtain a high-quality effective sequence set, providing reliable input for downstream analysis; (4) Sequence alignment and microbial identification: an efficient alignment algorithm is used to align the sequencing sequence to the pathogenic microorganism database, and abundance is estimated by combining the expectation-maximization (EM) model. Species-level detection and determination are achieved through specific thresholds and unique alignment rules. (5) Quality and contamination control: Filter background interference at the sample quality control index and batch level to ensure the reliability and repeatability of the results.
[0031] This invention also discloses the following method for analyzing drug resistance genes: (1) Sequence alignment and screening: The sequencing sequence is compared with the drug resistance gene database, and positive results are screened based on the preset similarity, coverage and length thresholds; (2) Mutation detection and annotation: The screening results are analyzed for drug resistance mutation sites, and drug resistance genes and their mutation sites are labeled; (3) Results integration and output: Generate a comprehensive test result table, including the name of the pathogenic microorganism, the number of sequence supports, the drug resistance gene and its mutation site, for clinical analysis and application.
[0032] The primers used in the data analysis methods of this application are compatible with general primer design methods, and their parameters are as follows: 1. Primer length: minimum 20 bases, maximum 27 bases; 2. Primer melting temperature: minimum melting temperature 58℃, maximum melting temperature 62℃; 3. Primer GC content: minimum GC content 40%, maximum GC content 60%.
[0033] Beneficial effects of the present invention Compared with the prior art, the present invention has the following beneficial effects: (1) Fast and efficient: The method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data in this application can complete the data analysis in just 6 hours under the same hardware conditions. The time required is significantly shorter than that of second-generation metagenomic sequencing and nanopore metagenomic sequencing methods, demonstrating fast and efficient performance.
[0034] (2) High accuracy: In species detection and drug resistance gene detection, targeted nanopore sequencing shows higher sensitivity and specificity, both reaching 100%. It can accurately detect low-abundance microorganisms and drug resistance genes, reduce false negatives, and ensure the reliability of detection results.
[0035] (3) Strict specificity control: By optimizing the comparison analysis algorithm, determining unique comparison reads, and batch contamination identification and filtering strategies, false positives are effectively avoided, and the specificity of microbial and drug resistance gene detection is improved.
[0036] In summary, the sequencing analysis method of this application, combined with proprietary analysis algorithms, enables rapid identification of potentially infectious pathogens and their drug resistance genes in samples to be tested with lower cost, shorter detection time, and higher sensitivity and specificity. Detailed Implementation
[0037] The following examples are used to illustrate preferred embodiments of the invention. Those skilled in the art will understand that the techniques disclosed in the examples represent techniques discovered by the inventors that can be used to implement the invention, and therefore can be considered preferred embodiments for implementing the invention. However, those skilled in the art should understand from this specification that many modifications can be made to the specific embodiments disclosed herein, still yielding the same or similar results, without departing from the spirit or scope of the invention.
[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains, and all materials disclosed herein and cited therein are incorporated herein by reference. Many equivalent techniques of specific embodiments of the invention described herein will be recognized or can be understood by ordinary experimentation by those skilled in the art. These equivalents will be included in the claims.
[0039] The technical solution of this application will be further described in detail below with reference to specific embodiments.
[0040] Example 1: System and Method for Detecting Microorganisms and Drug Resistance Genes in Samples Based on Targeted Nanopore Sequencing I. Database Construction 1. Construction of a microbial reference database (1) Microbial comparison database (a) Bacterial 16S Database: mainly derived from NCBI nt (ftp: / / ftp.ncbi.nlm.nih.gov / blast / db / FASTA / nt.gz) and RefSeq, used for bacterial classification and species identification; (b) Fungal ITS Database: Data from UNITE (ftp: / / ftp.unite.ut.ee) and NCBI nt, used for rapid identification of fungal species; (c) Species-Specific Genome Bank: Contains high-quality strain sequences of species or genes related to the reproductive tract, derived from the complete genome or target gene sequences published in GenBank (ftp: / / ftp.ncbi.nlm.nih.gov / genomes).
[0041] During the construction process, species names are mapped to NCBI Taxonomy, and unmatched species are labeled using the parent genus taxid. Redundant sequences are identified and removed at the species level, improving the coverage and comparison efficiency of microbial classification analysis.
[0042] (2) Microbial annotation database This database is used to annotate identified microbial species. It is categorized by bacteria, fungi, mycoplasma, parasites, and viruses, integrating information such as family, genus, species, Chinese name, Latin name, colonization and infection sites, pathogenicity, associated diseases, and transmission methods. The database focuses on common colonizing and pathogenic bacteria of the reproductive tract, containing approximately 600 microorganisms.
[0043] 2. Construction of a database of microbial resistance genes (1) Drug resistance gene comparison database This database is used for the alignment and identification of drug resistance genes in reproductive tract microbial samples obtained through targeted nanopore sequencing, integrating drug resistance gene sequences associated with common reproductive tract pathogens. Reference sequences are primarily derived from authoritative databases CARD (https: / / card.mcmaster.ca) and ARG-ANNOT (https: / / github.com / katholt / srst2 / blob / master / data / ARGannot_r3.fasta), supplemented with information on common clinical drug resistance mechanisms. Redundant and low-confidence sequences were removed during the construction process to ensure the integrity and accuracy of the gene sequences, thereby improving the sensitivity and specificity of drug resistance gene detection.
[0044] (2) Antidote resistance gene annotation database This database provides systematic annotation of drug-resistant genes identified by targeted nanopore sequencing, including gene name, mutation site, nucleotide / amino acid changes, drug resistance category, associated microorganism and its Chinese name, infection type, and drug-resistant antibiotic.
[0045] II. Sample Sequencing For details regarding the species range, primer sets, and experimental detection methods involved in sample sequencing, please refer to the patent with publication number CN120818616A: Primer Combinations, Kits, and Applications for Identification of Six Reproductive Tract Pathogens and Detection of Drug Resistance Genes.
[0046] III. Sequencing Data Analysis (1) Sequencing data preprocessing Quality control was performed on the raw sequencing data. The fastplong (v0.2.2) tool was used to remove base fragments with an average quality value of less than 15 using a sliding window method, while reads with an overall sequence length of less than 100 bp were also removed.
[0047] (2) Target sequence identification (a) Primer alignment: The sequencing sequences were aligned using vsearch (v2.27.0). The sequence similarity threshold was set to ≥75%, and only the best matching result for each sequence was retained to exclude interference from low-quality or multiple alignments.
[0048] (b) Target sequence identification: Candidate sequences are further strictly filtered using a self-developed algorithm. The specific rules are: primer coverage ≥90%; primer mismatch number ≤3 bp; primer matching sites exist within 15 bp at both ends of the amplified fragment; amplified fragment length is in the range of 200–2000 bp.
[0049] Only sequences that meet the directionality and site requirements are retained, while sequences with internal matching or abnormal amplification are removed to minimize false positives caused by nonspecific amplification.
[0050] (c) Target sequence extraction: The selected sequences are uniformly integrated and output as a clean.fasta file, which serves as input data for downstream analysis. In actual operation, the retention rate of the selected sequences remained stable at over 95%, ensuring the representativeness and completeness of the analysis results.
[0051] (3) Microbial comparison and detection (a) Species comparison and abundance estimation Sequencing sequences, after quality control and effective sequence screening, were aligned to a pre-constructed reference database using minimap2 software to obtain preliminary alignment results. The alignment results were then input into an Expectation-Maximization (EM) algorithm for abundance estimation to achieve quantitative analysis of different microorganisms in the mixed sample. To reduce low-level noise interference, an abundance threshold of 0.1% was set; species below this threshold were considered background signals and not included in subsequent statistical analysis.
[0052] (b) unique (species-unique comparison) read calculation To ensure specificity of the detection, each sequencing alignment result was analyzed and its unique attribution was determined in conjunction with species annotations: ① Extract candidate alignments for each read and calculate the alignment score. ② Among the candidate alignments, the best and second-best alignments are selected. If the ratio of the second-best to the best score is less than 0.9, the reads are determined to belong uniquely to this species. ③ If there is only a single credible alignment or no competing species, it is also recorded as a unique alignment; ④ Finally, the number of unique alignment reads for each species is counted and used for species abundance correction and threshold determination; ⑤ Based on the number of uniquely matched reads, set species positive determination thresholds for different pathogenic microorganisms (bacteria ≥10, parasites ≥30, fungi ≥30, viruses ≥3).
[0053] A species is considered positive only if the test result meets the above threshold requirements; results below the threshold will not be reported.
[0054] (c) Quality and batch contamination control To ensure the reliability of the results, the following quality standards were set: Q20 > 85%, Q30 > 75%, and the proportion of effective microbial sequences ≥ 10%. At the batch level, if the proportion of reads of a certain species in a single sample to the total reads of that species in the batch is < 1% (batch_ratio < 0.01), it is judged as contamination or background and is removed.
[0055] (d) Output of results After completing quality and contamination control, a final results table is generated, which includes species name, unique reads, relative abundance, batch filtering status, and clinical annotations, for result interpretation and application.
[0056] (4) Methods for detecting drug resistance genes (a) Sequence alignment and threshold setting After removing human reads, the sequencing sequences were aligned to the drug resistance gene database using minimap2 (v2.28). Alignment results had to meet the following criteria: sequence identity ≥ 90%, coverage ≥ 40%, and effective alignment length ≥ 450 bp. Only reads meeting these criteria were used for subsequent statistical analysis.
[0057] (b) Mutation detection and annotation Based on the detection of drug resistance genes, the generated VCF file is analyzed and compared. Combined with the drug resistance site annotation table, the total coverage depth, mutation depth and mutation frequency of the site are output, and the site is labeled as uncovered, wild type or drug resistance according to the coverage and mutation characteristics.
[0058] (c) Result filtering To ensure the reliability of the detection results, a two-layer filtering standard is set: Gene presence level: only positive results with ≥5 supporting reads are retained; SNP level: the mutation site must simultaneously meet the requirements of coverage depth ≥10 and mutation frequency ≥25%.
[0059] (d) Combined with microbial detection results The relevant drug resistance genes and drug resistance SNP sites will only be reported if the corresponding microorganism is detected in the sample; if the corresponding host microorganism is not detected, it will not be included in the report.
[0060] (e) Output of Results The final results table includes gene name, mutation site, number of supporting reads, coverage depth, mutation frequency, and drug resistance determination.
[0061] Example 2: Timeliness Detection of the Method I. Experimental Data To evaluate the differences in detection cycle time among different detection methods, 15 clinical samples from the reproductive tract were collected. Each sample was divided into three subsets: one subset was detected using the targeted nanopore sequencing method of this invention, with 0.6M reads per sample; another subset was detected using nanopore metagenomic sequencing, with 20M reads per sample; and the last subset was detected using next-generation metagenomic sequencing, with 20M reads per sample. All tests were conducted under identical hardware conditions (12 CPUs). The differences in detection cycle time among the three methods were compared, and the results are shown in Table 1.
[0062] Table 1 Comparison of the timeliness of the method of the present invention in Example 2
[0063] II. Experimental Results As shown in Table 1, the entire process of second-generation metagenomic sequencing takes about 16 hours, nanopore metagenomic sequencing takes about 12 hours, while the targeted nanopore sequencing of this invention can complete the entire detection process in only about 6 hours, shortening the overall detection cycle by more than 2 times.
[0064] Example 3: Accuracy Testing of the Method I. Experimental Data Fifteen samples of clinically diagnosed and PCR-confirmed genital tract infections were collected, including five negative controls (S1-S5) and ten positive samples (S6-S15). Each sample was simultaneously detected using the method of this invention (targeted nanopore sequencing), nanopore metagenomic sequencing, and next-generation metagenomic sequencing. The detection rates of species and drug resistance genes were compared among the three methods to evaluate the performance of the method of this invention (targeted nanopore sequencing) in terms of detection sensitivity and specificity. The results are shown in Table 2.
[0065] Table 2 Comparison of detection accuracy of three different methods
[0066] Table Notes: (1) + indicates detected (positive); - indicates not detected / found (negative).
[0067] (2) Normalized Reads: refers to the value after standardization per 100,000 sequencing reads. The calculation formula is: (Number of unique reads for a species / Total number of reads for a sample) × 10 5 .
[0068] II. Experimental Results (1) Species detection accuracy results As shown in Table 2, the method of the present invention achieved accurate detection in all 15 samples: all target microorganisms were successfully identified in 10 positive samples (S6–S15), and no false positive signals were found in 5 negative samples (S1–S5). The sensitivity and specificity of species detection were both 100% (10 / 10, 5 / 5).
[0069] In contrast, nanopore metagenomic sequencing missed two low-load samples (S12, S13), with a sensitivity of 80% (8 / 10) and a specificity of 100% (5 / 5); next-generation metagenomic sequencing also failed to detect effectively in two low-load samples (S12, S13), with a sensitivity of 80% (8 / 10) and a specificity of 80% (4 / 5), one of which, a negative sample (S3), was reported as a false positive.
[0070] (2) Results of drug resistance gene detection As shown in the table, the method of this invention accurately identified all preset drug resistance genes and mutation sites in all 10 positive samples, and no non-specific detections were observed in the negative samples. The sensitivity and specificity were both 100% (10 / 10, 5 / 5). Nanopore metagenomic sequencing and next-generation metagenomic sequencing missed detections at drug resistance sites (S12, S13, S14, S15) in 4 low-load samples, with a sensitivity of 60% (6 / 10) and a specificity of 100% (5 / 5).
[0071] III. Analysis of Experimental Results In summary, the targeted nanopore sequencing of this invention exhibits optimal performance in both species and drug resistance gene detection, achieving 100% sensitivity and specificity, and maintaining stable detection capability even in low-load / difficult-to-culture samples. Nanopore metagenomic sequencing and next-generation metagenomic sequencing each showed two false negatives in some low-load samples, indicating a slight decrease in sensitivity. Furthermore, next-generation metagenomic sequencing produced a false positive in one negative sample, suggesting potential limitations in specificity. Both nanopore metagenomic sequencing and next-generation metagenomic sequencing showed poor performance in detecting drug resistance genes in some low-load samples, with a sensitivity of only 60%.
[0072] The above description is merely a preferred embodiment and example of the present invention, and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments and examples, any person skilled in the art can make some modifications or alterations to the methods and techniques disclosed above without departing from the scope of the present invention to create equivalent embodiments. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data, characterized in that, Includes the following steps: S1. Construct a microbial reference database, including a microbial comparison database, a microbial annotation database, a drug resistance gene comparison database, and a drug resistance gene annotation database; S2. Obtain the raw sequencing data of the targeted nanopore sequencing of the sample; S3. Raw sequencing data preprocessing: First, the raw sequencing data is subjected to quality control to obtain a high-quality effective sequence set; then, the target sequence is identified and extracted after primer alignment. S4. Using the target sequence extracted in step S2 and the microbial reference database in step S1, perform species comparison and abundance estimation of microorganisms, calculate species unique comparison reads, and output microbial detection results after quality and batch contamination control. S5. After removing human reads from the target sequence extracted in step S2, the sequence of the drug resistance gene is compared with the drug resistance gene reference database in step S1. The sequence of the drug resistance gene is then compared, mutation is detected and annotated, and the results are filtered. Finally, the results are combined with the microbial detection results in step S4 to output the final result.
2. The method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data according to claim 1, characterized in that, The target sequence identification method in step S3 is as follows: S3-1. Set a sequence similarity threshold of ≥75%, perform primer sequence alignment on the sequencing sequences, and retain only the best matching result for each sequence; S3-2, Set primer coverage ≥90%; primer mismatch number ≤3 bp; primer matching sites exist within 15 bp at both ends of the amplified fragment; amplified fragment length is in the range of 200–2000 bp, and only sequences that meet the directionality and site requirements are retained as target sequences.
3. The method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data according to claim 1, characterized in that, In step S4, the abundance threshold for abundance estimation is set to 0.1%.
4. The method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data according to claim 1, characterized in that, In step S4, the method for calculating the species unique alignment read is as follows; S4-1. Extract candidate alignments for each read and calculate the alignment results; S4-2. Select the best and second-best alignments from the candidate alignments. If the ratio of the second-best to the second-best score is less than 0.9, then the reads are determined to belong uniquely to this species. S4-3. If there is only a single reliable alignment or no competing species, it is also recorded as a unique alignment. S4-4. Count the number of unique alignment reads for each species, which is used for species abundance correction and threshold determination. S4-5. Based on the number of uniquely matched reads, set a species positive determination threshold for different pathogenic microorganisms: bacteria ≥10, parasites ≥30, fungi ≥30, and viruses ≥3. Only when the test result meets the above threshold requirements is it determined to be a positive species; results below the threshold will not be reported.
5. The method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data according to claim 1, characterized in that, In the quality and batch contamination control of step S4, the quality standards are set as follows: Q20 > 85%, Q30 > 75%, and the proportion of effective microbial sequences ≥ 10%; the batch standards are set as follows: if the proportion of reads of a certain species in a single sample to the total reads of that species in the batch is < 1%, it is judged as contamination or background and is removed.
6. The method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data according to claim 1, characterized in that, In step S4, the microbial detection results include species name, number of uniquely matched reads, relative abundance, batch filtering status, and clinical annotations.
7. The method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data according to claim 1, characterized in that, In step S5, the sequence alignment of the drug resistance gene must meet the following requirements: sequence consistency ≥ 90%, coverage ≥ 40%, and effective alignment length ≥ 450 bp.
8. The method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data according to claim 7, characterized in that, The result filtering in step S5 adopts a two-layer filtering standard: gene presence / absence level: only positive results supporting read count ≥ 5 are retained; SNP level: the mutation site must simultaneously meet the requirements of coverage depth ≥ 10 and mutation frequency ≥ 25%.
9. The method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data according to claim 1, characterized in that, In the final output of step S5, the relevant drug resistance genes and drug resistance SNP sites are reported only if the corresponding microorganism is detected in the sample; if the corresponding host microorganism is not detected, it is not included in the report.
10. The method for analyzing reproductive tract microorganisms and drug resistance genes based on targeted nanopore sequencing data according to claim 1, characterized in that, In step S5, the drug resistance gene detection results include gene name, mutation site, number of homogenized reads, coverage depth, mutation frequency, and drug resistance determination.