Nanopore sequencing-based identification and analysis system and method for unknown pathogenic microorganisms

By constructing an unknown pathogenic microorganism identification and analysis system based on nanopore sequencing, the challenge of rapid and accurate identification of unknown pathogenic microorganisms in existing technologies has been solved, and efficient and accurate microbial identification and recognition of unknown pathogens have been achieved.

CN119811490BActive Publication Date: 2025-09-16THE NAVAL MEDICAL UNIV OF PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411856323.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-09-16
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

Existing technologies face challenges in quickly and accurately identifying unknown pathogenic microorganisms, especially the high sequencing error rate and data analysis complexity in nanopore sequencing technology.

Method used

A nanopore sequencing-based identification and analysis system for unknown pathogenic microorganisms was designed, including modules such as microbial database construction and update, information standardization and annotation, nanopore sequencing analysis, sequencing data quality control, host sequencing read removal, microbial database alignment and annotation, assembly of sequencing reads of the same genus, and microbial database alignment calculation.

Benefits of technology

Through multi-level data processing and analysis, the impact of nanopore sequencing error rate is reduced, the accuracy and efficiency of microbial identification are improved, and unknown pathogenic microorganisms can be quickly identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811490B_ABST
    Figure CN119811490B_ABST
Patent Text Reader

Abstract

The present invention provides a system and method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing, including: microbial database construction and updating, standardized annotation of microbial information, nanopore sequencing analysis, sequencing data quality control, host sequencing read removal, microbial database alignment annotation, assembly of same-genus sequencing reads, and microbial database alignment calculation; for each representative sequencing read, the representative sequencing read is computationally aligned with any genome sequence in the microbial database; when a first computational alignment condition is met, an output is output that the representative sequencing read is consistent with the species to which the genome sequence meeting the first computational alignment condition belongs; when a second computational alignment condition is met, an output is output that the representative sequencing read is homologous to the species to which the sequence meeting the second computational alignment condition belongs; when the representative sequencing read does not meet both conditions with any sequence in the microbial database, an output is output that the representative sequencing read is not aligned in the microbial database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gene sequencing analysis, and in particular to a system and method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing. Background Art

[0002] With the acceleration of globalization and the increased movement of people and goods across the globe, the disease burden caused by many emerging and unidentified pathogens continues to increase, posing unprecedented challenges to human health. For example, a range of diseases, including SARS, avian influenza, influenza A (H1N1), influenza H7N9, Ebola, and monkeypox, all demonstrate the potential for rapid spread and severe consequences. These diseases not only pose a threat to public health but can also cause economic losses and social panic, prompting public health systems around the world to continuously strengthen their capabilities for pathogen surveillance, rapid identification, and response. In the face of infectious disease threats, rapid and accurate pathogen detection is crucial, providing a crucial basis for effective prevention and control measures.

[0003] Currently, traditional pathogen identification technologies mainly include morphological observation, cell physiological and biochemical characteristics, bacterial culture typing, gene chips and automated microbial analysis systems. These methods usually rely on the cultivation of microorganisms, have a long cycle, and are extremely sensitive to culture conditions, and are prone to false negative or false positive results. In addition, detection methods based on specific primers, probes or antibodies (such as antigen-antibody reactions and PCR detection) rely on prior knowledge of the sequences of known microorganisms when identifying microorganisms, which makes them limited in effectiveness when dealing with unknown or mutated pathogens. The limitations of these technologies make it extremely difficult to quickly and accurately identify pathogens during infectious disease outbreaks, thereby affecting disease control and the timeliness of intervention measures.

[0004] The emergence of next-generation sequencing technologies, particularly nanopore sequencing, has provided a novel solution for microbial identification. Nanopore sequencing offers significant advantages: it requires no culture, shortens detection time, and is suitable for clinical scenarios where rapid response is crucial. It requires no prior knowledge and can identify unknown and mutated pathogens, expanding the scope of detection and enhancing response capabilities to emerging pathogens. Its long sequencing reads allow for the acquisition of complete genomic information, making subsequent genomic analysis and functional studies more accurate.

[0005] Metagenomic technology uses high-throughput sequencing to directly sequence all nucleic acids in a sample, revealing the genetic information of all microorganisms in the sample, including DNA viruses, bacteria, fungi, parasites, etc. It does not rely on prior culture of microorganisms or known sequence information for specific pathogens, making it particularly suitable for the rapid identification of unknown or rare pathogens.

[0006] Therefore, nanopore sequencing technology has demonstrated important value in many fields such as metagenomic sequencing, pathogen monitoring, genome sequencing of new species, and epigenetic research, becoming a powerful tool for rapid identification of microorganisms.

[0007] Despite its significant theoretical advantages, nanopore sequencing technology faces a series of challenges in practical application. First, nanopore sequencing accuracy is only around 95%, resulting in a high error rate, posing significant challenges for species identification and homology analysis. Second, the need to quickly and accurately identify pathogens from complex biological samples demands efficient data analysis algorithms and tools to generate actionable results in a short period of time. Finally, identifying unknown pathogens never seen in nature or pathogens not included in databases, particularly those causing infectious diseases overseas, remains a major challenge in pathogen data analysis. Summary of the Invention

[0008] In response to the problems and shortcomings of the prior art, the present invention provides a system and method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing.

[0009] The present invention solves the above technical problems through the following technical solutions:

[0010] The present invention provides an unknown pathogenic microorganism identification and analysis system based on nanopore sequencing, which is characterized in that it includes a microorganism database construction and update module, a microorganism information standardization and annotation module, a nanopore sequencing analysis module, a sequencing data quality control module, a host sequencing read removal module, a microorganism database comparison and annotation module, a same-genus sequencing read assembly module and a microorganism database comparison and calculation module;

[0011] The microbial database construction and update module is used to obtain multiple microbial species information based on a public database, each microbial species information includes at least a genome sequence, perform integrity assessment on each genome sequence to remove genome sequences of poor quality, remove redundant genome sequences based on the similarity principle for each retained genome sequence, identify and merge multiple genome sequences with a similarity greater than a set threshold, retain only one genome sequence of the same species or similarity after de-redundancy, construct a microbial database based on the microbial species information to which the retained genome sequences belong, and regularly update the microbial database;

[0012] The microbial information standardization and annotation module is used to classify the microbial species information in the microbial database according to the level attributes of kingdom, phylum, class, order, family, genus, and species in sequence, and perform standardized classification and annotation of species. After traversing the information of each microbial species, it issues a manual correction reminder message to remind people to manually correct the automatically annotated microbial database;

[0013] The nanopore sequencing analysis module is used to extract all nucleic acids in the sample to be tested, where the nucleic acid is DNA or RNA. DNA undergoes end-repair or RNA undergoes reverse transcription and end-repair to obtain products. The obtained products are used to build libraries and sequence using a nanopore sequencing platform. The sequencing directory of the nanopore sequencing process is monitored in real time. Whenever a new sequencing sequence file is generated, it is used as the sequencing sequence file to be analyzed until the nanopore sequencing completion condition is met or no new sequencing sequence file is generated within a set time. The sequencing sequence file to be analyzed includes multiple sequencing read segments.

[0014] The sequencing data quality control module is used to perform operations of removing read adapter contamination, removing low-quality reads, and removing short reads on each sequencing read in each sequencing sequence file to be analyzed, thereby obtaining a plurality of sequencing reads retained in the sequencing sequence file to be analyzed;

[0015] The host sequencing read removal module is used to align each sequencing read retained in each sequencing sequence file to be analyzed with the host genome sequence, and extract the sequencing reads that are not aligned to the host genome sequence;

[0016] The microbial database comparison and annotation module is used to compare each extracted sequencing read with the microbial database, and when the similarity between the sequencing read and the read in the genome sequence of a species reaches a set similarity, the sequencing read is annotated with the species level and genus level of the species, and when the similarity between the sequencing read and the read in the genome sequence of a species does not reach the set similarity, the sequencing read is removed, thereby obtaining each retained sequencing read and its species level and genus level species annotation;

[0017] The same-genus sequencing read assembly module is used to treat all sequencing reads annotated to the same genus as a read set for the retained sequencing reads, and to perform consensus sequence assembly on all sequencing reads in the read set of the same genus to obtain several representative sequencing reads, thereby obtaining representative sequencing reads under each genus;

[0018] The microbial database comparison calculation module is used to perform calculation comparison on each representative sequencing read segment with any genome sequence in the microbial database. When a first calculation comparison condition is met, it is output that the species to which the representative sequencing read segment belongs is consistent with the genome sequence that meets the first calculation comparison condition. When a second calculation comparison condition is met, it is output that the species to which the representative sequencing read segment belongs is homologous to the genome sequence that meets the second calculation comparison condition. When the representative sequencing read segment and each genome sequence in the microbial database do not meet the first calculation comparison condition and the second calculation comparison condition, it is output that the representative sequencing read segment is not compared in the microbial database.

[0019] The present invention also provides a method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing, which is characterized in that it includes the following steps:

[0020] S1. Obtain multiple microbial species information based on public databases, each microbial species information including at least a genome sequence, perform integrity assessment on each genome sequence to remove genome sequences of poor quality, remove redundant genome sequences based on the similarity principle for each retained genome sequence, identify and merge multiple genome sequences with a similarity greater than a set threshold, retain only one genome sequence of the same species or similarity after de-redundancy, and construct a microbial database based on the microbial species information to which the retained genome sequences belong;

[0021] S2. Classify each microbial species in the microbial database according to the attributes of kingdom, phylum, class, order, family, genus, and species, and perform standardized classification annotation on each microbial species information. After traversing each microbial species information, issue a manual correction reminder message to remind people to manually correct the automatically annotated microbial database;

[0022] S3. Extract all nucleic acids from the sample to be tested, where the nucleic acid is DNA or RNA. End-repair the DNA or reverse transcribe and end-repair the RNA to obtain products. Use a nanopore sequencing platform to build a library and sequence the obtained products. Monitor the sequencing directory of the nanopore sequencing process in real time. Whenever a new sequencing sequence file is generated, it is used as the sequencing sequence file to be analyzed until the nanopore sequencing completion condition is met or no new sequencing sequence file is generated within a set time. The sequencing sequence file to be analyzed includes multiple sequencing reads.

[0023] S4. For each sequencing sequence file to be analyzed, operations of removing read adapter contamination, removing low-quality reads, and removing short reads are performed on each sequencing read in the sequencing sequence file to be analyzed, thereby obtaining a plurality of sequencing reads retained in the sequencing sequence file to be analyzed;

[0024] S5. For each sequencing read retained in each sequencing sequence file to be analyzed, each sequencing read is aligned with the host genome sequence, and sequencing reads that are not aligned to the host genome sequence are extracted;

[0025] S6. Compare each extracted sequencing read with the microbial database. When the similarity between the sequencing read and the read in the genome sequence of a species reaches a set similarity, annotate the sequencing read with the species level and genus level of the species. When the similarity between the sequencing read and the read in the genome sequence of a species does not reach the set similarity, remove the sequencing read, thereby obtaining each retained sequencing read and its species level and genus level species annotation;

[0026] S7. For the retained sequencing reads, all sequencing reads annotated to the same genus are taken as a read set, and consensus sequence assembly is performed on all sequencing reads in the read set of the same genus to obtain several representative sequencing reads, thereby obtaining representative sequencing reads under each genus;

[0027] S8. For each representative sequencing read, the representative sequencing read is computationally aligned with any genome sequence in the microbial database. When the first computational alignment condition is met, the representative sequencing read is output to be consistent with the species to which the genome sequence that meets the first computational alignment condition belongs. When the second computational alignment condition is met, the representative sequencing read is output to be homologous to the species to which the genome sequence that meets the second computational alignment condition belongs. When the representative sequencing read and each genome sequence in the microbial database do not meet both the first computational alignment condition and the second computational alignment condition, the representative sequencing read is output to be not aligned in the microbial database.

[0028] The positive progress effect of the present invention is:

[0029] This invention provides a system and method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing. The system extracts all nucleic acids from a sample, then performs end-repair on the DNA and reverse transcription and end-repair on the RNA. The resulting products are then used to build libraries and sequence them on a nanopore sequencing platform. The sequencing data is then used to identify the unknown pathogenic microorganisms. This system, which takes advantage of nanopore sequencing's high error rate, can accurately identify microorganisms and facilitate the discovery of unknown pathogens.

[0030] The present invention has the following features: 1. Nanopore sequencing data is compared twice against microbial databases and assembled once at the genus level, thereby reducing the impact of the high error rate of nanopore sequencing and more accurately identifying homologous microorganisms; 2. The analysis cycle is short. After the entire sequencing is completed, only an additional 3-5 hours are required to complete the entire analysis process; 3. Since there are many unknown pathogens in nature or microorganisms not included in the microbial genome database, this system can prompt the presence of unknown pathogenic microorganisms in the sample. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a structural block diagram of an unknown pathogenic microorganism identification and analysis system based on nanopore sequencing according to a preferred embodiment of the present invention.

[0032] Figure 2 Flowchart of the method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing according to a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0034] like Figure 1 As shown, this embodiment provides an unknown pathogenic microorganism identification and analysis system based on nanopore sequencing, which includes a microbial database construction and update module 1, a microbial information standardization and annotation module 2, a nanopore sequencing analysis module 3, a sequencing data quality control module 4, a host sequencing read removal module 5, a microbial database comparison and annotation module 6, a same-genus sequencing read assembly module 7, a microbial database comparison and calculation module 8 and a microbial abundance statistics module 9.

[0035] The microbial database construction and update module 1 is used to obtain multiple microbial species information based on public databases (including NCBI, BV-BRC, EMBL, DDB, etc.), each microbial species information includes at least a genome sequence, and an integrity assessment is performed on each genome sequence to remove genome sequences of poor quality. For each retained genome sequence, redundant genome sequences are removed according to the similarity principle, and multiple genome sequences with a similarity greater than a set threshold (such as 90%) are identified and merged. After de-redundancy, only one of the same species or similar genome sequences is retained, and a microbial database is constructed based on the microbial species information to which the retained genome sequences belong; thereafter, after detecting that the public database has been updated, the updated public database is used to reconstruct the microbial database, and the microbial information standardization and annotation module is called to perform standardized classification annotation on the species of each microbial species information in the reconstructed microbial database.

[0036] In this example, QUAST software (a software used to evaluate the quality of genome assemblies, analyzing the completeness, accuracy, and consistency of genome assemblies) was used to perform an integrity assessment on each genome sequence to remove genome sequences of poor quality, such as genome sequences with an N base ratio greater than 5% or sequences shorter than 1000 bp.

[0037] In this example, genomic sequence de-redundancy was achieved using CD-HIT software (a tool for efficient clustering and de-redundant sequences, used to process large-scale nucleic acid or protein sequence data to identify and merge similar sequences, thereby reducing redundancy in the dataset). Redundant sequences were removed based on a 90% similarity threshold, ensuring that only one representative sequence was retained for the same species or similar genomic sequences, thereby improving the uniqueness and efficiency of the microbial database. The constructed microbial database will be updated based on updates to public databases to ensure that downloaded genomic sequences are consistent with the latest research results and to promptly remove outdated or duplicate genomic sequences with changes in the genomic sequence within the microbial database.

[0038] The microbial information standardization and annotation module 2 is used to classify the information of each microbial species in the microbial database according to the level attributes of kingdom, phylum, class, order, family, genus and species in turn, and perform standardized classification and annotation of species. After traversing the information of each microbial species, it issues a manual correction reminder message to remind people to manually correct the automatically annotated microbial database.

[0039] In this embodiment, due to the inconsistency of microbial annotation information between different public databases, such as a database that only includes the name of the microbial species but lacks analytical information, it is necessary to standardize the species classification information. The NCBIT axonomy classification system is used to standardize microbial species information, including kingdom, phylum, class, order, family, genus, and species. For example, the standardized classification information for Escherichia coli is Bacteria; Proteobacteria; Gammaproteobacteria; Enterobacterales; Enterobacteriaceae; Escherichia; and Escherichia coli.

[0040] In this embodiment, a script can be written for automatic processing, and corresponding annotations can be made according to the genome sequence of the species; if the species name is inconsistent with the database, manual processing is required; the microbial database is automatically annotated and manually corrected by reference to the NCBI database.

[0041] The nanopore sequencing analysis module 3 is used to extract all nucleic acids (DNA or RNA) in the sample to be tested. DNA is end-repaired or RNA is reverse transcribed and end-repaired to obtain products. The obtained products are then used to build libraries (adding adapters required for nanopore sequencing on both sides of the molecules) and sequence them using the nanopore sequencing platform. The sequencing directory of the nanopore sequencing process is monitored in real time. Whenever a new sequencing sequence file is generated, it is used as the sequencing sequence file to be analyzed until the conditions for nanopore sequencing completion are met or no new sequencing sequence files are generated within the set time. Each sequencing sequence file to be analyzed includes multiple sequencing reads.

[0042] In the present embodiment, the nanopore sequencing platform monitors the current changes during the sequencing process as DNA or RNA molecules pass through the nanopore, and converts the current into sequence information in real time and outputs it in the fastq file format. The watchdog package (watchdog package) of python (Python is a widely used high-level programming language) is used to monitor the sequencing directory in real time (sequencing will generate files, and the sequencing directory is the path for saving sequencing files). Whenever a new sequencing sequence file is generated, the system will automatically call the subsequent analysis process (analyzed at any time during the sequencing process, there is only sequence in the sequencing file - at this time, it will also contain many segments of sequence and unique numbers of sequencing sequences), and the sequence is annotated with microorganisms. Until sequencing is completed (there will be special information when normal sequencing is completed, which can be captured) or no new sequencing file is generated for a long time, the nanopore sequencing analysis module automatically exits.

[0043] The sequencing data quality control module 4 is used to perform operations of removing read adapter contamination, removing low-quality reads, and removing short reads for each sequencing read in each sequencing sequence file to be analyzed, thereby obtaining a number of sequencing reads retained in the sequencing sequence file to be analyzed, and performing quality checks on the sequencing reads retained in the sequencing sequence file to be analyzed to confirm that adapter contamination, low-quality reads, and short reads have been effectively removed.

[0044] In this example, during the quality control of nanopore sequencing data, Porechop (a software for processing nanopore sequencing data and removing adapter sequences from raw sequencing reads) was used to remove adapter contamination, and NanoFilt (a software for filtering and quality control of raw nanopore sequencing data) was used to remove low-quality reads and short reads. FastQC (a software for assessing the quality of high-throughput sequencing data) was then used to perform a quality check on the sequenced reads retained after quality control to confirm that adapter contamination, low-quality reads, and short reads had been effectively removed.

[0045] The host sequencing read removal module 5 is used to align each sequencing read retained in each sequencing sequence file to be analyzed with the host genome sequence, and extract the sequencing reads that are not aligned to the host genome sequence.

[0046] Because sequencing reads often contain host-derived reads, which often account for a large proportion, if not removed, microbial signals will be overwhelmed, leading to inaccurate analysis of microbial community diversity and composition. Removing host reads can significantly reduce the amount of data, thereby simplifying subsequent analysis, reducing the computational burden, and improving analysis efficiency and accuracy.

[0047] In this example, minimap2 (a software for fast and efficient alignment of long sequences to reference genomes or other sequences) and samtools (a software for processing alignment files such as SAM and BAM) were used in a sequential order to remove host sequencing reads. First, minimap2 was used to align sequencing reads to the host genome sequence (same species or homologous). Then, samtools was used to extract sequencing reads that did not align to the host genome sequence for subsequent microbial database analysis.

[0048] The microbial database comparison and annotation module 6 is used to compare each extracted sequencing read with the microbial database. When the similarity between the sequencing read and the read in the genome sequence of a species reaches a set similarity, the species level and genus level of the species are used to annotate the sequencing read. When the similarity between the sequencing read and the read in the genome sequence of a species does not reach the set similarity, the sequencing read is removed, so that each retained sequencing read and its species level and genus level species annotation can be obtained.

[0049] In this example, kraken2 (a fast and accurate microbial classification tool) and bracken (a tool for microbial classification and abundance estimation, used in conjunction with kraken2) software were used for microbial database comparison and species annotation. By directly comparing host-depleted sequencing reads with the microbial database, species annotation of bacteria, archaea, eukaryotes, and viruses in the test samples was performed at the species and genus levels. This read-based metagenomic species annotation method is more comprehensive and accurate, eliminating the resource consumption of the assembly and gene prediction processes, enabling the annotation of more low-abundance species, and significantly improving the accuracy of species annotation and relative species abundance.

[0050] The same genus sequencing read assembly module 7 is used to treat all sequencing reads that are aligned and annotated to the same genus as a read set for the retained sequencing reads, and to perform consensus sequence assembly on all sequencing reads in the read set of the same genus to obtain several representative sequencing reads, thereby obtaining representative sequencing reads under each genus.

[0051] Since nanopore sequencing has a high error rate, with an average sequencing error rate of up to 5%, there is a possibility of species annotation errors when sequencing reads are used directly for species annotation. For example, sequencing reads from a certain species are annotated to the genome of a homologous species with a similar genome, resulting in errors in subsequent species analysis. Therefore, to solve the problem of high error rates in nanopore sequencing, this system innovatively introduces the operation of assembling sequencing reads of the same genus. That is, based on the species annotation results at the genus level in the previous step, the sequencing reads of all species aligned to the same genus are extracted separately, and then the canu software (a software specially designed to process long read data and can efficiently generate high-quality genome assemblies) is used to perform consensus sequence assembly to obtain all representative sequencing reads under the genus reads. The representative reads are sequences after assembly, which eliminate random errors in sequencing and better represent the true genome sequence of the species. Subsequent use of representative sequences for microbial species annotation can obtain more accurate results.

[0052] The microbial database comparison calculation module 8 is used to perform calculation comparison on each representative sequencing read segment with any genome sequence in the microbial database. When the first calculation comparison condition is met, it is output that the species to which the representative sequencing read segment belongs is consistent with the genome sequence that meets the first calculation comparison condition. When the second calculation comparison condition is met, it is output that the species to which the representative sequencing read segment belongs is homologous to the genome sequence that meets the second calculation comparison condition. When the representative sequencing read segment and each genome sequence in the microbial database do not meet the first calculation comparison condition and the second calculation comparison condition, it is output that the representative sequencing read segment is not compared in the microbial database, prompting the user that an unknown pathogen may exist.

[0053] The first calculation and comparison condition is: similarity > first similarity threshold, and coverage > first coverage threshold; the second calculation and comparison condition is: second similarity threshold < similarity ≤ first similarity threshold, and second coverage threshold < coverage ≤ first coverage threshold. In this embodiment, the first similarity threshold is 97%, and the first coverage threshold is 95%.

[0054] Similarity = the number of bases that a representative sequencing read has in common with a genome sequence in the microbial database / the total length of the representative sequencing read.

[0055] Coverage = length of the region where the representative sequencing read is mapped to a genome sequence in the microbial database / total length of the representative sequencing read.

[0056] The microbial abundance statistics module 9 is used to count the number of representative sequencing reads of each species under species consistency, and calculate the microbial abundance of a species under species consistency = the number of representative sequencing reads corresponding to the species / the sum of the number of representative sequencing reads of each species under species consistency.

[0057] like Figure 2 As shown, this embodiment also provides a method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing, which includes the following steps:

[0058] Step 101, microbial database construction: Based on a public database, multiple microbial species information is obtained, each microbial species information includes at least a genome sequence, and each genome sequence is evaluated for integrity to remove genome sequences of poor quality. For each retained genome sequence, redundant genome sequences are removed based on the similarity principle, and multiple genome sequences with similarities greater than a set threshold are identified and merged. After redundancy removal, only one genome sequence of the same species or similar is retained, and a microbial database is constructed based on the microbial species information to which the retained genome sequences belong.

[0059] Step 102, standardized annotation of microbial information: Classify each microbial species in the microbial database according to the level attributes of kingdom, phylum, class, order, family, genus, and species in turn, and perform standardized classification annotation of species. After traversing each microbial species information, a manual correction reminder message is issued to remind people to manually correct the automatically annotated microbial database.

[0060] After steps 101 and 102, the microbial database is updated and information is standardized and annotated: after detecting that the public database has been updated, the microbial database is reconstructed using the updated public database, and the information of each microbial species in the reconstructed microbial database is standardized and classified.

[0061] Step 103, nanopore sequencing analysis: All nucleic acids in the sample to be tested are extracted. The nucleic acid is DNA or RNA. DNA is end-repaired or RNA is reverse transcribed and end-repaired to obtain products. The obtained products are used to build libraries and sequence using a nanopore sequencing platform. The sequencing directory of the nanopore sequencing process is monitored in real time. Whenever a new sequencing sequence file is generated, it is used as the sequencing sequence file to be analyzed until the nanopore sequencing completion condition is met or no new sequencing sequence file is generated within a set time. The sequencing sequence file to be analyzed includes multiple sequencing reads.

[0062] Step 104, sequencing data quality control: for each sequencing sequence file to be analyzed, operations of removing read adapter contamination, removing low-quality reads, and removing short reads are performed on each sequencing read in the sequencing sequence file to be analyzed, thereby obtaining a number of sequencing reads retained in the sequencing sequence file to be analyzed; the sequencing reads retained in the sequencing sequence file to be analyzed are quality checked to confirm that adapter contamination, low-quality reads, and short reads have been effectively removed.

[0063] Step 105, host sequencing read removal: for each sequencing read retained in each sequencing sequence file to be analyzed, each sequencing read is aligned with the host genome sequence, and sequencing reads that are not aligned to the host genome sequence are extracted.

[0064] Step 106, microbial database comparison and annotation: Each extracted sequencing read is compared with the microbial database. When the similarity between the sequencing read and the read in the genome sequence of a species reaches a set similarity, the sequencing read is annotated with the species level and genus level of the species. When the similarity between the sequencing read and the read in the genome sequence of a species does not reach the set similarity, the sequencing read is removed, thereby obtaining each retained sequencing read and its species level and genus level species annotation.

[0065] Step 107, assembly of sequencing reads of the same genus: for the retained sequencing reads, all sequencing reads annotated to the same genus are taken as a read set, and consensus sequence assembly is performed on all sequencing reads in the read set of the same genus to obtain several representative sequencing reads, thereby obtaining representative sequencing reads under each genus.

[0066] Step 108, microbial database comparison calculation: For each representative sequencing read, the representative sequencing read is compared with any genome sequence in the microbial database. When the first comparison condition is met, the representative sequencing read is output as being consistent with the species to which the genome sequence that meets the first comparison condition belongs. When the second comparison condition is met, the representative sequencing read is output as being homologous to the species to which the genome sequence that meets the second comparison condition belongs. When the representative sequencing read does not meet both the first and second comparison conditions with each genome sequence in the microbial database, the representative sequencing read is output as not being compared in the microbial database.

[0067] Among them, the first calculation and comparison condition is: similarity > first similarity threshold, and coverage > first coverage threshold; the second calculation and comparison condition is: second similarity threshold < similarity ≤ first similarity threshold, and second coverage threshold < coverage ≤ first coverage threshold.

[0068] Similarity = the number of bases that a representative sequencing read has in common with a genome sequence in the microbial database / the total length of the representative sequencing reads; Coverage = the length of the region where the representative sequencing read is mapped to a genome sequence in the microbial database / the total length of the representative sequencing reads.

[0069] Step 109, microbial abundance statistics: Count the number of representative sequencing reads of each species under species consistency, and calculate the microbial abundance of a species under species consistency = the number of representative sequencing reads corresponding to the species / the cumulative sum of the number of representative sequencing reads of each species under species consistency.

[0070] This analysis system can be used to quickly and accurately identify pathogenic microorganisms, including bacteria, fungi, and viruses. First, nanopore sequencing has a high random error rate. When performing species identification, it is easy to produce erroneous comparison results for homologous microorganisms, leading to subsequent errors in the identification of microbial species. Therefore, we have specially developed an analysis process to resolve species comparison errors. Secondly, the nanopore sequencing platform can sequence in real time, and this analysis system can capture and analyze nanopore sequencing data in real time, shortening the waiting period for sequencing and analysis; there are various algorithms and software for microbial database comparison. We have optimized the analysis process and can quickly obtain analysis results in a short time. Finally, the analysis system will perform microbial analysis at the genus level, analyzing the comparison results of reads in the species included in the genus. If an unknown pathogen exists, its reads will be compared to other species in the same genus, but the comparison results will be relatively poor, thus prompting the user that an unknown pathogen may exist.

[0071] Although specific embodiments of the present invention have been described above, those skilled in the art will appreciate that these are merely illustrative and that the scope of the present invention is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, and such changes and modifications are intended to fall within the scope of the present invention.

Claims

1. A nanopore sequencing-based identification and analysis system for unknown pathogenic microorganisms, characterized in that: It includes a microbial database construction and update module, a microbial information standardization and annotation module, a nanopore sequencing analysis module, a sequencing data quality control module, a host sequencing read removal module, a microbial database alignment and annotation module, a same-genus sequencing read assembly module, and a microbial database alignment calculation module; The microbial database construction and update module is used to obtain multiple microbial species information based on a public database, each microbial species information includes at least a genome sequence, perform integrity assessment on each genome sequence to remove genome sequences of poor quality, remove redundant genome sequences based on the similarity principle for each retained genome sequence, identify and merge multiple genome sequences with a similarity greater than a set threshold, retain only one genome sequence of the same species or similarity after de-redundancy, construct a microbial database based on the microbial species information to which the retained genome sequences belong, and regularly update the microbial database; The microbial information standardization and annotation module is used to classify the microbial species information in the microbial database according to the level attributes of kingdom, phylum, class, order, family, genus, and species in sequence, and perform standardized classification and annotation of species. After traversing the information of each microbial species, it issues a manual correction reminder message to remind people to manually correct the automatically annotated microbial database; The nanopore sequencing analysis module is used to extract all nucleic acids in the sample to be tested, where the nucleic acid is DNA or RNA. DNA undergoes end-repair or RNA undergoes reverse transcription and end-repair to obtain products. The obtained products are used to build libraries and sequence using a nanopore sequencing platform. The sequencing directory of the nanopore sequencing process is monitored in real time. Whenever a new sequencing sequence file is generated, it is used as the sequencing sequence file to be analyzed until the nanopore sequencing completion condition is met or no new sequencing sequence file is generated within a set time. The sequencing sequence file to be analyzed includes multiple sequencing read segments. The sequencing data quality control module is used to perform operations of removing read adapter contamination, removing low-quality reads, and removing short reads on each sequencing read in each sequencing sequence file to be analyzed, thereby obtaining a plurality of sequencing reads retained in the sequencing sequence file to be analyzed; The host sequencing read removal module is used to align each sequencing read retained in each sequencing sequence file to be analyzed with the host genome sequence, and extract the sequencing reads that are not aligned to the host genome sequence; The microbial database comparison and annotation module is used to compare each extracted sequencing read with the microbial database, and when the similarity between the sequencing read and the read in the genome sequence of a species reaches a set similarity, the sequencing read is annotated with the species level and genus level of the species, and when the similarity between the sequencing read and the read in the genome sequence of a species does not reach the set similarity, the sequencing read is removed, thereby obtaining each retained sequencing read and its species level and genus level species annotation; The same-genus sequencing read assembly module is used to treat all sequencing reads annotated to the same genus as a read set for the retained sequencing reads, and to perform consensus sequence assembly on all sequencing reads in the read set of the same genus to obtain several representative sequencing reads, thereby obtaining representative sequencing reads under each genus; The microbial database comparison calculation module is used to perform calculation comparison on each representative sequencing read segment with any genome sequence in the microbial database. When a first calculation comparison condition is met, it is output that the species to which the representative sequencing read segment belongs is consistent with the genome sequence that meets the first calculation comparison condition. When a second calculation comparison condition is met, it is output that the species to which the representative sequencing read segment belongs is homologous to the genome sequence that meets the second calculation comparison condition. When the representative sequencing read segment and each genome sequence in the microbial database do not meet the first calculation comparison condition and the second calculation comparison condition, it is output that the representative sequencing read segment is not compared in the microbial database.

2. The unknown pathogenic microorganism identification and analysis system based on nanopore sequencing according to claim 1, characterized in that: The first calculation comparison condition is: similarity > first similarity threshold, and coverage > first coverage threshold; Second calculation and comparison condition: second similarity threshold < similarity ≤ first similarity threshold, and second phase coverage threshold < coverage ≤ first coverage threshold; Similarity = the number of bases that a representative sequencing read has in common with a genome sequence in the microbial database / the total length of the representative sequencing read; Coverage = length of the region where the representative sequencing read is mapped to a genome sequence in the microbial database / total length of the representative sequencing read.

3. The unknown pathogenic microorganism identification and analysis system based on nanopore sequencing according to claim 1, characterized in that: The system also includes a microbial abundance statistics module, which is used to count the number of representative sequencing reads of each species under species consistency, and calculate the microbial abundance of a species under species consistency = the number of representative sequencing reads corresponding to the species / the cumulative sum of the number of representative sequencing reads of each species under species consistency.

4. The unknown pathogenic microorganism identification and analysis system based on nanopore sequencing according to claim 1, characterized in that: The sequencing data quality control module is used to perform quality checks on the sequencing reads retained in the sequencing sequence file to be analyzed to confirm that adapter contamination, low-quality reads, and short reads have been effectively removed.

5. The nanopore sequencing-based unknown pathogenic microorganism identification and analysis system according to claim 1, wherein: The microbial database construction and update module is used to reconstruct the microbial database using the updated public database after detecting that the public database has been updated, and call the microbial information standardization and annotation module to perform standardized classification annotation on the information of each microbial species in the reconstructed microbial database.

6. A method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing, characterized in that: It includes the following steps: S1. Obtain multiple microbial species information based on public databases, each microbial species information including at least a genome sequence, perform integrity assessment on each genome sequence to remove genome sequences of poor quality, remove redundant genome sequences based on the similarity principle for each retained genome sequence, identify and merge multiple genome sequences with a similarity greater than a set threshold, retain only one genome sequence of the same species or similarity after de-redundancy, and construct a microbial database based on the microbial species information to which the retained genome sequences belong; S2. Classify each microbial species in the microbial database according to the attributes of kingdom, phylum, class, order, family, genus, and species, and perform standardized classification annotation on each microbial species information. After traversing each microbial species information, issue a manual correction reminder message to remind people to manually correct the automatically annotated microbial database; S3. Extract all nucleic acids from the sample to be tested, where the nucleic acid is DNA or RNA. End-repair the DNA or reverse transcribe and end-repair the RNA to obtain products. Use a nanopore sequencing platform to build a library and sequence the obtained products. Monitor the sequencing directory of the nanopore sequencing process in real time. Whenever a new sequencing sequence file is generated, it is used as the sequencing sequence file to be analyzed until the nanopore sequencing completion condition is met or no new sequencing sequence file is generated within a set time. The sequencing sequence file to be analyzed includes multiple sequencing reads. S4. For each sequencing sequence file to be analyzed, operations of removing read adapter contamination, removing low-quality reads, and removing short reads are performed on each sequencing read in the sequencing sequence file to be analyzed, thereby obtaining a plurality of sequencing reads retained in the sequencing sequence file to be analyzed; S5. For each sequencing read segment retained in each sequencing sequence file to be analyzed, each sequencing read segment is aligned with the host genome sequence, and the sequencing read segments that are not aligned to the host genome sequence are extracted; S6. Compare each extracted sequencing read with the microbial database. When the similarity between the sequencing read and the read in the genome sequence of a species reaches a set similarity, annotate the sequencing read with the species level and genus level of the species. When the similarity between the sequencing read and the read in the genome sequence of a species does not reach the set similarity, remove the sequencing read, thereby obtaining each retained sequencing read and its species level and genus level species annotation; S7. For the retained sequencing reads, all sequencing reads annotated to the same genus are taken as a read set, and consensus sequence assembly is performed on all sequencing reads in the read set of the same genus to obtain several representative sequencing reads, thereby obtaining representative sequencing reads under each genus; S8. For each representative sequencing read, the representative sequencing read is computationally aligned with any genome sequence in the microbial database. When the first computational alignment condition is met, the representative sequencing read is output to be consistent with the species to which the genome sequence that meets the first computational alignment condition belongs. When the second computational alignment condition is met, the representative sequencing read is output to be homologous to the species to which the genome sequence that meets the second computational alignment condition belongs. When the representative sequencing read does not meet both the first computational alignment condition and the second computational alignment condition with each genome sequence in the microbial database, the representative sequencing read is output to be not aligned in the microbial database.

7. The method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing according to claim 6, wherein: The first calculation comparison condition is: similarity > first similarity threshold, and coverage > first coverage threshold; Second calculation and comparison condition: second similarity threshold < similarity ≤ first similarity threshold, and second phase coverage threshold < coverage ≤ first coverage threshold; Similarity = the number of bases that a representative sequencing read has in common with a genome sequence in the microbial database / the total length of the representative sequencing read; Coverage = length of the region where the representative sequencing read is mapped to a genome sequence in the microbial database / total length of the representative sequencing read.

8. The method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing according to claim 6, wherein: S9. Count the number of representative sequencing reads of each species under species consistency, and calculate the microbial abundance of a species under species consistency = the number of representative sequencing reads corresponding to the species / the sum of the number of representative sequencing reads of each species under species consistency.

9. The method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing according to claim 6, wherein: In step S4, the quality of the sequencing reads retained in the sequencing sequence file to be analyzed is checked to confirm that adapter contamination, low-quality reads, and short reads have been effectively removed.

10. The method for identifying and analyzing unknown pathogenic microorganisms based on nanopore sequencing according to claim 6, wherein: After steps S1 and S2, when it is detected that the public database has been updated, the updated public database is used to reconstruct the microbial database, and the information of each microbial species in the reconstructed microbial database is labeled with standardized species classification.

Citation Information

Patent Citations

  • Pathogen metagenome analysis method based on nanopore sequencing data

    CN116705160A

  • Targeted pathogen nanopore sequencing rapid analysis method based on multiple PCR

    CN116779036A