A method for identifying and analyzing unknown pathogenic microorganisms
By constructing a high-quality database for virus and bacteria identification, and combining it with virus host prediction models and pathogenicity prediction, the problem of accuracy in identifying unknown pathogenic microorganisms has been solved, enabling rapid identification of unknown pathogenic viruses and bacteria, and supporting epidemic prevention and control and vaccine development.
Patent Information
- Application Number
- CN202511648428.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-11-11
AI Technical Summary
Existing technologies are insufficient for the rapid and accurate identification of unknown pathogenic microorganisms. Public databases suffer from low integrity, high contamination, and naming and labeling errors, and it is difficult to distinguish the genomic sequence differences between unknown microorganisms and closely related species.
We constructed a high-quality virus and bacteria identification database, combined with machine learning and deep learning models, and used metagenomic sequencing data to predict virus hosts and identify bacteria. We then used the virus identification database and host prediction model to identify potential unknown viruses and bacteria, and combined pathogenicity prediction to screen out unknown pathogens.
It improves the accuracy of identifying unknown pathogenic microorganisms, reduces the risk of errors caused by low-quality sequences in the database, provides a method for rapid identification of unknown pathogenic viruses and bacteria, and provides a data foundation for epidemic prevention and control and vaccine development.
Smart Images

Figure CN121096438B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of metagenomic bioinformatics analysis technology, specifically relating to a method for identifying and analyzing unknown pathogenic microorganisms. Background Technology
[0002] Since the beginning of the 21st century, outbreaks caused by emerging pathogens such as Severe Acute Respiratory Syndrome Coronavirus (SARS-CoV), Middle East Respiratory Syndrome Coronavirus (MERS-CoV), and Severe Acute Respiratory Syndrome Coronavirus Type 2 (SARS-CoV-2) have seriously threatened human life and health. Rapid identification of emerging, unknown pathogens is crucial for containing their spread and is essential for public health and clinical diagnosis and treatment.
[0003] Traditional methods for identifying pathogenic microorganisms include morphological observation, physiological and biochemical identification, and serological testing, which mainly rely on microbial isolation and culture. The accuracy of the results is easily affected by culture conditions. Modern molecular techniques have developed rapidly, with 16S RNA sequencing and real-time quantitative PCR commonly used for pathogen detection. However, these methods are only applicable to known species and cannot detect unknown pathogenic microorganisms.
[0004] Metagenomics can directly analyze the genetic information of all microorganisms in a sample, including viruses, bacteria, and fungi. It does not rely on culture or known species sequence information, making it suitable for identifying unknown pathogens. In recent years, metagenomics has been widely used in clinical pathogen detection.
[0005] In practical applications, metagenomics technology still faces challenges in the detection of unknown pathogens. First, published public virus databases suffer from low completeness, high contamination, and naming / labeling errors, affecting the accuracy of virus identification. Second, the ICTV, as the authoritative body for virus classification and nomenclature, has a much smaller database than the one covering viral diversity; its recognized viral species do not include unverified, unclassified, or provisional viruses in other databases, highlighting the urgent need to establish a high-quality, high-coverage viral identification genome database. Third, accurately obtaining sequence information of unknown microorganisms and rapidly identifying unknown pathogens based on metagenomic sequencing data to differentiate them from closely related species remains a significant challenge in data analysis. Summary of the Invention
[0006] To overcome the shortcomings of existing analytical methods, this invention provides a method for constructing a microbial identification genome database and a method for identifying and analyzing unknown pathogenic microorganisms, which can identify unknown pathogenic microorganisms based on metagenomic high-throughput sequencing data.
[0007] The technical solution adopted in this invention is:
[0008] A method for identifying and analyzing unknown pathogenic microorganisms, characterized in that the method includes the following steps:
[0009] S1: Constructing a virus identification database: Obtain and integrate viral genome sequences from ICTV, BV-BRC, and NCBI Virus RefSeq databases, remove redundant and low-quality sequences, and screen high-quality genomes to establish a virus identification database;
[0010] S2: Constructing a virus host prediction model: Download and integrate virus-host relationship data from the ICTV, NCBI, VRION, and VirhostDB databases, extract single nucleotide frequency, dinucleotide frequency, and codon bias of virus genome sequences whose hosts are vertebrates as feature values, and train and construct a virus host prediction model.
[0011] S3: Constructing a bacterial identification database:
[0012] Bacterial genome information is collected from NCBI, and bacterial nomenclature and type strains are cross-validated using LPSN and IJSEM. Qualified bacterial genomes are obtained, and low-quality and erroneous genomes are removed through quality control. 16S rRNA gene sequences are extracted, and a bacterial identification database is established. The execution order of S3 is any one of the following: before S1, or after S1 and before S4.
[0013] S4: Conduct identification and analysis of unknown pathogenic microorganisms, perform quality control and assembly of metagenomic sequencing data, and obtain preliminary assembled sequences;
[0014] The initial assembly sequence uses machine learning or deep learning models to identify potential viral sequences, and then performs quality assessment. It uses a virus identification database to perform genome analysis of the most similar species to screen out potential unknown viral genome sequences. The host prediction model is used to predict the host of potential unknown viral genome sequences. When primates appear in the top three predicted hosts, it is judged as an unknown pathogenic virus.
[0015] The preliminary assembled sequences were binned and assembled to obtain metagenomic assembled genomes (MAGs). High-quality MAG sequences were screened for bacterial species identification analysis to identify potential unknown bacterial genome sequences. The pathogenicity of the potential unknown bacterial genome sequences was predicted, and the prediction result was that they were pathogenic to humans, thus identifying them as unknown pathogenic bacteria.
[0016] Furthermore, the method is preferably performed according to the following steps:
[0017] S1: Constructing a virus identification database:
[0018] Viral genome sequence data were downloaded from the International Committee on Taxonomy of Viruses (ICTV), BV-BRC database, and NCBI Virus RefSeq database. After merging and removing redundancy, the viral genome sequences were quality-assessed, and low-completeness, highly contaminated, and misclassified viral genome sequences were removed. High-quality genomes were screened to construct a virus identification database.
[0019] Furthermore, the quality assessment includes calculating genome sequence integrity, contamination rate, and maximum similarity viral sequence, removing sequences with integrity <90%, contamination rate >5%, or maximum similarity viral species that are inconsistent with the genome sequence tag, screening for high-quality genomes, retaining only one representative genome sequence for each virus or viral subtype, and constructing a virus identification database.
[0020] Furthermore, step S1 is preferably performed according to the following steps:
[0021] S1-1: Obtain all virus classification information and pattern representative genome sequence numbers from the International Committee on Taxonomy of Viruses (ICTV) website;
[0022] S1-2: Download the virus list and genome sequence numbers from the BV-BRC database;
[0023] S1-3: Download the virus list and genome sequence numbers from the NCBI Virus RefSeq database;
[0024] S1-4: Merge and remove redundancy from the virus data downloaded from ICTV, NCBI Virus RefSeq and BV-BRC, and download the complete viral genome sequence from the NCBI Genbank and RefSeq databases according to the genome sequence number;
[0025] S1-5: Quality assessment of viral genome sequences, removal of sequences with integrity <90%, contamination rate >5%, or inconsistencies between the most similar viral species and the genome sequence tag, screening for high-quality genomes, retaining only one representative genome sequence for each virus or viral subtype, and constructing a virus identification database.
[0026] Furthermore, in step S1-1, all virus classification information and pattern representative genome sequence numbers are obtained from ICTV, specifically including classification information of virus species kingdom, phylum, class, order, family, genus and species, as well as virus name, virus name abbreviation, virus Genbank sequence number, genome coverage, genome type and host information;
[0027] In steps S1-2, viral genomes and genomic metadata are downloaded from the BV-BRC database. The obtained viral data includes genome sequence numbers, genome information, isolate information, host information, phenotypic information, etc. from various sources.
[0028] In steps S1-3, the viral genome sequence number, viral species or strain name, species name, genome sequence integrity, geographical location, host, collection time, and genome type information are obtained from the NCBI Virus RefSeq database.
[0029] In steps S1-4, based on virus name, virus name abbreviation, and virus classification information, virus species information from ICTV, NCBIVirusRefSeq, and BV-BRC is integrated, and virus species from NCBI and BV-BRC that are redundant with those from ICTV are removed, while non-repeating virus species information is retained; for the de-redundant virus species information, complete viral genome sequences are downloaded from the NCBI Genbank and RefSeq databases according to the corresponding genome sequence numbers.
[0030] In steps S1-5, the viral genome sequences are quality-assessed, and high-quality genome sequences with integrity ≥90%, contamination rate ≤5%, and maximum similarity species consistent with the tag are selected. Only one representative genome sequence is retained for each virus or viral subtype, and a viral identification genome database is constructed.
[0031] The CheckV tool can be used for quality assessment.
[0032] Only one representative genome sequence is retained for each virus or viral subtype. The model representative genome sequence of the viral species published by ICTV is preferred. For viral species that have not been published by ICTV or for low-quality viral species whose model representative genome sequences do not meet the above quality standards, the genome sequences of the viral species from the NCBI and BV-BRC databases are evaluated for quality. High-quality genome sequences with integrity ≥90%, contamination rate ≤5%, and maximum similarity species consistent with the label are selected as the representative genome sequence.
[0033] The virus identification database constructed in step S1 includes data from ICTV, NCBI, and BV-BRC databases. The collected information includes virus classification (kingdom, phylum, class, order, family, genus, species), virus name, virus name abbreviation, virus Genbank sequence number, genome coverage, genome type, and host information.
[0034] S2: Constructing a virus host prediction model:
[0035] S2-1: Download and integrate virus-host relationship data from the ICTV, NCBI, VRION, and VirhostDB databases, select hosts that are vertebrates, and include the virus genome in the virus-host relationship data in the virus identification database constructed in step S1, integrate the genome sequences of virus species in the virus identification database, and generate a training dataset.
[0036] S2-2: Based on the training dataset, extract the single nucleotide frequency, dinucleotide frequency, and codon bias of the viral genome sequence as viral feature data;
[0037] S2-3: Classify the vertebrate hosts in the training dataset and use them as labels for the training dataset;
[0038] Furthermore, the classification of vertebrate hosts is generally based on the host classification name and NCBI classification status. They can be classified according to class, order, etc., into primates, bats, birds, fish, reptiles, etc., and can be divided into non-overlapping labels.
[0039] S2-4: Based on the feature value data and labels of the training dataset, the Gradient Boosting Machine (GBM) algorithm is used to train and build a virus host prediction model.
[0040] Furthermore, when constructing the prediction model, the holdout test is used to initially evaluate the accuracy of the prediction model, generally requiring an accuracy greater than 85%.
[0041] S2-5: Using the EID2 (ENHanCEd Infectious Diseases) database as the test set, test the accuracy of the virus host prediction model constructed in S2-4. The prediction accuracy for the viral genome sequence of primate hosts should be >90%.
[0042] Furthermore, step S2-5 is preferably performed as follows: using the EID2 database as the test set, extracting the feature values of the viral genome sequence according to step S2-2, using the viral host prediction model constructed in S2-4, predicting the host based on the feature values, and calculating the prediction accuracy. For viral genome sequences whose host is a primate, the host prediction model accuracy is required to be greater than 90%.
[0043] If the accuracy does not meet the requirements, the prediction model will be further trained and optimized until the accuracy meets the requirements.
[0044] S3: Constructing a bacterial identification database:
[0045] Bacterial genomes were collected from NCBI, and bacterial nomenclature and type strains were cross-validated using LPSN, Bergey's Manual of Bacterial Systematics, 2nd Edition, and IJSEM. Qualified bacterial genes were screened, and highly contaminated and low-complete genomes were removed. 16S rRNA gene sequences were extracted, and erroneously named and misidentified genomes were removed to establish a bacterial identification database.
[0046] The method for constructing a bacterial identification database is disclosed in CN 112863606 A, and can be constructed according to the method in this patent.
[0047] Specific methods include:
[0048] S3-1: Collection of microbial information.
[0049] S3-1-1: Collect bacterial genomes from NCBI (National Center for Biotechnology Information), including bacterial genome sequences, “strains,” “culture collections,” “clones,” and “annotations” metadata, and screen for model strains from them;
[0050] S3-1-2: Obtain a valid published list of bacterial names and strain types from the LPSN (The List of Prokaryotic names with Standing in Nomenclature);
[0051] S3-1-3: Consult the articles in Bergey's Manual of Bacterial Systematics, 2nd Edition and IJSEM (International Journal of Systematic and Evolutionary Microbiology) to determine the eligible species names and corresponding type strain strain numbers;
[0052] S3-1-4: Based on the species name and strain number in S3-1-3, the genome sequences and meta-information obtained in S3-1-1 are screened, and qualified bacterial genomes are entered into the database for management.
[0053] S3-2: Quality control of genome sequences in the database.
[0054] S3-2-1: Use CheckM (v1.0.18) based on pedigree marker gene sets to assess the integrity and contamination rate of each genome, and delete genomes with a contamination rate >5% or integrity less than 90%;
[0055] S3-2-2: For annotated genomes, the 16S rRNA gene sequence is directly extracted. For unannotated genomes, RNAmmer (v1.2) is used for extraction. These 16S rRNA gene sequences are compared with the LTP database (version: LTPs132\u SSU) to check for consistency. If there is inconsistency at the genus level, the IJSEM and related literature for this species are consulted to investigate whether the name has been changed. If the name has not been changed, it is determined to be contamination, and the contaminating genome is removed.
[0056] S3-2-3: Perform pairwise ANI (Average Nucleotide Identity) calculations between any two genomes to infer mislabeled genomes. Investigate whether renaming has been performed by reviewing IJSEM and related literature for this species. If not, identify the genomes as mislabeled and remove them. Delete genomes with incorrect clustering.
[0057] S4: Pathogenic microorganism identification and analysis:
[0058] S4-1: Perform quality control on metagenomic sequencing data to remove low-quality sequencing data and host contamination data, and obtain the quality-controlled data;
[0059] In step S4-1, the sequencing data undergoes quality control, preferably performed according to the following steps:
[0060] Quality control was performed on metagenomic sequencing data to remove sequencing adapter sequence contamination, low-quality and short sequence data, and host sequence contamination data to obtain quality-controlled data.
[0061] Generally, Trimmomatic and Fastqc tools can be used for quality control statistics to remove sequencing data contaminated with sequencing adapter sequences, low-quality sequences (sequences with an average sliding window quality value of less than 15) and short sequences (<75bp).
[0062] Data to remove host sequence contamination is typically obtained by using the BWA tool to compare sequencing data with host species genome databases.
[0063] S4-2: After quality control, the data is initially assembled and spliced to obtain the initial assembled sequence (contig).
[0064] Generally, based on the de Bruijn graph algorithm, the data after quality control is initially assembled and spliced to generate a preliminary assembled sequence (contig).
[0065] The MetaWrap (SPAdes) tool can be used for initial assembly and splicing;
[0066] S4-3: Identify unknown pathogenic virus species, wherein the execution order of S4-4 is any of the following: before S4-3, after S4-2, or after S4-3;
[0067] S4-3-1: Perform viral sequence identification analysis on the preliminary assembled sequence (contig) to identify potential viral sequences;
[0068] In step S4-3-1, a viral gene recognition tool is used to identify potential viral sequences (including DNA viruses and RNA viruses) from the preliminary assembled sequence. Generally, Virsorter, DeepVirFinder, and geNomad tools can be used to identify potential viral sequences. If any tool identifies a virus, it is considered a potential viral sequence.
[0069] S4-3-2: Perform quality assessment and maximum similarity species genome analysis on potential viral sequences, calculate the average amino acid similarity (AAI) and average nucleotide similarity (ANI) of potential viral sequences and all genomes in the virus identification database established in step S1; screen potential viral genome sequences with integrity ≥50%, contamination rate <5%, and maximum AAI and maximum ANI both less than 95% as potential unknown viral genome sequences;
[0070] Single-sequence viral genome assessment tools, such as CheckV, are generally used for quality assessment and preliminary identification of closely related species.
[0071] S4-3-3: Using the ViralHostPredictor tool and the virus host prediction model constructed in step S2, predict the host of the potential unknown virus genome sequence obtained in step S3-3-2. When the first three hosts in the prediction results are primates, it is judged as an unknown pathogenic virus.
[0072] S4-4: Identify unknown pathogenic bacterial species. The execution order of S4-4 and S4-3 can be interchanged or performed simultaneously.
[0073] S4-4-1: The preliminary assembled sequence is binned and assembled to obtain the metagenomic assembled genome MAG sequence;
[0074] The hybrid binning algorithm is generally used, and the MetaWARP binning assembly tool can be used for binning assembly.
[0075] S4-4-2: Redundancy removal and species classification annotation are performed on MAG sequences, followed by genome quality assessment. MAG sequences with integrity ≥50% and contamination rate <5% are selected as high-quality MAG sequences.
[0076] Microbial genome redundancy removal tools, such as the dRep tool, are generally used to remove redundancy from MAG sequences;
[0077] Species-level metagenomic analysis tools, such as Sylph, are used for species classification annotation.
[0078] The CheckM tool can be used to assess the genomic quality of MAG sequences;
[0079] S4-4-3: Use bacterial identification databases to perform bacterial species identification analysis, calculate the ANI of high-quality MAG sequences, and screen sequences with a maximum ANI of less than 95% as potential unknown bacterial genome sequences;
[0080] The ANI of high-quality MAG sequences can be calculated using the fastANI and MIGA tools.
[0081] Furthermore, calculating the ANI of a high-quality MAG sequence involves comparing the high-quality MAG sequence with all genomes in the bacterial identification and typing analysis genome database to calculate the ANI.
[0082] S4-4-4: Using machine learning methods and tools, the pathogenicity of potential unknown bacterial genome sequences is predicted. If the prediction result is "Human Pathogenic", it is identified as an unknown pathogenic bacterium.
[0083] The PathogenFinder tool is preferred for predicting the pathogenicity of potentially unknown bacterial genome sequences.
[0084] The method of the present invention may further include step S5: outputting results:
[0085] Report the genome sequence, AAI / ANI value, host or pathogenicity prediction results of unknown pathogenic viruses and bacteria.
[0086] The beneficial effects of this invention are as follows:
[0087] (1) The present invention constructs a high-quality virus identification database, which can reduce the risk of errors in the identification and analysis of viral pathogens caused by low-quality sequences in public databases.
[0088] (2) The present invention constructs a virus host prediction model, which can identify and analyze unknown viruses and quickly predict whether they are potential unknown pathogenic viruses based on machine learning methods. It can be applied to the analysis of rapid unknown viral pathogens and provides technical foundation and methodological support for epidemic prevention and control.
[0089] (3) The present invention can utilize a high-quality bacterial identification database to detect unknown bacteria and identify whether they are unknown pathogenic bacterial pathogens, thereby reducing the cost of manual analysis.
[0090] (4) Based on metagenomic sequencing data, this invention can obtain the genome sequence of unknown pathogenic microorganisms, providing a data foundation for further vaccine research and development and detection method development. Attached Figure Description
[0091] The accompanying drawings form part of the specification and are used together with the embodiments of the present invention to explain the invention, but do not constitute a limitation thereof.
[0092] Figure 1 This invention relates to a method for constructing a viral identification genome database.
[0093] Figure 2 This is a flowchart of the method for identifying and analyzing unknown pathogenic microorganisms according to the present invention. Detailed Implementation
[0094] To understand and implement this invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, but the scope of protection of this invention is not limited thereto.
[0095] Example 1
[0096] like Figure 1 As shown, the steps for constructing the virus identification genome database of the present invention are as follows:
[0097] Acquire virus classification data and genome sequence data
[0098] (1) Obtain all virus species information from the official website of the International Committee on Taxonomy of Viruses (ICTV) (https: / / ictv.global / ), including the classification information of virus species kingdom, phylum, class, order, family, genus and species, as well as virus name, virus name abbreviation, virus Genbank sequence number, genome coverage, genome type and host information;
[0099] (2) Select the VIRUSES module from the BV-BRC (Bacterial and Viral Bioinformatics Resource Center) database to download viral genomes and genomic metadata. The obtained viral data includes genome sequence numbers from multiple sources, genome information, isolate information, host information, phenotypic information, etc.
[0100] (3) Obtain the viral genome sequence number, viral species or strain name, species name, genome sequence integrity, geographical location, host, collection time, and genome type information from the NCBI (National Center for Biotechnology Information) Virus database (RefSeq);
[0101] (4) Based on virus name, virus abbreviation, and virus classification information, integrate the virus species information from ICTV and NCBI BV-BRC sources, remove virus species from NCBI and BV-BRC sources that are duplicated and redundant with ICTV sources, and retain non-duplicate virus species information.
[0102] (5) Based on the deredundant virus species and corresponding genome sequence number information obtained from ICTV, NCBI and BV-BRC, download the corresponding genome sequence of the species from the NCBI Genbank and RefSeq databases;
[0103] (6) The quality of the integrated and deduplicated viral genome sequences was assessed using the CheckV tool. High-quality genome sequences with a completeness greater than 90%, a contamination rate ≤ 5%, and the greatest similarity species consistent with the label were selected. For the same viral species or viral subtype, only one viral genome sequence was retained as the representative genome sequence of the viral species. The model representative genome sequence specified by ICTV was given priority. For low-quality viral species whose model representative genome sequences did not meet the above quality standards but were not published by ICTV, the quality of the genome sequences from NCBI and BV-BRC sources of the viral species was assessed. The quality of the genome sequences from NCBI and BV-BRC databases of the viral species was assessed, and high-quality genome sequences with a completeness ≥ 90%, a contamination rate ≤ 5%, and the greatest similarity species consistent with the label were selected as the representative genome sequences.
[0104] (7) Based on high-quality representative viral genome sequences, combined with information such as viral classification and host information, construct a viral identification genome database.
[0105] In this embodiment, a total of 18,747 (Refseq) and 16,236 (ICTV) viral genome sequences were downloaded. 14,466 redundant sequences were removed, and 2,692 low-quality viral genomes were removed, including 1,947 genomes with a completeness of less than 90%, 19 genomes with a contamination rate of more than 5%, and 726 genomes with species naming errors. This resulted in a virus identification database containing 8,908 viral species and 17,825 genome sequences.
[0106] Example 2
[0107] Constructing a virus host prediction model includes the following steps:
[0108] 1) Download and integrate virus-host relationship data from the ICTV, NCBI, VRION, and VirhostDB databases, select hosts that are vertebrates, and ensure that the virus genome is included in the virus-host relationship data in the virus identification database constructed in step S1. Integrate the genome sequences of virus species in the virus identification database to generate a training dataset.
[0109] 2) Based on the training dataset, extract the single nucleotide frequency, dinucleotide frequency, and codon bias of the viral genome sequence as viral feature data;
[0110] 3) Classify all vertebrate hosts in the training dataset (primates, bats, birds, fish, reptiles, etc.) according to the host classification name and NCBI classification data, and use them as labels for the training dataset;
[0111] 4) Based on the feature values and labels of the training dataset, a virus host prediction model is trained and constructed using the Gradient Boosting Machine (GBM) algorithm. The holdout test is then used to conduct a preliminary evaluation of the accuracy of the host prediction model, requiring an accuracy greater than 85%.
[0112] 5) The accuracy of the viral host prediction model was evaluated using viral genome sequences from EID2 (ENHanCEd Infectious Diseases) as the test set. The results showed that the prediction model had an accuracy of over 90% for viral sequences whose hosts are primates.
[0113] In this embodiment, a training dataset was constructed based on 12,795 virus-vertebrate host relationships from databases including ICTV, NCBI, VRION, and VirhostDB, along with 8,331 high-quality genome sequences and 5,442 virus species from the corresponding virus identification database. Based on the genome sequences in the training dataset, 3,950 feature values were extracted. The vertebrate host species classification labels in the training dataset comprise 14 categories, including Afrotheria, Amphibians, Carnivores, Eulipotyphla and Scandentia, Galloanserae, Marsupials, Primates, Ungulates, Bats, Birds, Fishes, Reptiles, Rodents and Lagomorphs. Rabbits and Xenarthrans were identified. The Gradient Boosting Machine (GBM) algorithm was used to build the prediction model. The host prediction model was tested based on 3,591 EID2 (ENHanCEd Infectious Diseases) virus sequences. For virus sequences whose hosts are primates, the host prediction model accuracy was 91.67%.
[0114] Example 3: Construction of a bacterial identification database
[0115] The bacterial identification database was constructed according to the method disclosed in CN 112863606 A, specifically including the following steps:
[0116] 1) Collect bacterial genome information from NCBI, and at the same time collect meta-information including “strains”, “culture collection”, “clones” and “annotations”, establish a correspondence table between genome information and meta-information, and clarify the source of each genome.
[0117] 2) Obtain a list of validly published bacterial names and types from LPSN, and consult Bergey's Handbook of Archaea and Bacteria Systematics and articles from IJSEM. After screening, obtain the genomes of qualified strains and add them to the database for management.
[0118] 3) After screening the bacterial genome sequences entered into the database, the bacterial identification database was constructed. The self-developed Python program MDBacQCTools was used to filter out erroneous and low-quality bacterial genomes.
[0119] The quality control steps using MDBacQCTools are as follows: First, CheckM (v1.0.18) based on phylogenetic marker gene sets is used to assess the integrity and contamination rate of each genome. Genomes with contamination exceeding 5% or integrity below 90% are removed from the database. Second, 16S rRNA gene sequences extracted from the genomes are extracted using RNAmmer (v1.2) and compared with the LTP database (version: LTPs132\u SSU) to check for consistency. Genome sequences with any divergence at the genus level are removed. Finally, pairwise ANI calculations are performed between any two genomes to infer mislabeled genomes. Based on the ANI values, clustering tree analysis is performed on different genome sources (≥2 genomes) with the same species name in the context of their respective genus, removing obvious outliers. For genomes with only a single genome, aberrant genomes are identified by constructing a clustering tree with the entire genus as the background.
[0120] The bacterial identification database extracts 16S rRNA sequences from the genomes of model strains in the genome database, extracting 25,209 16S rRNA gene sequences of model strains with a length greater than 300. These sequences are then merged with the 16S rRNA gene sequences contained in the LTP to form a 16S rRNA sequence database that can be used for identification.
[0121] Statistical analysis was performed on bacterial genomes in the constructed bacterial identification database.
[0122] Based on the strain lists retrieved from LPSN and IJSEM, 13,161 bacterial genome sets were collected from NCBI. 331 genomes were removed (mainly those with integrity less than 90% or contamination rates greater than 5%). Twelve assembly results with abnormal GC content or genome size were excluded. Genomes showing significant differences in 16S rRNA were also removed. Based on paired ANI value clustering plots and IJSEM literature, when investigating mislabeled genomes, for species with multiple genome assembly results in the type strain, clustering was conducted, resulting in the removal of 36 mislabeled genomes; for species containing only one type genome, genus-level clustering was used to identify and eliminate 22 cross-genus mislabeled genomes. The labels of 485 genomes were corrected according to updated nomenclature. After removing low-quality and mislabeled genomes, a final database of 12,745 genomes was included, covering 9,810 species and 2,448 genera. The average integrity reached 99.14%, and the average contamination rate was less than 0.79%.
[0123] Example 4
[0124] like Figure 2 As shown, the steps of the method for identifying and analyzing unknown pathogenic microorganisms of the present invention are as follows:
[0125] 1) Perform quality control on raw metagenomic sequencing data. Use Trimmomatic and Fastqc tools to remove sequencing adapter sequence contamination, low-quality (sequences with an average sliding window quality value of less than 15) and short sequences (<75bp). Use BWA tool to compare the sequencing data with the host species genome database to remove sequencing data contamination by host sequences, and obtain the quality-controlled data.
[0126] 2) The Kraken tool was used to perform species annotation and abundance analysis on the quality-controlled data to quickly obtain species annotation information for viruses, bacteria and fungi. At the same time, the MetaWrap (SPAdes) tool was used for preliminary assembly and splicing to generate preliminary assembled sequences (contigs).
[0127] 3) Identification of unknown viruses (pathogenic microorganisms) based on preliminary assembled sequences:
[0128] 3.1) Virsorter, DeepVirFinder, and geNomad tools were used to identify potential viral sequences in the preliminary assembled sequences. These tools are based on machine learning or deep learning models and do not compare against known virus databases. They can identify potential unknown viral sequences. If any tool identifies a sequence as a virus, it is considered a potential viral sequence.
[0129] 3.2) For the identified potential viral sequences, the CheckV tool was used for quality assessment, and the integrity and contamination rate of the sequence were calculated; Vclust was used to calculate the sequence similarity of closely related species, and the average amino acid similarity (AAI) and average nucleotide similarity (ANI) of the potential viral sequences and all genomes in the virus identification database established in Example 1 were calculated.
[0130] 3.3) Based on the calculation results of the above steps, viral genome sequences with integrity greater than or equal to 50%, contamination rate less than 5%, and maximum AAI and maximum ANI both less than 95% are selected as potential unknown viral genome sequences.
[0131] 3.4) Based on the genome sequence of potential unknown viruses, the ViralHostPredictor tool and the virus host prediction model constructed in Example 2 were used to predict the host and analyze whether the potential host is a primate. When the top three hosts in the prediction results are primates, they are judged to be unknown pathogenic viruses (pathogenic microorganisms).
[0132] 4) Identification of unknown pathogenic bacteria:
[0133] 4.1) Based on the preliminary assembled sequence, the MetaWARP binning assembly tool was used to obtain the metagenomic assembled genome (MAG) sequence;
[0134] 4.2) The dRep tool was used to remove redundancy from the MAG sequence, and the Sylph tool was used for species classification annotation.
[0135] 4.3) The CheckM tool was used to assess the genome quality of MAG sequences, and sequences with an integrity of ≥50% and a contamination rate of ≤5% were selected as high-quality MAG sequences;
[0136] 4.4) The fastANI tool was used to calculate the average nucleotide similarity (ANI) between high-quality MAG sequences and all genomes in the bacterial identification and typing analysis genome database for bacterial species identification. MAG sequences with a maximum ANI of less than 95% were identified as potentially unknown bacterial genome sequences. The bacterial identification and typing analysis genome database was constructed according to the method disclosed in CN 112863606 A.
[0137] 4.5) The PathogenFinder tool was used to predict the pathogenicity of potential unknown bacterial genome sequences. The prediction result was "Human Pathogenic", and the bacteria were identified as unknown pathogenic bacteria.
[0138] Example 5
[0139] The identification and analysis of metagenomic sequencing data of simulated unknown pathogenic viral species in this invention are as follows:
[0140] 1) The genome sequence of the severe acute respiratory syndrome coronavirus 2 (SARS-COV-2, taxid 694009) species was removed from the viral identification genome database;
[0141] 2) Simulate four metagenomic data samples containing SARS-CoV-2 (sample 1, sample 2, sample 3, sample 4), which contain high-throughput metagenomic sequencing data of seven viral species (human host) (Table 1).
[0142] 3) The method of Example 4 was used to identify and analyze unknown pathogenic microorganisms. The maximum AAI and maximum ANI and the corresponding nearest source species name are shown in Table 2. A viral contig sequence can be identified in each of the samples 1, 2, 3 and 4. The AAI and ANI of all sequences in the viral identification genome database are all below 95%, that is, no known viral species were identified. Therefore, it is determined that there is a potential unknown viral sequence.
[0143] 4) The host prediction results are shown in Table 3. The host of the unknown virus sequence is primates. The host prediction scores of the four samples are all greater than 99% and are the TOP1 prediction results. Therefore, they can be identified as unknown pathogenic viruses.
[0144] Table 1. Sample information of simulated metagenomic data
[0145]
[0146] Table 2. Similarity values for simulated metagenomic data samples
[0147]
[0148]
[0149]
[0150] Table 3. Host prediction results of potential new viral species sequences
[0151]
[0152] Example 6
[0153] The identification and analysis of metagenomic sequencing data of simulated unknown pathogenic bacterial species are as follows:
[0154] 1) Staphylococcus aureus was removed from the bacterial identification and typing analysis genomic database. Staphylococcus aureus txid1280) species genome sequence;
[0155] 2) Simulate four metagenomic data samples containing Staphylococcus aureus (SA1, SA2, SA3, SA4) (Table 4) to obtain metagenomic high-throughput sequencing data;
[0156] 3) The method of Example 4 was used to perform the identification and analysis process for unknown pathogenic microorganisms. The maximum ANI between the bacterial genome sequences in SA1, SA2, SA3 and SA4 samples and all genome sequences in the database was less than 95% of the MAG sequences (as shown in Table 5). That is, no known bacterial species were identified, and it was determined that there were potential unknown bacterial species.
[0157] 4) The pathogenicity prediction analysis results are shown in Table 6. The pathogenicity prediction results of the unknown bacterial species sequences are all "Human Pathogenic", that is, pathogenic to humans, and it is determined that there are unknown pathogenic bacteria.
[0158] Table 4. Sample Information of Simulated Metagenomic Data
[0159]
[0160] Table 5. ANI calculation results for potentially unknown bacterial sequences
[0161]
[0162] Table 6. Pathogenicity prediction results of unknown bacterial sequences
[0163]
Claims
1. A method for the identification analysis of unknown pathogenic microorganisms, characterized in that The method comprises the following steps: S1: Constructing a virus identification database: Obtain and integrate virus genome sequences from ICTV, BV-BRC, NCBI Virus RefSeq database, remove redundant and low-quality sequences, and screen high-quality genomes to establish a virus identification database; S2: Constructing a virus host prediction model: Download and integrate virus-host relationship data from ICTV, NCBI, VRION, and VirhostDB databases, extract single nucleotide frequency, double nucleotide frequency, and codon bias of virus genome sequences with vertebrate hosts as characteristic values, train and construct a virus host prediction model; S3: Constructing a bacterial identification database: Collect bacterial genome information from NCBI, cross-verify bacterial nomenclature and type strains with LPSN and IJSEM, obtain qualified strain genomes, remove low-quality and incorrect genomes, extract 16S rRNA gene sequences, and establish a bacterial identification database; the execution order of S3 is any one of the following: before S1, or after S1 and before S4; S4: Pathogenic microorganism identification analysis: Quality control and assembly of metagenomic sequencing data to obtain preliminary assembly sequences; The preliminary assembly sequences are identified by a machine learning or deep learning model to recognize potential virus sequences, and the potential virus sequences are subjected to quality assessment and maximum similarity species genome analysis to calculate the average amino acid similarity AAI and the average nucleotide similarity ANI of the potential virus sequences and all genomes in the virus identification database established in step S1; potential virus genome sequences with a completeness of ≥50%, a contamination rate of <5%, and maximum AAI and maximum ANI of <95% are selected as potential unknown virus genome sequences; the host of the potential unknown virus genome sequence is predicted by the virus host prediction model, and when the top three hosts in the prediction result are primates, it is judged as an unknown pathogenic virus; The preliminary assembly sequences are subjected to binning assembly to obtain metagenome assembly genomes MAG; high-quality MAG sequences are screened for bacterial species identification analysis, and potential unknown bacterial genome sequences are screened out; the pathogenicity of the potential unknown bacterial genome sequence is predicted, and the prediction result is pathogenic to humans, which is judged as an unknown pathogenic bacteria.
2. The method of claim 1, wherein The method comprises the following steps: S1: Constructing a virus identification database: Download virus genome sequence data from the International Committee on Taxonomy of Viruses ICTV, BV-BRC database, and NCBI Virus RefSeq database, combine, remove redundancies, and perform quality assessment on the virus genome sequences to remove low-completeness, high-contamination, and classification-error virus genome sequences, screen high-quality genomes, and construct a virus identification database; S2: Constructing a virus host prediction model: S2-1: Download and integrate virus-host relationship data in ICTV, NCBI, VRION, VirhostDB databases, filter the host as vertebrates, and the virus genome is included in the virus identification database constructed in step S1, integrate the genome sequence of the virus species in the virus identification database, and generate a training data set; S2-2: Based on the training data set, extract the single nucleotide frequency, double nucleotide frequency and codon bias of the virus genome sequence as the virus characteristic value data; S2-3: Classify the vertebrate hosts in the training data set as training data set labels; S2-4: Based on the characteristic value data and labels of the training data set, use the gradient boosting tree algorithm to train and construct a virus host prediction model; S3: Construct a bacterial identification database: Collect bacterial genomes from NCBI, use LPSN, Bergey's Manual of Systematic Bacteriology 2nd Edition and IJSEM to cross-verify bacterial nomenclature and model strains, filter qualified strain bacteria genes, remove high-pollution and low-integrity genomes, extract 16S rRNA gene sequences, remove incorrectly named and incorrectly identified genomes, and establish a bacterial identification database; S4: Pathogenic microorganism identification and analysis: S4-1: Quality control of metagenomic sequencing data, remove low-quality sequencing data and host contamination data to obtain quality-controlled data; S4-2: Preliminary assembly of the quality-controlled data to obtain preliminary assembly sequences; S4-3: Identify unknown pathogenic virus species: S4-3-1: Perform virus sequence recognition analysis on the preliminary assembly sequences to identify potential virus sequences; S4-3-2: Perform quality assessment and maximum similarity species analysis on the potential virus sequences, calculate the average amino acid similarity AAI and average nucleotide similarity ANI of the potential virus sequences and all genomes in the virus identification database established in step S1; filter potential virus genome sequences with integrity ≥ 50%, contamination rate < 5% and maximum AAI and maximum ANI less than 95% as potential unknown virus genome sequences; S4-3-3: Use the ViralHostPredictor tool and the virus host prediction model constructed in step S2 to predict the host of the potential unknown virus genome sequence obtained in step S3-3-2. When the top three hosts in the prediction result are primates, it is judged as an unknown pathogenic virus; S4-4: Identify unknown pathogenic bacterial species, and the execution order of S4-4 and S4-3 can be exchanged or performed simultaneously: S4-4-1: Box assembly of preliminary assembly sequences to obtain metagenome assembly genome MAG sequences; S4-4-2: Remove redundancy from the MAG sequence, perform species classification annotation, and then perform genome quality assessment, and filter MAG sequences with integrity ≥ 50% and contamination rate < 5% as high-quality MAG sequences; S4-4-3: Use the bacterial identification database to perform bacterial species identification analysis, calculate the ANI of the high-quality MAG sequence, and filter sequences with maximum ANI less than 95% as potential unknown bacterial genome sequences; S4-4-4: Using machine learning method tools to predict the pathogenicity of potential unknown bacterial genome sequences, and the prediction result is pathogenic to humans, which is judged as unknown pathogenic bacteria.
3. The method of claim 2, wherein The step S1 is performed as follows: S1-1: Obtain all virus classification information and model representative genome sequence numbers from the International Committee on Taxonomy of Viruses (ICTV) official website; S1-2: Download virus list and genome sequence number from BV-BRC database; S1-3: Download virus list and genome sequence number from NCBI Virus RefSeq database; S1-4: Merge and de-duplicate the virus data downloaded from ICTV, NCBI Virus RefSeq and BV-BRC, and download complete virus genome sequences from NCBI Genbank and RefSeq database according to genome sequence number; S1-5: Quality assessment of virus genome sequence, remove sequences with completeness < 90%, contamination rate > 5% or maximum similarity virus species and genome sequence label inconsistent, screen high-quality genomes, and only keep one representative genome sequence for each virus or virus subtype, to construct a virus identification database.
4. The method of claim 3, wherein The step S1 is performed as follows:
5. The method of claim 2, wherein In the step S2-4, when constructing the prediction model, the accuracy of the prediction model is preliminarily evaluated by using the holdout test, and the accuracy is required to be greater than 85%.
6. The method of claim 5, wherein The step S2 further includes the step S2-5: S2-5: Use EID2: database as test set to test the accuracy of the virus host prediction model constructed in S2-4, and the prediction accuracy for virus genome sequences of primate hosts is required to be greater than 90%.
7. The method of claim 2, wherein In the step S4-1, the sequencing data is quality controlled as follows: Quality control of metagenomic sequencing data, remove sequencing adapter sequence contamination, low quality and short sequence data, remove host sequence contaminated data, and obtain quality controlled data.
8. The method of claim 2, wherein In the step S4-3-1, Virsorter, DeepVirFinder and geNomad tools are used for potential virus sequence identification, and any tool identified as virus is a potential virus sequence.
9. The method of claim 2, wherein In the step S4-4-2, dRep tool is used to remove redundancy of MAG sequence; Sylph tool is used for species classification annotation; and CheckM tool is used for genome quality assessment of MAG sequence.
10. The method of claim 2, wherein In the step S4-4-4, PathogenFinder tool is used to predict the pathogenicity of potential unknown bacterial genome sequences.
Citation Information
Patent Citations
Bacteria identification and typing analysis genome database and identification and typing analysis method
CN112863606A
Virus gene recognition and host prediction method and system
CN114512182A
Pathogenic microorganism genome database and construction method and application thereof
CN116682496A
Metagenome-based pathogen identification method, metagenome-based pathogen identification device, medium and product
CN118072831A