Metagenome virus classification method based on weighted graph network and label propagation

Through the method of weighted graph network and label propagation, the problems of insufficient flexibility and data imbalance in virus classification methods are solved, and flexible classification of metagenomic viruses and effective exploration of unknown viruses are achieved.

CN120748490APending Publication Date: 2025-10-03HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510848320.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing virus classification methods lack flexibility and are difficult to adapt to the dynamic changes of ICTV classification standards, resulting in data class imbalance and inability to effectively analyze unknown viruses ('dark matter').

Method used

A method based on weighted graph networks and label propagation is adopted to achieve flexible classification of metagenomic viruses by updating the database, constructing a weighted graph network and a three-stage label propagation process.

Benefits of technology

It realizes the automatic update of the virus database, alleviates the problem of data imbalance, effectively explores unknown viruses, and improves the classification ability of "dark matter".

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748490A_ABST
    Figure CN120748490A_ABST
Patent Text Reader

Abstract

The invention discloses a metagenome virus classification method based on a weighted graph network and label propagation, and relates to a metagenome virus classification method. The method aims at solving the problems that in the prior art, the dynamic adaptability of ICTV classification standards is insufficient, data categories are unbalanced, and dark matter classification is difficult. The method comprises the following steps: in a database updating and constructing process, acquiring latest virus classification sample data, and constructing a latest virus protein database; in the weighted graph network construction process, firstly, a to-be-detected virus nucleotide sequence is translated into a protein sequence, then a protein cluster is obtained through Markov clustering, finally, the weight of an edge between the to-be-detected virus sequence and a reference sequence in a database is calculated through an edge weight calculation algorithm, and a weighted graph network is constructed; in a label propagation virus classification process, comprehensive virus species classification is carried out on metagenome viruses in three stages through a category weight normalization label propagation algorithm. The invention belongs to the technical field of gene engineering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for classifying macrogenomic viruses and belongs to the technical field of genetic engineering. Background Art

[0002] Metagenomic virus classification is the process of accurately classifying complex metagenomic viruses, thereby supporting disease diagnosis, antimicrobial drug development, and phage therapy. While current virus classification methods have some classification capabilities, they still face three major challenges. First, the official virus classification system is maintained by the International Cybersecurity and Telecommunications Commission (ICTV) and is continuously updated as research progresses. However, most existing virus classification methods lack flexibility, making it difficult to timely update the database to adapt to the dynamic changes in the ICTV classification standards, resulting in insufficient dynamic adaptability of the ICTV classification standards. Second, as the level of virus classification decreases, the number of minority classes gradually increases, leading to an increasingly unbalanced distribution of viral data. Due to the lack of information about minority classes, classification algorithms tend to favor the majority class, significantly reducing their predictive power for the minority class and causing data class imbalance. Third, metagenomic sequencing can sequence the uncultured "dark matter" of the microbial biosphere, which may contain a large number of unknown viruses. However, the known taxonomic genomes in the ICTV database only account for a small fraction of the total viral population. Most current methods can only classify known viruses and are unable to effectively analyze the "dark matter," making it difficult to classify this "dark matter." Summary of the Invention

[0003] In order to solve the problems of insufficient dynamic adaptability, imbalanced data categories and difficulty in classifying "dark matter" in the existing ICTV classification standard, the present invention proposes a metagenomic virus classification method based on weighted graph network and label propagation.

[0004] The technical solution adopted by the present invention to solve the above problems is: the steps of the present invention include: Step 1: Update and build the database; Step 2: Build a weighted graph network; Step 3: Label propagation virus classification.

[0005] Furthermore, the process of updating and constructing the database in step 1 is as follows: Step 101: Download the latest reference virus sequence and its classification label from the website; Step 102: Using software to translate the downloaded viral DNA sequence into the corresponding amino acid sequence, and further constructing a viral protein database based on the translated amino acid sequence; Step 103: Perform a Blastp alignment on the amino acid sequence in the viral protein database to obtain a Blastp alignment result of the reference database.

[0006] Furthermore, the process of constructing the weighted graph network in step 2 is: Step 201: Translate the macrogene viral sequence to be tested into six reading frames to extract all possible amino acid sequences; Step 202: Filter out amino acid sequences with a length of less than 12 and retain amino acid sequences with biological significance; Step 203: The filtered amino acid sequence is subjected to a Blastp comparison with a viral protein database to obtain a Blastp comparison result of the virus to be tested; Step 204: Markov clustering is performed by combining the Blastp comparison results of the reference database with the Blastp comparison results of the virus to be tested to obtain virus protein clusters; Step 205: construct a weighted graph neural network based on the protein clustering results and the Blastp comparison results of the virus to be tested.

[0007] Furthermore, the process of label propagation virus classification in step 3 is as follows: Step 301: In the first stage, the viruses to be tested that are directly connected to the reference virus nodes are classified into direct categories through category normalization label propagation; Step 302: In the second stage, based on the viruses classified in the first stage, the viruses to be tested that are indirectly connected to the reference virus nodes are classified into possible categories through category normalization label propagation; Step 303, the third stage is to classify the independent and mutually connected virus groups to be tested that have no connection with the reference virus node into an unknown category.

[0008] Furthermore, in step 205, a weighted graph neural network is constructed based on the protein clustering results and the Blastp comparison results of the virus to be tested, which specifically includes: Step 20501: Calculate the probability that two viral sequences share a certain number of protein clusters: , in, represents the total number of protein clusters, Representation sequence The number of protein clusters included, Representation sequence The number of protein clusters included, Representation sequence and sequence At least share The probability of protein clustering; Step 20502: According to Calculate the relationship between two viral sequences Weight: , in, Indicates the total number of virus sequences to be tested and the latest reference virus sequences downloaded from the ICTV website. represents the total number of sequence pairs, represents the weight of the edge constructed based on protein clustering; Step 20503: Calculate the inter-virus sequence distances based on the Blastp comparison results of the virus to be tested. Weight: , in, Representation sequence with sequence The number of Blastp comparison results between Representation sequence with sequence Between The comparison results Fraction, represents the weight of the edge constructed based on the protein Blastp alignment results; Step 20504: Establish the edges and thresholds constructed between virus sequences based on the thresholds: , in, express Constructing the threshold of edges in weighted graph networks, express Constructing the threshold of edges in weighted graph networks, Indicates building an edge between two virus sequences and assigning a weight of , Indicates building an edge between two virus sequences and assigning a weight of , Indicates that no edge is constructed between the two viral sequences.

[0009] Furthermore, the label propagation through category normalization in step 301 specifically includes: Step 30101: normalize the reference virus sequence to which the virus sequence to be tested is connected: , in, represents the virus sequence to be classified, represents the neighbor reference virus sequence, Indicates the virus sequence to be classified with neighbor reference virus sequences The weight of the edge between Indicates that the neighbor nodes belong to the category the number of Represents the node to be classified Belong to category probability; Step 30102: According to Assign virus classification labels to the virus sequences to be classified: .

[0010] The beneficial effects of the present invention are: 1. The virus classification database in the present invention can be flexibly created and automatically updated as ICTV standards are updated, overcoming the problem of insufficient flexibility in database updates in traditional virus classification methods. 2. The weighted graph-based category normalization label propagation algorithm in this invention normalizes the weights of the edges connected to each virus sequence, thereby increasing the focus on minority species in virus classification and alleviating the data imbalance problem in virus classification. 3. The weighted graph-based three-stage label propagation process in the present invention proposes the existing ICTV virus categories, the approximate unknown categories of the existing ICTV virus categories, and the independent virus groups to be tested as unknown categories in batches through three-stage label propagation, thereby achieving a full exploration of the "dark matter" in the metagenomic virus and alleviating the difficulty of classifying "dark matter". BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a flow chart of the present invention; Figure 2 It is a flowchart of the database update and construction process; Figure 3 It is a flowchart of the weighted graph network construction process; Figure 4 It is a schematic diagram of the label propagation virus classification process. DETAILED DESCRIPTION

[0012] Specific implementation method 1: Figures 1 to 4 As shown in FIG, a metagenomic virus classification method based on weighted graph network and label propagation includes the following steps: Step 1: Update and build the database; the specific process is: Step 101: Download the latest reference virus sequence and its classification label from the website; Step 102: Using software to translate the downloaded viral DNA sequence into the corresponding amino acid sequence, and further constructing a viral protein database based on the translated amino acid sequence; Step 103: Perform a Blastp alignment on the amino acid sequence in the viral protein database to obtain a Blastp alignment result of the reference database; Step 2: Build a weighted graph network; the specific process is: Step 201: Translate the macrogene viral sequence to be tested into six reading frames to extract all possible amino acid sequences; Step 202: Filter out amino acid sequences with a length of less than 12 and retain amino acid sequences with biological significance; Step 203: The filtered amino acid sequence is subjected to a Blastp comparison with a viral protein database to obtain a Blastp comparison result of the virus to be tested; Step 204: Markov clustering is performed by combining the Blastp comparison results of the reference database with the Blastp comparison results of the virus to be tested to obtain virus protein clusters; Step 205: constructing a weighted graph neural network based on the protein clustering results and the Blastp comparison results of the virus to be tested; specifically comprising: Step 20501: Calculate the probability that two viral sequences share a certain number of protein clusters: , in, represents the total number of protein clusters, Representation sequence The number of protein clusters included, Representation sequence The number of protein clusters included, Representation sequence and sequence At least share The probability of protein clustering; Step 20502: According to Calculate the relationship between two viral sequences Weight: , in, Indicates the total number of virus sequences to be tested and the latest reference virus sequences downloaded from the ICTV website. represents the total number of sequence pairs, represents the weight of the edge constructed based on protein clustering; Step 20503: Calculate the inter-virus sequence distances based on the Blastp comparison results of the virus to be tested. Weight: , in, Representation sequence with sequence The number of Blastp comparison results between Representation sequence with sequence Between The comparison results Fraction, represents the weight of the edge constructed based on the protein Blastp alignment results; Step 20504: Establish the edges and thresholds constructed between virus sequences based on the thresholds: , in, express Constructing the threshold of edges in weighted graph networks, express Constructing the threshold of edges in weighted graph networks, Indicates building an edge between two virus sequences and assigning a weight of , Indicates building an edge between two virus sequences and assigning a weight of , Indicates that no edge is constructed between the two viral sequences; Step 3: Label propagation virus classification; the specific process is as follows: Step 301: The first stage is to classify the viruses to be tested that are directly connected to the reference virus nodes into direct categories through category normalization label propagation. Specifically, the following steps are performed: Step 30101: normalize the reference virus sequence to which the virus sequence to be tested is connected: , in, represents the virus sequence to be classified, represents the neighbor reference virus sequence, Indicates the virus sequence to be classified with neighbor reference virus sequences The weight of the edge between Indicates that the neighbor nodes belong to the category the number of Represents the node to be classified Belong to category probability; Step 30102: According to Assign virus classification labels to the virus sequences to be classified: ; Step 302: In the second stage, based on the viruses classified in the first stage, the viruses to be tested that are indirectly connected to the reference virus nodes are classified into possible categories through category normalization label propagation; Step 303, the third stage is to classify the independent and mutually connected virus groups to be tested that have no connection with the reference virus node into an unknown category.

[0013] The website mentioned in step 101 refers to the website of the International Committee on Taxonomy of Viruses (ICTV).

[0014] The software in step 102 refers to Prodigal software.

[0015] The label propagation virus classification process comprehensively classifies viruses in the metagenome through three stages. In the first stage, test virus sequences directly linked to reference virus sequences are classified into existing ICTV virus categories due to their high confidence. In the second stage, test virus sequences indirectly linked to reference virus sequences are classified as approximate unknown categories of existing ICTV virus categories due to their low similarity to the reference virus sequences but with a certain degree of confidence. In the third stage, test virus groups that have no connection to reference virus nodes but are connected to each other are classified as unknown categories because they have no similarity to existing reference virus sequences but have a certain degree of similarity to each other.

[0016] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement of the above embodiments made according to the technical essence of the present invention, within the spirit and principles of the present invention, without departing from the content of the technical solution of the present invention, shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A metagenomic virus classification method based on weighted graph networks and label propagation, characterized in that: The specific steps include: Step 1: Update and build the database; Step 2: Build a weighted graph network; Step 3: Label propagation virus classification.

2. A method for metagenomic virus classification based on weighted graph network and label propagation according to claim 1, characterized in that: The process of updating and constructing the database in step 1 is: Step 101: Download the latest reference virus sequence and its classification label from the website; Step 102: Using software to translate the downloaded viral DNA sequence into the corresponding amino acid sequence, and further constructing a viral protein database based on the translated amino acid sequence; Step 103: Perform a Blastp alignment on the amino acid sequence in the viral protein database to obtain a Blastp alignment result of the reference database.

3. The method for metagenomic virus classification based on weighted graph network and label propagation according to claim 1, characterized in that: The process of constructing the weighted graph network in step 2 is: Step 201: Translate the macrogene viral sequence to be tested into six reading frames to extract all possible amino acid sequences; Step 202: Filter out amino acid sequences with a length of less than 12 and retain amino acid sequences with biological significance; Step 203: The filtered amino acid sequence is subjected to a Blastp comparison with a viral protein database to obtain a Blastp comparison result of the virus to be tested; Step 204: Markov clustering is performed by combining the Blastp comparison results of the reference database with the Blastp comparison results of the virus to be tested to obtain virus protein clusters; Step 205: construct a weighted graph neural network based on the protein clustering results and the Blastp comparison results of the virus to be tested.

4. The method for metagenomic virus classification based on weighted graph network and label propagation according to claim 1, characterized in that: The process of label propagation virus classification in step 3 is as follows: Step 301: In the first stage, the viruses to be tested that are directly connected to the reference virus nodes are classified into direct categories through category normalization label propagation; Step 302: In the second stage, based on the viruses classified in the first stage, the viruses to be tested that are indirectly connected to the reference virus nodes are classified into possible categories through category normalization label propagation; Step 303, the third stage is to classify the independent and mutually connected virus groups to be tested that have no connection with the reference virus node into an unknown category.

5. The method for metagenomic virus classification based on weighted graph network and label propagation according to claim 3, characterized in that: In step 205, a weighted graph neural network is constructed based on the protein clustering results and the Blastp comparison results of the virus to be tested, which specifically includes: Step 20501: Calculate the probability that two viral sequences share a certain number of protein clusters: , in, represents the total number of protein clusters, Representation sequence The number of protein clusters included, Representation sequence The number of protein clusters included, Representation sequence and sequence At least share The probability of protein clustering; Step 20502: According to Calculate the relationship between two viral sequences Weight: , in, Indicates the total number of virus sequences to be tested and the latest reference virus sequences downloaded from the ICTV website. represents the total number of sequence pairs, represents the weight of the edge constructed based on protein clustering; Step 20503: Calculate the inter-virus sequence distances based on the Blastp comparison results of the virus to be tested. Weight: , in, Representation sequence with sequence The number of Blastp comparison results between Representation sequence with sequence Between The comparison results Fraction, represents the weight of the edge constructed based on the protein Blastp alignment results; Step 20504: Establish the edges and thresholds constructed between virus sequences based on the thresholds: , in, express Constructing the threshold of edges in weighted graph networks, express Constructing the threshold of edges in weighted graph networks, Indicates building an edge between two virus sequences and assigning a weight of , Indicates building an edge between two virus sequences and assigning a weight of , Indicates that no edge is constructed between the two viral sequences.

6. The method for metagenomic virus classification based on weighted graph network and label propagation according to claim 1, characterized in that: The label propagation through category normalization in step 301 specifically includes: Step 30101: normalize the reference virus sequence to which the virus sequence to be tested is connected: , in, represents the virus sequence to be classified, represents the neighbor reference virus sequence, Indicates the virus sequence to be classified with neighbor reference virus sequences The weight of the edge between Indicates that the neighbor nodes belong to the category the number of Represents the node to be classified Belong to category probability; Step 30102: According to Assign virus classification labels to the virus sequences to be classified: 。