Classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics
By combining statistical methods and deep learning technology, using semantic information and statistical characteristics of viral DNA sequences, a classification algorithm based on viral genomics is constructed, which solves the accuracy and efficiency of identifying prokaryotic and eukaryotic virus sequences in the existing technology, and achieves high-precision virus sequence classification.
Patent Information
- Application Number
- CN202510281596.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has limitations in efficiently and accurately identifying prokaryotic and eukaryotic virus sequences from massive metagenomic data, and it is difficult to meet the needs of high-precision classification.
By combining statistical methods and deep learning technology, using semantic information and statistical characteristics of viral DNA sequences, a classification algorithm based on viral genomics is proposed. The specific steps include 3-mers encoding of the viral DNA sequence and generating a statistical feature matrix; pre-training using a stacked multi-layer Transformer encoder structure, learning semantic information in the sequence, and generating low-dimensional embedding vectors; fusing the statistical feature matrix with the embedding vectors to form a comprehensive feature representation; building a classification model based on a multi-layer self-attention network, extracting features layer by layer, and completing the classification of the virus sequence.
High-precision classification of prokaryotic and eukaryotic viruses was achieved, and experiments showed that the classification accuracy on the standard data set reached 96.7%, which was significantly better than the existing algorithms.
Smart Images

Figure CN120220829A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of bioinformatics and artificial intelligence, and particularly to a classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics. Background Art
[0002] Viruses are an important part of the biosphere, and the characteristics of their genomes are of great significance in molecular biology research and virology research. However, due to the high diversity and complexity of the viral genome structure, traditional virus classification methods have certain limitations in terms of accuracy and efficiency.
[0003] In recent years, with the popularization of high-throughput sequencing technology, the acquisition of metagenomic data has become more convenient, providing rich data resources for viral sequence classification research. However, how to efficiently and accurately identify prokaryotic and eukaryotic viral sequences from massive metagenomic data remains a hot and difficult issue in current research. For this reason, the present invention combines deep learning technology with metagenomics to propose an efficient viral sequence classification algorithm. Summary of the Invention
[0004] The purpose of the present invention is to provide a classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics. By combining statistical methods and deep learning technology, and using the semantic information and statistical features of viral DNA sequences, the accuracy and efficiency of the classification model are improved. To achieve the above purpose, the specific scheme is as follows:
[0005] S1. Encode the viral DNA sequence with 3-mers, extract all possible triplet combinations, and generate a statistical feature matrix;
[0006] S2. Use a stacked multi-layer Transformer encoder structure, model the nucleotides at different positions in the viral sequence through the multi-head self-attention mechanism, pre-train the DNA sequence, learn the semantic information contained in the sequence, and generate a low-dimensional embedding vector;
[0007] S3. Fuse the statistical feature matrix and the embedding vector to form a comprehensive feature representation;
[0008] S4. Build a classification model based on a multi-layer self-attention network, extract features layer by layer, and complete the classification of viral sequences;
[0009] S5. Use a multi-task learning strategy to optimize the model, and evaluate the classification performance through accuracy, recall rate, and F1-score metrics.
[0010] Preferably, the specific process of step S1 is as follows:
[0011] S11. Standardize the input viral DNA sequence, remove redundant sequences and non-nucleic acid characters;
[0012] S12. Perform 3-mers encoding on the standardized DNA sequence, and decompose the sequence into 64 possible triplet combinations;
[0013] S13. Extract all 3-mers combinations based on the sliding window technique, count their joint occurrence frequencies in the sequence, and construct a frequency matrix. The specific formula is as follows:
[0014]
[0015] where count(k i ,k j ) represents the joint occurrence times of triplets k i and k j , and N is the total number of triplets in the sequence;
[0016] S14. Record the relative positions of each 3-mers combination in the sequence, construct a position matrix, and capture the local context relationship. The specific formula is as follows:
[0017]
[0018] where pos(k i ) represents the position of triplet k i in the sequence.
[0019] Preferably, the specific process of step S2 is as follows:
[0020] S21. Randomly mask the bases in the DNA sequence to generate a partially masked sequence;
[0021] S22. Use a stacked multi-layer Transformer encoder structure to pre-train the masked sequence through the multi-head self-attention mechanism to capture the long-range dependence relationship between bases. The objective function formula is as follows:
[0022]
[0023] where M is the number of masked bases, and k \m represents the context of the remaining sequence after removing the m-th base;
[0024] S23. Encode the DNA sequence into a low-dimensional embedding vector through a deep neural network to extract the implicit semantic information of the sequence.
[0025] Preferably, the specific process of step S3 is as follows:
[0026] S31. Standardize the frequency matrix and the position matrix to ensure consistent numerical distributions;
[0027] S32. Integrate the embedding vector with the standardized statistical feature matrix to form a comprehensive feature representation. The specific integration formula is as follows:
[0028] F combined = α·F stat + β·F embed
[0029] where F stat is the statistical feature matrix, F embed is the semantic embedding matrix, and α and β are adjustable weight parameters.
[0030] Preferably, the specific process of step S4 is as follows:
[0031] S41. Construct a classification model based on a multi-layer self-attention mechanism. The model includes the following components:
[0032] (1) Input layer: Receive the comprehensive feature representation;
[0033] (2) Self-attention layer: Extract local features and global dependencies layer by layer. The attention weight calculation formula is as follows:
[0034]
[0035] where d k is the dimension of the key vector, QK T represents the dot product similarity between the query and the key, and the normalized attention weight is obtained through the softmax function.
[0036] (3) Fully connected layer: Map the extracted features to the classification results;
[0037] (4) Output layer: Output the classification category (prokaryotic or eukaryotic virus) of the virus sequence.
[0038] S42. To improve the stability of deep network training, residual connections and layer normalization operations are introduced into the model.
[0039] Preferably, the specific process of step S5 is as follows:
[0040] S51. Divide the encoded comprehensive feature dataset into a training set and a test set, where the training set accounts for 80% and the test set accounts for 20%;
[0041] S52. Use a multi-task learning strategy to optimize the model, including a main classification task and an auxiliary feature prediction task;
[0042] S53. Screen the classification model with the optimal performance by adjusting the hyperparameters of the model, such as the learning rate, the number of training epochs, and the batch size;
[0043] S54. Evaluate the model using precision, recall, and F1-score metrics.
[0044] A classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics proposed by the present invention has the following beneficial effects:
[0045] The present invention captures the global dependencies of viral sequences through a multi-head attention mechanism and extracts local features by combining multi-scale convolutions, achieving high-precision classification of prokaryotic and eukaryotic viruses. Experiments show that the classification accuracy of this algorithm reaches 96.7% on the standard dataset downloaded from the official website of Virus-Host DB, significantly superior to existing algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0047] Figure 1 It is a flowchart of a classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0049] As Figure 1 shown, a flowchart of a classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics, and the steps of this architecture are as follows:
[0050] S1: 3-mers encoding and statistical feature extraction of viral DNA sequences
[0051] S11. Standardization processing
[0052] First, standardize the input viral DNA sequences. This includes removing redundant parts in the sequences, such as low-quality sequences, repetitive sequences, and non-nucleic acid characters (such as N or other impurity characters). In specific implementation, use the Trimmomatic tool for sequence cleaning, set a quality threshold (such as Q30) to remove low-quality bases, and use a deduplication tool (such as cd-hit) for filtering repetitive sequences. Ensure that all sequences only contain the standard four bases (A, T, C, G), thereby improving the accuracy and consistency of subsequent analysis.
[0053] S12. 3-mers encoding
[0054] Perform 3-mers encoding on the standardized DNA sequences, breaking the sequences into all possible triplet combinations. For example, use a Python script to traverse the sequences by the sliding window method, moving one base each time to generate 3-mers. For example, the sequence "ATCGG" will be broken into "ATC", "TCG", and "CGG". When writing the script, ensure that the window slides seamlessly to capture each 3-mer. Applicable code frameworks such as Biopython. S13. Frequency matrix construction
[0055] Traverse the entire sequence based on the sliding window to count the occurrence frequency of each 3-mer. In specific implementation, use the collections.Counter class in Python to quickly count the frequency of each 3-mer and calculate the relative frequency. With a sliding window step size of 1, ensure that all possible 3-mers can be counted. Divide the occurrence times of each 3-mer by the total sequence length minus 2 to obtain the relative frequency, in order to construct the frequency matrix.
[0056] S14. Position matrix construction
[0057] To construct the position matrix, it is necessary to capture the relative positions of each 3-mer in the sequence. In specific implementation, use position indices to record each 3-mer, use the sliding window to record the initial position of each 3-mer, and normalize this position value (divide by the total sequence length) to form the relative position matrix. This can provide more position information for subsequent analysis and enhance the richness of feature representation.
[0058] S2: Masked Language Model (MLM) pre-training and embedding vector generation
[0059] S21. Masking process
[0060] In the masking process, 15% to 20% of the bases in each DNA sequence are randomly selected for masking. Specifically, the random module in Python can be used to determine the positions to be masked and replace them with the special character "X". The masking ratio needs to be adjusted according to the experiment to ensure that the model can effectively learn the context relationships between sequences.
[0061] S22. MLM Pretraining
[0062] Use the BERT architecture for pre-training of the masked language model, adopting the TensorFlow or PyTorch deep learning framework. Input the processed sequences into the BERT model, set hyperparameters such as the learning rate (e.g., 1e-4) and batch size (e.g., 64), and use the Adam optimizer for training. The model learns the context relationships between bases in the sequence by minimizing the prediction error of the masked bases. Pre-training usually needs to run on a GPU or TPU to accelerate the calculation process, and the training time depends on the size of the dataset, generally ranging from several hours to several days.
[0063] S23. Embedding Vector Generation
[0064] The pre-trained masked language model can encode the sequences into low-dimensional embedding vectors. Input each DNA sequence into the pre-trained model and use the output of the hidden layer of the model as the embedding vector. By saving these embedding vectors, they can be used for subsequent feature fusion and classification models. The embedding dimension is usually selected as 128 or 256, and the specific selection needs to be based on the experiment to ensure sufficient expression of information.
[0065] S3: Fusion of Statistical Features and Embedding Vectors
[0066] S31. Embedding Vector Generation
[0067] Normalize the frequency matrix and position matrix to ensure that each feature has a consistent numerical distribution. In specific implementation, use the StandardScaler in the sklearn library to normalize each feature, set the mean to zero and the variance to one, to prevent the numerical differences between features from affecting the performance of the model.
[0068] S32. Embedding Vector Generation
[0069] Fuse the statistical feature matrix and the embedding vector. The specific implementation method is to combine them through a concatenation operation, and use the concatenate function of NumPy for feature concatenation. To further improve the fusion effect, weighted combination can be introduced, set the initial weights, and automatically adjust them through the subsequent training process to optimally integrate statistical features and deep learning features.
[0070] S4: Construction of a Classification Model Based on a Multi-Layer Self-Attention Network
[0071] S41. Construct a classification model based on a multi-layer self-attention mechanism
[0072] Take the fused comprehensive feature representation as the input. The input layer uses a Dense layer (implemented in Keras) to map the high-dimensional features to a unified dimension, providing an adapted feature space for the subsequent self-attention mechanism. Use a multi-layer self-attention mechanism to extract features layer by layer, implemented using the Transformers library. Calculate the attention weights of the input features in each layer, and extract global and local information through matrix multiplication. Each layer of self-attention contains a multi-head attention mechanism, enabling the model to extract diverse information from different subspaces. Send the features extracted by the self-attention layer into a fully connected layer, and use a multi-layer perceptron (MLP) for classification. The ReLU activation function is used in the fully connected layer to enhance the non-linear expression ability of the model, and the softmax function is used in the last layer to output the classification result (prokaryotic or eukaryotic). The output layer converts the output of the final fully connected layer into class probability values through the softmax activation function. Select the class with the highest probability as the final classification result.
[0073] S42. To improve the stability of deep network training, residual connections and layer normalization operations are introduced into the model.
[0074] After each layer of self-attention, add residual connections and layer normalization (LayerNorm). Residual connections can be implemented through the Add layer in Keras, adding the input and output element-wise. Layer normalization uses the LayerNormalization class to ensure the consistent input distribution of each layer and stabilize model training.
[0075] S5: Model Optimization and Performance Evaluation
[0076] S51. Dataset Division
[0077] Divide the dataset. Use the train_test_split function in sklearn to divide the dataset into a training set (80%) and a test set (20%), ensuring that the training and test data distributions are consistent and ensuring the reliability of model evaluation.
[0078] S52. Multi-Task Learning Strategy
[0079] Optimize the main classification task and auxiliary tasks simultaneously during model training. Auxiliary tasks such as 3-mers frequency prediction, predict these features by adding an additional output layer, and use a joint loss function (such as the weighted sum of the main task cross-entropy loss and the auxiliary task mean squared error loss) to jointly optimize the model.
[0080] S53. Hyperparameter Tuning
[0081] Use grid search (GridSearchCV) or random search (RandomizedSearchCV) to tune the hyperparameters of the model. The tuned parameters include the learning rate, batch size, and the number of attention heads. Combine cross-validation techniques to select the optimal hyperparameter combination to achieve optimal performance.
[0082] S54. Performance evaluation
[0083] Evaluate the model using precision, recall, and F1-score metrics. Use the classification_report function in sklearn to evaluate the classification results of the test set, calculate the values of each metric, and observe the classification effect of the model through the confusion matrix to ensure the reliability of the classification task.
[0084] As described above, it is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. According to the technical essence of the present invention, any simple modifications, equivalent replacements, and improvements made to the above embodiments within the spirit and principles of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics, characterized in that: The following steps are involved: S1. Encode the viral DNA sequence by 3-mers, extract all possible triplet combinations, and generate a statistical feature matrix; S2. Use a stacked multi-layer Transformer encoder structure to model nucleotides at different positions in the virus sequence through a multi-head self-attention mechanism, pre-train the DNA sequence, learn the semantic information contained in the sequence, and generate a low-dimensional embedding vector; S3, fuse the statistical feature matrix with the embedding vector to form a comprehensive feature representation; S4. Build a classification model based on a multi-layer self-attention network, extract features layer by layer, and complete the classification of virus sequences; S5. Use a multi-task learning strategy to optimize the model and use precision, recall, and F1 score as indicators to evaluate the classification performance.
2. The classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics according to claim 1, characterized in that: The specific process of step S1 is as follows: S11, standardize the input viral DNA sequence to remove redundant sequences and non-nucleic acid characters; S12, perform 3-mers encoding on the standardized DNA sequence and decompose the sequence into 64 possible triplet combinations; S13. Extract all 3-mers combinations based on the sliding window technology, count their joint occurrence frequencies in the sequence, and construct a frequency matrix. The specific formula is as follows: Among them, count(k i ,k j ) represents triplet k i and k j The number of joint occurrences of , N is the total number of triplets in the sequence; S14. Record the relative position of each 3-mers combination in the sequence, construct a position matrix, and capture the local context relationship. The specific formula is as follows: Among them, pos(k i ) represents triplet k i Position in the sequence.
3. The classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics according to claim 1, characterized in that: The specific process of step S2 is as follows: S21, randomly performing masking processing on the bases in the DNA sequence to generate a partial masked sequence; S22. Using the stacked multi-layer Transformer encoder structure, the mask sequence is pre-trained through the multi-head self-attention mechanism to capture the long-range dependencies between bases. The objective function formula is as follows: Where M is the number of bases to be masked, k \m represents the context of the remaining sequence after removing the mth base; S23. Encode the DNA sequence into a low-dimensional embedding vector through a deep neural network to extract the implicit semantic information of the sequence.
4. The classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics according to claim 1, characterized in that: The specific process of step S3 is as follows: S31, standardizing the frequency matrix and the position matrix to ensure that the numerical distribution is consistent; S32. Fuse the embedded vector with the standardized statistical feature matrix to form a comprehensive feature representation. The specific fusion formula is as follows: F combined =α·F stat +β·F embed Among them, F stat is the statistical feature matrix, F embed is the semantic embedding matrix, and α and β are adjustable weight parameters.
5. The classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics according to claim 1, characterized in that: The specific process of step S4 is as follows: S41. Construct a classification model based on a multi-layer self-attention mechanism, wherein the model includes the following components: (1) Input layer: receives comprehensive feature representation; (2) Self-attention layer: extract local features and global dependencies layer by layer. The attention weight calculation formula is as follows: Among them, d k is the dimension of the key vector, QK T Represents the dot product similarity between the query and the key, and the normalized attention weight is obtained through the softmax function; (3) Fully connected layer: maps the extracted features into classification results; (4) Output layer: Output the classification category of the virus sequence as prokaryotic or eukaryotic virus; S42. In order to improve the stability of deep network training, residual connection and layer normalization operations are introduced into the model.
6. The classification algorithm for identifying prokaryotic and eukaryotic viruses based on viral genomics according to claim 1, characterized in that: The specific process of step S5 is as follows: S51. Divide the encoded comprehensive feature data set into a training set and a test set, where the training set accounts for 80% and the test set accounts for 20%; S52. Use a multi-task learning strategy to optimize the model, including the main classification task and the auxiliary feature prediction task; S53. Filter the classification model with the best performance by adjusting the model's learning rate, number of training rounds, and batch size hyperparameters; S54. Use precision, recall and F1 score indicators to evaluate model performance.