Phage virus multi-host prediction method, system, equipment and medium

By using a pre-trained BERT model and deep neural networks to process phage gene sequences, the problems of host label dependence and strict data format requirements in existing technologies are solved, enabling multi-host prediction of unlabeled gene sequences and improving prediction accuracy and model adaptability.

CN121034419APending Publication Date: 2025-11-28SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410664982.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing phage virus host prediction methods rely on data with clearly labeled hosts and have difficulty handling unlabeled gene sequences. Furthermore, existing models have strict requirements on the input data format and are difficult to adapt to new data.

Method used

A pre-trained BERT model was used to process phage gene sequences. The high-dimensional embedding vectors were reduced to low dimensions through self-attention mechanism and average pooling method. A multi-host prediction model was constructed by combining deep neural network with this model.

Benefits of technology

It significantly improves the accuracy of feature extraction and the versatility of the model, can handle gene sequences of different lengths, achieves effective prediction of unlabeled data, and improves prediction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034419A_ABST
    Figure CN121034419A_ABST
Patent Text Reader

Abstract

The invention discloses a bacteriophage virus multi-host prediction method, system and device and a medium, and the method comprises the steps: receiving knee joint data through a sliding window, building a joint physiological kinematics constraint, and building a joint physiological motion optimization equation based on the joint physiological kinematics constraint; solving the joint physiological motion optimization equation to obtain a joint axis and a joint position vector; based on the joint axis and the joint position vector, two groups of joint angles are obtained through acceleration information and angular velocity integration; and solving the weighted average of the two groups of joint angles through complementary filtering to obtain the joint angle. According to the method, the problem of disturbance caused by unstable wearing or external disturbance in the measurement process of the inertial sensing wearable inertial node is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bio-information analysis, and particularly relates to a bacteriophage virus multi-host prediction, a system, a device and a medium. BACKGROUND

[0002] With the rapid progress of viral genetic sequencing technology, the research on such small life forms is no longer simply dependent on traditional laboratory cultivation means. With the establishment of a large number of viral genetic sequence, protein sequence and other various information databases, the bioinformatics strategy of using computer technology for biological analysis is also increasing. These computer-aided analyses enable research to directly predict the biological properties of viruses, thereby providing a new perspective for research work. However, due to the unstructured nature of viral genetic information and the complexity of gene expression mechanisms, traditional computational statistical analysis methods often have little effect on gene data. In contrast, machine learning technology is data-driven, enabling the computing program to automatically identify the biological patterns sought from massive data.

[0003] And with the continuous development of globalization, viruses have become the most threatening pathogens to human health. In the field of virology research, the analysis and prediction of potential hosts of bacteriophage viruses is an important research direction and has great significance for biological research. However, the current method of extracting features from bacteriophage gene sequences and then predicting hosts is often based on supervised machine learning methods.

[0004] Prior art one: ([1] Ahlgren, Nathan A, et al. "alignment-free d2 oligonucleotide frequency dissimilarity measure improves prediction of hosts from metagenomically-derived viral sequences." (2019).) and others developed a method of predicting the host of a virus by using similar oligonucleotide frequency patterns between the virus and the host genome, by comparing the genomes of about 32,000 prokaryotic organisms with the genomes of 1427 known real hosts, comprehensively comparing 11 oligonucleotide frequency indicators and several k-mer lengths, and achieving a prediction accuracy (genus level 64%) using threshold and consensus methods that exceeds oligonucleotide frequency based on Euclidean distance (32%) or homology-based methods (22-62%).

[0005] Existing technology two: (Lu, Congyu et al. "Prokaryotic virus host predictor: a Gaussian model for host prediction of prokaryotic viruses in metagenomics", BMC biology 19.1(2021):5.) et al. developed a Gaussian model to predict the host of prokaryotic viruses, which has better performance than previous computational methods. They ultimately implemented a prokaryotic virus host predictor (PHP), a software tool that uses a Gaussian model to predict the host of prokaryotic viruses by utilizing the difference in k-mer frequencies between viral and host genome sequences as features.

[0006] Existing technology three: (Shang, Jiayu, and Yanni Sun. "CHERRY: a Computational Method for Accurate Prediction of Virus-Prokaryotic Interactions Using a Graphencoder-Decoder Model", Briefings in Bioinformatics 23.5 (2022)) et al. set host prediction as link prediction in a knowledge graph that integrates multiple protein- and DNA-based sequence features. They ultimately implemented a tool called CHERRY, which can be used to predict the host of newly discovered viruses and identify viruses infecting target bacteria. Its performance was compared with 11 popular host prediction methods, and it outperformed all existing methods at the species level, improving accuracy by 37%.

[0007] The main challenges currently facing phage virus feature extraction and host prediction methods include: 1. Over-reliance on well-labeled phage host data, which is extremely rare in reality, and the labels in databases may not be accurate enough, greatly limiting the application of supervised learning methods in phage host prediction. 2. The lack of application of advanced artificial intelligence models, such as the self-attention-based transformer model used in natural language processing, to phage host prediction tasks; previous methods have mainly been limited to using simple machine learning models such as Gaussian models. 3. Existing phage host prediction methods typically have strict requirements on the format of input data, which cannot be too long or too short, and struggle to handle new data not present in the model training database. Summary of the Invention

[0008] In order to solve the problem of predicting bacteriophage host data depending on the clear host marker in the prior art, the present application provides a bacteriophage virus multi-host prediction, system, device and medium, which can perform pre-training on a large amount of unlabeled genetic sequence data, thereby constructing a feature extraction framework of genetic sequence, and significantly reducing the dependence of the embodiment of the present application on labeled virus and genetic data.

[0009] In order to solve the above technical problems, the technical scheme adopted by the present application is:

[0010] In a first aspect, the present application provides a bacteriophage virus multi-host prediction method, comprising:

[0011] The bacteriophage genetic sequence to be predicted is divided into an ordered k-mer word sequence, and the k-mer word sequence is processed by a pre-trained BERT model to obtain an embedding vector to be predicted;

[0012] The embedding vector to be predicted is mapped to a fixed-dimensional low-dimensional space by using an average pooling method, and a feature vector to be predicted representing the original genetic sequence is obtained after reducing the dimension of the embedding;

[0013] The feature vector to be predicted is input into a pre-trained bacteriophage multi-host prediction model to predict the infection probability of the multi-host; wherein the bacteriophage multi-host prediction model is used to identify the correlation between the bacteriophage genetic sequence and the specific host bacteria.

[0014] As a further improvement of the present application, the training method of the pre-trained classification network model comprises:

[0015] The bacteriophage genetic sequence in the database is divided into an ordered k-mer word sequence, and the k-mer word sequence is processed by a pre-trained BERT model to obtain an embedding vector; the database is selected from a standard dataset VHM dataset, and the complete genome containing bacteriophage genetic sequence and host infection information

[0016] The embedding vector in the high-dimensional space is mapped to a fixed-dimensional low-dimensional space by using an average pooling method, and a feature vector representing the original genetic sequence is obtained after reducing the dimension of the embedding;

[0017] The feature vector is input into a pre-trained classification network model to predict the infection probability of the multi-host.

[0018] As a further improvement of the present application, the bacteriophage genetic sequence is divided into an ordered k-mer word sequence, which means that the long genetic sequence is divided into a plurality of fixed-length and non-overlapping fragments.

[0019] As a further improvement of the present application, the BERT model comprises a Transformer network employing a self-attention mechanism; the processing of the k-mer word sequence by the pre-trained BERT model to obtain the embedding vector also introduces markers, specifically:

[0020] The [BAR] marker at the beginning of the sequence serves as an identifier of the original genome of the fragment, [PAD] is used for sequence padding; [UNK] represents an unknown unit; [SEP] identifies the end of the sequence; [MASK] is used for masking operation;

[0021] The long gene sequence is segmented and tokenized, the k-mer word sequence is converted into a tokenized sentence, and [BAR] and [SEP] markers are added at the beginning and end of the sequence respectively; the embedding representation of each word is obtained by adding the position embedding vector and the paragraph embedding vector, and the continuous sequence is divided into a series of fixed-length fragments;

[0022] The fragment is converted into a k-mer word sequence, and [BAR] and [SEP] markers are introduced at the beginning and end of the sequence respectively, and the embedding of each word is obtained by combining the position embedding vector and the paragraph embedding vector, and the sum of the two constitutes the vector representation of the corresponding word.

[0023] As a further improvement of the present application, a random word masking mechanism is introduced in the pre-training of the BERT model, specifically:

[0024] According to a normal distribution with a mean of k and a standard deviation of k-3, a series of k-mer word sequences of different lengths are continuously masked;

[0025] In the early stage of training, the masked k-mer accounts for an initial proportion p of the total input k-mer; as the training progresses, the masking proportion will gradually increase by a certain percentage: from the second training cycle, the masking proportion increases by a fixed percentage q every cycle, and in the i-th training cycle, the masking proportion reaches p+q(i-1);

[0026] During the pre-training process, the model is optimized using the cross-entropy loss function, and an adaptive adjustment strategy is used for training.

[0027] As a further improvement of the present application, the average pooling method is used to map the embedding vectors in the high-dimensional space to the fixed-dimensional low-dimensional space, and the feature vector representing the original gene sequence is obtained after reducing the dimension of the embedding, including:

[0028] The gene sequence is segmented into fixed-length fragments, and then decomposed into multiple overlapping k-mer words; and the k-mer word sequence with a frequency lower than a threshold value in the entire corpus is removed;

[0029] The embedding vectors of different subsequences from the same virus sequence are combined to obtain a feature vector representing the entire virus sequence.

[0030] As a further improvement of the application, the phage multi-host prediction model adopts a deep neural network classification model, and based on the BERT model parameters, a mapping relationship between gene sequences and human infectivity is trained.

[0031] A multi-classification task is performed, and a classification network based on a multi-layer perception, a classification network based on a Transformer, and a classification network of a convolutional neural network are built to predict the multi-host label of the phage virus, and then the original virus feature vector is mapped to the multi-host class label.

[0032] In a second aspect, the application provides a phage virus multi-host prediction system, comprising:

[0033] An embedding vector unit is configured to divide the phage gene sequence to be predicted into an ordered k-mer word sequence, and process the k-mer word sequence through a pre-trained BERT model to obtain an embedding vector to be predicted.

[0034] A feature vector unit is configured to map the embedding vector to be predicted to a low-dimensional space of a fixed dimension by using an average pooling method, and obtain a feature vector to be predicted representing the original gene sequence after reducing the dimension of the embedding.

[0035] A classification prediction unit is configured to input the feature vector to be predicted into a pre-trained phage multi-host prediction model to predict the infection probability of the multi-host, wherein the phage multi-host prediction model is used to identify the correlation between the phage gene sequence and a specific host bacterium.

[0036] In a third aspect, the application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the phage virus multi-host prediction method when executing the computer program.

[0037] In a fourth aspect, the application provides a computer-readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the phage virus multi-host prediction method.

[0038] In a fifth aspect, the application provides a computer program product, which comprises computer instructions for instructing a computer to execute the phage virus multi-host prediction method.

[0039] The application has the following beneficial effects compared with the prior art:

[0040] The present application directly takes the complete phage virus sequence as input, avoids information loss caused by reliance on statistical properties, and significantly enhances the accuracy of feature extraction. The system uses the attention mechanism based on the BERT pre-training model to efficiently convert the virus sequence into a numerical vector, laying a solid foundation for the multi-host prediction task. The phage multi-host prediction model trained by the classification system based on deep neural networks realizes seamless migration of the model to other gene sequence datasets by clearly distinguishing between the pre-training and fine-tuning stages. The system can also process gene sequences of several hundred to tens of thousands of base pairs in length, greatly improving its versatility and flexibility. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the drawings of the related technical solutions in the embodiments of the present application or the prior art. It should be understood that the drawings in the following introduction are only for the convenience of clearly describing part of the embodiments of the technical solutions of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0042] Figure 1 A phage virus multi-host prediction method flowchart is given for the present application;

[0043] Figure 2 A general workflow diagram is given for the embodiments of the present application;

[0044] Figure 3 An embedding, pre-training and feature extraction flowchart is given for the present application;

[0045] Figure 4 A performance comparison diagram of different methods on the VHM baseline dataset is given.

[0046] Figure 5 A phage virus multi-host prediction system is provided for the present application;

[0047] Figure 6 An electronic device schematic diagram is provided for the present application. DETAILED DESCRIPTION

[0048] Embodiments of the present application are described below in detail with reference to the accompanying drawings, wherein the same or similar numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as a limitation of the present application. For the step numbers in the following embodiments, they are only set for the convenience of the description of the explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0049] In the description of the present application, the words such as setting, installing, connecting and the like should be understood in a broad sense unless otherwise explicitly limited, and those skilled in the art can reasonably determine the specific meaning of the above words in the present application in combination with the specific content of the technical solution.

[0050] As shown in Figure 1 The first object of the present application is to provide a phage virus multi-host prediction method, comprising:

[0051] S1, dividing the phage gene sequence to be predicted into an ordered k-mer word sequence, and processing the k-mer word sequence through a pre-trained BERT model to obtain a to-be-predicted embedding vector;

[0052] S2, mapping the to-be-predicted embedding vector to a low-dimensional space with a fixed dimension by using an average pooling method, and obtaining a to-be-predicted feature vector representing the original gene sequence after reducing the dimension of the embedding;

[0053] S3, inputting the to-be-predicted feature vector into a pre-trained phage multi-host prediction model to predict the infection probability of the multi-host; wherein the phage multi-host prediction model is used to identify the correlation between the phage gene sequence and the specific host bacteria.

[0054] The present application constructs a system and its architecture for automatically extracting features from phage gene sequences and predicting multiple hosts. The architecture uses deep learning technology for sequence analysis. Directly using the complete phage virus sequence as input avoids information loss caused by relying on statistical properties and significantly enhances the accuracy of feature extraction. The present system uses the attention mechanism based on the BERT pre-training model to efficiently convert the virus sequence into a numerical vector, laying a solid foundation for the multi-host prediction task. In addition, the embodiment of the present application also uses a classification network model based on a deep neural network, which realizes the seamless migration of the model to other gene sequence data sets by clearly distinguishing between the pre-training and fine-tuning stages. The system can also process gene sequences of different lengths ranging from a few hundred to tens of thousands of base pairs, greatly improving its versatility and flexibility.

[0055] As a preferred scheme, the training method of the pre-trained classification network model comprises:

[0056] S11, divide the bacteriophage gene sequence in the database into an ordered k-mer word sequence, and process the k-mer word sequence through a pre-trained BERT model to obtain an embedding vector; the database is selected from a standard dataset VHM dataset, a complete genome containing bacteriophage gene sequence and host infection information, and interaction data between bacteriophages and their hosts are recorded;

[0057] S12, using an average pooling method to map the embedding vector in a high-dimensional space to a fixed dimension in a low-dimensional space, and obtaining a feature vector representing the original gene sequence after reducing the dimension of the embedding;

[0058] S13, inputting the feature vector into a pre-trained classification network model to predict the infection probability of multiple hosts.

[0059] The present application can perform pre-training on a large amount of unlabeled gene sequence data, thereby constructing a feature extraction framework for gene sequences, and significantly reducing the dependence of the present application on labeled virus and gene data. Compared with the prior art, the bacteriophage virus multi-host prediction method developed by the present application greatly improves the prediction performance. The BERT model is applied to the prediction task of bacteriophage hosts, and remarkable results are achieved.

[0060] The present application is a deep learning method for predicting the multi-host of the gene sequence of bacteriophage virus. The main function is to use a deep learning feature extraction model to extract features from highly diverse and unstructured bacteriophage gene sequences, and obtain a feature vector that can represent the characteristics of the original bacteriophage gene sequence. Then use these feature vectors that show excellent separability in the feature space to perform subsequent bacteriophage virus multi-host prediction tasks; finally, the feature vector obtained by the feature extraction model of the present application realizes a bacteriophage virus multi-host prediction tool, which achieves good prediction results on known datasets, and can use a small part of bacteriophage gene data in the publicly published multi-modal bacteriophage virus database to outperform other methods using all the data in performance.

[0061] The training method of the bacteriophage multi-host prediction model of the present application will be further described in detail below in combination with the drawings and specific embodiments.

[0062] Figure 2The overall workflow diagram for the embodiment of the present application is shown in the figure. The method flow framework outline includes the following parts: (A) phage gene sequence embedding and BERT model pre-training: first, long phage gene sequences are processed by dividing them into shorter fragments of fixed length and marking the same source of these fragments. Then, the pre-embedded k-mers sequence is used to pre-train the BERT model. (B) Dimension reduction processing: the feature vectors obtained by the pre-trained BERT model belonging to the same original sequence are subjected to average pooling processing to reduce the dimension of the embedding and generate a feature vector representing the original phage gene sequence. (C) Classification network training: the feature vectors obtained in the foregoing steps are used for classification training of the neural network, and finally a classification network capable of predicting the infectivity of phage virus sequences to specific hosts, i.e. a phage multi-host prediction model, is constructed.

[0063] Figure 3 The embedding, pre-training and feature extraction flowchart of the present application is shown in the figure. The flow framework outline includes the following steps: (A) embedding processing: long-chain phage gene sequences are converted into short sequence fragments in the form of k-mers, such as 6-mer fragments. (B) model pre-training: the k-mers masked in the training process are recovered and correctly predicted by calculating and optimizing the cross-entropy loss. (C) feature vector extraction: in the pre-trained model, the embedding sequence is directly input without the mask step, so as to obtain the embedding vector of the [CLS] marker in each short sequence fragment. Then, the average value of these vectors is taken to form a feature vector representing the complete original phage gene sequence.

[0064] The following is a detailed description of each step:

[0065] (i) Sequence tokenization and embedding: as an initial step, the virus original sequence is converted into a format that can be effectively processed by machine learning algorithms. The method of dividing the virus sequence into an ordered sequence of k-mer words is adopted, and the BERT language model is used to vectorize these k-mer word sequences to learn their embedding representation.

[0066] (ii) Dimension reduction processing: since the BERT model generates embeddings in a high-dimensional space for each sequence, and the vector dimension varies due to the inconsistency of virus sequence lengths, the average pooling method is applied to project these high-dimensional vectors into a low-dimensional space with fixed dimensions, thereby obtaining a unified representation of the virus sequence.

[0067] (iii) Training the classification network: after dimension reduction, the vector representing the gene sequence is input into a deep neural network to predict the infection probability of the sequence to a specific number of hosts.

[0068] I. Gene sequence data preprocessing and pre-training model feature extraction specifically includes the following steps:

[0069] The embodiments of the present application adopt the BERT model, which contains a Transformer network using self-attention mechanism. This Transformer network performs well in capturing sparse correlations in long gene sequences, and is particularly suitable for processing accurate embedding representation learning of ordered data such as text or video. Recently, researchers have attempted to use BERT for DNA sequence encoding (e.g. DNABERT), mainly for solving classification problems, which usually involves truncating long gene sequences to adapt to the processing capacity of BERT.

[0070] Further, in order to solve the long gene sequence processing challenge in maternal prediction, the embodiments of the present application adopt the strategy of average pooling. Specifically, the long gene sequence is divided into multiple fixed-length and non-overlapping fragments, and these fragments are processed separately by the BERT model to obtain embedding vectors. Subsequently, these vectors are processed by average pooling to synthesize a fixed-dimensional vector.

[0071] For example, in the operation process, the gene sequence is divided into fragments with a length of 500 nt, and these fragments are further decomposed into (500-k+1) overlapping k-mer words, and the value of k is set to 6. These k-mer word sequences are used to train the BERT embedding model, and each k-mer is processed as an independent unit in the sentence. In order to remove sequencing noise, k-mers with very low frequency (e.g., less than 100 times can be specified as very low frequency) in the entire corpus are considered as false insertions and removed.

[0072] In addition, the embodiments of the present application also introduce five special markers, including [BAR] markers at the beginning of the sequence as the identification of the original genome of the fragment; [PAD] for sequence padding; [UNK] represents unknown units; [SEP] identifies the end of the sequence; [MASK] for masking operation. As shown in Figure 3 A, the long gene sequence is segmented, tokenized, and converted into a tokenized sentence using this method, and [BAR] and [SEP] markers are added at the beginning and end of the sequence, respectively. In the technology of the embodiments of the present application, the embedding representation of each word is obtained by adding the position embedding vector and the paragraph embedding vector, as shown in Figure 3 A. The data preprocessing strategy of the embodiments of the present application is to divide the longer continuous sequence into a series of fixed-length short sequence fragments. Then, these fragments are converted into k-mer word sequences, and [BAR] (as a unique identifier of the sequence) and [SEP] (representing the termination of the sequence) are introduced at the beginning and end of the sequence, respectively. In the proposed model, the embedding of each word is combined by the position embedding vector and the paragraph embedding vector, and the sum of the two constitutes the vector representation of the corresponding word, as shown in Figure 3 A.

[0073] Optionally, in this embodiment of the invention, the BERT-based framework has been adjusted and optimized to more efficiently pre-train the model for processing DNA sequence data. Specifically, a random word masking mechanism is introduced, forcing the model to predict these masked words based on contextual cues. This process involves analyzing neighboring k-mers to infer the target k-mer, thereby promoting deeper contextual understanding.

[0074] The main features of the pre-trained model given in this embodiment of the invention include:

[0075] First, a sequential masking strategy was implemented, which continuously masked k-mer word sequences of different lengths according to a normal distribution with a mean of k and a standard deviation of k-3.

[0076] Secondly, a dynamic k-mer masking approach was adopted. In the early stages of training, the masked k-mers accounted for an initial proportion p (usually set to 10%) of the total input k-mers. As training progressed, this masking proportion gradually increased by a fixed percentage. Specifically, starting from the second training cycle, the masking proportion increased by a fixed percentage q per cycle, meaning that in the i-th training cycle, the proportion of masked k-mers reached (p+q(i-1)). These strategies were designed to continuously increase the challenge of training, enabling the model to capture richer information.

[0077] During pre-training, the cross-entropy loss function can be used to optimize the model, thereby ensuring the efficiency and effectiveness of pre-training.

[0078] Where, y′ i y represents the true probability. i This represents the probability of each class predicted by the model. In this embodiment, the model was pre-trained for approximately 100k steps, covering five training epochs, and implemented on a large unlabeled dataset. The initial batch size was set to 64, and an adaptive adjustment strategy was employed to optimize the training process. To effectively manage GPU memory, a cumulative gradient method was performed, i.e., calculating and updating network parameters after a set number of pre-training steps, followed by gradient clearing, and then starting the next training epoch, with an initial cumulative gradient step count of 48. Furthermore, this embodiment uses the Adam optimizer, and the learning rate is adjusted according to a trapezoidal period to optimize performance, with an initial learning rate set to 5e-4.

[0079] To prevent overfitting while saving computational resources, the early stopping strategy is adopted in the embodiments of the present application. The structure of the model is inspired by BERT base, including 12 Transformer layers, each layer with 512 hidden units and 8 attention heads. All models use the same parameter configuration in the pre-training stage. Model training is carried out on a system equipped with 4 NVIDIA TITAN Xp GPUs, using single-precision floating-point numbers. By pre-training on a wide range of unlabeled corpora, the embodiments of the present application can combine embedding vectors from different subsequences of the same viral sequence through the average pooling strategy, thereby obtaining a vector representing the entire viral sequence, which provides strong support for subsequent tasks.

[0080] II. Constructing a phage multi-host classification network

[0081] In the pre-training stage of the model, the embodiments of the present application process all available data sets. In the fine-tuning training stage, the embodiments of the present application retain the BERT model parameters trained previously, and only adjust the parameters of the network part used for classification. Although both the application of deep neural network classification model in gene sequence human infectivity prediction and phage multi-host prediction rely on deep learning technology, there are some key differences in network structure. In the prediction of human infectivity of gene sequence, the deep neural network performs a binary classification task. When training the network structure, only the classification network model needs to learn the mapping relationship between the gene sequence and human infectivity. In contrast, the deep neural network used for phage multi-host prediction model focuses on identifying the relationship between phage gene sequence and specific host bacteria. This model uses a more complex network structure and performs a multi-classification task to map the original viral feature vector to the multi-host class label. This makes the deep learning classification network structures of the two different. Specifically, the present application builds a classification network based on multilayer perceptron, a classification network based on Transformer and a convolutional neural network to predict the multi-host label of phage virus.

[0082] Next, the embodiments of the present application divide the data into training set and test set, and carry out 10 training cycles to complete the training of the classification network. This process aims to optimize the performance of the network and ensure that it achieves high accuracy in predicting whether the phage can infect multiple hosts.

[0083] The advantages of the embodiments of the present application are as follows:

[0084] Firstly, the attention-based language model concept is introduced into the representation learning of phage gene sequences, and the deep feature representation of DNA sequences is extracted using self-supervised learning technology, which can reveal the deep semantic information in phage gene encoding. Secondly, the effective dimension reduction of gene sequence representation vector is realized, which ensures that the model can process sequences of different lengths, thereby adapting to different analysis scenarios of short sequences and long gene sequences. Finally, the attention-based language model technology is applied to the multi-host prediction problem of phages for the first time. The model can effectively analyze the potential infection risk of unknown virus sequences to multiple hosts. The method of the embodiment of the present application performs better than the prior art in the multi-host prediction task of phages, proving the wide applicability and outstanding effectiveness of the method proposed in the embodiment of the present application in this field.

[0085] The method proposed in the embodiment of the present application performs unsupervised feature extraction and multi-host prediction on phage gene sequences, and experiments are performed on a benchmark dataset. In order to verify the effect of the phage gene multi-host prediction method of the embodiment of the present application, a standard dataset VHM dataset is used for experiment.

[0086] These datasets mainly collect and provide complete genomes containing host infection information based on the NCBI database (website: https: / / www.ncbi.nlm.nih.gov / labs / virus / vssi / # / ). The VHM dataset was first created in 2017 by (Prior Art 1), and then updated by (Prior Art 2) and (Prior Art 3).

[0087] The dataset contains the interaction data between phages and their hosts recorded in the NCBI RefSeq database before 2020. The training set consists of 1,306 pairs of phage-host interactions discovered before 2015, covering 187 bacterial species; while the test set consists of 634 pairs of newly discovered phage-host interactions from 2015 to 2020, covering 95 bacterial species. The performances of various models trained and tested based on the VHM dataset are evaluated by integrating the methods mentioned in (Prior Art 1) (Prior Art 2) (Prior Art 3).

[0088] Figure 4 The comparison figure of different methods on the VHM baseline dataset. The performance of various methods compared in the embodiment of the present application on the VHM benchmark dataset is shown. The vertical axis shows the prediction accuracy, and the horizontal axis represents different classification levels in order from left to right, including species, genus, family, order, class, and phylum. Different patterns are used to distinguish different prediction methods.

[0089] As Figure 4As shown in the figure, the figure lists the accuracy of 11 models on the test set at different classification levels (door, class, order, family, genus, species). At the genus classification level, the accuracy of the host prediction is between 30%-60%, and only four models, PHIAF, DeepHost, VHM-net and vHULK, can predict at the species classification level, with a prediction accuracy of about 40%.

[0090] The present application also uses the prediction accuracy as an indicator to evaluate the feasibility of the method. The method of the present application is compared with the latest phage virus DNA sequence multi-host prediction classification tool and the prior art method. In the experiment, the method of the present application follows the above experimental idea, and the model is pre-trained using the unlabeled data set. After obtaining the pre-trained model, the VHM data set is feature extracted, and the virus sequence data is vectorized. Then, the DNA fragments are trained for phage virus DNA sequence multi-host prediction classification on the same classification network. The final experimental results show that the phage host prediction method reaches a prediction accuracy of 91% at the genus level, surpassing all the methods. The experiment verifies the feasibility of the method.

[0091] As Figure 5 shown, the second object of the present application is to provide a phage virus multi-host prediction system, comprising:

[0092] An embedding vector unit is configured to divide the to-be-predicted phage gene sequence into an ordered k-mer word sequence, and process the k-mer word sequence through a pre-trained BERT model to obtain a to-be-predicted embedding vector;

[0093] A feature vector unit is configured to map the to-be-predicted embedding vector into a fixed-dimensional low-dimensional space by using an average pooling method, and obtain a to-be-predicted feature vector representing the original gene sequence after reducing the dimension of the embedding;

[0094] A classification prediction unit is configured to input the to-be-predicted feature vector into a pre-trained phage multi-host prediction model to predict the infection probability of the multi-host; wherein the phage multi-host prediction model is configured to identify the association between the phage gene sequence and the specific host bacteria.

[0095] As Figure 6 shown, the third object of the present application is to provide an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned phage virus multi-host prediction method.

[0096] A fourth object of the embodiments of the present application is to provide a computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the above-mentioned bacteriophage virus multi-host prediction method.

[0097] A fifth object of the embodiments of the present application is to provide a computer program product comprising computer instructions, characterized in that the computer instructions instruct a computer to execute the above-mentioned bacteriophage virus multi-host prediction method.

[0098] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the function specified in the flow Figure 1 one or more flows and / or blocks in the flow Figure 1 one or more blocks or multiple blocks.

[0099] These computer program instructions can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the function specified in the flow Figure 1 one or more flows and / or blocks in the flow Figure 1 one or more blocks or multiple blocks.

[0100] The present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) having computer usable program code embodied therein.

[0101] The present application is described with reference to the flowcharts and / or block diagrams according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general purpose computer, a special purpose computer, an embedded processor or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices produce an apparatus for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks in the flow Figure 1 one or more blocks or multiple blocks.

[0102] Obviously, the described embodiments are only part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor should belong to the protection scope of the present application.

[0103] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them. Although the present application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that the specific embodiments of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered in the protection scope of the claims of the present application.

Claims

1. A method for predicting multiple hosts of bacteriophage viruses, characterized in that, include: The phage gene sequence to be predicted is divided into ordered k-mer word sequences, and the k-mer word sequences are processed by a pre-trained BERT model to obtain the embedding vector to be predicted. The average pooling method is used to map the embedding vector to be predicted into a low-dimensional space with a fixed dimension, thereby reducing the dimension of the embedding and obtaining the feature vector to be predicted that represents the original gene sequence. The feature vector to be predicted is input into a pre-trained phage multi-host prediction model to predict the probability of infection to multiple hosts; wherein, the phage multi-host prediction model is used to identify the association between phage gene sequences and specific host bacteria.

2. The phage virus multi-host prediction method according to claim 1, characterized in that, Training methods for pre-trained classification network models include: The phage gene sequences in the database are divided into ordered k-mer word sequences, and the k-mer word sequences are processed by a pre-trained BERT model to obtain embedding vectors; the database is selected from the standard VHM dataset, which contains the complete genome of phage gene sequences and host infection information, as well as data recording the interaction between phages and their hosts; The average pooling method is used to map the high-dimensional embedding vector to a fixed-dimensional low-dimensional space, thereby reducing the dimension of the embedding and obtaining the feature vector representing the original gene sequence. The feature vectors are input into a pre-trained classification network model to predict the probability of infection in multiple hosts.

3. The phage virus multi-host prediction method according to claim 2, characterized in that, The process of dividing the bacteriophage gene sequence into ordered k-mer word sequences refers to dividing a long gene sequence into multiple fixed-length, non-overlapping segments.

4. The phage virus multi-host prediction method according to claim 2, characterized in that, The BERT model includes a Transformer network employing a self-attention mechanism; the pre-trained BERT model processes k-mer word sequences to obtain embedding vectors and also introduces markers, specifically: The [BAR] marker at the beginning of the sequence serves as an identifier for the original genome fragment; [PAD] is used for sequence filling; [UNK] represents an unknown unit; [SEP] indicates the end of the sequence; and [MASK] is used for masking operations. Long gene sequences are segmented and labeled, k-mer word sequences are converted into labeled sentences, and [BAR] and [SEP] markers are added at the beginning and end of the sequence, respectively; the embedding representation of each word is obtained by adding the position embedding vector and the paragraph embedding vector, and the continuous sequence is segmented into a series of fixed-length segments; The fragments are converted into k-mer word sequences, and [BAR] and [SEP] are introduced at the beginning and end of the sequence, respectively. The embedding of each word is formed by combining the position embedding vector and the paragraph embedding vector, and the sum of the two constitutes the vector representation of the corresponding word.

5. The phage virus multi-host prediction method according to claim 2, characterized in that, The BERT model incorporates a random word masking mechanism during pre-training, specifically: Following a normal distribution with mean k and standard deviation k-3, continuously mask k-mer word sequences of different lengths; In the early stages of training, the masked k-mers account for an initial proportion p of the total input k-mers; As training progresses, the occlusion ratio will gradually increase by a fixed percentage: starting from the second training cycle, the occlusion ratio increases by a fixed percentage q in each cycle, and in the i-th training cycle, the k-mer occlusion ratio reaches p+q(i-1). During pre-training, the cross-entropy loss function is used to optimize the model, and an adaptive adjustment strategy is employed for training.

6. The phage virus multi-host prediction method according to claim 2, characterized in that, The method employs average pooling to map the high-dimensional embedding vector to a fixed-dimensional low-dimensional space, reducing the dimensionality of the embedding to obtain a feature vector representing the original gene sequence, including: The gene sequence was segmented into fixed-length fragments, then further decomposed into multiple overlapping k-mer words; and k-mer word sequences that appeared less than a threshold in the entire corpus were removed. Embedding vectors from different subsequences of the same viral sequence are merged to obtain a feature vector representing the entire viral sequence.

7. The phage virus multi-host prediction method according to claim 2, characterized in that, The phage multi-host prediction model employs a deep neural network classification model and, based on the BERT model parameters, trains a mapping relationship between gene sequences and human infectivity. We perform multi-classification tasks and build classification networks based on multilayer perceptrons, Transformers, and convolutional neural networks to predict multi-host labels for bacteriophage viruses. Then, we map the original virus feature vectors onto the multi-host labels.

8. A bacteriophage virus multi-host prediction system, characterized in that, include: Embedding vector units are used to divide the phage gene sequence to be predicted into ordered k-mer word sequences, and the k-mer word sequences are processed by a pre-trained BERT model to obtain the embedding vector to be predicted. The feature vector unit is used to map the embedding vector to be predicted into a low-dimensional space with a fixed dimension using the average pooling method, thereby reducing the dimension of the embedding and obtaining the feature vector to be predicted that represents the original gene sequence. A classification prediction unit is used to input the feature vector to be predicted into a pre-trained phage multi-host prediction model to predict the probability of infection to multiple hosts; wherein the phage multi-host prediction model is used to identify the association between phage gene sequences and specific host bacteria.

9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the bacteriophage multi-host prediction method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the bacteriophage virus multi-host prediction method according to any one of claims 1-7.