Marine bacteriophage antibacterial lysin prediction method, device, medium and equipment

Through the ProtT5 and ProtBERT deep learning framework combined with natural language processing technology, a marine phage antibiotic lysin prediction model was constructed, global marine microbiome data was integrated, and a systematic database was established, which solved the problems of scarcity and low efficiency in lysozyme development, and achieved high-precision and low-cost antibiotic lysin screening and verification.

CN120260696AActive Publication Date: 2025-07-04OCEAN UNIV OF CHINA

Patent Information

Application Number
CN202510327813.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

In the prior art, lysozyme development methods are scarce, the model is robust and reliable, the prediction results are high false positive rates, the verification method is single, and the screening and optimization efficiency are low, resulting in limited antibiotic lysozyme development efficiency and wide application.

Method used

The ProtT5 and ProtBERT deep learning framework combined with natural language processing technology is used to build a marine phage antibiotic lysin prediction model. Through multi-dimensional verification methods, global marine microbiome and phage group data are integrated, a systematic antibiotic lysin database is established, the data integration platform and algorithm design are optimized, and the technical threshold is lowered.

Benefits of technology

It significantly improves the accuracy and robustness of antibiotic lysin prediction, improves screening and verification efficiency, expands the source of antibiotic lysin, reduces development costs and technical thresholds, promotes diversified innovation, and solves the problems of insufficient resources and low efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260696A_ABST
    Figure CN120260696A_ABST
Patent Text Reader

Abstract

The invention discloses a marine bacteriophage antibacterial lysin prediction method, device, medium and equipment, and relates to the field of biological medicine and computational biology. Comprising the following steps: constructing a training data set based on a globally known antibacterial protein sequence; and training a protein sequence classification model based on deep learning by using the antibacterial protein data set to obtain a marine antibacterial lysin prediction model. The method comprises the following steps: carrying out data standardization processing on a to-be-predicted bacteriophage protein sequence, inputting the to-be-predicted bacteriophage protein sequence into a protein language model, carrying out self-attention modeling on the input sequence through a ProtT5 module, and extracting global structure and function related features to obtain a first feature vector; extracting a short sequence local dependency feature through a ProtBERT module to obtain a second feature vector; performing weighted summation on the first feature vector and the second feature vector to generate a fusion feature vector; and carrying out probability prediction on the fusion feature vector, and outputting the antibacterial tag and the prediction confidence of the marine phage sequence to be predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of biomedicine and computational biology, and particularly to a method, device, medium, and equipment for predicting marine phage antibacterial lysins. Background Art

[0002] In the field of antibacterial drug development, traditional methods mainly rely on screening antibacterial components from nature or obtaining antibacterial drugs through chemical synthesis. Lysozyme, as a natural antibacterial substance, has been widely used to combat bacterial infections, especially against certain drug-resistant bacteria.

[0003] However, traditional methods for developing lysozyme cannot overcome the problem of high false positives: In the past, traditional machine learning methods (such as support vector machines, random forests, etc.) were mostly used for predicting antibacterial proteins. However, the logic frameworks of these algorithms are simple, making it difficult to process massive and complex data, and the prediction results often have a high false positive rate, resulting in low robustness and reliability of the model. Summary of the Invention

[0004] Based on this, in order to solve the technical problems in the prior art, the present invention provides a method, device, medium, and equipment for predicting marine phage antibacterial lysins.

[0005] The present invention provides a method for predicting marine phage antibacterial lysins, including:

[0006] Construct a deep learning network including a ProtT5 module, a ProtBERT module, a weighting module, and a classification module. Based on the collected antibacterial protein sequence data, construct a data set, and use the data set to train the constructed deep learning network to obtain a marine phage antibacterial lysin prediction model;

[0007] Input the sequence of the marine phage antibacterial lysin to be predicted into the prediction model. Through the ProtT5 module, perform encoding and decoding operations on the sequence of the marine phage antibacterial lysin to be predicted to capture the global features in the sequence of the marine phage antibacterial lysin to be predicted, and obtain a first feature vector; through the ProtBERT module, perform encoding operations on the sequence of the marine phage antibacterial lysin to be predicted to capture the sequence local dependence relationship in the sequence of the marine phage antibacterial lysin to be predicted, and obtain a second feature vector; through the weighting module, perform weighted summation on the first feature vector and the second feature vector to obtain a fused feature vector; through the classification module, perform probability prediction on the fused feature vector, and output the confidence that the sequence of the marine phage protein to be predicted is an antibacterial protein sequence.

[0008] Further, the performing encoding and decoding operations on the sequence of the marine phage antibacterial lysin to be predicted through the ProtT5 module specifically includes:

[0009] In the encoding stage of ProtT5: The sequence of the marine phage antibacterial lysin to be predicted is position-encoded through the position encoding layer in the ProtT5 module. Combining with the protein sequence embedding operation, a protein sequence embedding vector is obtained. The global dependencies of the protein sequence embedding vector are captured through the multi-head self-attention mechanism layer in the ProtT5 module to generate a context-weighted protein sequence vector. The weighted protein sequence vector is non-linearly transformed through the first feed-forward neural network layer in the ProtT5 module, and combined with residual connection and normalization operations to obtain protein encoding features.

[0010] In the decoding stage of ProtT5, the protein encoding features are sequence-reconstructed through the sequence reconstruction layer in the ProtT5 module to output a predicted amino acid sequence vector. The amino acid sequence vector is masked through the masked multi-head self-attention mechanism layer in the ProtT5 module to capture the local dependencies of amino acids and obtain an amino acid masked weighted vector. Through the interactive attention mechanism layer in the ProtT5 module, the global dependencies between the protein encoding features and the amino acid masked weighted vector are calculated to obtain a fused weighted amino acid sequence vector. The fused weighted amino acid sequence vector is non-linearly transformed through the second feed-forward neural network layer in the ProtT5 module, and combined with residual connection and normalization operations to obtain a first feature vector.

[0011] Furthermore, the encoding operation of the marine phage antibacterial lysin sequence to be predicted through the ProtBERT module specifically includes:

[0012] The sequence of the marine phage antibacterial lysin to be predicted is position-encoded after being aligned to a preset fixed length through the position encoding layer in the ProtBERT module. Combining with the protein sequence embedding operation, a protein sequence embedding vector is obtained.

[0013] The protein sequence embedding vector is weighted through the multi-head self-attention layer in the ProtBERT module to capture the global dependencies of amino acids and obtain a weighted protein sequence vector.

[0014] The weighted protein sequence vector is randomly masked through the masked language modeling layer in the ProtBERT module to generate amino acid masked features.

[0015] The amino acid masked features are feature-aggregated through the pooling layer in the ProtBERT module to obtain a second feature vector.

[0016] Furthermore, the weighted summation of the first feature vector and the second feature vector through the weighting module to obtain a fused feature vector specifically includes:

[0017] E = α·E ProtT5 +(1 - α)·EProtBERT

[0018] Among them, E ProtT5 is the first eigenvector, and E ProtBERT is the second eigenvector, and α is the adaptive weight coefficient.

[0019] Furthermore, the probability prediction of the fused feature vector by the classification module specifically includes:

[0020] Receiving the normalized fused feature vector through the input layer in the classification module and passing it to the hidden layer in the classification module to extract the deep non-linear relationship in the fused feature vector:

[0021] h i = ReLU(W hid,i ·h i-1 + b hid,i )

[0022] Among them, h i and h i-1 are the output feature vectors of the i-th layer and the (i - 1)-th layer hidden layers respectively, i ∈ {1, …, n}, and n is the total number of hidden layer layers; W hid,i is the weight matrix of the i-th layer hidden layer; b hid,i is the bias vector of the i-th layer hidden layer;

[0023] Outputting the probability value through the output layer in the classification module:

[0024] P = σ(W out ·h n + b out )

[0025] Among them, σ(·) represents the Sigmoid function; h n is the output feature vector of the n-th layer hidden layer; W out is the weight matrix of the output layer; b out is the bias vector of the output layer; P represents the confidence that the output marine phage protein sequence to be predicted is an antibacterial protein sequence.

[0026] Furthermore, the construction of the data set based on the collected antibacterial protein sequence data specifically includes:

[0027] Obtain relevant sequence data from the global known phage genome database, which includes DDBJ, CHVD, EMBL, Genebank, GOV2, GPD, GVD, IGVD, IMG_VR, MGV, PhagesDB, RefSeq, STV, and TemPhD; obtain marine metagenomic sequencing data from the marine metagenomic database, which includes ENA, SRA, and ProGenomes3; extract antibacterial protein sequence data from the positive antibacterial protein sequence database, which includes NCBI, UniprotKB / Swiss-Prot, PDB, RefSeq, EMBL, GenBank, and PIR; integrate the known phage whole-genome data, marine metagenomic sequencing data, and antibacterial protein sequence data to obtain the original antibacterial protein sequence dataset;

[0028] Preprocess the original antibacterial protein sequence dataset to obtain the processed antibacterial protein sequence dataset; the preprocessing includes data integrity verification, standardizing the format, file segmentation, removing empty sequences, cleaning abnormal symbols, trimming sequences, deduplication, redundancy removal, quality control, gene annotation, and species classification;

[0029] Use different bioinformatics tools to identify phage sequences in the processed antibacterial protein sequence dataset, and take the union of the identification results to obtain the final dataset; the bioinformatics tools include PhiSpy, Iphop, virsorter2, DeepVirFinder, VIBRANT, kraken2, viralVerify, Metaphinder, geNomad, virfinder, PPR-Meta, and Phigaro.

[0030] The present invention provides a device for predicting marine phage antibacterial lysins, comprising:

[0031] A model construction module for constructing a deep learning network including a ProtT5 module, a ProtBERT module, a weighting module, and a classification module, constructing a dataset based on the collected antibacterial protein sequence data, and using the dataset to train the constructed deep learning network to obtain a marine phage antibacterial lysin prediction model;

[0032] A prediction module is used to input the sequence of the marine phage antibacterial lysin to be predicted into a prediction model, perform encoding and decoding operations on the sequence of the marine phage antibacterial lysin to be predicted through the ProtT5 module to capture the global features in the sequence of the marine phage antibacterial lysin to be predicted, and obtain a first feature vector; perform encoding operations on the sequence of the marine phage antibacterial lysin to be predicted through the ProtBERT module to capture the sequence local dependence relationship in the sequence of the marine phage antibacterial lysin to be predicted, and obtain a second feature vector; perform weighted summation on the first feature vector and the second feature vector through a weighting module to obtain a fused feature vector; perform probability prediction on the fused feature vector through a classification module, and output the confidence that the sequence of the marine phage protein to be predicted is an antibacterial protein sequence.

[0033] The present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned method for predicting marine phage antibacterial lysin is implemented.

[0034] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned method for predicting marine phage antibacterial lysin is implemented.

[0035] The above at least one technical solution adopted by the present invention can achieve the following beneficial effects:

[0036] In the method for predicting marine phage antibacterial lysin provided by the present invention, aiming at the problems of high false positive rate of traditional algorithms and single model optimization method, the present invention adopts a deep learning framework based on ProtT5 and ProtBERT to extract context features from protein sequences and combines multi-dimensional verification means, effectively improving the prediction accuracy and robustness of the model. Specifically: first, the ProtT5 module is used to encode and decode the sequence to generate a first feature vector. Based on the global feature extraction ability of ProtT5, the first feature vector reflects the overall folding pattern, long-range amino acid interaction and evolutionary information of the protein; the ProtBERT module extracts the local dependence relationship of the sequence and obtains a second feature vector, and the second feature vector reflects the interaction of short-distance amino acids, co-occurrence patterns and local structural stability; the combination of these two feature vectors obtains a comprehensive fused feature vector through a weighting module. This fused feature vector contains the global information of the protein and the local analysis of amino acids, and through the classification module, a high-confidence prediction of whether the marine phage protein sequence is an antibacterial protein is realized. Compared with traditional algorithms, the present invention enhances the optimization process of the model through multi-dimensional verification means, thereby significantly improving the accuracy and robustness of the prediction model, and providing more reliable technical support for the research and application of marine phage antibacterial proteins. Description of the Drawings

[0037] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0038] Figure 1 It is a schematic diagram of the complete workflow provided by the present invention;

[0039] Figure 2 It is a schematic diagram of the encoder and decoder based on the Transformer architecture provided by the present invention. Detailed implementation manners

[0040] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0041] The traditional methods for developing lysozyme have the following main problems:

[0042] Limited biological sources: Existing lysozymes mainly come from specific bacteria or phages. Although these sources can provide a certain amount of lysozyme, their resource scarcity and limitations in biological acquisition significantly limit large-scale production and wide application. At the same time, with the exponential growth of omics data, the maintenance and update of existing models and biological protein databases lag behind, making it difficult to efficiently process and analyze new and complex antibacterial lysozyme data, resulting in the inability to fully explore the potential antibacterial protein sequence features, which further restricts the development and utilization of antibacterial lysozymes.

[0043] Traditional algorithms cannot overcome the high false positive problem: In the past, traditional machine learning methods (such as support vector machines, random forests, etc.) were mostly used for antibacterial protein prediction. However, these algorithmic logic frameworks are simple, difficult to process massive and complex data, and the prediction results often have a high false positive rate, resulting in low robustness and reliability of the models. Among existing artificial intelligence methods, some natural language processing technologies (such as Word2Vec, FastText, GloVe, etc.) have been applied to the feature extraction of target protein sequences. However, these methods do not combine the latest natural language technologies and deep learning frameworks specifically for processing protein sequences, limiting the prediction accuracy and the exploration of antibacterial potential.

[0044] Single model optimization and validation methods: Existing methods have significant deficiencies in model optimization and validation. The model optimization process lacks systematicness. Usually, it only targets limited parameter adjustments and small-scale dataset training, making it difficult to comprehensively improve model performance. In addition, the validation means are simple and unreliable, mostly relying on low-throughput and single-dimensional experimental validation, and unable to fully evaluate the accuracy of prediction results and the actual activity of antibacterial lysozyme. This limitation not only reduces the validation efficiency but also makes it difficult to translate the model prediction results into actual medical applications and drug development achievements, significantly restricting the process of lysozyme development.

[0045] Low efficiency in screening and optimization processes: Currently, the screening and optimization of most lysozymes rely on traditional laboratory methods such as clone screening and directed evolution. These methods are usually time-consuming, laborious, and costly, unable to quickly mine potential lysozyme sequences from large-scale biological resources, further restricting the development efficiency and wide application of lysozymes.

[0046] The potential of antibacterial lysins in clinical treatment has not been fully explored: As emerging drugs to replace antibiotics against drug resistance, antibacterial lysins have the advantages of broad-spectrum antibacterial activity and low production costs. However, the current research status of antibacterial lysins is significantly insufficient compared to other antibacterial proteins (such as antibacterial peptides) and antibacterial prediction models. Since their first discovery in 1971, less than 2% of lysins have been successfully identified and experimentally verified, and only five lysins have entered the clinical trial stage.

[0047] Narrow antibacterial spectrum: Existing first-line antibacterial proteins (such as antibacterial peptides) mainly against bacterial infections are mostly only effective against specific types of bacteria, with a relatively limited antibacterial spectrum and difficult to cover diverse pathogens. Traditional screening methods are less efficient in discovering broad-spectrum antibacterial proteins, restricting their wide application in clinical practice. In contrast, the antibacterial lysins developed in this project not only have a broader antibacterial spectrum but also can provide effective solutions for different types of drug-resistant pathogens, expanding the application scope for antibacterial drug development.

[0048] Incomplete reference databases: Existing databases are mostly based on terrestrial microbial resources, lacking systematic integration of marine microorganisms and phage lysins, resulting in a single data source and insufficient samples for training models. This project mines data from the global marine microbiome and phageome to establish the first systematic marine antibacterial lysin database, providing strong support for the comprehensive development of antibacterial lysins and promoting the in-depth utilization of the unexploited resource treasure trove of the ocean.

[0049] Computer prediction problems and cost control defects in pharmaceutical company drug design: Pharmaceutical companies widely use computer-aided design (CADD) technology. However, existing tools mainly focus on small molecule drug development and lack dedicated prediction methods suitable for macromolecules (such as antibacterial lysins), with insufficient prediction accuracy and high computational costs. This project combines deep learning and natural language processing models specifically for protein sequences, not only significantly improving the prediction accuracy of antibacterial lysins but also reducing computational costs through efficient model optimization, providing a cost-effective new method for pharmaceutical company drug development.

[0050] Lack of a standardized functional verification system: The functional verification of existing antibacterial lysins mostly relies on low-throughput experiments and lacks a unified standardized process, resulting in difficult-to-compare results. This project established a standardized antibacterial lysin activity verification system by introducing high-throughput screening technology and multi-dimensional verification means, significantly improving the verification efficiency and result reliability, providing strong support for the clinical development of lysins.

[0051] Insufficient data integration and interoperability: Existing data are scattered across multiple omics platforms, lacking a unified format and standard, restricting the development and exploration of antibacterial lysins. This project integrated metagenomic data, protein sequence data, and phage data, developed a unified data processing and analysis platform, improved the efficiency of data integration and interoperability, and achieved in-depth exploration of antibacterial lysin resources.

[0052] Research on phages and lysins is limited to specific species: Current research mostly focuses on a few phage and lysin sources, ignoring the globally diverse microbial resources. This project focuses on exploring high-value lysin resources from the unexploited ecosystem of the entire ocean, providing a rich set of new candidates for research in the field.

[0053] Technical barriers limit the participation of small and medium-sized research institutions: The development technology of antibacterial lysins is complex and has a high threshold, restricting the participation of small and medium-sized institutions. The artificial intelligence-assisted method developed in this project is suitable for large-scale data processing, while optimizing the generality and operational convenience of the algorithm, reducing the research and development threshold, and promoting diverse innovation in the field.

[0054] To address these existing problems, the research of this invention has significantly improved the efficiency and application potential of antibacterial lysin development, providing a new technical path and solution for dealing with the drug resistance crisis. The technical solution involved in this invention belongs to the fields of biomedicine and computational biology and is specifically applied to the development of antibacterial drugs. This technical solution uses an artificial intelligence-assisted method to screen and optimize lysozymes from marine microorganisms (such as phages) with the aim of developing new antibacterial lysozymes for the research and development of antibacterial drugs. This invention can be widely applied in the field of anti-infection, especially in dealing with drug-resistant bacteria and developing innovative antibacterial therapies.

[0055] The present invention creates a method for mining and developing antibacterial lysins based on artificial intelligence, which systematically solves problems such as data scarcity, insufficient model optimization, and simple verification systems in the development of existing antibacterial lysins by integrating global marine microbiome and phageome data.

[0056] The following describes in detail the technical solutions provided by each embodiment of the present application with reference to the accompanying drawings.

[0057] Figure 1 It is a schematic flowchart of a method for predicting marine phage antibacterial lysins in this embodiment, which specifically includes the following steps

[0058] S1: Data collection and preprocessing.

[0059] Collect known phage sequences from databases such as DDBJ, CHVD, EMBL, Genebank, GOV2, GPD, GVD, IGVD, IMG_VR, MGV, PhagesDB, RefSeq, STV, TemPhD, etc. These databases cover the genomic information of various phages globally and can provide data on aspects such as phage species, genomic structure, and functional genes. Secondly, to expand the application scope of the data and cover more marine ecological backgrounds, collect marine-related metagenomic sequencing data from databases such as ENA, DDBJ, SRA, ProGenomes3, etc., and predict potential phage sequences from it. Finally, collect positive antibacterial protein sequences from databases such as past verified literature, NCBI, UniprotKB / Swiss-Prot, PDB, RefSeq, EMBL, GenBank, PIR, etc. These databases contain a large amount of verified antibacterial protein information, providing important references for subsequent model construction, feature extraction, and lysin activity prediction.

[0060] The original data was preprocessed, and the steps included verifying data integrity, standardizing the format, splitting files, removing empty sequences, cleaning abnormal symbols, trimming sequences, removing duplicates, removing redundancy, quality control, gene annotation, and species classification, etc., to ensure the accuracy and integrity of the data. Tools such as Globus, Entrez Direct, and iSeq were used for data download. Tools such as Fastp, Cutadapt, and Trimmomatic were used for cleaning sequencing data; Chopper was used for quality control of third-generation sequencing data; Shasta and wtdbg2 were used for the assembly of Nanopore and PacBio data respectively; Unicycler and Flye were used for long-read and hybrid assembly respectively. Seqkit was used for cleaning sequence data, and QUAST tools and self-built libraries were used for quality control and alignment; Bakta and Prokka were used for gene annotation. Other alignments, binning, etc. used traditional bioinformatics tools.

[0061] Use tools such as PhiSpy, Iphop, virsorter2, DeepVirFinder, VIBRANT, kraken2, viralVerify, Metaphinder, geNomad, virfinder, PPR-Meta, and Phigaro to predict and identify phage sequences in the assembled data, and take the union of the identification results.

[0062] S2: Model construction.

[0063] The core of the present invention is to embed and extract features from protein sequences through deep learning models and natural language processing techniques, and construct a classifier based on this for efficient identification and prediction of antibacterial lysins from marine phage data. The following are the specific technical solutions and principles:

[0064] Protein sequence embedding is a key step in constructing the model. Its core goal is to transform the amino acid sequence into a high-dimensional vector representation that can be processed by a computer. In this study, two protein language models (pLMs) based on the ProtTrans framework were used: ProtT5-XXL-BFD and ProtBERT-BFD. Based on the powerful context modeling ability of the Transformer architecture and combined with natural language processing (NLP) techniques, they are trained in a dedicated large protein database to efficiently extract global and local features in large-scale protein data and provide high-quality embeddings for downstream tasks.

[0065] S201: Advantages of ProtT5 and ProtBERT

[0066] ProtT5 models protein sequences using a bidirectional Transformer architecture, capturing global context information through the multi-head self-attention mechanism to reveal complex semantic relationships between amino acids. This global modeling approach can not only resolve long-range dependencies in the sequence but also combine local features to accurately characterize the structure and function of proteins. Additionally, the highly optimized pre-training and fine-tuning strategies in the T5 architecture further enhance the model's generalization ability. It is trained on a super-large protein sequence database such as BFD (Big Fantastic Database), learning the evolutionary relationships and structural and functional characteristics of various proteins, enabling it to extract rich feature representations from large-scale protein data. ProtT5 can also handle protein sequences that exceed the capabilities of ordinary machine learning models (such as 1024 or more), capturing long-range dependencies in the sequence through segmented modeling, which is particularly suitable for biological data of phages with long gene sequences. At the same time, ProtT5 uses Masked Language Modeling (MLM), randomly masking 15% of the amino acids in the input sequence during training and allowing the model to predict the masked positions based on the context. This mechanism simulates real mutations in protein sequences while incorporating biological characteristics. For example, the model learns the physicochemical properties (such as hydrophobicity, polarity, and spatial conformation) of amino acids and the evolutionary background (such as conservation among homologous sequences) through the attention mechanism, thereby enhancing the biological relevance of sequence embeddings. This mechanism enables ProtT5 to more accurately capture the biological background of protein sequences, thus improving the accuracy of the model in capturing protein sequence features. The embedded vector is then added with position encoding information, and finally, a reliable protein embedding vector is output. Input into the multi-head attention mechanism of the Transformer, it is used to capture the context relationship of each amino acid position in the protein sequence, understanding how the function of an amino acid in the sequence is affected by other amino acids:

[0067]

[0068] Among them, Q represents the query vector of the current amino acid, used to retrieve context information related to its function. It can be understood as a question vector, such as "What context information does the current amino acid need to determine its function". K represents the information features provided by all positions in the sequence, that is, the key values of each amino acid position determine how other positions interact with it. For example, an amino acid at an active site may attract other amino acid positions related to its function. V contains the specific feature information of amino acids, and the output features are obtained by the weighted value matrix of the mutual relationship between the query matrix and the key matrix. Represents the attention weight, which shows how each amino acid interacts with other amino acids in the sequence. The larger the weight, the stronger the association between two amino acids. The correlation of each amino acid with other amino acids in the sequence is quantified. It is used to prevent the attention weight from being too large when the dimension is large, which may lead to unstable gradients.

[0069] Finally, the context-aware vector output by the multi-head attention mechanism can generate a feature-enhanced representation by calculating the global dependencies within the amino acid sequence, which is used to describe the structural and functional relationships of proteins. Then it is stacked with a feed-forward neural network (FFN). The principle of the FNN is as follows. It is used to perform independent non-linear transformations on the features at each position to extract deeper features:

[0070] t i = Layer normalization(z i + FFN(z i ))

[0071] Here, FFN(z i ) represents performing a non-linear transformation on the context-aware feature z i through the feed-forward neural network. The feed-forward network itself usually consists of two fully connected layers and a non-linear activation function (such as ReLU):

[0072] FFN(z i ) = ReLU(W ff1 z i + b ff1 )W ff2 + b ff2

[0073] Among them, W1 and W2 map the features to high-dimensional and low-dimensional spaces respectively. ff1 and ff2 represent the first and second layers of the feed-forward network respectively. The input features are linearly transformed through W1z i + b1 and mapped to a higher-dimensional feature space. Then, the non-linear modeling ability of the model is increased through the ReLU activation function to capture the complex relationships in the protein sequence. Subsequently, the high-dimensional features are remapped back to the original dimension through W2. This ensures that the input and output feature dimensions are consistent, facilitating subsequent residual connections, that is, the z i + FFN(z i ) part, making the original features and the transformed features added together. In this way, both the original input information is retained and deep features are added. Finally, layer normalization is applied to the residual result. By normalizing the mean and variance of the features of each layer, the problem of gradient disappearance is alleviated and the training process is stabilized.

[0074] After the above part is completed, the encoder of the protein sequence is completed. At this time, the position embedding of each amino acid not only reflects its own information, but also combines the information of other positions in the sequence. Subsequently, a decoder is established, whose function is to extract task-specific features from the context representation generated by the encoder and generate task outputs through additional layers, such as sequence classification or sequence annotation. The decoder consists of an Encoder-Decoder Attention and an output generation layer, and generates specific outputs for the context features according to the task requirements. In this invention, the task is sequence classification, that is, whether it is an antibacterial lysin. The decoder generates global features by taking the average pooling of the features t at all positions i to generate global features:

[0075]

[0076] Then, the classification result is generated through a fully connected layer and an activation function:

[0077] P(y = 1) = σ(W cls ·T global + b cls )

[0078] The advantage of the encoder-decoder architecture is that the context representation it generates can flexibly adapt to protein classification tasks. The decoder precisely transforms the context features generated by the encoder into task-specific outputs through a specially designed output module, ensuring that the functional characteristics of proteins are fully expressed. At the same time, this architecture uses the multi-head self-attention mechanism and multi-layer stacking structure, which can efficiently capture long-range dependencies and complex patterns in protein sequences, and is especially suitable for processing long sequence data such as phages. At the same time, ProtT5 is trained on the BFD database, which contains more than 200 million protein sequences, covering a variety of biological functions and structural information. It significantly improves the model's ability to model sequence complexity and prediction accuracy.

[0079] Another language model, ProtBERT, is a pre-trained language model based on BERT (Bidirectional Encoder Representations from Transformers) and is suitable for extracting local and global features from protein sequences. Similar to ProtT5, ProtBERT is also an architecture of Transformer. The difference is that ProtT5 contains an encoder and a decoder, while ProtBERT does not contain a decoder part (the encoder and decoder are as Figure 2Rather than outputting context-aware features of protein sequences directly through an encoder stack as shown, it focuses on deeply capturing the local dependencies of amino acids in protein sequences. Additionally, although both ProtT5 and ProtBERT are based on MLM, there are the following significant differences in implementation and characteristics. For ProtT5, MLM is a subtask, and its core design is a Text-to-Text task, with MLM being just one of the pre-training objectives. The model can be extended for generation tasks such as sequence translation, sequence filling, reconstructing the complete sequence, etc. This design is more suitable for handling complex context modeling tasks such as genomic-level sequence reconstruction, and the ultimate goal in the present invention is to generate sequences. For ProtBERT, it is a pure MLM model that focuses on the masked language model task without considering context, so it is more efficient in modeling short sequences and local features. For example, for a typical antibacterial lysin sequence (usually in the range of 20 - 200 amino acid lengths), the embedding efficiency of ProtBERT is better than that of ProtT5. Although ProtT5 can handle long sequences (>1024 amino acids), ProtBERT is clearly more advantageous in modeling local features of short sequences.

[0080] ProtBERT adopts the BERT structure, relying only on the encoder part without a decoder. The input sequence needs to be aligned to a fixed length to avoid affecting the calculation efficiency due to dynamic padding; it uses bidirectional attention to calculate the global amino acid dependencies without a unidirectional mask, with a relatively large amount of calculation, and the number of attention heads can be reduced or the Dropout rate can be adjusted to prevent overfitting; ProtBERT uses masked language modeling, randomly masking 15% of the amino acids, and only predicting the masked sites based on the context during the training process without sequence reconstruction. Different from ProtT5, ProtBERT has no decoding process, and the MLM training strategy maintains a fixed attention pattern during inference, and layer normalization or batch normalization is needed to stabilize the feature distribution; since the ProtBERT structure does not contain Encoder-Decoder Attention, it needs to be changed to bidirectional self-attention calculation, and the batchsize should be fixed during inference to avoid the normalization layer being sensitive to distribution changes; ProtBERT uses a pooling layer for feature aggregation, and a global average pooling or max pooling layer needs to be added in the classification task to adapt the output structure, rather than generating prediction results through a decoder.

[0081] S202: The combination strategy of ProtT5 and ProtBERT.

[0082] The embedding features of ProtBERT enhance the model's prediction ability for short sequence or local feature-dominated tasks, while ProtT5 makes up for the deficiencies in long sequence modeling. This synergy achieves higher accuracy and more comprehensive functional analysis in the antibacterial lysin prediction task. To effectively combine the embedding features of ProtT5 and ProtBERT, the present invention designs specific fusion strategies and feature extraction processes to generate high-quality feature vectors suitable for classification tasks.

[0083] First, the method of weighted voting is adopted to maximize the utilization of the prediction results of ProtT5 and ProtBERT. The present invention designs a weighted voting mechanism based on the performance on the validation set. Specifically, the two models are used to independently predict the same sequence, and weights w1 and w2 (the sum is 1) are assigned according to their performance on the validation set and the length ratio of the input data. This method can effectively reduce the possible prediction bias of a single model and enhance the robustness of the classification results. The final fused prediction probability is calculated by the following formula:

[0084] P = w1·P ProtT5 + w2·P ProtBERT

[0085] Subsequently, a weighted average strategy is adopted to integrate the high-dimensional embedding vectors generated by ProtT5 and ProtBERT. The specific calculation formula is:

[0086] E = α·E ProtT5 +(1 - α)·E ProtBERT

[0087] E ProtT5 and E ProtBERT are the embedding vectors output by the two models respectively, and α is the weighting coefficient. α is optimized through cross-validation to ensure that the fused features are more biologically meaningful, taking into account both global context and local fine features. After completing the embedding fusion, to further improve the performance of the classifier, the present invention processes and optimizes the generated feature vectors, mainly including the following steps:

[0088] First, feature normalization is performed. Since the embedding features generated by different models may have differences in numerical ranges, directly inputting them into the classifier may lead to unstable training. Therefore, the present invention performs standardization processing on the fused embedding features, and the formula is as follows:

[0089]

[0090] μ and σ are the mean and standard deviation of the features respectively. Normalization can eliminate the bias of the feature distribution and ensure that the model converges faster during the gradient descent process. During the embedding process, the high-dimensional embedded features may contain redundant information, increasing the computational complexity and potentially introducing noise. The present invention uses principal component analysis (PCA) to reduce the dimension of the features, reduce redundant information, and retain key features at the same time. After normalization and dimensionality reduction are completed, the finally extracted feature matrix has a description of the data dimension and biological significance. The output features retain the key functional regions and evolutionary background of the protein sequence, and are particularly suitable for the prediction and analysis of antibacterial lysins.

[0091] ProtT5 encodes and decodes the marine phage antibacterial lysin sequence to capture the global structural and functional features in the marine phage antibacterial lysin sequence to be predicted, and obtains the first feature vector; ProtBERT encodes the marine phage antibacterial lysin sequence to capture the sequence local dependencies in the marine phage antibacterial lysin sequence to be predicted, and obtains the second feature vector.

[0092] The global information of the first feature vector can define the action mode of the whole protein, such as the folding pattern and sequence function of the protein. It may contain information such as the transmembrane region of the protein, specific domains of lysozyme, and enzyme activity, and tends to define the characteristics of long-sequence proteins; the local features of the second feature vector mainly describe the mutual influence in the primary structure of the protein and reflect the co-occurrence mode of amino acids within a short distance, such as the formation trend of certain conserved secondary structures (α-helix, β-sheet), or the catalytic active sites and receptor binding regions between certain amino acids. However, it cannot reflect the spatial action of the protein (for example, the active site of a protein may be determined by two amino acids that are far apart in sequence but adjacent in three-dimensional structure), mainly the mutual relationship between short-distance amino acids. Generally speaking, if a person wants to learn a new language, the first feature is like learning grammar and sentence structure, and the second feature is like learning phrases and words. If you want to express "Apples are very delicious", using only the first feature to describe, it will be expressed as "I ate an apple today" or "I ate something that is very delicious", and it doesn't know which words are key. If using only the second feature to describe, it will be expressed as "apple", "eat", "delicious", but it won't be strung together into a complete sentence. Therefore, for a huge protein language model, it is necessary to consider both the local and global aspects, that is, the first feature and the second feature.

[0093] S3: Classifier construction

[0094] Based on the extracted protein sequence embedding features, the present invention integrates the global and local information of protein embedding features and constructs a deep neural network (DNN) classifier for identifying antibacterial lysins from phage protein sequences predicted from marine metagenomes. The choice of DNN as the classifier is based on its significant advantages in processing high-dimensional numerical features and complex relationship modeling. The protein feature vectors generated by ProtT5 and ProtBERT embedding models usually have high dimensions, and the fully connected layers of DNN can naturally adapt to this structured numerical data and flexibly model non-linear relationships in the feature space.

[0095] Compared with convolutional neural networks (CNNs) or recurrent neural networks (RNNs), DNN is more suitable for the data characteristics of this project. CNNs are usually applicable to data with spatial local correlations and are suitable for medium and small-scale datasets in protein applications, and rely on convolutional kernels to extract local features. However, the protein embedding features have captured global semantic representations through pre-trained models, so there is no need to design additional local feature extraction modules, which enables DNN to avoid unnecessary computational overhead. In addition, RNNs are usually prone to problems of gradient disappearance or explosion when processing long sequence data, while the fully connected structure of DNN combined with normalization (such as Batch Normalization) and residual connection (Residual Connection) is not only more stable during the training process but also can efficiently process the embedding features generated by pre-trained models. At the same time, DNN has high flexibility and can adapt to different task requirements by increasing the network depth or adjusting the number of layers. In the present invention, the output of each layer of DNN can be regarded as a non-linear transformation in the feature space, and this hierarchical representation method can fully capture the complex patterns in the protein sequence, including global dependencies (such as the interaction between active sites and remote regulatory regions) and local patterns (such as the functional distribution of amino acids around active sites). This property makes DNN an efficient and compatible classifier choice.

[0096] In summary, the present invention adopts a DNN classifier to accurately associate the embedding features of protein sequences with antibacterial activity labels through deep non-linear modeling, thereby achieving high-precision prediction of antibacterial lysins. Its main architecture includes the following parts:

[0097] First, after the input layer receives the normalized embedding features and the dimension of each sample vector, the embedding features are passed to the hidden layer while maintaining the integrity of the sequence context information. The hidden layer consists of multiple fully connected layers (FC Layers) and is used to extract the deep non-linear relationships in the embedding features. Each layer realizes efficient modeling of the input features through the combination of linear transformation and non-linear activation functions. The output of each layer can be expressed as:

[0098] hi = ReLU(W hid,i ·h i-1 + b hid,i )

[0099] h i is the output feature vector of the i-th hidden layer; W i is the weight matrix of the current layer, representing the weight relationship between features, which is initialized to a Gaussian distribution; b i is the bias vector of the current layer; The ReLU activation function maps the input features to a higher-dimensional feature space and can introduce non-linear modeling capabilities to avoid the vanishing gradient problem, thereby enhancing the model's ability to express complex patterns of proteins. In terms of parameter settings, the number of layers is 3 - 5 layers, and the number of neurons in a single layer is: 512, 256, and 128, decreasing layer by layer; Dropout is used to randomly mask lysin features during iteration (the probability is relatively low at 0.1. This is because the data involved in training is large enough, so it doesn't matter if there is an overfitting phenomenon in predicting classification sequences), which also avoids losing too many important features during model training. The output layer uses the Sigmoid activation function to predict whether the protein has antibacterial lysin activity. The output result is a probability value, representing the confidence that the sequence belongs to the positive class (antibacterial lysin), and is defined as follows:

[0100] P = σ(W out ·h n + b out )

[0101] The probability value P output by the Sigmoid function represents the confidence that the input sequence is an antibacterial lysin. To improve the credibility of the prediction, the present invention sets a dynamic threshold T, and only when P > T, the sequence is marked as an antibacterial lysin. The classifier training uses the cross-entropy loss function to measure the error between the prediction result and the true label (1 represents antibacterial lysin, 0 represents non-antibacterial lysin), and is defined as:

[0102]

[0103] y i is the true label of the sample, and P i is the predicted probability. To further improve the model performance, the present invention adopts the following optimization strategies during training: using the Adam optimizer (Adaptive Moment Estimation), combined with the learning rate decay mechanism, that is, the initial value of the learning rate is set to 0.001, and it decays exponentially to 0.0001 every 10 epochs. At the same time, since antibacterial lysins are rare sequences in phage protein sequences, during training, the non-antibacterial sequences are undersampled or the antibacterial lysin sequences are data-augmented to keep the sample distribution balanced.

[0104] This classification method can reliably identify potential antibacterial lysins from large-scale marine metagenomic data, providing precise and efficient technical support for the development of antibacterial drugs.

[0105] S4: To ensure the reliability and applicability of model predictions, the present invention adopts different performance evaluation strategies to verify the performance of the model in the antibacterial lysin prediction task from two aspects: evaluation metrics and model comparison.

[0106] First, the method of Matthews Correlation Coefficient (MCC) is adopted. This is a balanced metric suitable for imbalanced datasets, that is, in the proportion of positive and negative antibacterial proteins, usually the number of positive samples is much smaller than that of negative samples in the general case. In the screening of antibacterial lysin sequences, antibacterial lysins (positive class) also account for a very small part of the total samples, while non-antibacterial lysins (negative class) account for the majority. MCC provides a comprehensive performance summary for the classification model by considering true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN) simultaneously, avoiding biases that may be caused by simply relying on precision or recall. Compared with other metrics (such as accuracy), the significant advantage of MCC is that it is insensitive to changes in the proportion of positive and negative samples in the dataset. The formula is as follows:

[0107]

[0108] Where:

[0109] TP: True positives, that is, the number of samples predicted as positive and actually positive;

[0110] TN: True negatives, that is, the number of samples predicted as negative and actually negative;

[0111] FP: False positives, that is, the number of samples predicted as positive but actually negative;

[0112] FN: False negatives, that is, the number of samples predicted as negative but actually positive.

[0113] The value range of MCC is between -1 and 1:

[0114] 1: Indicates that the model's prediction is completely consistent with the true label (perfect classification).

[0115] 0: Indicates that the model's prediction is equivalent to random prediction and has no classification ability.

[0116] -1: Indicates that the model's prediction is completely opposite to the true label (the prediction result is completely wrong).

[0117] It does not rely on a specific threshold. Different from AUC-ROC or PR-AUC, MCC directly evaluates the entire confusion matrix, avoiding the fluctuations caused by threshold selection. At the same time, it has high stability. Even when the data distribution is extremely imbalanced, MCC can still provide a reliable evaluation of model performance. Compared with traditional evaluation metrics, such as Accuracy, it may produce misleading results when the data is imbalanced. For example, in the case where 95% of the samples are negative classes, simply predicting the negative class can achieve an accuracy of 95%, while MCC can reflect the bias brought by this imbalance. Another example is the F1 score, which only focuses on the performance of positive samples, while MCC evaluates the performance of both positive and negative samples simultaneously.

[0118] In the antibacterial lysin prediction task, due to the scarcity of positive samples and the high cost of detection experiments, MCC can provide a comprehensive evaluation of positive and negative samples, helping to screen high-confidence candidate lysins. At the same time, MCC has great optimization space. In terms of balancing the sample distribution, the data distribution can be balanced by undersampling negative samples or oversampling positive samples (such as the SMOTE method). The generalization ability of the model can also be improved through regularization methods. It is an evaluation method particularly suitable for protein sequence prediction tasks.

[0119] S5. Antibacterial Lysin Prediction

[0120] After the model is completed, by inputting the phage protein sequence generated by preprocessing and embedding, it is predicted whether it is an antibacterial lysin. The model calculates the antibacterial lysin probability of each input sequence through a classifier, and the output result is between 0 and 1. A binary classification result (whether it is an antibacterial lysin) is output through the threshold T. In this process, to ensure the accuracy of the prediction result, the output result undergoes a series of filtering and parameter screening. First, the dynamic threshold T is defaulted to 0.8, and the specific value is determined according to the characteristics of the user's dataset. For example, reducing T increases the recall rate, or increasing T increases the precision. Since the present invention is applicable to large-scale dataset training and prediction, this value is not recommended to be too low. Subsequently, a comparison tool is used to filter redundant sequences to ensure the uniqueness of candidate lysins. If there is active site data, the embedding features can be used to mark the possible functional regions in the sequence, and low-quality sequences lacking functional regions can be excluded.

[0121] The final output result is presented in a structured form, including the following information: sequence ID, sequence content, antibacterial lysin probability, whether it is filtered, structural features (optional), data source, and similarity alignment result.

[0122]

[0123] S6. Database Construction

[0124] After completing the performance evaluation and screening out highly confident antibacterial lysins, the present invention constructs a comprehensive antibacterial lysin database platform for storing, managing, and querying all prediction and verification results. The database design is divided into multiple functional modules, including a lysin sequence information table (recording sequence ID, amino acid sequence, and data source), a functional annotation table (recording active sites, structural features, and mechanism of action), a prediction result table (recording the confidence level and classification results of the model), and a verification experiment data table (containing antibacterial activity and cytotoxicity information verified by experiments). The database uses MySQL or PostgreSQL to store structured data, and at the same time introduces MongoDB for flexible management of large-scale sequences or unstructured files (such as 3D structures). In terms of technical architecture, the backend is built based on FastAPI, Django, or Flask, supporting data storage, user authentication, and API interface functions; the frontend uses React or Vue.js to implement an interactive interface, providing user-friendly data retrieval and display functions. In addition, full-text search and advanced query capabilities are provided through Elasticsearch and GraphQL, and visualization functions such as antibacterial activity distribution and confidence analysis are completed using D3.js or Plotly. The database is updated dynamically on a regular basis, adding new screening results and experimental data, and is deployed on a cloud service platform (such as AWS or Azure), combined with Docker and Kubernetes container technologies, supporting distributed computing and high-concurrency access. At the same time, the platform is seamlessly integrated with the AI model, supporting real-time prediction and automatic storage of results after users upload sequences, providing a complete technical chain for the recommendation and research of antibacterial lysins. Finally, the database will serve as a set of dynamically updated and efficient tools, providing convenient antibacterial lysin screening, characteristic analysis, and model optimization support for researchers, pharmaceutical companies, and AI developers, accelerating the development and transformation of new antibacterial drugs.

[0125] The present invention solves the problem of limited biological sources in the development of existing antibacterial lysozymes through artificial intelligence technology. Traditional lysozymes are mostly derived from specific bacteria or phages, and their resource scarcity limits large-scale applications. The present invention integrates global marine microbiome and phageome data to construct the first systematic marine antibacterial lysin database, greatly expanding the sources of potential antibacterial lysins. At the same time, the database integrates various omics data, providing rich resources and laying a foundation for the extensive development of future lysins.

[0126] Aiming at the problems of high false positive rate of traditional algorithms and single model optimization methods, the present invention adopts deep learning frameworks (such as ProtT5, ProtBERT) and natural language processing technologies to extract context features from protein sequences, and combines multi-dimensional verification means to effectively improve the prediction accuracy and robustness of the model. In addition, by optimizing the algorithm and introducing high-throughput screening technology, the present invention significantly improves the efficiency of antibacterial lysin screening and verification, solves the problems of long time consumption and high cost of existing methods, and provides an efficient solution for large-scale screening and development.

[0127] The present invention also significantly reduces the technical threshold for the development of antibacterial lysins. By optimizing the data integration platform and algorithm design, the interoperability of data processing is improved, and the in-depth mining of lysin resources is promoted. Aiming at the problem of limited participation of small R & D institutions, the method developed by the present invention is generally suitable for datasets of different scales, different platforms, and different omics, reducing the technical barriers and promoting diversified innovation in the field. In addition, by simulating diverse marine environments, lysins adapted to special environments are screened, providing support for their application under extreme conditions.

[0128] In summary, the present invention has achieved comprehensive innovation in data integration and model prediction. It effectively solves the problems of insufficient resources, low accuracy, and low efficiency in the existing lysozyme development, provides a powerful solution for the antibiotic resistance crisis, and promotes the development and wide application of new antibacterial drugs.

[0129] The above is the method for predicting marine phage antibacterial lysins provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding device for predicting marine phage antibacterial lysins, including:

[0130] A model construction module, configured to construct a deep learning network including a ProtT5 module, a ProtBERT module, a weighting module, and a classification module, construct a dataset based on the collected protein sequence data of marine phage antibacterial lysins, and use the dataset to train the constructed deep learning network to obtain a marine phage antibacterial lysin prediction model.

[0131] A prediction module, configured to input the protein sequence of the marine phage antibacterial lysin to be predicted into the prediction model, encode and decode the protein sequence to be predicted through the ProtT5 module to obtain a first feature vector; encode and decode the protein sequence to be predicted through the ProtBERT module to obtain a second feature vector; perform weighted summation on the first feature vector and the second feature vector through the weighting module to obtain a fused feature vector; and perform probability prediction on the fused feature vector through the classification module to output the class label of the protein sequence to be predicted.

[0132] For the specific limitations of the marine phage antibacterial lysin prediction device, reference may be made to the limitations of the marine phage antibacterial lysin prediction method in the above text, which will not be elaborated here. Each module in the above marine phage antibacterial lysin prediction device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0133] The present invention also provides a computer-readable storage medium storing a computer program, which can be used to execute the above-provided marine phage antibacterial lysin prediction method.

[0134] The present invention also provides the structure of a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, other hardware required for other services may also be included. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-provided marine phage antibacterial lysin prediction method.

[0135] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-described method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include at least one of non-volatile and volatile memories. The non-volatile memory can include a read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory. By way of illustration and not limitation, the RAM can be in various forms, such as a static random access memory (SRAM) or a dynamic random access memory (DRAM), etc.

[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded by the present invention.

Claims

1. A method for predicting marine phage antibacterial lysins, characterized in that, Including: Construct a deep learning network including a ProtT5 module, a ProtBERT module, a weighting module, and a classification module. Based on the collected antibacterial protein sequence data, construct a data set, and use the data set to train the constructed deep learning network to obtain a prediction model for marine phage antibacterial lysins; Input the sequence of the marine phage antibacterial lysin to be predicted into the prediction model. Through the ProtT5 module, perform encoding and decoding operations on the sequence of the marine phage antibacterial lysin to be predicted to capture the global features in the sequence of the marine phage antibacterial lysin to be predicted, and obtain a first feature vector; Through the ProtBERT module, perform encoding operations on the sequence of the marine phage antibacterial lysin to be predicted to capture the sequence local dependencies in the sequence of the marine phage antibacterial lysin to be predicted, and obtain a second feature vector; through the weighting module, perform weighted summation on the first feature vector and the second feature vector to obtain a fused feature vector; through the classification module, perform probability prediction on the fused feature vector, and output the confidence that the protein sequence of the marine phage to be predicted is an antibacterial protein sequence.

2. The method for predicting marine phage bacteriolytic lysin according to claim 1, wherein The encoding and decoding operations performed on the sequence of the marine phage antibacterial lysin to be predicted by the ProtT5 module specifically include: In the encoding stage of ProtT5: Through the position encoding layer in the ProtT5 module, perform position encoding on the sequence of the marine phage antibacterial lysin to be predicted, and combine protein sequence embedding operations to obtain a protein sequence embedding vector; through the multi-head self-attention mechanism layer in the ProtT5 module, capture the global dependencies of the protein sequence embedding vector to generate a context-weighted protein sequence vector; through the first feed-forward neural network layer in the ProtT5 module, perform non-linear transformation on the weighted protein sequence vector, and combine residual connection and normalization operations to obtain protein encoding features; In the decoding stage of ProtT5, through the sequence reconstruction layer in the ProtT5 module, perform sequence reconstruction on the protein encoding features, and output a predicted amino acid sequence vector; through the masked multi-head self-attention mechanism layer in the ProtT5 module, perform masking operations on the amino acid sequence vector to capture the local dependencies of the amino acids to obtain an amino acid masked weighted vector; through the interactive attention mechanism layer in the ProtT5 module, calculate the global dependencies between the protein encoding features and the amino acid masked weighted vector to obtain a fused weighted amino acid sequence vector; through the second feed-forward neural network layer in the ProtT5 module, perform non-linear transformation on the fused weighted amino acid sequence vector, and combine residual connection and normalization operations to obtain a first feature vector.

3. The method for predicting marine phage antibacterial lysin according to claim 1, wherein The encoding operations performed on the sequence of the marine phage antibacterial lysin to be predicted by the ProtBERT module specifically include: Through the position encoding layer in the ProtBERT module, perform position encoding on the sequence of the marine phage antibacterial lysin to be predicted after aligning it to a preset fixed length, and combine protein sequence embedding operations to obtain a protein sequence embedding vector; The protein sequence embedding vectors are weighted by the multi-head self-attention layer in the ProtBERT module to capture the global dependencies of amino acids and obtain weighted protein sequence vectors; The weighted protein sequence vectors are randomly masked by the masked language modeling layer in the ProtBERT module to generate amino acid mask features; The amino acid mask features are aggregated by the pooling layer in the ProtBERT module to obtain the second feature vector.

4. The method for predicting marine phage antibacterial lysin according to claim 1, wherein The weighted sum of the first feature vector and the second feature vector is obtained by the weighting module, specifically including: E = α·E ProtT5 +(1 - α)·E ProtBERT Among them, E ProtT5 is the first eigenvector, E ProtBERT is the second eigenvector, and α is the adaptive weight coefficient.

5. The method for predicting marine phage antibacterial lysin according to claim 1, characterized in that, The probability prediction of the fused feature vector is performed by the classification module, specifically including: The normalized fused feature vector is received by the input layer in the classification module and passed to the hidden layer in the classification module to extract the deep non-linear relationships in the fused feature vector: h i = ReLU(W hid,i ·h i-1 + b hid,i ) where h i and h i-1 are the output feature vectors of the i-th and (i - 1)-th hidden layers respectively, where i ∈ {1, …, n} and n is the total number of hidden layers; W hid,i is the weight matrix of the i-th hidden layer; b hid,i is the bias vector of the i-th hidden layer; The probability value is output by the output layer in the classification module: P = σ(W out ·h n +b out ) Among them, σ(·) represents the Sigmoid function; h n is the output feature vector of the n-th hidden layer; W out is the weight matrix of the output layer; b out is the bias vector of the output layer; P represents the confidence that the output marine phage protein sequence to be predicted is an antibacterial protein sequence.

6. The method for predicting marine phage antibacterial lysin according to claim 1, wherein The dataset is constructed based on the collected antimicrobial protein sequence data, specifically including: Relevant sequence data are obtained from the global known phage genome databases, which include DDBJ, CHVD, EMBL, Genebank, GOV2, GPD, GVD, IGVD, IMG_VR, MGV, PhagesDB, RefSeq, STV, and TemPhD; marine metagenomic sequencing data are obtained from the marine metagenomic databases, which include ENA, SRA, and ProGenomes3; antimicrobial protein sequence data are extracted from the positive antimicrobial protein sequence databases, which include NCBI, UniprotKB / Swiss-Prot, PDB, RefSeq, EMBL, GenBank, and PIR; the known phage whole-genome data, marine metagenomic sequencing data, and antimicrobial protein sequence data are integrated to obtain the original antimicrobial protein sequence dataset; The original antimicrobial protein sequence dataset is preprocessed to obtain the processed antimicrobial protein sequence dataset; the preprocessing includes data integrity verification, standardizing the format, file splitting, removing empty sequences, cleaning abnormal symbols, trimming sequences, deduplication, redundancy removal, quality control, gene annotation, and species classification; Phage sequences are identified in the data of the processed antimicrobial protein sequence dataset using different bioinformatics tools, and the union of the identification results is taken to obtain the final dataset; the bioinformatics tools include PhiSpy, Iphop, virsorter2, DeepVirFinder, VIBRANT, kraken2, viralVerify, Metaphinder, geNomad, virfinder, PPR-Meta, and Phigaro.

7. An apparatus for predicting marine phage antibacterial lysins, characterized in that, Including: A model construction module, configured to construct a deep learning network including a ProtT5 module, a ProtBERT module, a weighting module, and a classification module, construct a data set based on the collected antibacterial protein sequence data, and use the data set to train the constructed deep learning network to obtain a prediction model for marine phage antibacterial lysin; A prediction module, configured to input the sequence of the marine phage antibacterial lysin to be predicted into the prediction model, perform encoding and decoding operations on the sequence of the marine phage antibacterial lysin to be predicted through the ProtT5 module to capture the global features in the sequence of the marine phage antibacterial lysin to be predicted, and obtain a first feature vector; Perform encoding operations on the sequence of the marine phage antibacterial lysin to be predicted through the ProtBERT module to capture the sequence local dependencies in the sequence of the marine phage antibacterial lysin to be predicted, and obtain a second feature vector; perform weighted summation on the first feature vector and the second feature vector through the weighting module to obtain a fused feature vector; perform probability prediction on the fused feature vector through the classification module, and output the confidence that the sequence of the marine phage protein to be predicted is an antibacterial protein sequence.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 6 above is implemented.

9. A computer device, characterized in that, Comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, when the processor executes the program, the method described in any one of claims 1 to 6 above is implemented.

Citation Information

Patent Citations

  • Antibacterial peptide prediction method and device based on protein pre-training representation learning

    CN112614538A

  • Prediction method for protein binding nucleotide sites on full-length circular RNA

    CN114187963A

  • Antibacterial peptide recognition method and application of antibacterial peptide recognition method in inhibition of multi-drug-resistant bacteria

    CN118136121A

  • Resistance polypeptide recognition method based on deep learning

    CN119007829A

  • Identification method and system of short antibacterial peptide sequence, terminal and storage medium

    CN119541641A

Cited By

  • 6-HB targeted membrane fusion inhibitory peptide prediction method, device, equipment and medium

    CN120913657A

  • Phage endolysin prediction method and system based on machine learning, storage medium and electronic equipment

    CN121260266A

  • Drug sensitivity phenotype prediction method, device and equipment

    CN121768517A