Anti-virus peptide prediction method and system based on queue negative sampling strategy and contrast learning
Through cohort negative sampling and contrast learning methods, combined with the protein language model ESM2 and feature encoding, an antiviral peptide prediction model was constructed, which solved the problems of high cost and low efficiency in the existing technology, and achieved efficient and accurate antiviral peptide prediction.
Patent Information
- Application Number
- CN202510490646.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-05
AI Technical Summary
The existing antiviral peptide prediction methods are expensive, time-consuming and inefficient, the model performance is unstable, making it difficult to achieve efficient and low-cost antiviral peptide screening and prediction.
The cohort negative sampling strategy and contrast learning method were used to capture context features through the protein language model ESM2, and negative samples were selected based on feature encoding and cosine similarity. The antiviral peptide prediction model was constructed using contrast learning definition loss function.
It improves the accuracy and efficiency of antiviral peptide prediction, reduces costs, enhances the accuracy of the model when distinguishing complex samples, optimizes the construction of feature space, and improves the ability to capture key features of antiviral peptides.
Smart Images

Figure CN120432019A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biological peptide identification, and in particular to an antiviral peptide prediction method and system based on a queue negative sampling strategy and comparative learning. Background Art
[0002] With the advancement of biomedicine, the limitations of current antiviral therapies and the lack of research on viral pathogens have posed significant challenges to the prevention of viral diseases. Antiviral peptides (AVPs) are a class of small peptide molecules that have been shown to exert antiviral effects by inhibiting viral attachment and replication. Compared with traditional antiviral drugs, AVPs offer significant advantages, including high specificity, low side effects, cost-effectiveness, ease of synthesis and modification, and strong sensitivity to viral resistance. These properties make AVPs ideal candidates for the development of novel antiviral therapies. Currently, some AVPs have been successfully applied in clinical practice. However, the screening process for natural AVPs remains time-consuming and costly, and more efficient, low-toxic, and highly selective AVP screening methods are urgently needed. Furthermore, because AVPs can target specific viruses, predicting their antiviral activity requires not only determining whether a peptide has antiviral potential but also accurately predicting its efficacy against specific viral families or species. This dual predictive capability is crucial for advancing AVP-based therapies and has both scientific value and practical significance.
[0003] Traditional antiviral peptide prediction methods mainly include biological experimental methods, machine learning methods, and deep learning methods. Among them, biological experimental methods such as chemical analysis and immunoassay are widely used in laboratories and can effectively detect the antiviral peptide (AVP) activity of specific viruses with high sensitivity and specificity, but the process is time-consuming and costly; high-throughput computing technology has been widely used in AVP research, especially AVP prediction models based on machine learning, such as SVM, random forest and ensemble learning methods, which use the physical and chemical properties of peptide sequences for feature extraction and classification, thereby improving the efficiency and accuracy of AVP screening; deep learning methods such as contrastive learning, protein language models and transfer learning strategies have effectively improved the ability to identify and classify AVP. Computer-aided analysis methods are combined with machine learning algorithms to model peptide sequence features, improve the degree of automation of AVP prediction, and can perform more accurate classification for different viruses.
[0004] However, among traditional antiviral peptide prediction methods, biological experimental methods are costly and time-consuming, which limits the feasibility of large-scale screening. In addition, the operations are complex, and the experimental conditions have a great impact on sensitivity and specificity, making it difficult to ensure stability and repeatability. Machine learning methods often rely on manually designed feature extraction methods and heavily rely on traditional feature encoding technologies. They require complex feature engineering steps, which increases the difficulty and computational cost of modeling, and the prediction performance is limited to the selected feature set, and has poor versatility. Deep learning methods rely heavily on a large amount of high-quality labeled data, and are prone to overfitting problems when data is insufficient. They also require high computing resources and are more expensive to train than traditional methods. Existing models still lack a comprehensive understanding of sequence context semantic relationships when identifying AVPs, which affects the accuracy of predictions.
[0005] Therefore, traditional antiviral peptide prediction methods often have the problems of high cost and low efficiency in antiviral peptide prediction due to complicated operation, unstable model performance and high resource requirements. Summary of the Invention
[0006] Based on this, in order to solve the above technical problems, a method and system for predicting antiviral peptides based on queue negative sampling strategy and contrastive learning are provided, which can predict antiviral peptides efficiently and at low cost, and improve the accuracy of antiviral peptide prediction.
[0007] A method for predicting antiviral peptides based on a cohort negative sampling strategy and contrastive learning, the method comprising:
[0008] Collecting antiviral peptide data from various databases, and constructing an antiviral peptide dataset after preprocessing the antiviral peptide data;
[0009] Inputting the antiviral peptide dataset into the protein language model ESM2 to capture context features, extracting sequence features in the antiviral peptide dataset using a feature encoding method, and fusing the context features with the sequence features to obtain a fused feature sequence;
[0010] Storing different labels for each sample in the feature sequence, and dividing each sample into a first queue and a second queue according to each label;
[0011] Determine an anchor sample label, select a target queue from the first queue and the second queue according to the anchor sample label, and select a negative sample in the target queue using cosine similarity;
[0012] A contrastive learning method is used to define a contrast loss function and a cross entropy loss function based on the negative sample to obtain a final loss function, and an antiviral peptide prediction model is constructed based on the final loss function to complete the antiviral peptide prediction.
[0013] In one embodiment, the antiviral peptide data is preprocessed to construct an antiviral peptide dataset, including:
[0014] determining antiviral peptide attribute keywords, and screening the antiviral peptide data according to the attribute keywords to obtain screened antiviral peptide data;
[0015] The target similarity of each of the screened antiviral peptide data is calculated respectively, a similarity threshold is determined, and duplicate removal processing is performed based on the target similarity and the similarity threshold to obtain an antiviral peptide data set.
[0016] In one embodiment, the antiviral peptide dataset is input into the protein language model ESM2 to capture contextual features, including:
[0017] Inputting the antiviral peptide dataset into the protein language model ESM2 based on the Transformer architecture to determine the pre-training parameters of the protein language model;
[0018] Based on the pre-training parameters, the protein language model is lightweighted and the sequence encoding capability is balanced to obtain a processed protein language model;
[0019] The processed protein language model is used for feature extraction, and the context-aware sequence feature representation is output as the context feature.
[0020] In one embodiment, a feature encoding method is used to extract sequence features from the antiviral peptide dataset, including:
[0021] The antiviral peptide dataset was comprehensively extracted using binary coding, Zscale coding, distance pair coding, CKSAAGP coding, QSOrder coding, and DDE coding methods to obtain sequence features.
[0022] In one embodiment, selecting a target queue from the first queue and the second queue according to the anchor sample label includes:
[0023] Compare the anchor sample label with the labels of the first queue and the second queue respectively to obtain a comparison result;
[0024] If the comparison result shows that the label of the anchor sample is opposite to the label of the first queue, the first queue is used as the target queue;
[0025] If the comparison result shows that the label of the anchor sample is opposite to the label of the second queue, the second queue is used as the target queue.
[0026] In one embodiment, using cosine similarity to select negative samples in the target queue includes:
[0027] Calculate the cosine similarity between each sample in the target queue and the anchor sample respectively;
[0028] Sorting each sample in the target queue based on each cosine similarity to obtain a sorted target queue;
[0029] The number of samples is determined, and negative samples are determined from the sorted target queue according to the number of samples.
[0030] In one embodiment, a contrastive learning method is used to define a contrastive loss function and a cross entropy loss function based on the negative sample to obtain a final loss function, including:
[0031] Defining a contrastive loss function using a contrastive learning method according to each of the cosine similarities and each sample in the target queue;
[0032] Determine the category weights, and fine-tune the weights in the initial cross entropy loss function based on the category weights to obtain a final cross entropy loss function;
[0033] The contrast loss function and the cross entropy loss function are combined to obtain the final loss function.
[0034] In one embodiment, the method further comprises:
[0035] Obtaining an evaluation index, and using the evaluation index to evaluate the antiviral peptide prediction model to obtain an evaluation result;
[0036] The evaluation indicators include accuracy, sensitivity, specificity, Matthews correlation coefficient, geometric mean of sensitivity and specificity, area under the precision-recall curve, area under the receiver operating characteristic curve, and F1 score.
[0037] In one embodiment, the method further comprises:
[0038] collecting an antiviral peptide sequence to be predicted, and inputting the antiviral peptide sequence to be predicted into the antiviral peptide prediction model;
[0039] The antiviral peptide prediction model is used to output the antiviral peptide prediction results and display them on a page.
[0040] An antiviral peptide prediction system based on a queue negative sampling strategy and contrastive learning, the system comprising:
[0041] A data set construction module is used to collect antiviral peptide data from various databases, and construct an antiviral peptide data set after preprocessing the antiviral peptide data;
[0042] A feature extraction module is used to input the antiviral peptide dataset into the protein language model ESM2 to capture context features, and use a feature encoding method to extract sequence features in the antiviral peptide dataset, and fuse the context features with the sequence features to obtain a fused feature sequence;
[0043] a queue division module, configured to store different labels for each sample in the feature sequence, and to divide each sample into a first queue and a second queue according to each label;
[0044] A negative sample selection module is configured to determine an anchor sample label, select a target queue from the first queue and the second queue according to the anchor sample label, and select negative samples from the target queue using cosine similarity;
[0045] The antiviral peptide prediction module is used to define a contrast loss function and a cross entropy loss function based on the negative sample using a contrast learning method to obtain a final loss function, and to construct an antiviral peptide prediction model based on the final loss function to complete the antiviral peptide prediction.
[0046] The above-mentioned antiviral peptide prediction method and system based on queue negative sampling strategy and contrastive learning can capture the deep contextual information and evolutionary characteristics of peptide sequences through the protein language model ESM2, and extract features by combining feature encoding methods for fusion, so that the extracted features are more comprehensive; negative sample selection based on feature queues can enhance the accuracy of the model in distinguishing complex samples, and can also optimize the construction of feature space, so that the model can better capture the key features of antiviral peptides; the use of contrastive learning methods enables the model to not only learn the similarities and differences between samples, but also improve the accuracy of antiviral peptide prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 FIG2 is an application environment diagram of an antiviral peptide prediction method based on a queue negative sampling strategy and contrastive learning in one embodiment;
[0048] Figure 2 Schematic diagram of a process for predicting antiviral peptides based on a cohort negative sampling strategy and contrastive learning in one embodiment;
[0049] Figure 3 Schematic diagram of the model framework of AVP-HNCL in one embodiment;
[0050] Figure 4 An ablation experiment diagram comparing ESM model encoding with other feature encodings in one embodiment;
[0051] Figure 5 FIG1 is an ablation experiment diagram for verifying the effectiveness of the enhancement strategy in one embodiment;
[0052] Figure 6 FIG2 is an ablation experiment diagram for verifying the effectiveness of the ESM2 model in one embodiment;
[0053] Figure 7 FIG2 is an ablation experiment diagram for verifying the effectiveness of contrastive learning in one embodiment;
[0054] Figure 8 A schematic diagram showing a performance comparison between using pre-training and not using pre-training in one embodiment;
[0055] Figure 9 1 is a structural block diagram of an antiviral peptide prediction system based on a queue negative sampling strategy and contrastive learning in one embodiment;
[0056] Figure 10 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0058] It is understood that the terms "first," "second," and the like as used herein may be used herein to describe queues, but these queues are not limited by these terms. These terms are used solely to distinguish a first queue from another queue. For example, a first queue may be referred to as a second queue, and similarly, a second queue may be referred to as a first queue, without departing from the scope of this application. The first queue and the second queue are both queues, but they are not the same queue.
[0059] The antiviral peptide prediction method based on queue negative sampling strategy and contrastive learning provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Figure 1As shown, the application environment includes a computer device 110. The computer device 110 can collect antiviral peptide data from various databases, pre-process the antiviral peptide data, and construct an antiviral peptide dataset. The computer device 110 can input the antiviral peptide dataset into the protein language model ESM2 to capture contextual features, and use a feature encoding method to extract sequence features in the antiviral peptide dataset, fusing the contextual features with the sequence features to obtain a fused feature sequence. The computer device 110 can store different labels for each sample in the feature sequence and divide each sample into a first queue and a second queue based on the labels. The computer device 110 can determine anchor sample labels, select a target queue from the first queue and the second queue based on the anchor sample labels, and select negative samples in the target queue using cosine similarity. The computer device 110 can use a contrastive learning method to define a contrastive loss function and a cross-entropy loss function based on the negative samples to obtain a final loss function. Based on the final loss function, the computer device 110 constructs an antiviral peptide prediction model to complete the antiviral peptide prediction. The computer device 110 can be, but is not limited to, various personal computers, laptops, smartphones, robots, unmanned aerial vehicles, tablet computers, and other devices.
[0060] In one embodiment, Figure 2 As shown, a method for predicting antiviral peptides based on a queue negative sampling strategy and contrastive learning is provided, comprising the following steps:
[0061] Step 202 : collecting antiviral peptide data from various databases, and constructing an antiviral peptide dataset after preprocessing the antiviral peptide data.
[0062] Establishing a reliable benchmark dataset is the basis for developing powerful predictions of the potential mechanisms of antiviral peptides (AVPs). Among them, when collecting data, the antiviral peptide (AVP) dataset provided by Guan et al. can be used. The dataset is divided into a training set (2129 samples) and a test set (553 samples). In the training and test datasets, positive and negative samples are evenly distributed, ensuring the stability and generalization ability of the model. The dataset contains two sets of antiviral peptide data, non-AVPs and non-AMPs, sharing 2662 balanced positive and negative samples. Among them, these samples come from multiple databases, including AVPdb, dbAMP, DRAMP, DBAASP and HIPdb. Specifically, the sources of negative samples are different: the negative samples of the first set of data come from dbAMP, DBAASP and DRAMP, which do not contain virus-specific entries; the negative samples of the second set of data are obtained from the UniProt database.
[0063] In one embodiment, a provided antiviral peptide prediction method based on a queue negative sampling strategy and contrastive learning may also include a data preprocessing process, the specific process including: determining antiviral peptide attribute keywords, and screening antiviral peptide data according to the attribute keywords to obtain screened antiviral peptide data; calculating the target similarity of each screened antiviral peptide data, determining the similarity threshold, and performing deduplication processing based on the target similarity and the similarity threshold to obtain an antiviral peptide data set.
[0064] After obtaining the antiviral peptide data, it can be filtered according to attribute keywords such as "toxicity", "membrane", "secretion", "defense", "antibiotic", "anticancer", "antiviral" and "antifungal". Then, in order to reduce redundancy, the CD-HIT tool can be used to remove duplicate sequences with a similarity threshold of 40%, and finally a balanced data set is obtained, in which the number of positive samples and negative samples is equal, totaling 2662 sequences. In this embodiment, the two data sets are referred to as data set 1 and data set 2, respectively. In addition, two independent test data sets can be included to verify the generalization ability of the model. Among them, independent data set 1 includes 604 experimentally verified AVPs and 452 non-AVPs; independent data set 2 contains 604 verified AVPs and 604 unverified non-AVPs. Specifically, the specific number of samples in data set 1, data set 2, independent data set 1, and independent data set 2 are shown in the following table:
[0065]
[0066] In one embodiment, in order to further verify the generalization ability of the model and its performance in specific virus classification, a plurality of data sets for specific virus families and viruses were constructed. Specifically, according to the versatility of the antiviral peptides, the data can be classified into six major virus families, including Coronaviridae, Retroviridae, Flaviviridae, Orthomyxoviridae, Paramyxoviridae and Herpesviridae, and eight specific viruses, including feline immunodeficiency virus (FIV), hepatitis C virus (HCV), human immunodeficiency virus (HIV), human metapneumovirus type 3 (HPIV3), herpes simplex virus type 1 (HSV1), influenza A virus (INFVA), respiratory syncytial virus (RSV) and severe acute respiratory syndrome coronavirus (SARS-CoV). The selection of these virus families and target viruses is based on strict criteria: each virus family contains at least 100 sequence records and each target virus contains at least 80 sequence records to ensure the completeness of the training set and the feasibility of model development. The detailed specifications of the datasets for each virus family are shown in the following table:
[0067]
[0068] The detailed specifications of each specific virus dataset are shown in the following table:
[0069]
[0070] In step 204 , the antiviral peptide dataset is input into the protein language model ESM2 to capture context features, and sequence features in the antiviral peptide dataset are extracted using a feature encoding method. The context features are fused with the sequence features to obtain a fused feature sequence.
[0071] The Protein Language Model (ESM2) is a masked language model based on the Transformer architecture, trained using a large number of protein sequences. During training, amino acids are masked and mutations at specific positions are evaluated, enabling the model to capture contextual information and long-range dependencies in protein sequences.
[0072] In one embodiment, a method for predicting antiviral peptides based on a queue negative sampling strategy and contrastive learning is provided, which may also include a process of extracting features using ESM2. The specific process includes: inputting the antiviral peptide dataset into the protein language model ESM2 based on the Transformer architecture, and determining the pre-training parameters of the protein language model; based on the pre-training parameters, the protein language model is lightweighted and the sequence encoding capability is balanced to obtain a processed protein language model; the processed protein language model is used to extract features, and the sequence feature representation containing context awareness is output as a context feature.
[0073] The Protein Language Model (ESM2) offers pre-trained parameters of various sizes. In this example, an ESM2 model with 35M parameters was selected to achieve model lightweighting while improving sequence encoding capabilities. The use of protein pre-trained models provides a new perspective for sequence feature engineering. The Protein Language Model (ESM2) was used to obtain global contextual features for the antiviral peptide dataset.
[0074] In one embodiment, a method for predicting antiviral peptides based on a queue negative sampling strategy and contrastive learning is provided, which may also include a process of feature extraction using a feature encoding method. The specific process includes: using binary encoding, Zscale encoding, distance pair encoding, CKSAAGP encoding, QSOrder encoding, and DDE encoding methods to comprehensively extract features from the antiviral peptide dataset to obtain sequence features.
[0075] Among them, when using binary encoding, each amino acid is uniquely encoded using a 20-dimensional binary vector, which effectively distinguishes the other 19 standard amino acids; when using Zscale encoding, the physicochemical properties of amino acids in protein sequences can be measured, and various biophysical and chemical properties of amino acids can be compared by applying z-score normalization and standardization. In this system, each amino acid is represented by a five-dimensional vector; when using distance pair encoding, the distance pair and the pseudo amino acid composition of the simplified alphabet (PseAAC) encoding describe the pairwise distances between amino acids in the protein sequence, reflecting their relative positions and potential interactions. This encoding method is closely related to the structural properties of the protein; when using CKSAAGP encoding (CKSAAGPEncoding), the protein sequence is divided into overlapping amino acid groups and their frequencies are calculated, which effectively captures local patterns and helps to identify functional regions or active sites in peptides; using QSOrder encoding (QSOrder When using quasi-sequence order (QSOrder) encoding, the quasi-sequence descriptor approach analyzes the overall composition of amino acids while integrating information about their position in the sequence, effectively combining local and global sequence features. When using DDE encoding, the dipeptide deviation from expected mean (DDE) encoding assesses the difference between the actual frequency of dipeptides and their predicted frequency based on codon usage patterns. Significant changes in specific dipeptides may indicate underlying evolutionary, functional, or structural factors.
[0076] That is, in this embodiment, six traditional encoding methods are combined: Binary Encoding, Zscale Encoding, DistancePair Encoding, CKSAAGP Encoding, QSOrder Encoding and DDE Encoding, to extract rich features of AVP sequences from multiple perspectives such as physical and chemical properties and sequence order information.
[0077] Then, the features extracted by the ESM2 model can be combined with the features extracted by traditional encoding methods to construct a powerful feature representation that can fully reflect the local and global information of the AVP sequence and effectively enhance the feature expression ability of the model.
[0078] Step 206: Store different labels for each sample in the feature sequence, and divide each sample into a first queue and a second queue according to each label.
[0079] When forming negative sample pairs, data with different labels is stored in a queue. The top k most similar sequences from the queue with labels opposite to the anchor sample are selected as negative samples based on cosine similarity. This targeted negative sample selection method can effectively improve the model's ability to learn sequence representations, helping the model better distinguish between positive and negative samples, and enhancing the model's discrimination and generalization performance.
[0080] Specifically, during model training, feature engineering can be performed on the positive sample sequences extracted by the protein language model and the original sequences after sequence encoding. Each individual sequence learns the local context information and global features of the sequence during the feature engineering process.
[0081] In this embodiment, a queue-based negative sample pair selection strategy is also developed for contrastive learning. The goal of contrastive learning is to construct a latent space in which the distance between the anchor sample and the positive sample is minimized, while the distance between the anchor sample and the negative sample is maximized.
[0082] Step 208: Determine the anchor sample label, select a target queue from the first queue and the second queue according to the anchor sample label, and use cosine similarity to select negative samples in the target queue.
[0083] In the negative sample selection process, for each batch, the individual sample sequences are first divided into two queues according to the sample labels.
[0084] Specifically, in one embodiment, a method for predicting antiviral peptides based on a queue negative sampling strategy and contrastive learning is provided, which may also include a process of dividing queues. The specific process includes: comparing the anchor sample label with the labels of the first queue and the second queue respectively to obtain a comparison result; if the comparison result is that the anchor sample label is opposite to the label of the first queue, the first queue is used as the target queue; if the comparison result is that the anchor sample label is opposite to the label of the second queue, the second queue is used as the target queue.
[0085] When the label of the anchor sample is 1, the top k most similar samples can be retrieved from the target queue with label 0. These samples are highly similar to the anchor sample but have opposite labels. The main idea of this strategy is to select negative samples that are similar to the anchor sample but have opposite labels. This helps to effectively increase the distance between positive and negative samples in the latent space, which not only improves the quality of negative samples but also helps the model learn more discriminative features.
[0086] In one embodiment, a method for predicting antiviral peptides based on a queue negative sampling strategy and contrastive learning is provided, which may also include a process for selecting negative samples. The specific process includes: calculating the cosine similarities between each sample in the target queue and the anchor sample respectively; sorting each sample in the target queue based on each cosine similarity to obtain a sorted target queue; determining the number of samples, and determining negative samples from the sorted target queue based on the number of samples.
[0087] In this embodiment, cosine similarity can be used as an indicator to measure the distance between samples. The specific calculation formula can be expressed as: Here, x represents the embedding of the anchor sequence, and y represents the corresponding positive or negative sample.
[0088] Among them, the negative sample selection method enhances the construction of the latent space, optimizes model training, and improves learning efficiency and generalization ability. The mathematical formula for negative sample selection can be expressed as: in: represents the sample feature representation opposite to the x label, It means that the first k most similar samples are aggregated into a set, Q is a queue, and the embedding of the negative sample sequence is stored.
[0089] In step 210 , a contrastive loss function and a cross entropy loss function are defined based on negative samples using a contrastive learning method to obtain a final loss function, and an antiviral peptide prediction model is constructed based on the final loss function to complete the antiviral peptide prediction.
[0090] Among them, contrastive learning trains the model by minimizing the distance between anchor samples and positive samples, while maximizing the distance between anchor samples and negative samples.
[0091] In one embodiment, a method for predicting antiviral peptides based on a queue negative sampling strategy and contrastive learning is provided, which may also include a process of defining a loss function. The specific process includes: defining a contrast loss function using a contrastive learning method based on each cosine similarity and each sample in the target queue; determining the category weight, fine-tuning the weight in the initial cross-entropy loss function based on the category weight, obtaining the final cross-entropy loss function, and combining the contrast loss function and the cross-entropy loss function to obtain the final loss function.
[0092] Among them, the calculation formula of the contrast loss function can be expressed as: Among them, sim(x, y) represents the similarity between samples x and y, where x +represents the feature representation of the enhanced sample; margin is a key parameter in contrastive learning, which is used to control the degree of separation between positive and negative samples in the feature space; temperature T is another important hyperparameter used to scale the similarity distribution; |T| represents the absolute value of the temperature parameter, which penalizes the temperature value to prevent it from being too large or too small, thereby ensuring the stability of model training; λ is the regularization coefficient, which is used to adjust the influence of the regularization term in the loss function.
[0093] To enhance classification performance, class weights are also introduced into the cross-entropy loss function. By assigning different weights to each class, the model can be more effectively guided to focus on the minority class, thereby alleviating the class imbalance problem. This approach not only improves the overall performance of the model, but also ensures reliable predictions for all classes. By fine-tuning the class weights, the model's ability to handle imbalanced data is further improved, resulting in better performance on these tasks. The weighted cross-entropy loss function can be expressed as: Where N is the total number of samples, y i Represents sample x i The label, p i is the sample x i The probability of being predicted as a positive class, α1 is the weight of the positive class, and α0 is the weight of the negative class. By combining contrast loss and weighted cross entropy loss, the model can enhance feature representation and effectively solve the problem of class imbalance, ultimately improving the overall classification performance and enhancing the generalization ability of the model. The final loss function can be expressed as: L total =L CE-w +αL contrastive ; Among them, the α coefficient controls the proportion of contrast loss in determining the final loss.
[0094] In one embodiment, a provided antiviral peptide prediction method based on a cohort negative sampling strategy and contrastive learning may also include a model evaluation process, the specific process including: obtaining evaluation indicators, using the evaluation indicators to evaluate the antiviral peptide prediction model, and obtaining evaluation results; wherein the evaluation indicators include accuracy, sensitivity, specificity, Matthews correlation coefficient, geometric mean of sensitivity and specificity, area under the precision-recall curve, area under the receiver operating characteristic curve, and F1 score.
[0095] During the model evaluation process, an independent test set was used to comprehensively assess the model's performance. Evaluation metrics included accuracy (ACC), sensitivity (SN), specificity (SP), Matthews correlation coefficient (MCC), geometric mean (G-mean), area under the precision-recall curve (AUPRC), and area under the receiver operating characteristic curve (AUROC). These metrics comprehensively reflect the model's predictive performance and robustness from different perspectives, ensuring that the model not only performs well on the training set but also has excellent generalization capabilities and robust predictive performance on unseen data.
[0096] Specifically,
[0097] Among them: TP represents true positive examples, TN represents true negative examples, FP represents false positive examples, and FN represents false negative examples. The true positive rate (TPR) represents the proportion of actual positive samples correctly identified by the model, reflecting the model's sensitivity or recall ability to positive examples; the false positive rate (FPR) represents the proportion of actual negative samples that are mistakenly identified as positive samples by the model, reflecting the model's tendency to make incorrect judgments.
[0098] In one embodiment, a method for predicting antiviral peptides based on a queue negative sampling strategy and comparative learning is provided, which may also include a process for displaying a web page. The specific process includes: collecting the antiviral peptide sequence to be predicted, and inputting the antiviral peptide sequence to be predicted into the antiviral peptide prediction model; outputting the antiviral peptide prediction result through the antiviral peptide prediction model, and displaying it on the page.
[0099] By building a user-friendly online platform with a simple and intuitive interface, researchers can easily input AVP (antiviral peptide) sequences and quickly obtain predictions of their antiviral activity. The platform is designed to simplify the application process of AVP prediction technology, making it more convenient and accessible, effectively lowering the technical barriers to entry and enabling researchers and biopharmaceutical professionals to quickly obtain prediction information, thereby accelerating the progress of antiviral peptide research.
[0100] The present application provides an antiviral peptide prediction method based on a queue negative sampling strategy and contrastive learning, which can be applied to Figure 3The AVP-HNCL model, an antiviral peptide prediction framework shown in Figure 1, utilizes a unique contrastive learning approach to more effectively learn to distinguish subtle differences between antiviral and non-antiviral peptides, significantly improving prediction accuracy and generalization. In this study, a cohort-based negative sample sampling strategy was introduced. By carefully selecting negative samples that are highly similar to anchor samples but have opposite labels, the model's learning ability for difficult negative samples was enhanced. This not only improved the model's accuracy in distinguishing complex samples but also optimized the construction of the feature space, enabling the model to better capture key features of antiviral peptides. Furthermore, the model combined multiple feature encoding methods with advanced deep learning techniques, including the ESM2 model, a protein language model based on the Transformer architecture, which captures deep contextual information and evolutionary characteristics of peptide sequences. Furthermore, data augmentation strategies (such as mutations, insertions, and deletions) were used to generate positive sample pairs, further improving the model's generalization. This multi-dimensional feature fusion approach, combined with the innovative application of contrastive learning, has enabled AVP-HNCL to achieve breakthrough progress in the field of antiviral peptide prediction. In terms of model architecture, a convolutional neural network (CNN) combined with a bidirectional long short-term memory (BiLSTM) network was employed to effectively extract local and global features of the sequence. By combining contrastive learning with traditional supervised learning, the model not only learns similarities and differences between samples but also performs well in classification tasks. Furthermore, the concept of transfer learning was introduced, leveraging the pre-trained model parameters from the first phase to fine-tune the second phase's tasks, further enhancing the model's ability to classify specific virus families and viruses.
[0101] In one embodiment, to more comprehensively demonstrate the superiority of the antiviral peptide prediction framework AVP-HNCL model, it was compared in depth with several typical existing AVP prediction models, including AVP-IFT, AVPpred, AntiVPP1.0, Meta-iAVP, and DeepAVP. The performance comparison results of existing methods and the method in this application on independent dataset 1 are shown in the following table:
[0102]
[0103] The performance comparison results of different models on independent dataset tests are shown in the following table:
[0104]
[0105] Test results show that AVP-HNCL outperforms all existing models in multiple key performance indicators, including accuracy (ACC), sensitivity (SN), specificity (SP), and Matthews correlation coefficient (MCC). This demonstrates that AVP-HNCL not only significantly improves accuracy but also balances the predictive power of different categories while maintaining high sensitivity and specificity. Furthermore, AVP-HNCL combines the ESM2 pre-trained model with multiple traditional feature encoding methods to capture more comprehensive antiviral peptide sequence information, achieving a breakthrough in performance and further validating its effectiveness and reliability as an antiviral peptide prediction tool.
[0106] In one embodiment, an experiment can also be conducted to compare the performance of using only the ESM model encoding and the performance of ESM combined with the other six feature encoding methods. Figure 4 The experimental results show that the ESM model combined with other feature encoding methods outperforms the ESM model alone in all five evaluation indicators, with performance improved by 0.4 to 1.5 units. This indicates that the fusion of multiple feature encoding methods can enhance the model's ability to distinguish between positive and negative samples.
[0107] In one embodiment, Figure 5 As shown in the figure, by comparing the model performance using different enhancement strategies (mutation, insertion and deletion) in the multi-step enhancement process, according to Figure 5 Experimental results show that when these three augmentation strategies are introduced, the model shows significant improvements in four key metrics (ACC, MCC, SN, and F1), especially in MCC and SN. Further verification by removing each augmentation strategy individually shows that each strategy is effective, confirming their contribution to model performance. These strategies help the model better cope with data variability and complexity by generating diverse training samples, significantly improving the model's robustness and generalization capabilities.
[0108] In one embodiment, the performance of using the ESM2-35M model is compared with that of not using the ESM2 model. Figure 6 Experimental results show that the ESM2-35M model outperforms the model without ESM2 in all key performance indicators, which demonstrates the powerful feature extraction capability of the ESM2 model, which can capture deep information within the sequence, provide the model with richer feature representation, and effectively improve the accuracy and reliability of AVP prediction.
[0109] In one embodiment, the performance of the model using the proposed contrastive learning method, the traditional contrastive learning method, and the model without contrastive learning is also compared. Figure 7Experimental results show that the model using contrastive learning achieved an ACC of 0.9362 and an MCC of 0.8730, respectively. This compares to 0.9287 and 0.8574 for the traditional contrastive learning method, and 0.9268 and 0.8537 for the model without contrastive learning. This suggests that traditional contrastive learning has little effect on improving model performance, while contrastive learning effectively improves model performance by focusing on negative samples that are highly similar to anchor samples, enhancing the model's ability to learn key features and further improving classification performance.
[0110] like Figure 8 As shown, in one embodiment, the effectiveness experiment of pre-training is carried out using three data sets of the second stage. Figure 8 By comparing the performance of models using pre-training and those without it in the second phase, the authors found that pre-training helps to more clearly distinguish between positive and negative samples in high-dimensional space, forming more compact feature clusters, improving classification accuracy and the model's generalization ability. Furthermore, fine-tuning enables the model to quickly adapt to new data distributions, reducing the need for large amounts of training data. This demonstrates significant robustness and practicality in situations with limited datasets or complex sample distributions.
[0111] It should be understood that, although the various steps in the above flow chart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the above flow chart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0112] In one embodiment, Figure 9 As shown, an antiviral peptide prediction system based on a cohort negative sampling strategy and contrastive learning is provided, comprising: a data set construction module 910, a feature extraction module 920, a cohort division module 930, a negative sample selection module 940, and an antiviral peptide prediction module 950, wherein:
[0113] A data set construction module 910 is used to collect antiviral peptide data from various databases, and construct an antiviral peptide data set after preprocessing the antiviral peptide data;
[0114] Feature extraction module 920 is used to input the antiviral peptide dataset into the protein language model ESM2 to capture context features, extract sequence features from the antiviral peptide dataset using a feature encoding method, and fuse the context features with the sequence features to obtain a fused feature sequence;
[0115] A queue division module 930 is configured to store different labels for each sample in the feature sequence and to divide each sample into a first queue and a second queue according to each label;
[0116] A negative sample selection module 940 is configured to determine an anchor sample label, select a target queue from the first queue and the second queue based on the anchor sample label, and select negative samples in the target queue using cosine similarity;
[0117] The antiviral peptide prediction module 950 is used to define a contrast loss function and a cross entropy loss function based on negative samples using a contrast learning method to obtain a final loss function, and to construct an antiviral peptide prediction model based on the final loss function to complete the antiviral peptide prediction.
[0118] In one embodiment, the dataset construction module 910 is also used to determine antiviral peptide attribute keywords, and filter the antiviral peptide data according to the attribute keywords to obtain filtered antiviral peptide data; calculate the target similarity of each filtered antiviral peptide data, determine the similarity threshold, and perform deduplication processing based on the target similarity and the similarity threshold to obtain an antiviral peptide dataset.
[0119] In one embodiment, the feature extraction module 920 is also used to input the antiviral peptide dataset into the protein language model ESM2 based on the Transformer architecture to determine the pre-training parameters of the protein language model; based on the pre-training parameters, the protein language model is lightweighted and the sequence encoding capability is balanced to obtain a processed protein language model; the processed protein language model is used to perform feature extraction, and the sequence feature representation containing context awareness is output as a context feature.
[0120] In one embodiment, the feature extraction module 920 is further configured to perform comprehensive feature extraction on the antiviral peptide dataset using binary coding, Zscale coding, distance pair coding, CKSAAGP coding, QSOrder coding, and DDE coding to obtain sequence features.
[0121] In one embodiment, the negative sample selection module 940 is also used to compare the anchor sample label with the labels of the first queue and the second queue respectively to obtain a comparison result; if the comparison result is that the anchor sample label is opposite to the label of the first queue, the first queue is used as the target queue; if the comparison result is that the anchor sample label is opposite to the label of the second queue, the second queue is used as the target queue.
[0122] In one embodiment, the negative sample selection module 940 is further used to respectively calculate the cosine similarities between each sample in the target queue and the anchor sample; sort the samples in the target queue based on the cosine similarities to obtain a sorted target queue; determine the number of samples, and determine the negative samples from the sorted target queue based on the number of samples.
[0123] In one embodiment, the antiviral peptide prediction module 950 is also used to define a contrast loss function using a contrast learning method based on each cosine similarity and each sample in the target queue; determine the category weight, fine-tune the weight in the initial cross entropy loss function based on the category weight, obtain the final cross entropy loss function, and combine the contrast loss function and the cross entropy loss function to obtain the final loss function.
[0124] In one embodiment, the antiviral peptide prediction module 950 is also used to obtain evaluation indicators, use the evaluation indicators to evaluate the antiviral peptide prediction model, and obtain evaluation results; wherein the evaluation indicators include accuracy, sensitivity, specificity, Matthews correlation coefficient, geometric mean of sensitivity and specificity, area under the precision-recall curve, area under the receiver operating characteristic curve, and F1 score.
[0125] In one embodiment, the antiviral peptide prediction module 950 is also used to collect the antiviral peptide sequence to be predicted, and input the antiviral peptide sequence to be predicted into the antiviral peptide prediction model; output the antiviral peptide prediction result through the antiviral peptide prediction model, and display it on the page.
[0126] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for predicting antiviral peptides based on a queue negative sampling strategy and comparative learning is implemented. The display screen of the computer device can be a liquid crystal display or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.
[0127] Those skilled in the art will understand that Figure 10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0128] In one embodiment, a computer device is provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the antiviral peptide prediction method based on a queue negative sampling strategy and contrastive learning are implemented.
[0129] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the antiviral peptide prediction method based on queue negative sampling strategy and contrastive learning are implemented.
[0130] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0131] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0132] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for predicting antiviral peptides based on queue negative sampling strategy and contrastive learning, characterized in that: The method comprises: Collecting antiviral peptide data from various databases, and constructing an antiviral peptide dataset after preprocessing the antiviral peptide data; Inputting the antiviral peptide dataset into the protein language model ESM2 to capture context features, extracting sequence features in the antiviral peptide dataset using a feature encoding method, and fusing the context features with the sequence features to obtain a fused feature sequence; Storing different labels for each sample in the feature sequence, and dividing each sample into a first queue and a second queue according to each label; Determine an anchor sample label, select a target queue from the first queue and the second queue according to the anchor sample label, and select a negative sample in the target queue using cosine similarity; A contrastive learning method is used to define a contrast loss function and a cross entropy loss function based on the negative sample to obtain a final loss function, and an antiviral peptide prediction model is constructed based on the final loss function to complete the antiviral peptide prediction.
2. The antiviral peptide prediction method based on queue negative sampling strategy and contrastive learning according to claim 1, characterized in that: After preprocessing the antiviral peptide data, an antiviral peptide dataset is constructed, including: determining antiviral peptide attribute keywords, and screening the antiviral peptide data according to the attribute keywords to obtain screened antiviral peptide data; The target similarity of each of the screened antiviral peptide data is calculated respectively, a similarity threshold is determined, and duplicate removal processing is performed based on the target similarity and the similarity threshold to obtain an antiviral peptide data set.
3. The antiviral peptide prediction method based on queue negative sampling strategy and contrastive learning according to claim 1, characterized in that: The antiviral peptide dataset is input into the protein language model ESM2 to capture contextual features, including: Inputting the antiviral peptide dataset into the protein language model ESM2 based on the Transformer architecture to determine the pre-training parameters of the protein language model; Based on the pre-training parameters, the protein language model is lightweighted and the sequence encoding capability is balanced to obtain a processed protein language model; The processed protein language model is used for feature extraction, and the context-aware sequence feature representation is output as the context feature.
4. The antiviral peptide prediction method based on queue negative sampling strategy and contrastive learning according to claim 1, characterized in that: The sequence features in the antiviral peptide dataset were extracted using a feature encoding method, including: The antiviral peptide dataset was comprehensively extracted using binary coding, Zscale coding, distance pair coding, CKSAAGP coding, QSOrder coding, and DDE coding methods to obtain sequence features.
5. The antiviral peptide prediction method based on queue negative sampling strategy and contrastive learning according to claim 1, characterized in that: Selecting a target queue from the first queue and the second queue according to the anchor sample label includes: Compare the anchor sample label with the labels of the first queue and the second queue respectively to obtain a comparison result; If the comparison result shows that the label of the anchor sample is opposite to the label of the first queue, the first queue is used as the target queue; If the comparison result shows that the label of the anchor sample is opposite to the label of the second queue, the second queue is used as the target queue.
6. The antiviral peptide prediction method based on queue negative sampling strategy and contrastive learning according to claim 1, characterized in that: The negative samples in the target queue are selected using cosine similarity, including: Calculate the cosine similarity between each sample in the target queue and the anchor sample respectively; Sorting each sample in the target queue based on each cosine similarity to obtain a sorted target queue; The number of samples is determined, and negative samples are determined from the sorted target queue according to the number of samples.
7. The antiviral peptide prediction method based on queue negative sampling strategy and contrastive learning according to claim 6, characterized in that: A contrastive learning method is used to define a contrast loss function and a cross entropy loss function based on the negative sample to obtain a final loss function, including: Defining a contrastive loss function using a contrastive learning method according to each of the cosine similarities and each sample in the target queue; Determine the category weights, and fine-tune the weights in the initial cross entropy loss function based on the category weights to obtain a final cross entropy loss function; The contrast loss function and the cross entropy loss function are combined to obtain the final loss function.
8. The antiviral peptide prediction method based on queue negative sampling strategy and contrastive learning according to claim 1, characterized in that: The method further comprises: Obtaining an evaluation index, and using the evaluation index to evaluate the antiviral peptide prediction model to obtain an evaluation result; The evaluation indicators include accuracy, sensitivity, specificity, Matthews correlation coefficient, geometric mean of sensitivity and specificity, area under the precision-recall curve, area under the receiver operating characteristic curve, and F1 score.
9. The antiviral peptide prediction method based on queue negative sampling strategy and contrastive learning according to claim 1, characterized in that: The method further comprises: collecting an antiviral peptide sequence to be predicted, and inputting the antiviral peptide sequence to be predicted into the antiviral peptide prediction model; The antiviral peptide prediction model is used to output the antiviral peptide prediction results and display them on a page.
10. An antiviral peptide prediction system based on queue negative sampling strategy and contrastive learning, characterized in that: The system comprises: A data set construction module is used to collect antiviral peptide data from various databases, and construct an antiviral peptide data set after preprocessing the antiviral peptide data; A feature extraction module is used to input the antiviral peptide dataset into the protein language model ESM2 to capture context features, and use a feature encoding method to extract sequence features in the antiviral peptide dataset, and fuse the context features with the sequence features to obtain a fused feature sequence; a queue division module, configured to store different labels for each sample in the feature sequence, and to divide each sample into a first queue and a second queue according to each label; A negative sample selection module is configured to determine an anchor sample label, select a target queue from the first queue and the second queue according to the anchor sample label, and select negative samples from the target queue using cosine similarity; The antiviral peptide prediction module is used to define a contrast loss function and a cross entropy loss function based on the negative sample using a contrast learning method to obtain a final loss function, and to construct an antiviral peptide prediction model based on the final loss function to complete the antiviral peptide prediction.