Anticancer peptide prediction method and system based on feature fusion and cross attention mechanism
Through the method of feature fusion and cross-attention mechanism, combined with protein language model and traditional feature extraction model, the problems of low efficiency, low accuracy and high cost of anti-cancer peptide prediction in traditional methods are solved, and efficient and accurate anti-cancer peptide prediction is achieved.
Patent Information
- Application Number
- CN202510111098.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Traditional methods are inefficient, low-accurate, and cost-effective in identifying and discovering anticancer peptides, and involve ethical and ethical issues.
Using a method based on feature fusion and cross-attention mechanism, peptide structural features were extracted through protein language model ESM-2, and peptide physical and chemical features were extracted in combination with traditional feature extraction models, and feature fusion was performed through cross-attention mechanism. Finally, anti-cancer peptide prediction was used using multi-layer perceptron MLP.
Fast, efficient and accurate anti-cancer peptide prediction is achieved, avoiding the problems of low efficiency, low accuracy and high cost in traditional methods, and does not involve ethical and moral issues.
Smart Images

Figure CN120015123A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to an anticancer peptide prediction method and system based on feature fusion and cross-attention mechanism. Background Art
[0002] Cancer remains a huge challenge facing modern medicine, with its wide variety and strong metastatic potential, making it a malignant disease that humans have yet to fully conquer. Despite significant progress in the field of cancer treatment and multiple treatment attempts, a universally effective and patient-friendly treatment has yet to be found. Current cancer treatments—including chemotherapy, radiotherapy, surgery, and targeted therapy—each have their limitations. Chemotherapy and radiotherapy place a great burden on the body and are often accompanied by severe side effects, such as hair loss and vomiting. Surgical interventions are usually unable to eliminate cancer cells that have metastasized to other parts of the body; while the primary tumor can be removed, isolated tumor cells that remain in the surrounding tissues may not be detected. In addition, targeted therapies are expensive and only effective for specific types of cancer.
[0003] In recent years, anticancer peptides (ACPs) have attracted widespread attention from researchers. Anticancer peptides are widely present in various organisms, including mammals, amphibians, insects, plants and microorganisms, and can also be obtained synthetically. They interact with the phospholipid bilayer of cancer cell membranes, change the permeability of cell membranes, lead to leakage of cell contents, and ultimately cause cell death. Anticancer peptides have many advantages in tumor treatment: they have low molecular weight, simple structure, strong anticancer activity, high selectivity; they have few side effects, can be administered through a variety of routes, and are not prone to induce multidrug resistance.
[0004] However, the identification and discovery of anticancer peptides in a large number of extracted peptides usually rely on traditional methods, such as in vitro cell experiments or animal experiments. These methods are time-consuming, require careful design of experimental protocols, selection of appropriate control groups, and require a lot of financial support. In addition, animal experiments have gradually been questioned more and more due to the ethical and moral issues they involve. With the rapid development of artificial intelligence, a large number of methods based on machine learning or deep learning have been proposed, but these methods rely heavily on traditional feature encoding technology, require complex feature engineering steps, and there is not much correlation between the extracted features, and the focus is relatively one-sided.
[0005] It can be seen that the traditional biological experimental method has a large experimental scale and the machine learning method has low accuracy. Therefore, the traditional way of discovering anti-cancer peptides has the problems of low efficiency, low accuracy and high cost. Summary of the invention
[0006] Based on this, in order to solve the above technical problems, a method and system for predicting anticancer peptides based on feature fusion and cross-attention mechanism are provided, which can predict anticancer peptides quickly, efficiently and accurately.
[0007] A method for predicting anticancer peptides based on feature fusion and cross-attention mechanism, the method comprising:
[0008] Anticancer peptide sequences and non-anticancer peptide sequences are collected from the database to construct a data set containing various protein sequences;
[0009] Inputting the protein sequence into the protein language model ESM-2, and extracting peptide structural features in the protein sequence through the Transformer encoder;
[0010] Inputting the protein sequence into a feature extraction model, and extracting the physicochemical characteristics of peptides in the protein sequence through the feature extraction model;
[0011] Performing dimension transformation processing on the peptide structural features to obtain processed peptide structural features; using BiLSTM to make discrete peptide physicochemical features continuous to obtain continuous peptide physicochemical features;
[0012] Using a cross-attention mechanism, the processed peptide structural features and the continuous peptide physicochemical features are fused to obtain target features;
[0013] The target features are input into a multi-layer perceptron (MLP) to obtain the anticancer peptide prediction results.
[0014] In one embodiment, anti-cancer peptide sequences and non-anti-cancer peptide sequences are collected from a database to construct a data set containing various protein sequences, including:
[0015] Collecting experimentally verified anticancer peptide sequences from the database, using the CD-HIT tool to perform redundancy removal on the anticancer peptide sequences, and filtering to obtain the final anticancer peptide sequences;
[0016] randomly collecting non-anticancer peptide sequences, and extracting final non-anticancer peptide sequences from the non-anticancer peptide sequences using homology deviation removal and PSSM extraction criteria;
[0017] Anticancer peptide sequences and non-anticancer peptide sequences are randomly selected from the final anticancer peptide sequences and the final non-anticancer peptide sequences to construct a data set containing various protein sequences.
[0018] In one embodiment, the protein sequence is input into the protein language model ESM-2, and the peptide structural features in the protein sequence are extracted by the Transformer encoder, including:
[0019] Inputting the protein sequence into a protein language model ESM-2, and converting the protein sequence into a numerical vector representation through the protein language model;
[0020] Inputting the numerical vector representation into the Transformer encoder in the protein language model ESM-2;
[0021] The peptide structural features are obtained by performing calculations through the point-wise attention mechanism and the linear layer in the Transformer encoder.
[0022] In one embodiment, the protein sequence is input into a feature extraction model, and the physical and chemical characteristics of peptides in the protein sequence are extracted by the feature extraction model, including:
[0023] Inputting the protein sequence into a feature extraction model, using one-hot encoding in the feature extraction model to represent each amino acid in the protein sequence by one-hot encoding, and obtaining a binary vector corresponding to the protein sequence;
[0024] The feature extraction model is used to calculate the sum of all element masses of each amino acid in the protein sequence, and the sum of the element masses is used as the molecular weight;
[0025] Obtaining the pH values of amino acids in the protein sequence, and calculating the isoelectric point according to the pH values;
[0026] Calculating the number and properties of amino acid hydrophobic groups in the protein sequence, and determining the amino acid hydrophobicity based on the number and properties of the hydrophobic groups;
[0027] The binary vector, molecular weight, isoelectric point, and hydrophobicity are used as physicochemical characteristics of peptides in the protein sequence.
[0028] In one embodiment, the peptide structural features are dimensionally transformed to obtain processed peptide structural features; the discrete peptide physicochemical features are continuousized using BiLSTM to obtain continuous peptide physicochemical features, including:
[0029] Inputting the peptide structure features into a linear layer for dimensional transformation processing to obtain processed peptide structure features;
[0030] The physicochemical characteristics of the peptides are input into a bidirectional long short-term memory network BiLSTM, and the long-range dependencies of long sequences in the physicochemical characteristics of the peptides are captured by BiLSTM to complete the continuity of the physicochemical characteristics of the peptides, thereby obtaining continuous physicochemical characteristics of the peptides.
[0031] In one embodiment, the processed peptide structural features and the continuous peptide physicochemical features are fused using a cross-attention mechanism to obtain target features, including:
[0032] Using the processed peptide structural features and the continuous peptide physicochemical features through a cross-attention mechanism, a query matrix, a bond matrix, and a value matrix are generated respectively;
[0033] Determining the bond vector dimensions corresponding to the structural features of the processed peptides and the continuous physicochemical features of the peptides;
[0034] Calculate cross attention according to the query matrix, key matrix, value matrix, and key vector dimensions;
[0035] Based on the cross attention, the feature fusion of the processed peptide structural features and the continuous peptide physicochemical features is completed to obtain the target features.
[0036] In one embodiment, the method further comprises:
[0037] Input the target features into a Transformer architecture, and calculate a query vector, a key vector, and a value vector for each position through a multi-head self-attention mechanism in the Transformer architecture;
[0038] Performing weighted average calculation on the query vector, key vector, and value vector to obtain a weighted feature;
[0039] A feedforward neural network is used to perform position-by-position nonlinear transformation on the weighted features.
[0040] In one embodiment, the target feature is input into a multi-layer perceptron MLP to obtain an anticancer peptide prediction result, including:
[0041] Input the target feature into the input layer of the multi-layer perceptron MLP, and perform result prediction through the hidden layer and linear layer in the multi-layer perceptron MLP;
[0042] The prediction result is output from the output layer of the multi-layer perceptron MLP to obtain the anticancer peptide prediction result.
[0043] An anticancer peptide prediction system based on feature fusion and cross-attention mechanism, the system comprising:
[0044] A data collection module, used for collecting anticancer peptide sequences and non-anticancer peptide sequences from a database to construct a data set containing various protein sequences;
[0045] A structural feature extraction module, used to input the protein sequence into the protein language model ESM-2, and extract the peptide structural features in the protein sequence through the Transformer encoder;
[0046] Another feature extraction module is used to input the protein sequence into a feature extraction model, and extract the physical and chemical characteristics of peptides in the protein sequence through the feature extraction model;
[0047] A feature processing module is used to perform dimension transformation processing on the peptide structural features to obtain processed peptide structural features; use BiLSTM to make discrete peptide physicochemical features continuous to obtain continuous peptide physicochemical features;
[0048] A feature fusion module, used for fusing the processed peptide structural features and the continuous peptide physicochemical features using a cross-attention mechanism to obtain target features;
[0049] The prediction module is used to input the target features into a multi-layer perceptron MLP to obtain the anticancer peptide prediction results.
[0050] In one embodiment, the data collection module is also used to collect experimentally verified anticancer peptide sequences from a database, use the CD-HIT tool to perform redundancy processing on the anticancer peptide sequences, and filter them to obtain the final anticancer peptide sequences; randomly collect non-anticancer peptide sequences, and use homology deviation removal and PSSM extraction standards to extract the final non-anticancer peptide sequences from the non-anticancer peptide sequences; and randomly select anticancer peptide sequences and non-anticancer peptide sequences from the final anticancer peptide sequences and the final non-anticancer peptide sequences to construct a data set containing various protein sequences.
[0051] The above-mentioned anticancer peptide prediction method and system based on feature fusion and cross-attention mechanism uses a protein language model to extract peptide structural features, uses a traditional feature extraction model to extract peptide physicochemical features, uses a cross-attention mechanism for feature fusion, and finally obtains anticancer peptide prediction results based on a multi-layer perceptron MLP. It does not require time-consuming and high-cost, nor does it involve ethical issues. The extracted features are interrelated, and anticancer peptide predictions can be performed quickly, efficiently, and accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 This is an application environment diagram of an anticancer peptide prediction method based on feature fusion and cross-attention mechanism in one embodiment;
[0053] Figure 2 Schematic diagram of a process of an anticancer peptide prediction method based on feature fusion and cross-attention mechanism in one embodiment;
[0054] Figure 3 A schematic diagram of the model framework of ACP-ETBLCA in one embodiment;
[0055] Figure 4 A schematic diagram of a KDE graph, an AUC curve, and a pre-Recall curve of a model predicting probability distribution on ACP135 and ACP99 data sets in one embodiment;
[0056] Figure 5 It is a schematic diagram of measuring different ESM-2 parameters under six indicators, namely, pre, F1, recall, Sp, MCC, and ACC, in one embodiment;
[0057] Figure 6 A schematic diagram of exploring whether BiLSTM needs to be used in an embodiment;
[0058] Figure 7 A schematic diagram of exploring the origin of Q bonds in cross-attention in one embodiment;
[0059] Figure 8 A schematic diagram of exploring the number of Transformer encoder layers in one embodiment;
[0060] Fig. 9 It is a schematic diagram of the classification of samples in ACP135 after the data is extracted by the traditional feature extraction model, BiLSTM, ESM-2, linear layer, and cross-attention in one embodiment;
[0061] Fig.10 Schematic diagram of sample classification after data in ACP135 passes through 6 transformer encoder layers and average pooling layers in one embodiment;
[0062] Fig.11 is a structural block diagram of an anticancer peptide prediction system based on feature fusion and cross-attention mechanism in one embodiment;
[0063] Fig.12 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0065] The anticancer peptide prediction method based on feature fusion and cross-attention mechanism provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Figure 1 As shown, the application environment includes a computer device 110. The computer device 110 can collect anticancer peptide sequences and non-anticancer peptide sequences from a database to construct a data set containing various protein sequences; the computer device 110 can input the protein sequence into the protein language model ESM-2, and extract the peptide structural features in the protein sequence through the Transformer encoder; the computer device 110 can input the protein sequence into the feature extraction model, and extract the peptide physicochemical features in the protein sequence through the feature extraction model; the computer device 110 can perform dimensional transformation processing on the peptide structural features to obtain processed peptide structural features; use BiLSTM to make discrete peptide physicochemical features continuous to obtain continuous peptide physicochemical features; the computer device 110 can use the cross-attention mechanism to fuse the processed peptide structural features and continuous peptide physicochemical features to obtain target features; the computer device 110 can input the target features into the multi-layer perceptron MLP to obtain the anticancer peptide prediction results. Among them, the computer device 110 can be, but not limited to, various personal computers, laptops, smart phones, robots and other devices.
[0066] In one embodiment, Figure 2 As shown, a method for predicting anticancer peptides based on feature fusion and cross-attention mechanism is provided, comprising the following steps:
[0067] Step 202: collect anticancer peptide sequences and non-anticancer peptide sequences from a database to construct a data set containing various protein sequences.
[0068] The database may be a public database such as CancerPPD, APD3, SATPd, etc. The computer device may obtain 1350 experimentally verified anticancer peptide sequences through the public database, and simultaneously find 1839 non-anticancer peptide data in the literature, thereby constructing a data set containing various protein sequences.
[0069] In one embodiment, a method for predicting anticancer peptides based on feature fusion and cross-attention mechanism is provided, which may also include a process for constructing a data set, and the specific process includes: collecting experimentally verified anticancer peptide sequences from a database, using the CD-HIT tool to perform redundancy processing on the anticancer peptide sequences, and filtering to obtain the final anticancer peptide sequences; randomly collecting non-anticancer peptide sequences, and using homology deviation removal and PSSM extraction standards to extract the final non-anticancer peptide sequences from the non-anticancer peptide sequences; and randomly selecting anticancer peptide sequences and non-anticancer peptide sequences from the final anticancer peptide sequences and the final non-anticancer peptide sequences, respectively, to construct a data set containing various protein sequences.
[0070] The computer device can fuse anticancer peptides and non-anticancer peptides and use CD-HIT to remove anticancer peptide samples with more than 90% sequence similarity. Subsequently, sequences with a length between 5 and 50 amino acids were extracted using Seqkit, and only those sequences that could generate a position-specific scoring matrix (PSSM) profile through the PSI-BLAST tool were retained, and 1839 non-anticancer peptide samples were finally obtained. 487 positive sequences and 1479 negative sequences were randomly selected from these samples to construct a training data set, and the remaining sequences constituted an independent test set ACP135. At the same time, the widely used ACP99 was also used in this example.
[0071] Specifically, the training set and ACP135 were originally developed by Jilong et al., who collected 1,350 experimentally verified anticancer peptide sequences from the CancerPPD, APD3, and SATPdb databases. In order to mitigate homology bias and prevent artificial exaggeration of recognition accuracy, the filtering process in this embodiment ultimately obtained 622 positive anticancer peptide samples. For negative samples, Jilong et al. randomly selected non-anticancer peptide sequences and applied the same homology bias removal and PSSM extraction criteria, ultimately obtaining 1,839 non-anticancer peptide samples. 487 positive sequences and 1,479 negative sequences were randomly selected from these samples to construct a training data set.
[0072] ACP135 contains 135 anticancer peptide sequences and 360 non-anticancer peptide sequences, while the training set includes 487 anticancer peptide sequences and 1479 non-anticancer peptide sequences. ACP99 was developed by Agrawal et al. and contains 256 sequences, of which 99 are positive samples and 157 are negative samples. The data set has also been subjected to homology bias removal and PSSM matrix extraction. Using these strictly screened databases for training and testing can ensure that the classification performance of the model is accurately evaluated and the interference of homology-related artifacts is avoided.
[0073] Step 204: input the protein sequence into the protein language model ESM-2, and extract the peptide structural features in the protein sequence through the Transformer encoder.
[0074] The computer device can input the obtained protein sequence into the ESM-2 model for feature extraction. The ESM-2 model consists of 30 encoding layers, each of which contains a multi-head self-attention layer, a feedforward network layer, a residual connection layer, etc., to fully mine the peptide structure information.
[0075] Among them, ESM-2 is a deep learning pre-training model for protein sequence feature extraction. Similar to BERT, ESM-2 uses a self-attention mechanism to learn the representation of amino acid sequences to capture complex features and long-term dependencies in the sequences.
[0076] In one embodiment, a method for predicting anticancer peptides based on feature fusion and cross-attention mechanism is provided, which may also include a process for extracting peptide structural features, the specific process including: inputting a protein sequence into a protein language model ESM-2, converting the protein sequence into a numerical vector representation through the protein language model; inputting the numerical vector representation into a Transformer encoder in the protein language model ESM-2; and calculating through a point product attention mechanism and a linear layer in the Transformer encoder to obtain peptide structural features.
[0077] In this embodiment, in order to adapt to the context of protein sequences, ESM-2 converts the input protein sequence into a numerical vector representation, and then extracts sequence features through a 6-layer, 12-layer or higher-layer Transformer encoder. Specifically, each layer of the Transformer encoder contains multiple self-attention heads, which are calculated through the point multiplication attention mechanism and the linear layer to generate high-quality sequence features, which are then used for downstream tasks such as protein function prediction. Among them, the encoding process of ESM-2 can be expressed as: Among them, Q, K, and V represent query vector, key vector, and value vector respectively. k is the dimension of the key vector.
[0078] Step 206: Input the protein sequence into the feature extraction model, and extract the physicochemical characteristics of peptides in the protein sequence through the feature extraction model.
[0079] Among them, the feature extraction model can use traditional feature extraction methods to extract features. Specifically, the obtained protein sequence can be based on hydrophobicity, isopyclicity, one-hot encoding, molecular weight and other traditional feature extraction methods to obtain features based on peptide biochemical information.
[0080] In one embodiment, a method for predicting anticancer peptides based on feature fusion and cross-attention mechanism is provided, which may also include a process for extracting physicochemical features of peptides, and the specific process includes: inputting a protein sequence into a feature extraction model, and using the unique hot encoding in the feature extraction model to represent each amino acid in the protein sequence through one-hot encoding to obtain a binary vector corresponding to the protein sequence; calculating the sum of all element masses of each amino acid in the protein sequence through the feature extraction model, and taking the sum of the element masses as the molecular weight; obtaining the pH value of the amino acids in the protein sequence, and calculating the isoelectric point based on the pH value; calculating the number and properties of the hydrophobic groups of the amino acids in the protein sequence, and determining the hydrophobicity of the amino acids based on the number and properties of the hydrophobic groups; and taking the binary vector, molecular weight, isoelectric point, and hydrophobicity as the physicochemical features of peptides in the protein sequence.
[0081] One-hot encoding is a widely used method in machine learning and deep learning to convert categorical data into numerical vectors. In this example, each amino acid is represented by one-hot encoding, generating a binary vector in which a single position is set to 1 to indicate the presence of a specific category, and all other positions are 0. Given C categories, one-hot encoding converts each category C i Convert to a vector v of length C i , the definition formula can be expressed as:
[0082] Molecular weight (MW) refers to the total mass of all atoms in a molecule, usually in Daltons (Da). For an amino acid AAi, its molecular weight MWi is the sum of the masses of all its constituent elements.
[0083] The isoelectric point (pI) refers to the pH value at which the amino acid has no net charge in the solution. It is calculated based on the acidity and alkalinity of the amino acid side chain and involves the dissociation constants of the carboxyl and amino groups. The calculation formula can be expressed as: Among them, pK a1 , pK a2 represent the dissociation constants of carboxyl and amino groups, respectively.
[0084] Hydrophobicity is a measure of the number and nature of hydrophobic groups in amino acids, and is usually assessed using the Kyte-Doolittle scale. The hydrophobicity score HiH_iHi is defined as the value derived from the Kyte-Doolittle scale. The Kyte-Doolittle scale is a measure based on the hydrophobicity of amino acids to water, which divides the hydrophobicity of amino acids into positive values (indicating strong hydrophobicity) and negative values (indicating strong hydrophilicity). Each amino acid is assigned a score based on its hydrophobicity, and the higher the score, the more hydrophobic the amino acid.
[0085] Step 208, perform dimensionality transformation processing on the peptide structural features to obtain processed peptide structural features; use BiLSTM to make the discrete peptide physicochemical features continuous to obtain continuous peptide physicochemical features.
[0086] The computer device can use the linear layer to transform the dimension of the features extracted by ESM-2, and use BiLSTM to make the discrete data continuous.
[0087] Specifically, in one embodiment, a method for predicting anticancer peptides based on feature fusion and cross-attention mechanism is provided, which may also include a feature processing process, and the specific process includes: inputting the peptide structural features into the linear layer for dimensionality transformation processing to obtain the processed peptide structural features; inputting the peptide physicochemical features into the bidirectional long short-term memory network BiLSTM, capturing the long-distance dependencies of long sequences in the peptide physicochemical features through BiLSTM, completing the continuity of the peptide physicochemical features, and obtaining the continuous peptide physicochemical features.
[0088] Among them, the bidirectional long short-term memory network BiLSTM is an enhanced version of the traditional long short-term memory (LSTM) network. Its main advantage is that it can simultaneously capture the global contextual information of the sequence, thereby significantly improving the modeling of long-distance dependencies. Unlike the traditional LSTM model, which usually processes information in one direction, BiLSTM overcomes this limitation by allowing information to flow in two directions. For protein sequences, they can be regarded as a biological language, in which peptides correspond to sentences and single amino acid residues correspond to words. The contextual relationship between these residues is crucial for accurate prediction.
[0089] In this embodiment, although LSTM effectively alleviates the problem of gradient vanishing and gradient exploding, its inherent defect is that it cannot utilize future context information. Therefore, BiLSTM combines two independent LSTM layers: one in the forward processing sequence and the other in the reverse processing sequence. Mathematically, the forward and reverse transfer can be expressed as: h t → =LSTM forward (x t ,h t-1 → );h t ← =LSTM forward (x t ,h t+1 ← ), where h t → 、h t ←Respectively represent the hidden states produced by the forward and reverse LSTM layers at time step t. The final output of BiLSTM is obtained by concatenating or combining the forward and reverse hidden states:
[0090] This concatenated hidden state h t It provides a comprehensive contextual representation that covers both past and future state information. The bidirectional mechanism enables BiLSTM to more accurately capture long-distance dependencies in long sequences, making it particularly suitable for tasks that require extensive context understanding.
[0091] In this example, the features extracted by the traditional feature extraction method contain physicochemical properties. However, since the correlation between features from different methods is relatively weak, it may affect the quality of the synthesized features after fusion with the ESM-2 features. In order to retain the physicochemical information extracted by the traditional feature extraction model and enhance the relationship between features, BiLSTM is introduced to ensure that the physicochemical properties are effectively retained and the relationship between different feature dimensions is enhanced, thereby improving the overall feature representation and the prediction performance of the model.
[0092] In step 210, a cross-attention mechanism is used to fuse the processed peptide structural features and the continuous peptide physicochemical features to obtain target features.
[0093] Computer devices can use the cross-attention mechanism to fuse the processed features to obtain a structural information feature containing physical and chemical information. Among them, the cross-attention mechanism is a variant of the attention mechanism, which is particularly suitable for multimodal tasks, multi-sequence tasks, and scenarios involving upstream and downstream feature interactions. Its core idea is to extract contextual information from one input sequence and enhance or guide the representation of another input sequence through the attention mechanism.
[0094] In one embodiment, a method for predicting anticancer peptides based on feature fusion and cross-attention mechanism is provided, which may also include a feature fusion process, and the specific process includes: using the processed peptide structural features and the continuous peptide physicochemical features through the cross-attention mechanism to generate a query matrix, a bond matrix, and a value matrix respectively; determining the bond vector dimensions corresponding to the processed peptide structural features and the continuous peptide physicochemical features; calculating the cross-attention based on the query matrix, the bond matrix, the value matrix, and the bond vector dimensions; completing the feature fusion of the processed peptide structural features and the continuous peptide physicochemical features based on the cross-attention to obtain the target features.
[0095] The cross attention mechanism generates query, key, and value vectors through the input sequence itself. The cross attention uses two different sequences to generate query and key-value vectors respectively. The calculation formula of cross attention is: Where Q represents the query matrix obtained from one input sequence, K and V represent the key and value matrices obtained from another input sequence, respectively. k is the dimension of the key vector. Combining the features extracted from ESM-2 and fused through the cross-attention mechanism helps to effectively integrate the structural features obtained through ESM-2 with the traditional features (physicochemical features processed by BiLSTM). This configuration enables the cross-attention mechanism to effectively combine physicochemical information with structural features to generate a feature set that can capture the comprehensive properties of anticancer peptides, thereby improving the overall performance of the model.
[0096] In one embodiment, a method for predicting anticancer peptides based on feature fusion and cross-attention mechanism is provided, which may also include a feature processing process, and the specific process includes: inputting the target feature into the Transformer architecture, and calculating the query vector, key vector, and value vector for each position through the multi-head self-attention mechanism in the Transformer architecture; performing weighted average calculation on the query vector, key vector, and value vector to obtain weighted features; and using a feedforward neural network to perform nonlinear transformation on the weighted features position by position.
[0097] The computer device can use a 6-layer transformer encoder layer to further process the fused features to obtain processed features and pass them through the pooling layer.
[0098] Specifically, the Transformer encoder layer is a basic component in the Transformer architecture, which is used to encode the input sequence into a high-dimensional feature representation, which can effectively capture local and global dependencies in the data. Each encoder layer contains two main submodules: a multi-head self-attention mechanism and a feed-forward neural network (FFN). Among them, the multi-head self-attention mechanism allows each position in the sequence to pay attention to all other positions, thereby capturing long-range dependencies and contextual information. The multi-head self-attention mechanism includes calculating the query vector (Q), key vector (K), and value vector (V) for each position, and then applying weighted averaging to calculate the weighted representation, where multiple attention heads work in parallel to capture multiple features of the input data.
[0099] By using multiple attention heads, we can focus on different aspects of the input features, thereby enhancing the representation ability. After the self-attention mechanism, the feedforward neural network (FFN) performs a position-by-position nonlinear transformation on the features. FFN usually consists of two linear transformations and an activation function (such as ReLU or GELU). This structure enables the model to capture more complex feature interactions.
[0100] Step 212, input the target features into the multi-layer perceptron MLP to obtain the anticancer peptide prediction results.
[0101] The computer device can input the processed target features into the MLP to obtain the final prediction results, and classify whether it is an anti-cancer peptide according to whether the prediction probability is greater than 0.5.
[0102] Multi-Layer Perceptron (MLP) is a feed-forward artificial neural network that consists of multiple layers, each containing several neurons, capable of complex nonlinear mapping. MLP is often used to handle supervised learning problems, especially classification and regression tasks. In MLP, data is passed through the input layer, through several hidden layers (each layer consists of multiple neurons), and finally outputs the result. The hidden layer and the output layer introduce nonlinear characteristics through activation functions to improve the expressiveness of the model.
[0103] Specifically, in one embodiment, a method for predicting anticancer peptides based on feature fusion and cross-attention mechanism may also include a prediction process, which specifically includes: inputting target features into the input layer of a multi-layer perceptron MLP, and predicting results through hidden layers and linear layers in the multi-layer perceptron MLP; outputting the prediction results from the output layer of the multi-layer perceptron MLP to obtain anticancer peptide prediction results.
[0104] The application of the anticancer peptide prediction method based on feature fusion and cross-attention mechanism provided in this application is as follows: Figure 3As shown, it includes A, database data collection; B, feature extraction cross-attention mechanism for fusion and multi-layer perceptron MLP; C, practical application. Specifically, the ACP-ETBLCA prediction model consists of two stages. The first stage is used to extract features from the input protein sequence, and the second stage uses the constructed model for prediction. In the first stage, the input protein sequence is put into ESM-2 and the traditional feature extraction method that focuses on extracting the physicochemical properties of peptides, that is, the feature extraction model to extract features; the linear layer is used to transform the dimension of the features extracted by ESM-2; BiLSTM is used to further process the features extracted by the traditional feature extraction method, and the two processed features are fused using Cross-attention, and then the fused features are input into the 6-layer Transformer Encoder Layer to obtain new features; the second step is to input the obtained features into an MLP for prediction and obtain the results. Through the diversity of feature extraction and deep learning methods, the accuracy and generalization of the model are improved.
[0105] In one embodiment, in order to verify the effect of the application model of an anticancer peptide prediction method based on feature fusion and cross-attention mechanism provided in this application, five commonly used statistical indicators are used: accuracy (ACC), Matthews correlation coefficient (MCC), sensitivity (Sn) and specificity (Sp). In addition, considering the imbalance problem of the training and test data sets of the model in this application, the area under the curve (AUC) and F1 score are also included in the evaluation. The calculation formulas of these indicators are as follows:
[0106]
[0107] Among them, TP, TN, FP and FN represent the counts of true positives, true negatives, false positives and false negatives, respectively. In addition, in this embodiment, the precision-recall curve is also calculated to further evaluate the performance of the model. These indicators provide a comprehensive evaluation of the classification ability of the model, especially in the case of an unbalanced data set, where the traditional accuracy rate may not fully capture the nuances of the model performance.
[0108] In one embodiment, in order to demonstrate the superiority of the anticancer peptide prediction method application model ACP-ETBLCA based on feature fusion and cross-attention mechanism provided in this application, the ACP-ETBLCA model was compared with the existing models on the ACP135 and ACP99 datasets. The comparison between the constructed ACP-ETBLCA model and the existing model on the ACP135 dataset is as follows:
[0109] Comparison of the constructed model with existing models on the ACP135 dataset
[0110]
[0111] Among them, the comparison between the constructed ACP-ETBLCA model and the existing model on the ACP99 dataset is as follows:
[0112] Comparison of the constructed model with existing models on the ACP99 dataset
[0113]
[0114]
[0115] It can be seen that ACP-ETBLCA is basically superior to all previously constructed models in terms of indicators such as ACC (accuracy), Sp (specificity) and MCC (Matthews correlation coefficient), which shows that iBitter-GRE effectively mines the local and global information of peptide sequences and shows excellent ability in predicting bitter peptides.
[0116] Specifically, when analyzing the prediction results of ACP135 and ACP99 samples, we can observe the distribution characteristics of the two. Figure 4 As shown in Figure A, for the ACP135 sample, the predicted probabilities of negative samples are mostly concentrated in the area close to 0, and the probability of a small number of misclassified negative samples is slightly greater than 0.5. Although the positive samples are slightly scattered, they are mainly concentrated in the position close to 1, and the probability of misclassification is low. In contrast, Figure 4 As shown in Figure B, the predicted probability of negative samples of ACP99 samples is also concentrated near 0, but its misclassified negative samples are fewer, and the predicted probability of positive samples is significantly higher, with the distribution concentrated in the area close to 1. This explains the higher accuracy of ACP99 compared to ACP135.
[0117] In further performance evaluation, such as Figure 4 In C and D, the ROC curves of the models shown show that the AUC of ACP135 is 0.9589, while that of ACP99 is 0.9893, indicating that ACP99 is slightly better in model discrimination ability. The PR curve further reveals that ACP99 outperforms ACP135, with an average precision (AP) of 0.9905, higher than ACP135's 0.9380. Although both are significantly better than random classifiers, ACP99 shows a more stable advantage in balancing precision and recall.
[0118] Specifically, the measurements of different ESM-2 parameters under the six indicators of pre, F1, recall, Sp, MCC, and ACC are as follows: Figure 5As shown in the figure, using ESM-2 with 650M parameters can make the model have higher average pre, F1, recall, Sp, MCC, ACC and more concentrated data under 10-fold cross validation, better generalization ability, and better classification effect of the model.
[0119] like Figure 6 As shown in the figure, using BiLSTM to make the features extracted by the discrete traditional feature extraction method continuous can make the model have higher average pre, F1, recall, Sp, MCC, ACC and more concentrated data under 10-fold cross validation, better generalization ability, and better classification effect.
[0120] like Figure 7 As shown in the figure, using the features obtained after BiLSTM as Q in cross-attention, that is, obtaining structural features containing physical and chemical information, can enable the model to have higher average F1, recall, MCC, ACC under 10-fold cross validation, as well as more concentrated data and better generalization ability. This shows that using the features obtained after BiLSTM as cross-attention can improve the classification ability of the model.
[0121] pass Figure 8 It can be found that the number of transformer encoder layers has a significant impact on the generalization ability of the model. When using 6 layers (the same as the transformer encoder layer used in the ESM-28M parameter version), the results of the 10-fold cross validation are more concentrated, which indicates that the model has better generalization ability. In this embodiment, the T-SNE graph can also be used to visualize each block of the model.
[0122] After the data in ACP135 is extracted by traditional feature extraction models, BiLSTM, ESM-2, linear layer, and cross-attention, the classification of samples is as follows Fig. 9 As shown by Fig. 9 (A) and Fig. 9 (B) By comparison, we can find that the classification of samples has been significantly improved after using BiLSTM. Fig. 9 (C) and Fig. 9 (D) By comparison, we can find that using the linear layer only changes the dimension of the features extracted by ESM-2, and does not change the quality of the features. Fig. 9 (E) It can be seen that after using cross-attention, the samples have been significantly classified.
[0123] After the data in ACP135 passes through 6 layers of transformer encoder layer and average pooling layer, the classification of samples is as follows: Fig.10 As shown by Fig.10 By comparing (A), (B), (C), (D), (E), and (F), we can find that the classification effect is better without a transformer encoder layer. Fig.10 (G) It can be found that the use of average pooling further improves the classification effect.
[0124] It should be understood that, although the various steps in the above-mentioned flow chart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flow chart may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0125] In one embodiment, Fig.11 As shown, an anticancer peptide prediction system based on feature fusion and cross-attention mechanism is provided, including: a data collection module 1110, a structural feature extraction module 1120, other feature extraction modules 1130, a feature processing module 1140, a feature fusion module 1150 and a prediction module 1160, wherein:
[0126] A data collection module 1110 is used to collect anticancer peptide sequences and non-anticancer peptide sequences from a database to construct a data set containing various protein sequences;
[0127] The structural feature extraction module 1120 is used to input the protein sequence into the protein language model ESM-2 and extract the peptide structural features in the protein sequence through the Transformer encoder;
[0128] Other feature extraction modules 1130 are used to input the protein sequence into the feature extraction model, and extract the physical and chemical features of peptides in the protein sequence through the feature extraction model;
[0129] The feature processing module 1140 is used to perform dimensional transformation processing on the peptide structural features to obtain processed peptide structural features; use BiLSTM to make discrete peptide physicochemical features continuous to obtain continuous peptide physicochemical features;
[0130] A feature fusion module 1150 is used to fuse the processed peptide structural features and the continuous peptide physicochemical features using a cross attention mechanism to obtain target features;
[0131] The prediction module 1160 is used to input the target features into the multi-layer perceptron MLP to obtain the anticancer peptide prediction results.
[0132] In one embodiment, the data collection module 1110 is also used to collect experimentally verified anticancer peptide sequences from a database, use the CD-HIT tool to perform redundancy processing on the anticancer peptide sequences, and filter to obtain the final anticancer peptide sequences; randomly collect non-anticancer peptide sequences, and use homology deviation removal and PSSM extraction standards to extract the final non-anticancer peptide sequences from the non-anticancer peptide sequences; randomly select anticancer peptide sequences and non-anticancer peptide sequences from the final anticancer peptide sequences and the final non-anticancer peptide sequences, respectively, to construct a data set containing various protein sequences.
[0133] In one embodiment, the structural feature extraction module 1120 is also used to input the protein sequence into the protein language model ESM-2, convert the protein sequence into a numerical vector representation through the protein language model; input the numerical vector representation into the Transformer encoder in the protein language model ESM-2; and calculate through the point product attention mechanism and linear layer in the Transformer encoder to obtain the peptide structural features.
[0134] In one embodiment, other feature extraction modules 1130 are also used to input the protein sequence into the feature extraction model, use the one-hot encoding in the feature extraction model to represent each amino acid in the protein sequence through one-hot encoding, and obtain a binary vector corresponding to the protein sequence; calculate the sum of all element masses of each amino acid in the protein sequence through the feature extraction model, and use the sum of the element masses as the molecular weight; obtain the pH value of the amino acids in the protein sequence, and calculate the isoelectric point based on the pH value; calculate the number and properties of the hydrophobic groups of the amino acids in the protein sequence, and determine the hydrophobicity of the amino acids based on the number and properties of the hydrophobic groups; use the binary vector, molecular weight, isoelectric point, and hydrophobicity as the physicochemical characteristics of peptides in the protein sequence.
[0135] In one embodiment, the feature processing module 1140 is also used to input the peptide structural features into the linear layer for dimensional transformation processing to obtain the processed peptide structural features; input the peptide physicochemical features into the bidirectional long short-term memory network BiLSTM, capture the long-distance dependencies of long sequences in the peptide physicochemical features through BiLSTM, complete the continuity of the peptide physicochemical features, and obtain the continuous peptide physicochemical features.
[0136] In one embodiment, the feature fusion module 1150 is also used to generate a query matrix, a bond matrix, and a value matrix respectively using the processed peptide structural features and the continuous peptide physicochemical features through a cross-attention mechanism; determine the bond vector dimensions corresponding to the processed peptide structural features and the continuous peptide physicochemical features; calculate the cross-attention based on the query matrix, the bond matrix, the value matrix, and the bond vector dimensions; complete the feature fusion of the processed peptide structural features and the continuous peptide physicochemical features based on the cross-attention to obtain the target features.
[0137] In one embodiment, the feature fusion module 1150 is also used to input the target features into the Transformer architecture, calculate the query vector, key vector, and value vector for each position through the multi-head self-attention mechanism in the Transformer architecture; perform weighted average calculation on the query vector, key vector, and value vector to obtain weighted features; and use a feedforward neural network to perform nonlinear transformation on the weighted features position by position.
[0138] In one embodiment, the prediction module 1160 is also used to input the target features into the input layer of the multi-layer perceptron MLP, and perform result prediction through the hidden layer and linear layer in the multi-layer perceptron MLP; output the prediction results from the output layer of the multi-layer perceptron MLP to obtain the anti-cancer peptide prediction results.
[0139] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig.12 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for predicting anticancer peptides based on feature fusion and cross-attention mechanism is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse, etc.
[0140] Those skilled in the art will understand that Fig.12The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0141] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the anti-cancer peptide prediction method based on feature fusion and cross-attention mechanism are implemented.
[0142] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the anti-cancer peptide prediction method based on feature fusion and cross-attention mechanism are implemented.
[0143] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0144] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0145] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for predicting anticancer peptides based on feature fusion and cross-attention mechanism, characterized in that: The method comprises: Anticancer peptide sequences and non-anticancer peptide sequences are collected from the database to construct a data set containing various protein sequences; Inputting the protein sequence into the protein language model ESM-2, and extracting peptide structural features in the protein sequence through the Transformer encoder; Inputting the protein sequence into a feature extraction model, and extracting the physicochemical characteristics of peptides in the protein sequence through the feature extraction model; Performing dimension transformation processing on the peptide structural features to obtain processed peptide structural features; using BiLSTM to make discrete peptide physicochemical features continuous to obtain continuous peptide physicochemical features; Using a cross-attention mechanism, the processed peptide structural features and the continuous peptide physicochemical features are fused to obtain target features; The target features are input into a multi-layer perceptron (MLP) to obtain the anticancer peptide prediction results.
2. The anticancer peptide prediction method based on feature fusion and cross-attention mechanism according to claim 1, characterized in that: Anticancer peptide sequences and non-anticancer peptide sequences were collected from the database to construct a data set containing various protein sequences, including: Collecting experimentally verified anticancer peptide sequences from the database, using the CD-HIT tool to perform redundancy removal on the anticancer peptide sequences, and filtering to obtain the final anticancer peptide sequences; randomly collecting non-anticancer peptide sequences, and extracting final non-anticancer peptide sequences from the non-anticancer peptide sequences using homology deviation removal and PSSM extraction criteria; Anticancer peptide sequences and non-anticancer peptide sequences are randomly selected from the final anticancer peptide sequences and the final non-anticancer peptide sequences to construct a data set containing various protein sequences.
3. The anticancer peptide prediction method based on feature fusion and cross-attention mechanism according to claim 1, characterized in that: The protein sequence is input into the protein language model ESM-2, and the peptide structural features in the protein sequence are extracted through the Transformer encoder, including: Inputting the protein sequence into a protein language model ESM-2, and converting the protein sequence into a numerical vector representation through the protein language model; Inputting the numerical vector representation into the Transformer encoder in the protein language model ESM-2; The peptide structural features are obtained by performing calculations through the point-wise attention mechanism and the linear layer in the Transformer encoder.
4. The anticancer peptide prediction method based on feature fusion and cross-attention mechanism according to claim 1, characterized in that: The protein sequence is input into a feature extraction model, and the physical and chemical characteristics of peptides in the protein sequence are extracted by the feature extraction model, including: Inputting the protein sequence into a feature extraction model, using one-hot encoding in the feature extraction model to represent each amino acid in the protein sequence by one-hot encoding, and obtaining a binary vector corresponding to the protein sequence; The feature extraction model is used to calculate the sum of all element masses of each amino acid in the protein sequence, and the sum of the element masses is used as the molecular weight; Obtaining the pH values of amino acids in the protein sequence, and calculating the isoelectric point according to the pH values; Calculating the number and properties of amino acid hydrophobic groups in the protein sequence, and determining the amino acid hydrophobicity based on the number and properties of the hydrophobic groups; The binary vector, molecular weight, isoelectric point, and hydrophobicity are used as physicochemical characteristics of peptides in the protein sequence.
5. The anticancer peptide prediction method based on feature fusion and cross-attention mechanism according to claim 1, characterized in that: Performing dimensional transformation processing on the peptide structural features to obtain processed peptide structural features; BiLSTM is used to make discrete peptide physicochemical features continuous, and continuous peptide physicochemical features are obtained, including: Inputting the peptide structure features into a linear layer for dimensional transformation processing to obtain processed peptide structure features; The physicochemical characteristics of the peptides are input into a bidirectional long short-term memory network BiLSTM, and the long-range dependencies of long sequences in the physicochemical characteristics of the peptides are captured by BiLSTM to complete the continuity of the physicochemical characteristics of the peptides, thereby obtaining continuous physicochemical characteristics of the peptides.
6. The anticancer peptide prediction method based on feature fusion and cross-attention mechanism according to claim 1, characterized in that: The processed peptide structural features and the continuous peptide physicochemical features are fused using a cross-attention mechanism to obtain target features, including: Using the processed peptide structural features and the continuous peptide physicochemical features through a cross-attention mechanism, a query matrix, a bond matrix, and a value matrix are generated respectively; Determining the bond vector dimensions corresponding to the structural features of the processed peptides and the physicochemical features of the continuous peptides; Calculate cross attention according to the query matrix, key matrix, value matrix, and key vector dimensions; Based on the cross attention, the feature fusion of the processed peptide structural features and the continuous peptide physicochemical features is completed to obtain the target features.
7. The anticancer peptide prediction method based on feature fusion and cross-attention mechanism according to claim 6, characterized in that: The method further comprises: Input the target features into a Transformer architecture, and calculate a query vector, a key vector, and a value vector for each position through a multi-head self-attention mechanism in the Transformer architecture; Performing weighted average calculation on the query vector, key vector, and value vector to obtain a weighted feature; A feedforward neural network is used to perform position-by-position nonlinear transformation on the weighted features.
8. The anticancer peptide prediction method based on feature fusion and cross-attention mechanism according to claim 1, characterized in that: The target features are input into the multi-layer perceptron MLP to obtain the anticancer peptide prediction results, including: Input the target feature into the input layer of the multi-layer perceptron MLP, and perform result prediction through the hidden layer and linear layer in the multi-layer perceptron MLP; The prediction result is output from the output layer of the multi-layer perceptron MLP to obtain the anticancer peptide prediction result.
9. An anticancer peptide prediction system based on feature fusion and cross-attention mechanism, characterized in that: The system comprises: A data collection module, used for collecting anticancer peptide sequences and non-anticancer peptide sequences from a database to construct a data set containing various protein sequences; A structural feature extraction module, used to input the protein sequence into the protein language model ESM-2, and extract the peptide structural features in the protein sequence through the Transformer encoder; Another feature extraction module is used to input the protein sequence into a feature extraction model, and extract the physical and chemical characteristics of peptides in the protein sequence through the feature extraction model; A feature processing module is used to perform dimension transformation processing on the peptide structural features to obtain processed peptide structural features; use BiLSTM to make discrete peptide physicochemical features continuous to obtain continuous peptide physicochemical features; A feature fusion module, used for fusing the processed peptide structural features and the continuous peptide physicochemical features using a cross-attention mechanism to obtain target features; The prediction module is used to input the target features into a multi-layer perceptron MLP to obtain the anticancer peptide prediction results.
10. The anticancer peptide prediction system based on feature fusion and cross-attention mechanism according to claim 9, characterized in that: The data collection module is also used to collect experimentally verified anticancer peptide sequences from a database, use the CD-HIT tool to perform redundancy processing on the anticancer peptide sequences, and filter to obtain the final anticancer peptide sequences; randomly collect non-anticancer peptide sequences, and use homology deviation removal and PSSM extraction standards to extract the final non-anticancer peptide sequences from the non-anticancer peptide sequences; and randomly select anticancer peptide sequences and non-anticancer peptide sequences from the final anticancer peptide sequences and the final non-anticancer peptide sequences to construct a data set containing various protein sequences.
Citation Information
Patent Citations
Anticancer peptide recognition method and system based on attention mechanism and multi-granularity hierarchical characteristics
CN116935951A
Anticancer peptide prediction method and related device
CN118314949A
Natural language processing to predict properties of proteins
US20230386610A1
Cited By
Novel method for predicting therapeutic peptide by multi-kernel fuzzy system based on deep stacked encoder
CN120260693A
Method for classifying antihypertensive peptides by fusing sequence and structure multi-modal features and combining contrast-generative combined optimization
CN120766773A
Multi-model fusion lactic acid bacteria antibacterial peptide multi-activity prediction method and system
CN121438946A