Identification method of anticancer peptide

By employing a multimodal feature fusion method that combines amino acid sequence and physicochemical properties, dynamically adjusting weights, and utilizing deep learning technology, the accuracy of anticancer peptide identification is improved. This addresses the shortcomings of existing algorithms in feature extraction and fusion, achieving high-precision and efficient anticancer peptide identification.

CN121725879APending Publication Date: 2026-03-24HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing anticancer peptide recognition algorithms neglect the special physicochemical properties of anticancer peptides, such as hydrophilicity, hydrophobicity, and plasmon resonance, during the feature extraction stage. Furthermore, they lack deep fusion during feature fusion, resulting in insufficient recognition accuracy and susceptibility to overfitting.

Method used

A multimodal feature fusion method is adopted, which combines amino acid sequence features and physicochemical property features through natural language processing and deep learning techniques. The multimodal feature fusion algorithm with dynamic adaptive weights dynamically adjusts the feature weights, combines multi-head attention mechanism for information fusion, and uses convolutional neural network and multilayer perceptron for classification.

Benefits of technology

It improved the accuracy and interpretability of anticancer peptide identification, reduced the false positive rate, and enhanced the model's generalization ability and identification efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121725879A_ABST
    Figure CN121725879A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an anti-cancer peptide recognition method based on multi-modal feature fusion. The anti-cancer peptide recognition method is based on multi-modal feature fusion of a natural language processing technology and a deep learning technology, and a high-precision interpretable anti-cancer peptide recognition algorithm is constructed by analyzing amino acid composition and physicochemical properties of a peptide sequence and dynamically fusing multi-source information in combination with a multi-head attention mechanism. According to the method, seven types of key physicochemical indexes can be systematically extracted, multi-dimensional characteristics such as coverage charge distribution, hydrophobicity and structural stability can be systematically extracted, cross-modal fusion is carried out on sequence codes and seven-dimensional physicochemical characteristics through a sequence-physicochemical characteristic dynamic fusion strategy, weights of the two types of characteristics are dynamically adjusted through gating and a multi-head attention mechanism, and therefore, multi-modal fusion is realized. The potential ACPs can be rapidly screened from massive peptide sequences, and efficient calculation support is provided for cancer treatment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of biological information, and particularly relates to an anticancer peptide recognition method based on multi-modal feature fusion by applying natural language processing technology and deep learning technology. BACKGROUND

[0002] Cancer is the second leading cause of death worldwide after cardiovascular disease, and research and development of effective cancer treatment strategies have always been an important issue in the field of medical research. With the gradual deepening of cancer treatment research, researchers focus on a promising treatment method, anticancer peptide therapy.

[0003] Anticancer peptides (ACPs) are a kind of bioactive peptides with short amino acid sequences, usually composed of 2-50 amino acid residues, with a relative molecular mass of 50-5000. In terms of amino acid composition, ACPs are rich in positively charged lysine, arginine and histidine, and hydrophobic leucine. These residues can endow ACPs with cationic properties and hydrophobicity, and they can target and kill cancer cells without damaging normal cells. More and more ACPs are entering the clinic and becoming a potential alternative for cancer treatment. More and more ACPs have been identified and verified, and under this background, it is laborious, time-consuming and expensive to identify new ACPs from a large number of active peptides using traditional wet experimental methods. With the expansion of computer application technology, large-scale screening of sequences with potential to become ACPs through recognition algorithms to guide biological experimental verification has become a key means to improve ACP recognition efficiency and accelerate the clinical application of new ACPs.

[0004] Over the past decade, researchers have proposed a large number of ACP recognition algorithms that can screen sequences with the most anticancer peptide characteristics from a large number of peptide sequences. ACP recognition algorithms can be divided into two parts: feature extraction and classifier.

[0005] In the feature extraction part, researchers such as Tyagi, Wei Chen, Nalini Schaduangrat, Rao, Yi H-C, etc. focus on the sequence features of peptides, i.e. the composition and order of amino acids. Early Tyagi et al. used amino acid composition (ACC) and binary profile (BP) as main input features; thereafter, Agrawal et al. optimized the AntiCP model and proposed AntiCP2.0, which added mixed features and terminal composition to the original model to improve the prediction accuracy; Rao et al. proposed the predictor ACpred-Fuse, which used 29 different sequence-based feature encoding algorithms to further improve the feature extraction capability by combining multi-view information; inspired by natural language processing technology (NLP) and transfer learning, Zhibin Lv et al. further proposed an identification model iACP-DRLF based on deep representation learning, which used two pre-trained embedding models to encode amino acid sequences; recently, Yao et al. also used Transformer and BERT to extract sequence features of anticancer peptides and fused them with physicochemical properties to propose the identification model ACP-CapsPred, which has higher accuracy and stronger interpretability than the above models.

[0006] In the classifier part, researchers try to use deep learning algorithms for classification. The ACP-DL model proposed by Yi H-C et al. uses deep learning algorithms in the classifier part and uses LSTM for classification; the Con-ACP model proposed by Lee, Byungjo and Shin Dongkwan uses a contrast learning strategy to improve the prediction accuracy of the model; the identification model ACP-CapsPred proposed by Yao et al. uses a capsule network to realize the binary classification task.

[0007] However, in the field of ACP identification, existing algorithms still have shortcomings: (1) Based on the mechanism of anticancer peptides, their special physicochemical properties such as hydrophilicity, hydrophobicity and plasma are not paid attention to in the feature extraction stage; (2) When fusing sequence features and physicochemical property features, the matrix is simply spliced without deep fusion; (3) There are fewer known ACP types, and overfitting is prone to occur. In view of the shortcomings of existing ACP identification algorithms, it is a realistic technical need to develop new ACP identification technology that can overcome the above shortcomings. SUMMARY

[0008] In view of the problems in algorithm identification of ACPs in the prior art, the application provides an anti-cancer peptide identification method based on multi-modal feature fusion, which is based on multi-modal feature fusion of natural language processing technology and deep learning technology, analyzes the amino acid composition and physicochemical properties of the peptide sequence, dynamically fuses multi-source information by combining a multi-head attention mechanism, and constructs a high-precision and interpretable ACP identification algorithm.

[0009] To solve the above technical problems, the application provides an anti-cancer peptide identification method based on multi-modal feature fusion, which applies ACP sequence feature extraction algorithm and ACP physicochemical feature extraction algorithm, and ACP multi-modal feature fusion algorithm with dynamic adaptive weight.

[0010] Preferably, the anti-cancer peptide identification method comprises preparing original peptide series data first, and then the steps include:

[0011] Step 1: collecting anti-cancer peptide data;

[0012] Step 2: data preprocessing;

[0013] Step 3: performing ACP sequence feature extraction and ACP physicochemical feature extraction;

[0014] Step 4: performing ACP multi-modal feature fusion with dynamic adaptive weight;

[0015] Step 5: performing ACP classification calculation based on multi-layer perception;

[0016] Finally, the anti-cancer peptide identification result is obtained.

[0017] Preferably, the physicochemical features extracted by performing the ACP physicochemical feature extraction algorithm include molecular weight (Molecular Weight), isoelectric point (Isoelectric Point), GRAVY hydrophobic value (Grand Average of Hydropathy), instability index (Instability Index), extinction coefficient (Extinction Coefficient), aliphatic index (Aliphatic Index), and charge distribution (Charge Distribution), and the physicochemical properties of the anti-cancer peptide are analyzed by constructing a multi-dimensional feature vector.

[0018] Preferably, the ACP multi-modal feature fusion algorithm with dynamic adaptive weight comprises an algorithm based on a gating fusion mechanism.

[0019] Preferably, the ACPs multimodal feature fusion algorithm with dynamic adaptive weights maps the feature vectors of sequential modes and physicochemical modes to an isomorphic latent space through linear projection operations, thereby establishing a cross-modal geometric alignment relationship.

[0020] As a preferred embodiment of step 5, the classifier used in the ACPs classification calculation based on multilayer perceptron adopts a three-layer fully connected neural network architecture, which extracts high-order features step by step through linear transformation layers.

[0021] As a preferred embodiment of step 5, the ACPs classification calculation step based on multilayer perceptron includes: the network input is first projected to the latent space dimension through the first fully connected layer, and then nonlinearity is introduced through the ReLU activation function and Dropout regularization is applied to suppress overfitting; the second layer performs feature recombination in the latent space, and its output is fused with the original input features after linear transformation. This design retains key feature information through cross-layer connections; the final layer uses the Sigmoid function to constrain the output to the [0,1] probability space, and sets 0.5 as the decision threshold to achieve binary classification.

[0022] The ACPs identification method proposed in the above technical solution of this invention has the following advantages compared with similar methods in the prior art:

[0023] 1. To address the problem of insufficient mining of physicochemical features by traditional algorithms, the technical solution of this invention systematically extracts seven key physicochemical indicators, including molecular weight, isoelectric point, instability index, etc., covering multi-dimensional characteristics such as charge distribution, hydrophobicity, and structural stability, and fully extracts the global physicochemical features of multi-dimensional peptide chains;

[0024] 2. To address the shortcomings of traditional anticancer peptide identification methods that often employ simple fusion strategies such as static matrix splicing, this invention innovatively proposes a multimodal feature fusion method based on dynamic adaptive weights. It adopts a sequence-physicochemical feature dynamic fusion strategy to fuse sequence encoding with seven-dimensional physicochemical features across modalities. By dynamically adjusting the weights of the two types of features through gating and multi-head attention mechanisms, it effectively captures the dynamic interaction between different sequences and physicochemical features.

[0025] 3. The technical solution of this invention designs a high-precision hybrid model architecture. The proposed CNN+Gate model combination shows good performance in terms of efficiency and accuracy in the short peptide recognition task. Attached Figure Description

[0026] The following figures are provided to further illustrate the invention and form part of the specification. They are used together with the detailed embodiments to explain the invention, but do not constitute a limitation thereof. They include:

[0027] Figure 1A schematic flowchart illustrating the steps of the anticancer peptide identification method based on multimodal feature fusion provided in an embodiment of the present invention;

[0028] Figure 2 This is an architecture diagram of the ACPs physicochemical feature extraction algorithm provided in an embodiment of the present invention;

[0029] Figure 3 This is an architecture diagram of an algorithm based on a gating fusion mechanism provided in an embodiment of the present invention. Detailed Implementation

[0030] To make the technical problems, solutions, and advantages of this invention clearer, a detailed description will be provided below with reference to the accompanying drawings and specific embodiments. The examples given are for illustrative purposes only and are not intended to limit the scope of the invention.

[0031] To address the shortcomings of existing ACPs identification algorithms, this invention provides an anticancer peptide identification method based on multimodal feature fusion. Amino acids are treated as words, and peptide segments composed of amino acids are treated as sentences. Natural language processing techniques, such as pre-trained language models and further pre-training in the domain, are used to form the anticancer peptide identification method proposed in this invention.

[0032] To achieve the above technical objectives, such as Figure 1 As shown in the example, this embodiment provides a method for identifying anticancer peptides based on multimodal feature fusion. First, raw peptide series data is prepared, and the subsequent specific steps include:

[0033] Step 1: Collect anticancer peptide data;

[0034] Step 2: Data preprocessing;

[0035] Step 3: Perform ACP sequence feature extraction and ACP physicochemical feature extraction;

[0036] Step 4: Perform ACPs multimodal feature fusion with dynamic adaptive weights;

[0037] Step 5: Perform ACPs classification calculation based on multilayer perceptron;

[0038] The final result was the identification of anticancer peptides.

[0039] The specific implementation method of step 1 is as follows:

[0040] The dataset used by the AntiCP2.0 identification model was collected for model training and performance evaluation. The first 20 N-terminal amino acid residues were extracted from the AntiCP2.0 dataset.

[0041] ACP sequences exhibit high conservation of specific amino acids at certain positions. F (phenylalanine), G (glycine), and K (lysine) occur frequently at position 1 in ACP sequences, which may be related to their important roles in anticancer peptide function. Non-ACP sequences show a more dispersed amino acid distribution and lack a similar conservation pattern. AMP residues are relatively concentrated at certain key positions, showing a certain conservation trend. Random peptides show a more even distribution of residues across positions, lacking a clear concentration trend. Furthermore, L (leucine) occurs frequently in both ACP and non-ACP sequences, but at a higher frequency in non-ACP sequences, which may indicate that leucine has higher conservation or functional relevance in non-anticancer peptides.

[0042] The specific implementation method of step 2 is as follows:

[0043] This step aims to transform the raw character-based peptide sequence data into a numerical form that can be directly processed by deep learning models, and to standardize it to ensure the uniformity of the model input format, the appropriateness of the numerical range, and to improve the training efficiency and robustness of the model.

[0044] First, discretization encoding is performed on the character data of the peptide sequence. This process involves constructing an amino acid-integer mapping dictionary (char_to_int), converting each amino acid character (e.g., 'A', 'C') into a corresponding unique integer index (e.g., A→1, C→2), and assigning index 0 to the padding token. Subsequently, each amino acid sequence is converted into an integer sequence. This discretization encoding is the foundation of sequence data embedding; it transforms discrete symbolic information into a numerical form that the model can process, facilitating subsequent neural network learning of the sequence's semantic features.

[0045] Secondly, sequence length standardization is implemented to meet the requirements of deep learning models for fixed input dimensions and to focus on the functional core regions of peptides. Since the functional regions of anticancer peptides are typically concentrated in relatively short sequences, while the original peptide sequence lengths may vary, this embodiment limits the maximum length to 50 residues. For long sequences exceeding 50 residues, this invention truncates them to the first 50 residues to preserve their core functional information; for short sequences less than 50 residues, padding (corresponding to integer encoding 0) is used to ensure a uniform length of 50 bits. After this processing, all sequence data have a uniform tensor shape of [batch_size, 50] ([64, 50] when batch_size is 64), thereby reducing computational complexity and avoiding the challenges of handling variable-length sequences.

[0046] Finally, the extracted seven-dimensional physicochemical features are normalized. These seven features are molecular weight, isoelectric point, Gravy hydrophobicity, instability index, extinction coefficients 1 and 2, fat index, and charge distribution. This step is crucial because the numerical range and units of the physicochemical features vary significantly. Without processing, features with larger numerical values ​​may dominate during model training, leading to underfitting and prolonged convergence time. In this embodiment, RobustScaler is used for normalization. RobustScaler is a scaler with good robustness to outliers. By removing the median of the features and scaling according to the quartile range (IQR), it effectively reduces the impact of extreme values ​​on the scaling results, making it suitable for datasets containing outliers. The processed physicochemical feature data (train_pc_normalized and test_pc_normalized) are finally converted into PyTorch tensors of type torch.float32 for computation by the deep learning model.

[0047] The specific implementation method of step 3 is as follows:

[0048] (1) ACPs sequence feature extraction

[0049] This step aims to leverage the local pattern capture capabilities of convolutional neural networks (CNNs) on peptide sequences. Its core implementation follows the design of the CNNEncoder class, with details as follows:

[0050] In the initial stage, each preprocessed amino acid integer sequence (tensor shape [batch_size, seq_len], where batch_size is 64 and seq_len is normalized to 50 in this embodiment) first undergoes feature transformation through an embedding layer (nn.Embedding). This embedding layer is responsible for mapping discrete residue index representations into dense vector embeddings with continuous vector space projection properties. Specific parameters are set as vocab_size (vocabulary size, covering 20 standard amino acids and 1 padding character, totaling 21) and emb_dim (embedding dimension, set to 64). This step transforms the input sequence from discrete symbol space to continuous vector space, and the output embedded tensor shape is [64, 50, 64]. To adapt to the input specification of the nn.Conv1d layer (i.e., [batch_size, in_channels, seq_len]), this embedded tensor undergoes dimension transposition using the permute(0, 2, 1) operation, changing its shape to [64, 64, 50]. After this transformation, emb_dim (64) is treated as the number of input channels, while seq_len (50) represents the sequence length, consistent with the expected input format for the convolution operation.

[0051] During the feature extraction stage, multi-scale one-dimensional convolution operations are performed on the dimensionally transposed embedded representation. To this end, a set of parallel convolutional layers (nn.ModuleList) containing three different scale convolutional kernels (filter_sizes = [3, 4, 5]) is constructed. Each scale is configured with num_filters (set to 100 in this embodiment) of independent feature map channels, enabling parallel feature extraction through channel-wise operations. This means that for each kernel size (e.g., kernel_size=3), the model slides 100 different convolutional kernels in parallel across the sequence, each kernel designed to learn to detect a specific 3-mer pattern. The in_channels of each convolutional layer (nn.Conv1d) are set to emb_dim (64), while the out_channels are num_filters (100). After the convolution operation, each scale generates a feature map with a shape of [64, 100, 50 - kernel_size + 1]. Subsequently, the output of each convolutional layer is immediately nonlinearly activated using the ReLU (Rectified Linear Unit) activation function. The role of ReLU(x) = max(0, x) is to eliminate negative responses, ensuring that only positive feature signals with predictive validity of anticancer peptide function are retained, thereby enhancing the nonlinear expressive power of the model.

[0052] Next, for each ReLU-activated convolutional output, global max pooling is performed along the sequence axis (i.e., dim=2) (torch.max(conv_output, dim=2)[0]). The purpose of global max pooling is to extract the expression intensity spectrum peaks from each feature map, that is, to capture the most representative activation value of each convolutional kernel in the entire sequence. Among them, the maximum response value of each channel corresponds to the optimal matching degree of a specific functional mode in the sequence. This operation gives the model insensitivity to the position of a specific mode in the sequence (position invariance). After pooling, the output shape of each scale becomes [64, 100].

[0053] Finally, the max pooling results of all convolutional kernels of different scales (i.e., kernels of sizes 3, 4, and 5) are concatenated along the feature dimension (dim=1) (torch.cat(pooled_results, dim=1)). This concatenation operation effectively integrates local features captured from different receptive fields (i.e., N-grams of different lengths). Since num_filters is set to 100 and filter_sizes contains 3 different scales, the shape of the concatenated feature vector is [64, 100 * 3], i.e., [64, 300]. Finally, to suppress the risk of model overfitting and improve the model's generalization ability, a random weight decay strategy is applied, i.e., a Dropout layer is introduced, with the probability dropout_rate set to 0.2 in this embodiment.

[0054] (2) Extraction of physicochemical features of ACPs

[0055] This step aims to quantify the intrinsic biophysical and chemical properties of the peptide sequence, which have been widely proven to be closely related to anticancer activity. This embodiment achieves this goal through a series of systematic calculations and preprocessing steps.

[0056] Specifically, such as Figure 2As shown, in this embodiment, the Biopython bioinformatics toolkit was used to systematically calculate seven types of key physicochemical parameters. These parameters include: molecular weight, calculated based on the sum of the atomic masses of all amino acid residues and their side chains in the peptide; isoelectric point (pI), determined by predicting the net charge of the peptide at different pH values ​​to determine the pH value at which its net charge is zero; GRAVY (Grand Average of Hydrophobicity), calculated by averaging the hydrophobicity indices of all amino acids in the peptide to assess the overall hydrophobicity of the peptide; instability index, calculated based on the statistical weights of specific amino acid pairs to predict the stability of the peptide under in vitro conditions; extinction coefficient, which reflects the ability of the peptide to absorb ultraviolet light at a specific wavelength (usually 280 nm), mainly affected by the content of aromatic amino acids (tyrosine, tryptophan), and in this embodiment further subdivided into extinction coefficient 1 and extinction coefficient 2, which usually correspond to the calculation results of the reduced and oxidized states of cysteine; and aliphatic index. The index quantifies the relative content of aliphatic side chains (alanine, valine, isoleucine, and leucine) in the peptide, which is closely related to the peptide's thermal stability; and the charge distribution, which typically refers to the net charge of the peptide under physiological pH conditions (e.g., pH 7.4), a key factor in its interaction with negatively charged cancer cell membranes. Through the above comprehensive calculations, each peptide sequence is characterized as a seven-dimensional physicochemical feature vector. To eliminate model bias that may be caused by differences in dimensions and numerical ranges between different physicochemical features and to accelerate model convergence, the extracted seven-dimensional physicochemical feature vectors were normalized. In this embodiment, RobustScaler was used for processing. This scaler, by removing the median and scaling according to the interquartile range (IQR), has good robustness to outliers in the data. After RobustScaler processing, the original training set physicochemical feature train_pc was converted to train_pc_normalized, and the test set physicochemical feature test_pc was converted to test_pc_normalized. Finally, these normalized physicochemical features are converted into PyTorch tensors, with the data type uniformly set to torch.float32, to facilitate input to deep learning models, and form a feature matrix of shape [64, 7] in each batch (batch_size is 64).In the early stages of data preprocessing, the Mann-Whitney U nonparametric test verified that these seven physicochemical indicators showed statistically significant differences between the anticancer peptide and non-anticancer peptide groups (p<0.05), thus providing strong biological and statistical evidence for the effectiveness of these characteristics in distinguishing ACPs.

[0057] The specific implementation method of step 4 is as follows:

[0058] The Dynamic Adaptive Weighted Multimodal Feature Fusion of ACPs aims to overcome the limitations of static feature splicing in traditional anticancer peptide identification methods. It innovatively proposes a gated dynamic adaptive weighted multimodal feature fusion strategy to achieve deep and effective integration of sequence features and physicochemical features. The core of this fusion process lies in the implementation of the FeatureFusion_gate module, the specific technical details of which are as follows:

[0059] First, to establish cross-modal geometric alignment and map feature vectors from different modalities to a homogeneous latent space, linear projection operations are performed on the sequence features (shape [batch_size, 300], where batch_size is 64) from the ACPs sequence feature extraction module and the seven-dimensional physicochemical features (shape [64, 7]) from the ACPs physicochemical feature extraction module. Specifically, the sequence features are transformed using `self.seq_proj = nn.Linear(300, 64)`, reducing their dimension from 300 to 64; simultaneously, the physicochemical features are transformed using `self.pc_proj = nn.Linear(7, 64)`, increasing their dimension from 7 to 64. After these linear projections, both types of features are mapped to a unified latent space with a dimension set to `hidden_dim` (i.e., 64), thus eliminating the dimensional differences between modalities and laying the foundation for subsequent deep fusion.

[0060] Subsequently, a learnable gating mechanism (self.gate) is introduced to dynamically adjust the contribution weights of sequential and physicochemical modes in the fusion process, such as... Figure 3The algorithm is shown below, based on a gated fusion mechanism. First, the sequence features (shape [64, 64]) after linear projection are concatenated with the physicochemical features (shape [64, 64]) along the feature dimension (torch.cat([seq_features, pc_features], dim=1)) to form a joint feature vector with shape [64, 128]. Next, this joint feature vector is processed by a gated network composed of a Multilayer Perceptron (MLP). This gated network contains a linear layer (nn.Linear(128, 64)), followed by a ReLU activation function to introduce non-linearity, then another linear layer (nn.Linear(64, 1)), and finally a Sigmoid activation function. The output gate_value of the Sigmoid function is a scalar between 0 and 1 (shape [64, 1]), dynamically representing the importance weight of the sequence features relative to the physicochemical features in the current input batch.

[0061] The final gated fusion is achieved through the following weighted summation operation:

[0062] gated_fused_features = gate_value * seq_features + (1 - gate_value) *pc_features

[0063] In this model, `gate_value` directly affects the sequence features, while `(1 - gate_value)` affects the physicochemical features. This mechanism allows the model to adaptively allocate weights for different modalities based on the context of the input data, thereby achieving context-aware joint representation. For example, when `gate_value` is close to 1, the contribution of sequence features is enhanced; when `gate_value` is close to 0, the contribution of physicochemical features is more significant. This dynamic adjustment capability enables the model to integrate multi-source information more flexibly and effectively, avoiding information redundancy or insufficiency that may result from traditional static concatenation and fusion. This generates a 64-dimensional (i.e., `hidden_dim`) fused feature vector containing rich semantic and physicochemical information for subsequent classification tasks.

[0064] The specific implementation method of step 5 is as follows:

[0065] The ACPs classification computation based on a multilayer perceptron involves using a fused feature vector generated by a dynamically adaptive weighted multimodal feature fusion module (FeatureFusion_gate), and then employing a specially designed MLP classifier to achieve the final ACPs binary classification task. The architecture and operational details of this MLP classifier are as follows:

[0066] The classifier employs a three-layer fully connected neural network architecture, aiming to extract high-order discriminative information from the fused features through progressive linear transformations and nonlinear activations. Its input dimension, input_dim, is set to 64, consistent with the output dimension of the fused features; the intermediate hidden layer dimension, hidden_dim, is set to 32; and the final output dimension, output_dim, is set to 2, corresponding to the two categories (anticancer peptides / non-anticancer peptides) in the binary classification task.

[0067] The network input x (i.e., the fused feature vector, with shape [batch_size, 64]) is first projected onto the latent space dimension of 32 by the first fully connected layer self.fc1 = nn.Linear(64, 32). To enhance the model's generalization ability and suppress overfitting, a Dropout layer (self.dropout) is then applied, with the dropout_rate set to 0.1 in this embodiment. Next, a non-linearity is introduced through the ReLU activation function self.relu to capture the complex relationships between features. It is worth noting that this implementation includes a residual connection design (if input_dim != hidden_dim: self.shortcut = nn.Linear(input_dim, hidden_dim) else: self.shortcut = nn.Identity()), although commented out in the current forward propagation (# x = x + identity). Its original intention was to establish a skip connection between the original input and the first layer output to preserve key feature information and potentially help alleviate the vanishing gradient problem.

[0068] Subsequently, the processed feature vectors enter the second fully connected layer `self.fc2 = nn.Linear(32, 32)`, where further reorganization and abstraction of the features are performed within the latent space. After this layer, Dropout (`self.dropout`) is applied for regularization, and non-linearity is introduced again through the ReLU activation function.

[0069] Finally, the feature vectors, after two layers of feature extraction and nonlinear transformation, are passed through a third fully connected layer, `self.fc3 = nn.Linear(32, 2)`, to convert their dimension to 2, corresponding to the original scores (logits) of the two classes. Although the sigmoid activation function (`self.sigmoid`) is included in the class definition, it is commented out in the actual forward propagation (`# x = self.sigmoid(x)`), indicating that the model usually combines the loss function `nn.CrossEntropyLoss`, which already includes a softmax operation. Therefore, it is more in line with standard practice for the classifier to directly output logits. During prediction, these logits are converted into a probability distribution through softmax, and the final binary classification is achieved by setting a decision threshold of 0.5: when the probability of a certain class is higher than 0.5, the sample is classified as belonging to that class.

[0070] The final identification result of the anticancer peptide can be obtained, and the resulting binary identification result is: yes / no.

[0071] Experimental verification

[0072] The ACPs recognition algorithm based on the CNN+Gate model combination provided in the above embodiments is compared with the performance of mainstream existing algorithms. Following standard experimental procedures, the dataset was divided into training and test sets in an 8:2 ratio. The model was trained using the Adam optimizer (learning rate 0.00001, weight decay coefficient 1e-5) and cross-entropy loss function with a batch size of 64, and the number of training epochs (num_epochs) was set to 250. To improve the model's generalization performance, Dropout rates of 0.2 and 0.1 were used in sequence feature extraction and the classifier, respectively.

[0073] Validated on the ACP2alter dataset, the proposed CNN+Gate model combination exhibits good performance. The training loss eventually stabilized at 0.05, and the testing loss converged to 0.15. The final test accuracy reached 95.36%, indicating good generalization performance. In terms of key performance indicators, the specificity improved by 4.03% compared to the state-of-the-art (SOTA) model, reducing the false positive rate from 6.2% to 2.8%. This improvement can reduce unnecessary experimental validation costs in medical scenarios. Simultaneously, the model's actual sensitivity was 93.12%, essentially on par with the SOTA model's 93.05%, maintaining the ability to identify true anticancer peptides. Furthermore, the Matthews correlation coefficient (MCC) increased from 0.86 in the SOTA model to 0.90 in this model, indicating improvements in classification balance and robustness. The model's ROC curve is smooth, and the AUC value reaches 0.98, maintaining a high true positive rate even with a low false positive rate, demonstrating its good recognition ability. The performance characteristics of the above models are attributed to the local convolutional kernel design of CNNs, which helps to capture short peptide sequence patterns, and the gating mechanism, which dynamically fuses sequences with physicochemical features to suppress redundant information. L2 regularization and Dropout strategies are also used to balance model complexity and generalization performance.

[0074] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0075] For the preferred embodiments of the present invention described above, common knowledge such as specific structures and characteristics in the technical solutions are not described in detail; each embodiment is described in a progressive manner, and the technical features involved in each embodiment can be combined with each other without conflicting with each other. The same or similar parts between the embodiments can be referred to each other.

[0076] It should be noted that, for those skilled in the art, the above embodiments can be improved and modified in various ways without departing from the principles of the present invention, and such improvements and modifications should also be considered to fall within the protection scope of the present invention.

Claims

1. A method for identifying anticancer peptides, characterized in that, The algorithm employs ACP sequence feature extraction algorithm and ACP physicochemical feature extraction algorithm, as well as ACP multimodal feature fusion algorithm with dynamic adaptive weights.

2. The identification method according to claim 1, characterized in that, First, prepare the raw peptide series data. Subsequent steps include: Step 1: Collect anticancer peptide data; Step 2: Data preprocessing; Step 3: Perform ACP sequence feature extraction and ACP physicochemical feature extraction; Step 4: Perform ACPs multimodal feature fusion with dynamic adaptive weights; Step 5: Perform ACPs classification calculation based on multilayer perceptron; The final result was the identification of anticancer peptides.

3. The identification method according to any one of claims 1 to 2, characterized in that, The physicochemical features extracted by the ACPs physicochemical feature extraction algorithm include: molecular weight, isoelectric point, GRAVY hydrophobicity value, instability index, extinction coefficient, lipid index, and charge distribution. The physicochemical properties of the anticancer peptides are analyzed by constructing a multidimensional feature vector.

4. The identification method according to any one of claims 1 to 2, characterized in that, The ACPs multimodal feature fusion algorithm with dynamic adaptive weights includes an algorithm based on a gated fusion mechanism.

5. The identification method according to any one of claims 1 to 2, characterized in that, The ACPs multimodal feature fusion algorithm with dynamic adaptive weights maps the feature vectors of sequence modes and physicochemical modes to an isomorphic latent space through linear projection operations, establishing a cross-modal geometric alignment relationship.

6. The identification method according to claim 2, characterized in that, In step 5, the classifier used for the ACPs classification calculation based on the multilayer perceptron adopts a three-layer fully connected neural network architecture, and extracts high-order features step by step through the linear transformation layer.

7. The identification method according to claim 2, characterized in that, The steps for performing the ACPs classification calculation based on the multilayer perceptron include: The network input is first projected to the latent space dimension through the first fully connected layer, and then nonlinearity is introduced through the ReLU activation function and Dropout regularization is applied to suppress overfitting. The second layer performs feature recombination in the latent space, and its output is fused with the original input features after linear transformation. This design retains key feature information through cross-layer connections. The final layer uses the Sigmoid function to constrain the output to the [0,1] probability space, and sets 0.5 as the decision threshold to achieve binary classification.