Novel method for predicting therapeutic peptide by multi-kernel fuzzy system based on deep stacked encoder
Through the multi-core fuzzy system of deep stack encoder, combined with pre-trained language model and BiLSTM, the nonlinear correlation modeling problem in therapeutic peptide detection is solved, efficient and accurate peptide sequence detection is achieved, the number of fuzzy rules is reduced and the detection effect is improved.
Patent Information
- Application Number
- CN202510695184.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing therapeutic peptide detection methods lack the application of fuzzy systems, and traditional methods are time-consuming and labor-intensive. Although machine learning methods have been improved, they still need to be improved, especially in the nonlinear correlation modeling of peptide sequences.
A multi-core fuzzy system based on deep stack encoder is adopted, combining a pre-trained protein language model with a bidirectional long and short-term memory encoder, through hierarchical feature extraction and cross-modal attention fusion, a multi-core fuzzy inference module is constructed, fuzzing and kernel mapping is performed, and parameters are optimized using end-to-end training strategy.
It realizes efficient therapeutic peptide detection, improves detection accuracy, reduces the number of fuzzy rules, enhances nonlinear classification capabilities, and maintains the interpretability of the model.
Smart Images

Figure CN120260693A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer bioinformatics, and particularly relates to a new method for predicting therapeutic peptides based on a multi-core fuzzy system with a deep stacked encoder. Background Art
[0002] Currently, the prediction methods for therapeutic peptide sequences mainly include traditional methods (mass spectrometry-based techniques, bioinformatics techniques, cell experiments, etc.) and machine learning-based methods. Traditional biochemical experimental methods are time-consuming and laborious. In recent years, machine learning techniques have provided new solutions for this. Representative studies include: iACP-GAEnsC: constructing an ensemble model through genetic algorithms, combining random forests with polypeptide physicochemical properties to predict anti-cancer peptides; AntiAngioPred: an anti-angiogenic peptide prediction web server; ACPred-FL: a feature learning method based on word2vec sequence representation and convolutional networks; LSTM anti-cancer peptide predictor, an RF-based multi-feature CPP prediction model; MIMML: a meta-learning bioactive peptide recognition framework based on mutual information maximization; CNN-BiLSTM hybrid architecture and propensity score representation learning PSRQSP; However, there is still a lack of application of fuzzy systems in the field of therapeutic polypeptide detection. To supplement these methods, computational model methods based on machine learning have become effective tools for improving the detection of therapeutic peptides.
[0003] To solve the above problems, the present application proposes a new method for predicting therapeutic peptides based on a multi-core fuzzy system with a deep stacked encoder, which integrates a pre-trained protein language model with a stacked bidirectional long short-term memory (BiLSTM) encoder. This hybrid architecture can hierarchically extract global context patterns (through the language model) and local sequential dependencies (through BiLSTM), effectively modeling the non-linear correlations in peptide sequences. Summary of the Invention
[0004] The object of the present invention is to provide a new method for predicting therapeutic peptides based on a multi-core fuzzy system with a deep stacked encoder, which is reasonably designed and has good therapeutic peptide detection effect, aiming at the above problems.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A new method for predicting therapeutic peptides based on a multi-core fuzzy system with a deep stacked encoder, comprising the following steps: S1: Data preparation and preprocessing, performing pre-trained language model feature extraction and sequence tokenization and encoding; S2: Feature extraction module design, performing hierarchical feature extraction and cross-modal attention fusion; S3: Construction of a multi-core fuzzy inference module, performing fuzzification and kernel mapping, fuzzy rule generation and inference, and defuzzification output; S4: Joint training and optimization, adopting an end-to-end training strategy and joint parameter optimization; S5: Experiment and verification; S6: System deployment and application.
[0006] In the above new method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system, step S1 includes the following steps: S11: Input the amino acid sequence of the therapeutic polypeptide; S12: Use a pre-trained protein language model to perform embedding representation on the amino acid sequence, generate global semantic features, and output a high-dimensional semantic vector; S13: Divide the sequence into subsequences of a fixed length and convert them into a numerical representation.
[0007] In the above new method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system, the hierarchical feature extraction in step S2 includes the following steps: S21: The bidirectional LSTM layer captures local sequence dynamic features; S22: Stack multiple-level BiLSTM layers to extract high-order abstract features layer by layer.
[0008] In the above new method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system, the cross-modal attention fusion in step S2 includes the following steps: S23: Through the attention mechanism, fuse the global semantic features output by the pre-trained language model and the local sequence features output by the BiLSTM to generate a 128-dimensional fusion feature vector; S24: Dynamically adjust the contribution ratio of global and local features by attention weights.
[0009] In the above new method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system, the fuzzification and kernel mapping in step S3 include the following steps: S31: Input the 128-dimensional feature vector into the fuzzy subsystem, and through Gaussian membership function fuzzification processing, generate fuzzy membership degrees; S32: Use multi-core combination to map the fuzzified features to the reproducing kernel Hilbert space.
[0010] In the above new method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system, the fuzzy rule generation and inference in step S3 include the following steps: S33: Define the logical relationship between the antecedent and the consequent based on a small number of fuzzy rules; S34: Calculate the inner product through the kernel trick in the Hilbert space, implicitly construct a non-linear classification boundary, and distinguish therapeutic peptides from non-therapeutic peptides.
[0011] In the above-mentioned novel method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system, the defuzzification output in step S3 includes the following steps: S35: Defuzzify the fuzzy inference result through weighted average and output the final classification probability.
[0012] In the above-mentioned novel method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system, step S4 includes the following steps: S41: Adopt the cross-entropy loss function to jointly optimize the parameters of BiLSTM, the consequent parameters of fuzzy rules, and the parameters of kernel functions; S42: Dynamically adjust the fusion weight of global semantic features and local sequence features through gradient backpropagation.
[0013] In the above-mentioned novel method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system, step S5 includes the following steps: S51: Conduct a performance comparison experiment; S52: Interpretability analysis; S53: Hyperparameter tuning.
[0014] In the above-mentioned novel method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system, step S6 includes the following steps: S61: Model lightweighting, pruning redundant fuzzy rules, quantifying model parameters, and adapting to edge computing devices; S62: Prediction service interface, encapsulating the model as an API, and returning the probability of therapeutic polypeptides and the interpretation of key features when inputting the amino acid sequence.
[0015] Compared with the existing technologies, the advantages of the present invention are as follows: dual capture of local sequence dynamics and global semantic context is achieved through a stacked BiLSTM architecture, the transfer learning advantage of a pre-trained language model is fused, multi-granularity semantic representations of amino acid sequences are established, and cross-modal attention gates achieve hierarchical fusion of complementary features, effectively solving the problem of long-range dependence modeling of peptide sequences; the multi-core fuzzy system breaks through the linear separability limitation of traditional fuzzy models, maps 128-dimensional features to the RKHS space through a kernel method, and uses kernel tricks to achieve implicit calculation in a high-dimensional space, avoiding dimensional explosion, and only requiring a small number of fuzzy rules to achieve SOTA performance; the end-to-end training framework realizes the joint optimization of the parameters of the BiLSTM encoder and the fuzzy system, synchronously learns the Gaussian kernel parameter γ and the consequent parameters of fuzzy rules, constructs a closed-loop feedback between deep features and fuzzy inference, and an adaptive weight allocation mechanism dynamically balances the contribution degrees of local sequence features and global semantic information. Description of the Drawings
[0016] Figure 1 It is the implementation flowchart of the method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system of the present invention; Figure 2 It is the schematic diagram of the prediction method of the multi-core fuzzy system based on the deep stacked encoder of the present invention; Figure 3 It is the performance of the model under different parameters of the method of the present invention. Specific implementation manners
[0017] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0018] As Figures 1-3 shown, a new method for predicting therapeutic peptides based on a multi-core fuzzy system with a deep stacked encoder includes the following steps: S1: Data preparation and preprocessing, performing pre-trained language model feature extraction and sequence tokenization and encoding; S2: Feature extraction module design, performing hierarchical feature extraction and cross-modal attention fusion; the feature extraction module realizes the deep feature extraction of therapeutic polypeptides through a pre-trained language model, adopts a stacked BiLSTM architecture to enhance the sequence context modeling ability, integrates the transfer learning advantages of the pre-trained language model, and realizes the multi-level semantic representation of amino acid sequences; S3: Construction of a multi-core fuzzy inference module, performing fuzzification and kernel mapping, fuzzy rule generation and inference, and defuzzification output; the extracted features are input into the multi-core fuzzy system, which has three major technical advantages, maps high-dimensional features to the reproducing kernel Hilbert space (RKHS) through the kernel method, and improves the non-linear classification ability while maintaining the interpretability of the fuzzy system, and only requires a small number of fuzzy rules to achieve SOTA classification performance (experiments show that the accuracy is improved by 5.2% when the number of rules is reduced by 40%); S4: Joint training and optimization, adopting an end-to-end training strategy and parameter joint optimization; the two modules are deeply coupled through end-to-end training, and the 128-dimensional feature vector output by the amino acid feature extraction module is used as the input of the multi-core fuzzy system. The Gaussian kernel parameter γ and the fuzzy rule consequent parameters are jointly optimized, and adaptive weight allocation is used to balance local sequence features and global semantic information.
[0019] S5: Experiment and verification; S6: System deployment and application.
[0020] Specifically, step S1 includes the following steps: S11: Input the amino acid sequence of the therapeutic polypeptide; S12: Use a pre-trained protein language model (such as ProtBERT, ESM, etc.) to perform embedding representation on the amino acid sequence, generate global semantic features (such as residue interaction patterns, long-range dependencies), and the output is a high-dimensional semantic vector (for example, 1024 dimensions); S13: Split the sequence into subsequences of a fixed length (such as a sliding window) and convert them into a numerical representation (such as one-hot encoding or physicochemical property encoding).
[0021] Specifically, the hierarchical feature extraction in step S2 includes the following steps: S21: A bidirectional LSTM layer captures local sequence dynamic features (such as physicochemical properties, short-range folding patterns); S22: Multiple-level BiLSTM layers are stacked to extract high-order abstract features layer by layer (such as residue co-action, secondary structure tendency).
[0022] Furthermore, the cross-modal attention fusion in step S2 includes the following steps: S23: Through the attention mechanism, fuse the global semantic features output by the pre-trained language model and the local sequence features output by the BiLSTM to generate a 128-dimensional fusion feature vector; S24: The attention weights dynamically adjust the contribution ratio of global and local features.
[0023] In summary, the architecture of step S2 utilizes the ability of BiLSTM to encode global semantic context (e.g., residue interaction motifs), while the stacked BiLSTM layers capture local sequential dynamics, such as the physicochemical properties of residues and short-range folding patterns. By hierarchically fusing these complementary features through cross-modal attention gates, the inherent non-linear correlations in peptide sequences are addressed, including long-range dependencies that are crucial for stability and binding affinity.
[0024] Even further, the fuzzification and kernel mapping in step S3 include the following steps: S31: Input the 128-dimensional feature vector into the fuzzy subsystem and perform fuzzification processing through a Gaussian membership function to generate fuzzy membership degrees (such as the membership degree values of each feature); S32: Use a multi-kernel combination (such as Gaussian kernel, polynomial kernel) to map the fuzzified features to a reproducing kernel Hilbert space (RKHS) to solve the problem of linear inseparability in low-dimensional spaces.
[0025] In addition, the fuzzy rule generation and inference in step S3 include the following steps: S33: Define the logical relationship between the antecedent (feature membership degree) and the consequent (classification weight) based on a small number of fuzzy rules (such as 5 - 10 rules); S34: Calculate the inner product through the kernel trick in the Hilbert space to implicitly construct a non-linear classification boundary to distinguish therapeutic peptides from non-therapeutic peptides.
[0026] Meanwhile, the defuzzification output in step S3 includes the following steps: S35: Defuzzify the fuzzy inference result through weighted average and output the final classification probability.
[0027] Step S3 is a fuzzy system with a multi-core structure based on fuzzy rules, which integrates fuzzy logic and kernel feature mapping to enhance the identification of therapeutic polypeptides. The kernel function enables the fuzzy system to capture the complex non-linear feature relationships in the therapeutic peptide sequence and map the low-dimensional input to a high-dimensional Reproducing Kernel Hilbert Space (RKHS), breaking the traditional fuzzy system that relies on linearly separable data. The reproducing kernel trick allows the fuzzy system to directly calculate the inner product in the implicit high-dimensional space, avoiding the complexity explosion caused by the explicit high-dimensional mapping calculation.
[0028] Visibly, step S4 includes the following steps: S41: Adopt the cross-entropy loss function to jointly optimize the parameters of BiLSTM, the consequent parameters of fuzzy rules, and the parameters of the kernel function (such as γ of the Gaussian kernel); S42: Dynamically adjust the fusion weight of global semantic features and local sequence features through gradient backpropagation to improve the robustness of the model.
[0029] Obviously, step S5 includes the following steps: S51: Conduct performance comparison experiments and compare with SOTA models (such as CNN, RNN, traditional fuzzy systems) in terms of metrics such as accuracy and recall; S52: Interpretability analysis, visualize the fuzzy rule activation patterns, analyze key features (such as specific amino acid residues or physicochemical properties), count the rule usage frequency, and verify the hypothesis of "efficient classification with a small number of rules"; S53: Hyperparameter tuning, adjust the number of BiLSTM layers, attention dimension, kernel function type / number to optimize the model performance.
[0030] The experimental results show that the model of this method has good performance. The best accuracies for CPP, QSP, ACE inhibitory activity, DPPⅳ inhibitory activity, bitter peptides, and umami peptides are 95.7%, 97.5%, 90.1%, 85.7%, 94.5%, and 89.8% respectively.
[0031] Preferably, step S6 includes the following steps: S61: Model lightweighting, pruning redundant fuzzy rules, quantifying model parameters, and adapting to edge computing devices; S62: Prediction service interface, encapsulate the model as an API, and input the amino acid sequence to return the probability of therapeutic polypeptides and the interpretation of key features.
[0032] In summary, the principle of this embodiment is as follows: Stack BiLSTM + pre-trained model + attention mechanism, taking into account both local and global features, multi-core mapping to enhance non-linear classification ability, while reducing the number of rules by 40%, jointly optimizing the language model, sequence dynamics modeling, and fuzzy logic, breaking through the limitations of traditional module fragmentation.
[0033] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains may make various modifications or supplements to the described specific embodiments or use similar means for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
[0034] Although terms such as feature extraction and cross-modal attention fusion are used more frequently herein, the possibility of using other terms is not excluded. The use of these terms is only for more conveniently describing and explaining the essence of the present invention; interpreting them as any additional limitation is contrary to the spirit of the present invention.
Claims
1. A new method for predicting therapeutic peptides based on a deep stacked encoder multi-core fuzzy system, characterized in that, It includes the following steps: S1: Data preparation and preprocessing, performing pre-trained language model feature extraction and sequence tokenization and encoding; S2: Feature extraction module design, performing hierarchical feature extraction and cross-modal attention fusion; S3: Construction of a multi-core fuzzy inference module, performing fuzzification and kernel mapping, fuzzy rule generation and inference, and defuzzification output; S4: Joint training and optimization, adopting an end-to-end training strategy and parameter joint optimization; S5: Experiment and verification; S6: System deployment and application.
2. A new method for predicting therapeutic peptides based on a multi-core fuzzy system using a deep stacked encoder according to claim 1, characterized in that, The step S1 includes the following steps: S11: Input the amino acid sequence of the therapeutic polypeptide; S12: Use a pre-trained protein language model to perform embedding representation on the amino acid sequence, generate global semantic features, and output a high-dimensional semantic vector; S13: Divide the sequence into subsequences of a fixed length and convert them into a numerical representation.
3. A new method for predicting therapeutic peptides based on a multi-core fuzzy system with a deep stacked encoder according to claim 1, characterized in that, The hierarchical feature extraction in the step S2 includes the following steps: S21: The bidirectional LSTM layer captures local sequence dynamic features; S22: Stack multiple levels of BiLSTM layers to extract high-order abstract features layer by layer.
4. A new method for predicting therapeutic peptides by a multi-core fuzzy system based on a deep stacked encoder according to claim 3, characterized in that, The cross-modal attention fusion in the step S2 includes the following steps: S23: Through the attention mechanism, fuse the global semantic features output by the pre-trained language model and the local sequence features output by BiLSTM to generate a 128-dimensional fusion feature vector; S24: Dynamically adjust the contribution ratio of global and local features by attention weights.
5. A new method for predicting therapeutic peptides in a multi-core fuzzy system based on a deep stacked encoder according to claim 1, characterized in that, The fuzzification and kernel mapping in the step S3 includes the following steps: S31: Input the 128-dimensional feature vector into the fuzzy subsystem, perform fuzzification processing through the Gaussian membership function, and generate fuzzy membership degrees; S32: Use multi-core combination to map the fuzzified features to the reproducing kernel Hilbert space.
6. A new method for predicting therapeutic peptides based on a multi-core fuzzy system of a deep stacked encoder according to claim 5, characterized in that, The fuzzy rule generation and inference in the step S3 includes the following steps: S33: Define the logical relationship between the antecedent and the consequent based on a small number of fuzzy rules; S34: Calculate the inner product through the kernel trick in the Hilbert space, implicitly construct a non-linear classification boundary to distinguish therapeutic peptides from non-therapeutic peptides.
7. A new method for predicting therapeutic peptides in a multi-core fuzzy system based on a deep stacked encoder according to claim 6, characterized in that The defuzzification output in the step S3 includes the following steps: S35: Perform defuzzification on the fuzzy inference result through weighted average and output the final classification probability.
8. A novel method for predicting therapeutic peptides by a multi-core fuzzy system based on a deep stacked encoder according to claim 1, characterized in that, The step S4 includes the following steps: S41: Adopt the cross-entropy loss function to jointly optimize the BiLSTM parameters, the consequent parameters of the fuzzy rules, and the kernel function parameters; S42: Dynamically adjust the fusion weights of global semantic features and local sequence features through gradient backpropagation.
9. A new method for predicting therapeutic peptides by a multi-core fuzzy system based on a deep stacked encoder according to claim 1, characterized in that, The step S5 includes the following steps: S51: Conduct performance comparison experiments; S52: Interpretability analysis; S53: Hyperparameter tuning.
10. A new method for predicting therapeutic peptides by a multi-core fuzzy system based on a deep stacked encoder according to claim 1, characterized in that, The step S6 includes the following steps: S61: Model lightweighting, pruning redundant fuzzy rules, quantifying model parameters, and adapting to edge computing devices; S62: Prediction service interface, encapsulate the model as an API, and input the amino acid sequence to return the probability of therapeutic polypeptides and the interpretation of key features.
Citation Information
Patent Citations
Antibacterial peptide prediction method and device based on protein pre-training representation learning
CN112614538A
Specific biological sequence prediction method and system based on sequence homology
CN117953973A
Biological sequence prediction method and system based on correlation entropy kernel sparse representation model
CN118609644A
Peptide-HLA class I allele affinity prediction system based on protein language model and dual multi-instance attention mechanism
CN119132408A
Anticancer peptide prediction method and system based on feature fusion and cross attention mechanism
CN120015123A