Polypeptide sequence characterization, retrieval and editing-oriented integrated processing method
By combining the peptide controllable editing model with multimodal coding and Bayesian optimization, efficient alignment and generation of peptide sequences and text descriptions are achieved, solving the retrieval and editing problems of peptide sequences and text descriptions, and improving the property prediction and controllable editing capabilities of peptide sequences.
Patent Information
- Application Number
- CN202511059891.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies lack the ability to align polypeptide sequences with text descriptions in a multimodal manner, making it difficult to achieve high-quality representation, bidirectional retrieval, and text-guided generation. They also lack the ability to generate controllable sequences for complex functional objectives.
A peptide controllable editing model is used, combined with multimodal encoding, contrastive learning and Bayesian optimization, to achieve bidirectional retrieval and editing of peptide sequences and text descriptions. The deep representation is extracted from the target peptide sequence through the trained peptide controllable editing model, and the execution result is obtained based on the deep representation.
It improves the bidirectional retrieval capability between peptide sequences and text descriptions, enhances property prediction performance, and improves the controllable editing and optimization capabilities of peptide sequences.
Smart Images

Figure CN120808893A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of polypeptide sequence processing, and particularly relates to an integrated processing method for polypeptide sequence representation, retrieval and editing. BACKGROUND
[0002] Peptides are short-chain biomolecules composed of amino acid residues connected by peptide bonds, which are the basic building blocks of proteins. They are widely involved in life processes such as signal transduction, immune regulation, and antibacterial defense. Peptides have important applications in drug development, synthetic biology, and tissue engineering. Compared to small molecule drugs and intact proteins, short peptides have the advantages of strong targeting, high biocompatibility, easy synthesis and modification, and are the forefront of research in the field of biological drugs. However, traditional polypeptide function screening and design rely on biochemical experiments, which have the significant disadvantages of long cycle, high cost, and low efficiency, making it difficult to meet the needs of large-scale and customized development. With the rapid development of artificial intelligence, especially deep learning, new ideas have been provided for polypeptide structure modeling, function prediction, and sequence generation.
[0003] Protein structure prediction methods represented by AlphaFold have greatly promoted the development of structural biology, and large-scale protein language models based on Transformer have been widely applied to biological sequence modeling tasks. For example, ESM (Evolutionary Scale Modeling) extracts evolutionary semantics from massive protein sequences through multi-layer Transformer, supporting downstream prediction tasks. ProtBERT, ProtGPT, and other models also extract protein language features through self-supervised learning for sequence generation and classification. For polypeptide sequences, models such as PepBERT make lightweight adjustments in structure and take short sequences as the core of training, showing high performance in predicting properties such as hemolyticity and solubility. However, these methods generally rely on modeling functional information from sequences themselves and lack the ability to understand natural language descriptions, making it difficult to meet the needs of semantic-driven sequence retrieval or editing.
[0004] In the field of computer vision, models such as CLIP (Contrastive Language-Image Pre-training) achieve alignment of image and text representations through contrastive learning, opening up extensive exploration of cross-modal modeling. Preliminary efforts such as DrugCLIP have introduced this idea into the biomedical field, attempting to align small molecule structures with drug action description texts, providing a new paradigm for drug discovery. However, there is currently no systematic study of polypeptides, a special sequence category, to construct a unified sequence-text alignment representation space. Methods such as MoleculeSTM primarily serve the small molecule drug scene and lack adaptation mechanisms tailored to the characteristics of short polypeptide sequences and the need for multifunctional semantic expression. Therefore, there is currently no polypeptide multi-modal system that can achieve high-quality representation alignment, bidirectional retrieval, and text-guided generation.
[0005] A variety of machine learning and deep learning-based polypeptide property prediction methods have been proposed. For example, the CICERON system combines random forests, support vector machines (SVM), and ProtBERT representations to establish a baseline for polypeptide function classification and property prediction. Scholars such as Jiashun Mao use LSTM, Transformer, and other models to model antimicrobial peptides, achieving functional prediction. Although the above methods have achieved good results in specific tasks, they generally have the following problems: (a) the model only uses sequence information, lacking text semantic understanding ability; (b) the prediction results are not interpretable, making it difficult to guide subsequent design; (c) lack of controllable sequence generation capability under "complex functional goals."
[0006] In summary, there is currently no mature framework that can achieve multi-modal alignment and semantic mapping of "polypeptide sequence-text description," and no complete system to support the generation of a closed loop from natural language to structural optimization. SUMMARY
[0007] To solve the above problems in the prior art, the present application provides an integrated processing method for polypeptide sequence representation, retrieval, and editing. The technical problem to be solved by the present application is solved by the following technical scheme: An integrated processing method for polypeptide sequence representation, retrieval, and editing, comprising: S100, receiving a downstream task, a target polypeptide sequence corresponding to the downstream task, and a trained polypeptide controllable editing model corresponding to the downstream task; S200, executing the downstream task, thereby inputting the target polypeptide sequence into the corresponding trained polypeptide controllable editing model, so that the trained polypeptide controllable editing model extracts deep representations from the target polypeptide sequence, and obtains an execution result corresponding to the downstream task according to the deep representations.
[0008] Beneficial effects: The application provides a polypeptide sequence characterization, retrieval and editing integrated processing method, which comprises the following steps: BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 is a flowchart of the polypeptide sequence characterization, retrieval and editing integrated processing method provided by the application; Figure 2 is a structural diagram of the polypeptide controllable editing model provided by the application; Figure 3 is a retrieval effect diagram for a retrieval matching task provided by the application; Figure 4 is a solubility probability prediction diagram of a target polypeptide sequence before and after editing provided by the application. DETAILED DESCRIPTION
[0010] The application will be further described in detail below in combination with specific embodiments, but the embodiments of the application are not limited thereto.
[0011] As shown in the accompanying drawings, Figure 1 The application provides a polypeptide sequence characterization, retrieval and editing integrated processing method, which comprises the following steps: S100, receiving a downstream task, a target polypeptide sequence corresponding to the downstream task and a trained polypeptide controllable editing model corresponding to the downstream task; The downstream task is one or more of a polypeptide sequence and text retrieval matching task, a polypeptide sequence property prediction task and a polypeptide sequence optimization task.
[0012] S200, executing the downstream task, so as to input the target polypeptide sequence into the corresponding trained polypeptide controllable editing model, so that the trained polypeptide controllable editing model extracts deep features from the target polypeptide sequence, and obtains an execution result corresponding to the downstream task according to the deep features.
[0013] Reference Figure 2 , Figure 2 The logic framework diagram for performing downstream tasks of the present application. The internal structure of the polypeptide controllable editing model is related to the downstream tasks. If the downstream task is the retrieval matching task of the polypeptide sequence and the text, the trained polypeptide editable model includes a text encoder, a sequence encoder, a projection layer and a decoder; if the downstream task is the property prediction task of the polypeptide sequence, the trained polypeptide controllable editing model corresponding to the downstream task includes a sequence encoder and a decoder; if the downstream task is the optimization task of the polypeptide sequence, the trained polypeptide editable model includes a text encoder, a sequence encoder, a projection layer and a ProtGPT2 decoder.
[0014] In Figure 2 , for the retrieval matching task, as shown in part b of Figure 2 , the polypeptide sequence and the text description are mapped to a unified semantic space by a sequence encoder ( ) and a text encoder ( ), bidirectional retrieval (sequence to text / text to sequence) is realized by calculating similarity, and matching accuracy is verified in a multi-scale candidate set (4 / 10 / 20). Figure 2 For the property prediction task, as shown in part c of , the trained sequence encoder based on multi-modal alignment ( Figure 2 ) extracts sequence features, and a classifier is used to predict properties, experiments show that its accuracy and robustness exceed traditional models. For the optimization task, as shown in d1 and d2 of , in d1, the target polypeptide sequence is represented by the projection layer of the generation encoder and aligned with the representation of the sequence encoder ; in d2, based on the aligned representation, the optimal code balancing property retention (minimizing distance from the original sequence) and instruction following (maximizing similarity to the edited instruction text) is searched using Bayesian optimization (BO), and the edited sequence is output by the generation decoder
[0015] , and finally the function compliance is verified by the property module. In a specific embodiment of the present application, S200 includes:If the downstream task is a polypeptide sequence and text retrieval matching task, a polypeptide sequence and text retrieval matching task is performed, so that the target polypeptide sequence is input into the sequence encoder to obtain a first sequence representation, and a candidate description is selected from the text database and input into the text encoder to obtain a first text representation; the first sequence representation and the first text representation are projected to the same semantic space through a projection layer, and the similarity between the two in the same semantic space is calculated, the first text representation closest to the first sequence representation is obtained, and the first text representation is decoded through the decoder to obtain the text description closest to the target polypeptide sequence, and the text description is taken as the execution result of the retrieval matching task.
[0016] The present application designs a bidirectional retrieval task of polypeptide sequence and text description. For polypeptide sequence to text retrieval, the target polypeptide sequence is input, candidate descriptions are selected from the text database, the similarity between the first sequence representation and the first text representation is calculated, and the first text representation with the highest similarity is selected as the matching result; for text to polypeptide sequence retrieval, the target text description is input, candidate sequences are selected from the polypeptide sequence database, the similarity between the text representation and the polypeptide representation is calculated, and the polypeptide sequence with the highest similarity is selected. The matching accuracy of the two tasks is counted under different candidate scales. The first sequence representation and the first text representation are extracted by the trained encoder and projected to the shared semantic space, and the specific similarity calculation formula is as follows:
[0017] The projection layer and the encoder from PepCLIP are all frozen. is a negative text sequence from an alternative.
[0018] In a specific embodiment of the present application, S200 includes: If the downstream task is a polypeptide sequence property prediction task, a property prediction task is performed, so that the target polypeptide sequence is input into the sequence encoder to obtain a second sequence representation, and the second sequence representation is input into the encoder to obtain the property of the target polypeptide sequence.
[0019] The polypeptide sequence property prediction task adopts a multi-layer perceptron architecture, and a fully connected layer is stacked on the basis of the sequence representation , each layer containing a linear transformation and a ReLU activation function, and finally outputting a prediction result matched with a specific prediction task.
[0020] In one specific embodiment of the present application, S200 comprises: If the downstream task is an optimization task of a polypeptide sequence, the optimization task is performed, so as to input the target polypeptide sequence into a sequence encoder to obtain a third sequence representation, and input the task requirement of the optimization task into a text encoder in a text manner to obtain a latent code of the text; the third sequence representation and the latent code are projected to the same semantic space through a projection layer, and the optimal latent code is searched in the latent code space through a Bayesian optimization method with the latent code as a constraint condition, and the optimal latent code is decoded through a decoder to obtain an optimized polypeptide sequence; the optimized polypeptide sequence is taken as an execution result of the optimization task.
[0021] The searching of the optimal latent code in the latent code space through the Bayesian optimization method with the latent code as the constraint condition comprises: constructing a constraint target with the latent code, the constraint target comprising minimization of a Euclidean distance from the target polypeptide sequence and maximization of a similarity to the latent code; constructing a target function through the Bayesian optimization method with the constraint target as a constraint condition.
[0022] Given an input polypeptide sequence and a text description (such as “this polypeptide is soluble”), the latent codes of the input polypeptide and the text are first extracted through a sequence encoder and a text encoder respectively; a Bayesian optimization (BO) method is used to search for an optimal code in the latent space. The optimal code needs to satisfy the following two objectives at the same time: 1) minimization of the Euclidean distance from the input polypeptide code (to retain the key characteristics of the original polypeptide to the greatest extent); 2) maximization of the similarity to the text description code (to ensure that the edited result meets the requirements of the text instructions). The BO method efficiently finds the optimal latent point satisfying the balance of the double objectives by constructing a proxy model of the target function and iteratively sampling based on an acquisition function.
[0023] The optimal latent code is searched in the latent code with the maximization of the target function as an objective. The target function is expressed by a formula as follows:
[0024] wherein, is an optimized sequence representation, is a third sequence representation, is a latent code, is a Euclidean distance loss coefficient, is a padding code, is a polypeptide space mapping function, is a meanpooling function, is a Euclidean distance.
[0025] In a specific embodiment of the present application, the training process of the trained polypeptide controllable editing model comprises: In the pre-training stage, the predetermined polypeptide controllable editing model corresponding to the downstream task is pre-trained, and the parameters of the encoder in the predetermined polypeptide controllable editing model are adjusted according to the objective function until the training stopping condition is reached to obtain a pre-trained polypeptide controllable editing model. In the pre-training stage, the present application constructs a comprehensive and high-quality polypeptide-text data set based on the UniProt database. Specifically, polypeptide sequences with a length of 50 or less are collected from UniProt, and for each polypeptide, detailed species source information (including Latin name and taxonomic ID) and experimentally verified biological function annotations (such as antibacterial activity, signal transduction function, enzyme inhibition activity, etc.) are collected. Finally, 400,000 polypeptide-text data sets are obtained, which are used for pre-training and re-training.
[0026] In the pre-training stage, the present application adopts a contrastive learning strategy, and the core optimization objective is InfoNCE
[11] loss function (Noise Contrastive Estimation with Mutual Information). This function maximizes the mutual information of positive sample pairs while minimizing the similarity of negative sample pairs, effectively learning cross-modal alignment representation. The specific objective function is defined as:
[0027] wherein, is the corresponding sequence-text pair, and is a negative sample randomly sampled from noise, and the energy function which has a flexible form, and here adopts the dot product representation in the joint learning space, i.e.:
[0028] In the re-training stage, for downstream tasks other than the optimization task of polypeptide sequences, the parameters of the encoder in the pre-trained polypeptide controllable editing model corresponding to the downstream task are fixed, and the parameters of other modules except the encoder are adjusted according to the corresponding loss function to reach the training stopping condition, obtaining the trained polypeptide controllable editing model; for the downstream task of the optimization task of polypeptide sequences, the parameters of the encoder in the pre-trained polypeptide controllable editing model corresponding to the downstream task are fixed, and the parameters of the projection layer are adjusted according to the spatial alignment loss until the training stopping condition is reached, and then the ProtGPT2 decoder is used to replace the decoder in the polypeptide controllable editing model at the training stopping time, obtaining the trained polypeptide controllable editing model. Wherein, the sequence encoder adopts ESM encoder, and the text encoder adopts BERT encoder.
[0029] If the downstream task is the retrieval matching task of the polypeptide sequence and the text, the corresponding loss function is the similarity; if it is the property prediction task of the polypeptide sequence, the corresponding loss function is the classification loss, and if it is the optimization task of the polypeptide sequence, the projection layer parameters are adjusted according to the alignment loss.
[0030] UniRep and other lightweight CNN frameworks can also be used to replace the sequence encoder, and BioBERT and the like can also be used to replace the text encoder.
[0031] The parameters of the pre-trained text encoder and polypeptide sequence encoder are frozen, and the polypeptide representation output by the encoder is aligned with the representation of the sequence encoder by training a lightweight projection layer (such as a multilayer perception), so as to ensure that the encoding of the same polypeptide in the text, sequence and generation three modalities is located in the same semantic space. The loss function of the space alignment is:
[0032] , wherein, is the trained generation encoder projection layer, is the generation model encoder, is the projection layer and the encoder of ProtGPT2 (generation model).
[0033] In addition to ClipLoss, other losses can also be used to replace the loss function of the space alignment.
[0034] (1) Evaluation of bidirectional retrieval capability between polypeptide sequence and text description To evaluate the performance of the bidirectional retrieval task, the invention performs experiments in the text-to-sequence and sequence-to-text directions respectively. The experimental task is set as follows: given a query item (text or sequence) and a set containing T candidate options, the model needs to retrieve the corresponding item (sequence or text) most similar to it from the set, and calculate the retrieval accuracy. Specifically, in the task of retrieving polypeptide sequences from text descriptions: Direct similarity retrieval: for the given query text , calculate the representation of the query text, and calculate the similarity between the representation and all sequence representations in the candidate pool, and retrieve the sequence with the highest similarity. If the sequence is consistent with the true target sequence, it is considered to be retrieved correctly.
[0035] Fuzzy similarity retrieval: generate the representation of the true target sequence using the PepCLIP model. Calculate the similarity between and all sequence representations in the database, and retrieve the most similar sequence . This sequence defined as the query text "gold standard" sequence. Subsequently, a text to sequence retrieval task is performed, and if the highest similarity sequence retrieved is consistent with , it is considered correct retrieval.
[0036] The experimental results are shown in Figure 3 . In the task of retrieving polypeptide sequences given a text, the accuracy of the PepCLIP model is over 80%; in the task of retrieving polypeptide sequences given a text description, the accuracy is over 95%. These results fully demonstrate that the PepCLIP model has achieved excellent performance on the bidirectional retrieval task.
[0037] (2) Performance evaluation of polypeptide property prediction To evaluate the performance of the PepCLIP model in the peptide property prediction task, comparative experiments were conducted on two types of datasets: physicochemical property and bioactive peptide. The physicochemical property data comes from Peptide Dashboard
[12] , including: non-fouling (Nonfouling, 3200 positive samples / 13535 negative samples), solubility (Solubility, 8785 positive samples / 9668 negative samples), and hemolysis (Hemolysis, 1826 positive samples / 7490 negative samples). The bioactive peptide data uses the dataset used by the CICERON model, which contains 263 antidiabetic polypeptides (Antidiabetic), 240 celiac disease polypeptides (Celiac disease), 540 antimicrobial polypeptides (Antimicrobial), 140 opioid polypeptides (Opioid), 165 neuropeptides (Neuropeptides), and other types of polypeptides, totaling 3998 samples. All data are randomly divided into training set, validation set and test set according to the ratio of 0.8:0.1:0.1, and ensure class balance. In the binary classification task of physicochemical property prediction, the accuracy of PepCLIP and benchmark models such as LSTM, UniRep+Logistic, Onehot+RNN, and PiptideBERT is evaluated, and the results are shown in Table 1: Table 1 Comparison of physicochemical property prediction accuracy
[0038] In the bioactive peptide prediction task, for each target property, the present application sets its corresponding peptide sample as the positive sample, and sets all the remaining peptide samples as the negative sample to construct multiple binary classification tasks, and compares PepCLIP with the CICERON benchmark model in the Matthews Correlation Coefficient (MCC) index, and the results are shown in Table 2: Table 2 Comparison of MCC of bioactive peptide prediction
[0039] The experimental results show that the PepCLIP model has almost better results than the benchmark models in the above physicochemical property prediction and bioactive peptide prediction tasks.
[0040] (3) Evaluation of controllable editing and optimization ability of polypeptides To evaluate the performance of PepCLIP-BO in controllable editing tasks (optimization tasks), the present application carries out experiments taking solubility optimization as an example. The experiment is carried out under the guidance of the text instruction "this polypeptide is soluble", aiming to generate a new sequence that is highly similar to the original polypeptide sequence and significantly improves its solubility. Specifically, the present application applies the Bayesian Optimization (BO) strategy to edit and optimize 100 initial polypeptide sequences, and independently performs 5 rounds of optimization iterations for each sequence. The new sequence generated by optimization is finally evaluated for solubility probability by the aforementioned polypeptide property prediction model. As shown in Figure 4 , the solubility probability distribution of the polypeptide before editing is mainly concentrated around 0.5, as shown in Figure 4 part a, while the solubility probability of the polypeptide after editing is significantly improved, as shown in Figure 4 part b. This result verifies that the PepCLIP-BO model can effectively meet the specified property editing requirements while maintaining the key structural features of the original sequence to the greatest extent.
[0041] It should be noted that the terms "first", "second" in the present application are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0042] The above is further detailed description of the present application in combination with specific preferred embodiments, and cannot be deemed as limitation of the specific implementation of the present application to these descriptions. For those skilled in the art to which the present application belongs, without departing from the concept of the present application, a number of simple deductions or substitutions can be made, and all of them shall be deemed as falling within the protection scope of the present application.
Claims
1. An integrated processing method for polypeptide sequence characterization, retrieval and editing, characterized in that: include: S100, receiving a downstream task, a target polypeptide sequence corresponding to the downstream task, and a trained polypeptide controllable editing model corresponding to the downstream task; S200, execute the downstream task, thereby inputting the target polypeptide sequence into the corresponding trained polypeptide controllable editing model, so that the trained polypeptide controllable editing model extracts a deep representation from the target polypeptide sequence, and obtains an execution result corresponding to the downstream task based on the deep representation.
2. The integrated processing method for polypeptide sequence characterization, retrieval and editing according to claim 1, characterized in that: The downstream task is one or more of a polypeptide sequence and text retrieval and matching task, a polypeptide sequence property prediction task, and a polypeptide sequence optimization task.
3. The integrated processing method for polypeptide sequence characterization, retrieval and editing according to claim 2, characterized in that: If the downstream task is a retrieval and matching task between a peptide sequence and a text, the trained peptide editable model includes a text encoder, a sequence encoder, a projection layer, and a decoder; If the downstream task is a polypeptide sequence property prediction task, the trained polypeptide controllable editing model corresponding to the downstream task includes a sequence encoder and a decoder; If the downstream task is a polypeptide sequence optimization task, the trained polypeptide editable model includes a text encoder, a sequence encoder, a projection layer and a ProtGPT2 decoder.
4. The integrated processing method for polypeptide sequence characterization, retrieval and editing according to claim 3, characterized in that: S200 includes: If the downstream task is a retrieval and matching task of polypeptide sequence and text, the retrieval and matching task of polypeptide sequence and text is performed, so that the target polypeptide sequence is input into the sequence encoder to obtain a first sequence representation, and a candidate description is selected from the text database and input into the text encoder to obtain a first text representation; the first sequence representation and the first text representation are projected into the same semantic space through the projection layer, and the similarity between the two in the same semantic space is calculated to obtain a first text representation closest to the first sequence representation, and the first text representation is decoded by the decoder to obtain a text description closest to the target polypeptide sequence, and the text description is used as the execution result of the retrieval and matching task.
5. The integrated processing method for polypeptide sequence characterization, retrieval and editing according to claim 3, characterized in that: S200 includes: If the downstream task is a property prediction task of a polypeptide sequence, the property prediction task is performed, thereby inputting the target polypeptide sequence into the sequence encoder to obtain a second sequence representation, and passing the second sequence representation through the encoder to obtain the property of the target polypeptide sequence.
6. The integrated processing method for polypeptide sequence characterization, retrieval and editing according to claim 3, characterized in that: S200 includes: If the downstream task is a polypeptide sequence optimization task, the optimization task is performed, so that the target polypeptide sequence is input into the sequence encoder to obtain a third sequence representation, and the task requirements of the optimization task are input into the text encoder in text form to obtain a latent code of the text; the third sequence representation and the latent code are projected into the same semantic space through the projection layer, and with the latent code as a constraint condition, the optimal latent code is searched in the latent code space through the Bayesian optimization method, and the optimal latent code is decoded by the decoder to obtain an optimized polypeptide sequence; the optimized polypeptide sequence is used as the execution result of the optimization task.
7. The integrated processing method for polypeptide sequence characterization, retrieval and editing according to claim 6, characterized in that: The searching for the optimal latent code in the latent code space by the Bayesian optimization method with the latent code as a constraint condition includes: Constructing a constraint target with the potential code, wherein the constraint target includes minimizing the Euclidean distance with the target polypeptide sequence and maximizing the similarity with the potential code; Taking the constraint target as a constraint condition, constructing an objective function through a Bayesian optimization method; An optimal latent code is searched among the latent codes with the goal of maximizing the objective function.
8. The integrated processing method for polypeptide sequence characterization, retrieval and editing according to claim 7, characterized in that: The objective function is expressed as follows: in, For the optimized sequence representation, For the third order representation, is the latent code, is the Euclidean distance loss coefficient, Encode for padding, Peptide space mapping function, is the mean pooling function, is the Euclidean distance.
9. The integrated processing method for polypeptide sequence characterization, retrieval and editing according to claim 3, characterized in that: The training process of the trained polypeptide controllable editing model includes: In the pre-training stage, a predetermined peptide controllable editing model corresponding to the downstream task is pre-trained, and the parameters of the encoder in the predetermined peptide controllable editing model are adjusted according to the objective function until the training cutoff condition is reached to obtain the pre-trained peptide controllable editing model; In the retraining stage, for downstream tasks other than the peptide sequence optimization task, the parameters of the encoder in the pre-trained peptide controllable editing model corresponding to the downstream task are fixed, and then the parameters of other modules except the encoder are adjusted according to the corresponding loss function to reach the training cutoff condition, thereby obtaining a trained peptide controllable editing model; for downstream tasks of the peptide sequence optimization task, the parameters of the encoder in the pre-trained peptide controllable editing model corresponding to the downstream task are fixed, and then the parameters of the projection layer are adjusted according to the spatial alignment loss until the training cutoff condition is reached, and then the decoder in the peptide controllable editing model at the training cutoff is replaced with the ProtGPT2 decoder to obtain a trained peptide controllable editing model.
10. The integrated processing method for polypeptide sequence characterization, retrieval and editing according to claim 3, characterized in that: The sequence encoder adopts the ESM encoder, and the text encoder adopts the BERT encoder.