Protein language model alignment method and device based on RLAIF

By combining reinforcement learning technology and protein language model, the protein sequence generation process is optimized, and the application challenges of protein language model in the protein field are solved, more efficient protein design and prediction effects are achieved, and research in related fields is promoted.

CN120072066APending Publication Date: 2025-05-30INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510147280.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing protein language model has problems such as training data limitations, difficulty in processing long-distance dependencies, risk of data leakage, insufficient interpretability and high demand for computing resources when generating protein sequences, and it is difficult to effectively apply in the protein field.

Method used

Combining reinforcement learning technology and protein language model, by obtaining the target task data set, inputting pre-trained models, evaluating the stability of protein folding state, and adjusting model parameters, the reinforcement learning method is used to optimize the protein language model, define the protein sequence folding alignment standards, and iteratively optimize using the reward values of the scoring model and reference model.

Benefits of technology

It improves the performance of protein language models in different generation tasks, reduces the cost of protein design, improves the accuracy and effectiveness of protein sequence prediction, and promotes research and development in the fields of protein design and drug discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072066A_ABST
    Figure CN120072066A_ABST
Patent Text Reader

Abstract

The invention discloses a protein language model alignment method and device based on RLAIF. The training method comprises the following steps: acquiring a target task for predicting an amino acid sequence of protein; based on the target task, generating a first input data set corresponding to the target task; inputting the first input data set into a pre-trained RLAIF-based protein language model to obtain a first output data set; inputting the first output data set into a preset protein structure prediction model to obtain a second output data set; according to the second output data set, determining an optimal evaluation model for quantifying the stability of the protein folding state as a scoring model; inputting the first output data set into a scoring model to obtain corresponding reward value output data; and according to the reward value output data, adjusting each parameter of the pre-trained protein language model based on the RLAIF to obtain a trained protein language model based on the RLAIF.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the fields of artificial intelligence and biotechnology, and more specifically, to a method and apparatus for aligning a protein language model based on RLAIF. Background Art

[0002] Proteins are the basic executors of life activities. Traditional wet experimental methods for studying and analyzing proteins have obvious disadvantages, such as being time-consuming and laborious. In related fields, with the continuous development of deep learning technology, reinforcement learning, as a method for an agent to learn the optimal behavior strategy by interacting with the environment, has brought a significant improvement in the performance of large models in the natural language field. However, it is very challenging to transfer this paradigm from the natural language field to the protein field to obtain an effective protein language model. For example, there is a problem of difficulty in defining the alignment standard in the protein field.

[0003] In addition, the existing protein language models have the following deficiencies when used to generate protein sequences: for example, limitations of training data, difficulty in dealing with long-distance dependencies, risk of data leakage, lack of interpretability, high computational resource requirements, data bias, etc.

[0004] Therefore, in order to solve the problems existing in related fields, an improved protein language model is needed. Summary of the Invention

[0005] Embodiments of the present disclosure provide a training method and apparatus for a protein language model based on RLAIF and a prediction method and apparatus for protein data. By combining reinforcement learning technology and a protein language model, the performance of the protein language model in different generation tasks is improved, and the cost of using the protein language model to assist protein design is reduced.

[0006] In one general aspect, a training method for a protein language model based on RLAIF is provided. The training method includes: obtaining a target task for predicting the amino acid sequence of a protein; based on the target task, generating a first input data set corresponding to the target task, where the first input data set includes task description data related to the target task; inputting the first input data set into a pre-trained protein language model based on RLAIF to obtain a first output data set, where the first output data set includes the predicted amino acid sequence of the protein; inputting the first output data set into a preset protein structure prediction model to obtain a second output data set for indicating the folding result of the protein sequence; determining an optimal evaluation model for quantifying the stability of the protein folding state according to the second output data set as a scoring model; inputting the first output data set into the scoring model to obtain corresponding reward value output data; and adjusting each parameter of the pre-trained protein language model based on RLAIF according to the reward value output data to obtain a trained protein language model based on RLAIF.

[0007] Optionally, the pre-trained protein language model based on RLAIF can be obtained by the following method: obtaining a preset protein language model; using the first input data set to adjust each parameter of the preset protein language model through the gradient backpropagation algorithm to obtain an initial actor model as the pre-trained protein language model based on RLAIF.

[0008] Optionally, the step of adjusting each parameter of the pre-trained protein language model based on RLAIF according to the reward value output data to obtain a trained protein language model based on RLAIF may include: calculating an optimization objective value for adjusting each parameter according to the reward value output data through the following formula:

[0009]

[0010] where opt represents the optimization objective value, x represents the first input data set, V(x) represents the reward value of the prediction result of the trained protein language model based on RLAIF evaluated by the scoring model for x, β represents a preset parameter value, V ref (x) represents the reward value of the prediction result of the pre-trained protein language model based on RLAIF evaluated by the scoring model for x, pθ represents the probability distribution of the trained protein language model based on RLAIF, p ref represents the probability distribution of the pre-trained protein language model based on RLAIF, t represents the number of predicted amino acid sequences, a tDenotes the process of predicting the amino acid sequence of the next protein based on the amino acid sequence of the current protein, s t Denotes the state of the number of amino acid sequences of the proteins that already exist currently; according to the optimization target value, the corresponding loss value is calculated through the following formula: loss = -opt, where loss represents the loss value; based on the loss value, the various parameters are adjusted to obtain the trained protein language model based on RLAIF.

[0011] Optionally, after the step of obtaining the target task for predicting the amino acid sequence of the protein, the training method may further include: based on the target task, determining the target task input format and the target task output format, and the step of generating the first input data set corresponding to the target task based on the target task may include: based on the target task, generating the first input data set with the target task input format as the first input data set corresponding to the target task.

[0012] Optionally, the step of determining the optimal evaluation model for quantifying the stability of the protein folding state as the scoring model according to the second output data set may include: determining a plurality of evaluation models corresponding to a plurality of performance indicators for quantifying the stability of the protein folding state according to the second output data set; performing a weighted sum on the plurality of evaluation models according to the importance degrees of the plurality of performance indicators to obtain the optimal evaluation model as the scoring model.

[0013] In another general aspect, a prediction method for protein data is provided, the prediction method including: obtaining a target task for predicting the amino acid sequence of a protein; inputting the target task into a protein language model based on RLAIF to obtain the predicted amino acid sequence of the protein as the prediction result of the protein data, where the protein language model based on RLAIF is trained by using the training method as described above.

[0014] In another general aspect, a training device for a protein language model based on RLAIF is provided. The training device includes: a task acquisition module configured to acquire a target task for predicting an amino acid sequence of a protein; a data construction module configured to generate a first input data set corresponding to the target task based on the target task, the first input data set including task description data related to the target task; a prediction module configured to input the first input data set into a pre-trained protein language model based on RLAIF to obtain a first output data set, the first output data set including the predicted amino acid sequence of the protein; a reward value determination module configured to input the first output data set into a preset protein structure prediction model to obtain a second output data set for indicating the folding result of the protein sequence, determine an optimal evaluation model for quantifying the stability of the protein folding state as a scoring model according to the second output data set, and input the first output data set into the scoring model to obtain corresponding reward value output data; and an optimization module configured to adjust each parameter of the pre-trained protein language model based on RLAIF according to the reward value output data to obtain a trained protein language model based on RLAIF.

[0015] Optionally, the pre-trained protein language model based on RLAIF can be obtained by the following method: acquiring a preset protein language model; using the first input data set to adjust each parameter of the preset protein language model through the gradient backpropagation algorithm to obtain an initial actor model as the pre-trained protein language model based on RLAIF.

[0016] Optionally, the operation of the optimization module adjusting each parameter of the pre-trained protein language model based on RLAIF according to the reward value output data to obtain a trained protein language model based on RLAIF may include: calculating an optimization target value for adjusting each parameter through the following formula according to the reward value output data:

[0017]

[0018] where opt represents the optimization target value, x represents the first input data set, V(x) represents the reward value of the prediction result of the trained protein language model based on RLAIF for x evaluated by the scoring model, β represents a preset parameter value, V ref (x) represents the reward value of the prediction result of the pre-trained protein language model based on RLAIF for x evaluated by the scoring model, pθ represents the probability distribution of the trained protein language model based on RLAIF, p refrepresents the probability distribution of the pre-trained RLAIF-based protein language model, t represents the number of predicted amino acid sequences, and a t represents the process of predicting the amino acid sequence of the next protein based on the amino acid sequence of the current protein, s t represents the state of the number of amino acid sequences of the currently existing protein; according to the optimization target value, the corresponding loss value is calculated by the following formula: loss = -opt, where loss represents the loss value; based on the loss value, the respective parameters are adjusted to obtain the trained RLAIF-based protein language model.

[0019] Optionally, after the operation of obtaining the target task for predicting the amino acid sequence of the protein, the data construction module is further configured to: based on the target task, determine the target task input format and the target task output format, and the operation of the data construction module generating the first input data set corresponding to the target task based on the target task may include: generating the first input data set with the target task input format based on the target task as the first input data set corresponding to the target task.

[0020] Optionally, the operation of the reward value determination module determining the optimal evaluation model for quantifying the stability of the protein folding state based on the second output data set as the scoring model may include: determining multiple evaluation models corresponding to multiple performance indicators for quantifying the stability of the protein folding state based on the second output data set; performing weighted summation on the multiple evaluation models according to the importance degree of the multiple performance indicators to obtain the optimal evaluation model as the scoring model.

[0021] In another general aspect, a protein data prediction device is provided, where the prediction device includes: an acquisition module configured to acquire a target task for predicting the amino acid sequence of a protein; a prediction module configured to input the target task into an RLAIF-based protein language model to obtain the predicted amino acid sequence of the protein as the prediction result of the protein data, where the RLAIF-based protein language model is trained by the training method described above.

[0022] In another general aspect, a computer program product is provided, where the computer program product includes computer programs / instructions that, when executed by a processor, implement the training method of the RLAIF-based protein language model described above and the prediction method of protein data described above.

[0023] In another general aspect, there is provided a computer-readable storage medium, which, when the instructions stored therein are executed by a processor of an electronic device / server, enables the electronic device / server to execute the above-described training method of the RLAIF-based protein language model and the above-described prediction method of protein data.

[0024] In another general aspect, there is provided a computing device, including: at least one processor; at least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the above-described training method of the RLAIF-based protein language model and the above-described prediction method of protein data.

[0025] According to the training method and apparatus of the RLAIF-based protein language model and the prediction method and apparatus of protein data according to embodiments of the present disclosure, by combining reinforcement learning techniques and protein language models, the performance of the protein language model in different generation tasks is improved, and the cost of using the protein language model to assist in protein design is reduced. In addition, the training method and apparatus of the RLAIF-based protein language model proposed by the present disclosure contribute to promoting research and development in the fields of protein design, drug discovery, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and other objects and features of the embodiments of the present disclosure will become clearer through the following description with reference to the drawings showing embodiments, where:

[0027] Figure 1 is a flowchart showing the training method of the RLAIF-based protein language model according to an embodiment of the present disclosure;

[0028] Figure 2 is an exemplary flowchart showing the training method of the RLAIF-based protein language model according to an embodiment of the present disclosure;

[0029] Figure 3 is a flowchart showing the prediction method of protein data according to an embodiment of the present disclosure;

[0030] Figure 4 is a block diagram showing the training apparatus of the RLAIF-based protein language model according to an embodiment of the present disclosure;

[0031] Figure 5 is a block diagram showing the prediction apparatus of protein data according to an embodiment of the present disclosure;

[0032] Figure 6 is a block diagram showing the computing device according to an embodiment of the present disclosure. Detailed implementation manners

[0033] The following detailed implementation manners are provided to help readers obtain a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent. For example, the order of operations described herein is merely an example and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, descriptions of features known in the art may be omitted for greater clarity and conciseness.

[0034] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings, wherein like reference numerals always refer to like elements. The following embodiments will be described with reference to the accompanying drawings to explain the present disclosure.

[0035] As described above, in the related art, the process of migrating the natural language field to the protein field is full of challenges (such as difficulty in defining alignment standards in the protein field). To achieve the above migration process, protein language models have become a new tool and an important research direction in the fields of bioinformatics and artificial intelligence. Protein language models draw on the language model technology in natural language processing. By treating protein sequences as a kind of "language", deep learning models are used to capture patterns and regularities in protein sequences.

[0036] However, due to the disadvantages and deficiencies of existing protein language models as described above, it is very necessary but difficult to propose a protein language model alignment method combining reinforcement techniques.

[0037] To solve the above and / or other problems, the present disclosure proposes a training method and apparatus for a protein language model based on RLAIF (i.e., reinforcement learning based on AI feedback) and a prediction method and apparatus for protein data. By combining reinforcement learning technology and protein language models, the performance of protein language models in different generation tasks is improved, and the cost of protein language model-assisted protein design is reduced.

[0038] The following refers to Figures 1 to 6 Describe in detail a training method and apparatus for a protein language model based on RLAIF and a prediction method and apparatus for protein data according to embodiments of the present disclosure.

[0039] First, refer to Figures 1 to 2 Describe a training method for a protein language model based on RLAIF according to embodiments of the present disclosure.

[0040] Figure 1It is a flowchart showing a training method 100 of an RLAIF-based protein language model according to an embodiment of the present disclosure. Figure 2 It is an example flowchart showing a training method of an RLAIF-based protein language model according to an embodiment of the present disclosure.

[0041] Referring to Figure 1 , according to an embodiment of the present disclosure, in step S101, a target task for predicting the amino acid sequence of a protein is obtained. For example, the target task may be antibody design (e.g., the corresponding target task description may be to obtain the amino acid sequence of an antibody from an antigen), bending of a protein, etc., but the present disclosure is not limited thereto.

[0042] In addition, the scope of target tasks applicable to the training method of the present disclosure is very wide, that is, it is not limited to specific types of tasks, but is applicable to all task scenarios described by protein language models that can be expressed by amino acid sequences or other coding methods, including both conditional generation and unconditional generation protein language models.

[0043] For example, after step S101, the following steps may also be executed: Based on the target task, determine the target task input format and the target task output format.

[0044] According to an embodiment of the present disclosure, in step S102, based on the target task, a first input data set corresponding to the target task is generated, and the first input data set includes task description data related to the target task.

[0045] For example, step S102 may further include: Based on the target task, generate a first input data set having the target task input format as the first input data set corresponding to the target task.

[0046] By defining the target task input format and the target task output format, the generated first input data set can be more matched with the target task.

[0047] According to an embodiment of the present disclosure, in step S103, the first input data set is input into a pre-trained RLAIF-based protein language model to obtain a first output data set, and the first output data set includes the predicted amino acid sequence of the protein.

[0048] As an example, a pre-trained RLAIF-based protein language model can be obtained in the following way: Obtain a preset protein language model. For example, a specific protein language model can be selected or pre-trained to be used as a base model; Use the first input dataset (for example, fine-tune the preset protein language model using the first input dataset), and adjust the various parameters of the preset protein language model through the gradient backpropagation algorithm to obtain an initial actor model, which serves as the pre-trained RLAIF-based protein language model.

[0049] In addition, the pre-trained RLAIF-based protein language model can be additionally used as a reference model for subsequent comparison with the re-adjusted model.

[0050] By pre-training the protein speech model using the first input dataset associated with the target task, supervised fine-tuning of the initial protein language model can be achieved, which can serve as a more effective initial actor model, thus adapting to the specific task requirements of various different tasks.

[0051] According to an embodiment of the present disclosure, in step S104, the first output dataset is input into a preset protein structure prediction model to obtain a second output dataset for indicating the protein sequence folding result.

[0052] According to an embodiment of the present disclosure, in step S105, based on the second output dataset, an optimal evaluation model for quantifying the stability of the protein folding state is determined as a scoring model.

[0053] As an example, step S105 may further include: determining multiple evaluation models corresponding to multiple performance metrics for quantifying the stability of the protein folding state according to the second output dataset; performing weighted summation on the multiple evaluation models according to the importance of the multiple performance metrics to obtain an optimal evaluation model as the scoring model.

[0054] For example, a protein structure prediction model is used to fold the amino acid sequence / protein sequence generated by the actor model, and the stability of the generated protein folding state is comprehensively predicted according to at least one of the model perplexity, energy, hydrophilicity-hydrophobicity, positive-negative charge, etc. of the folding result.

[0055] By using the above evaluation model as the scoring model, the defect that the protein sequences generated by the protein language model in the prior art are difficult to correctly fold in water and further perform functions can be solved, thereby greatly improving the accuracy and effectiveness of the predicted amino acid sequences.

[0056] In addition, by using the above evaluation model as the scoring model, it is more conducive to predicting the protein structure or function by learning the patterns in the amino acid sequence and even generating brand-new proteins.

[0057] According to an embodiment of the present disclosure, in step S106, the first output data set with the target task input format is input into the scoring model to obtain the corresponding reward value output data.

[0058] According to an embodiment of the present disclosure, in step S107, according to the reward value output data, each parameter of the pre-trained RLAIF-based protein language model is adjusted to obtain the trained RLAIF-based protein language model.

[0059] According to an embodiment of the present disclosure, through the above training method of the RLAIF-based protein language model according to the embodiment of the present disclosure, it is possible to reduce costs and increase efficiency in the process of protein design assisted by the protein language model by migrating the paradigm aligned in the large model field.

[0060] As an example, step S107 may further include S1071 to S1073:

[0061] In step S1071, according to the reward value output data, the optimization target value for adjusting each parameter is calculated by the following formula (1):

[0062]

[0063] where opt represents the optimization target value, x represents the first input data set, V(x) represents the reward value of the prediction result of the trained RLAIF-based protein language model (actor model) for x evaluated by the scoring model, β represents the preset parameter value, V ref (x) represents the reward value of the prediction result of the pre-trained RLAIF-based protein language model (reference model) for x evaluated by the scoring model, pθ represents the probability distribution of the trained RLAIF-based protein language model (actor model), p ref represents the probability distribution of the pre-trained RLAIF-based protein language model (reference model), t represents the number of predicted amino acid sequences, a t represents the process of predicting the next amino acid sequence of the protein according to the current amino acid sequence of the protein (for example, it can correspond to the action in reinforcement learning), s t represents the state of the number of amino acid sequences of the currently existing protein (for example, it can correspond to the state in reinforcement learning).

[0064] Here, the adaptive factor according to the present disclosure (i.e., the second term in formula (1) ) reflects the difficulty of the input sample for the initial actor model. For example, for the input sample for a predetermined task description, if the V corresponding to the reference model itself refIf the score of (x) is relatively high, the weight of the input sample can be reduced; otherwise, the weight can be increased.

[0065] In addition, the KL divergence according to the present disclosure (i.e., the third term in formula (1) ) is a constraint part that can fully consider the KL divergence between the actor model and the reference model. For example, by using the clip function to take a value of 1.2 when the ratio part of the two probability distributions is too large, the overall optimization objective will be relatively small, and it can be gradually adjusted to be larger during subsequent optimization, so as to achieve the purpose that the actor model will not deviate too far from the reference model.

[0066] In addition, opt is a function of pθ, and the meaning of the clip(c, a, b) function can be expressed as pruning when c is outside the interval [a, b], and the specific expression is formula (2) below:

[0067]

[0068] In step S1072, according to the optimization objective value, the corresponding loss value is calculated through the following formula (3):

[0069] loss = -opt (3)

[0070] where loss represents the loss value.

[0071] In step S1073, based on the loss value, each parameter is adjusted to obtain a trained protein language model based on RLAIF.

[0072] According to the above embodiments of the present disclosure, by defining the alignment standard of the protein language model as whether the protein sequence can be correctly folded in an aqueous solution environment, combining the reward value of the scoring model and the constraints of the reference model, designing the optimization objective of the actor model, and using the reinforcement learning method to iteratively optimize the actor model, the prediction result can meet the standard for quantifying the stability of the protein folding state.

[0073] In addition, the protein language model alignment method based on RLAIF using the above training method not only considers maximizing the score of the output result of the actor model (for example, reflected by the first term V(x) in formula (1)), but also considers the difference between the actor model and the reference model. By calculating the KL divergence between the two to impose certain limiting conditions and using an adaptive factor, the optimization strategy for different difficulty samples is adjusted.

[0074] In addition, further, during the process of multiple iterative optimizations of the RLAIF-based protein language model alignment method of the present disclosure, the model can be repeatedly trained using the same input data or different input data. For example, inference is performed using the current actor model, and the obtained result is input into the scoring model for scoring, thereby obtaining a reward value. The current actor model is optimized according to the above objective function to obtain an updated actor model. Through the above multiple iterative processes, the prediction effect of the trained protein language model can be improved.

[0075] Hereinafter, with reference to Figure 2 an example will be given to illustrate the training method of the RLAIF-based protein language model according to an embodiment of the present disclosure.

[0076] Figure 2 FIG. is an example flowchart showing the training method of the RLAIF-based protein language model according to an embodiment of the present disclosure.

[0077] In step S21, the target task to be solved is obtained, and a corresponding data set is constructed according to the target task. At the same time, the target task input format and the target task output format are defined.

[0078] In step S221, a protein language model is selected or pre-trained as a base model and further fine-tuned on a data set related to a specific task to meet specific requirements, forming a preliminary actor model.

[0079] In step S222, a copy of the initial actor model is created as a reference model for subsequent comparison.

[0080] In step S23, the alignment criterion of the protein language model is defined as whether the protein sequence can be correctly folded in an aqueous solution environment, and a suitable evaluation model for quantifying the stability of the protein folding state is selected as the scoring model.

[0081] In step S24, the data in the target task input format is input into the actor model, and the output result obtained by inference is input into the scoring model to output the corresponding reward value. Based on the reward value output by the scoring model and the reference model, the optimization objective function of the actor model is obtained.

[0082] Further, box S5 represents: optimizing the current actor model according to the optimization objective function to obtain an actor model with updated parameters. After multiple iterations, an actor model with the desired optimal performance is obtained, that is, the trained RLAIF-based protein language model.

[0083] Here, it should be noted that the above descriptions of Figure 1 each step are merely exemplary, and the steps in the method according to the present disclosure are not limited thereto.

[0084] Next, refer to Figure 3 to describe the prediction method of protein data according to the present disclosure. Figure 3 FIG. 300 is a flowchart showing a prediction method of protein data according to an embodiment of the present disclosure.

[0085] According to an embodiment of the present disclosure, in step S301, a target task for predicting the amino acid sequence of a protein is obtained.

[0086] According to an embodiment of the present disclosure, in step S302, the target task is input into a protein language model based on RLAIF, and the predicted amino acid sequence of the protein is obtained as the prediction result of the protein data.

[0087] Here, the protein language model based on RLAIF is trained by using the training method described above.

[0088] Next, refer to Figure 4 to describe a training device for a protein language model based on RLAIF according to an embodiment of the present disclosure.

[0089] Figure 4 FIG. 400 is a block diagram showing a training device for a protein language model based on RLAIF according to an embodiment of the present disclosure.

[0090] Refer to Figure 4 , and according to an embodiment of the present disclosure, the training device 400 for a protein language model based on RLAIF may include a task acquisition module 410, a data construction module 420, a prediction module 430, a reward value determination module 440, and an optimization module 450.

[0091] According to an embodiment of the present disclosure, the task acquisition module 410 may execute: obtaining a target task for predicting the amino acid sequence of a protein.

[0092] According to an embodiment of the present disclosure, the data construction module 420 may execute: generating a first input data set corresponding to the target task based on the target task, and the first input data set includes task description data related to the target task.

[0093] As an example, after the operation of obtaining a target task for predicting the amino acid sequence of a protein, the data construction module 420 is further configured to: determine a target task input format and a target task output format based on the target task.

[0094] In addition, as an example, the operation of the data construction module 420 generating a first input data set corresponding to the target task based on the target task may include: generating a first input data set having a target task input format based on the target task as the first input data set corresponding to the target task.

[0095] According to an embodiment of the present disclosure, the prediction module 430 may execute: inputting the first input data set into a pre-trained RLAIF-based protein language model to obtain a first output data set, where the first output data set includes the amino acid sequence of the predicted protein.

[0096] As an example, the pre-trained RLAIF-based protein language model may be obtained by: acquiring a preset protein language model; fine-tuning the preset protein language model using the first input data set, and adjusting each parameter of the preset protein language model through the gradient backpropagation algorithm to obtain an initial actor model, which serves as the pre-trained RLAIF-based protein language model.

[0097] According to an embodiment of the present disclosure, the reward value determination module 440 may execute: inputting the first output data set into a preset protein structure prediction model to obtain a second output data set for indicating the protein sequence folding result, and determining, based on the second output data set, an optimal evaluation model for quantifying the stability of the protein folding state as a scoring model, and inputting the first output data set into the scoring model to obtain a corresponding reward value output data.

[0098] As an example, the operation of the reward value determination module 440 to determine an optimal evaluation model for quantifying the stability of the protein folding state as a scoring model based on the second output data set may include: determining multiple evaluation models corresponding to multiple performance indicators for quantifying the stability of the protein folding state according to the second output data set; performing a weighted sum on the multiple evaluation models according to the importance of the multiple performance indicators to obtain an optimal evaluation model as the scoring model.

[0099] According to an embodiment of the present disclosure, the optimization module 450 may execute: adjusting each parameter of the pre-trained RLAIF-based protein language model according to the reward value output data to obtain a trained RLAIF-based protein language model.

[0100] As an example, the operation of the optimization module 450 to adjust each parameter of the pre-trained RLAIF-based protein language model according to the reward value output data to obtain a trained RLAIF-based protein language model may include: calculating an optimization target value for adjusting each parameter through the above formula (1) according to the reward value output data; calculating a corresponding loss value through the above formula (3) according to the optimization target value; and adjusting each parameter based on the loss value to obtain a trained RLAIF-based protein language model.

[0101] It should be noted that the operations performed by the above respective structural blocks may be similar to the related content described with reference to Figure 1 and will not be elaborated here.

[0102] Next, with reference to Figure 5 a prediction device for protein data according to an embodiment of the present disclosure will be described. Figure 5 FIG. is a block diagram showing a prediction device 500 for protein data according to an embodiment of the present disclosure.

[0103] With reference to Figure 5 , the prediction device 500 for protein data according to an embodiment of the present disclosure may include an acquisition module 510 and a prediction module 520.

[0104] According to an embodiment of the present disclosure, the acquisition module 510 may perform: acquiring a target task for predicting an amino acid sequence of a protein.

[0105] According to an embodiment of the present disclosure, the prediction module 520 may perform: inputting the target task into a protein language model based on RLAIF to obtain a predicted amino acid sequence of the protein as a prediction result of the protein data.

[0106] Here, the protein language model based on RLAIF is trained by using the training method described above.

[0107] Figure 6 FIG. is a block diagram showing a computing device 600 according to an embodiment of the present disclosure.

[0108] With reference to Figure 6 , the computing device 600 according to an embodiment of the present disclosure may include a processor 610 and a memory 620. The processor 610 may include (but is not limited to) a central processing unit (CPU), a digital signal processor (DSP), a microcomputer, a field programmable gate array (FPGA), a system on chip (SoC), a microprocessor, an application specific integrated circuit (ASIC), etc. The memory 620 may store computer executable instructions to be executed by the processor 610. The memory 620 includes high-speed random access memory and / or non-volatile computer-readable storage media. When the processor 610 executes the computer executable instructions stored in the memory 620, the training method of the protein language model based on RLAIF and the prediction method of protein data described above can be implemented.

[0109] The training method of the RLAIF-based protein language model and the prediction method of protein data according to an embodiment of the present disclosure can be written as computer programs / instructions to form a computer program product and stored on a computer-readable storage medium. When the computer programs / instructions are executed by a processor, the training method of the RLAIF-based protein language model and the prediction method of protein data as described above can be implemented. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device / server, the electronic device / server is enabled to execute the training method of the RLAIF-based protein language model and the prediction method of protein data as described above. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. In one example, the computer program and any associated data, data files, and data structures are distributed across a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0110] The training method and device of the RLAIF-based protein language model and the prediction method and device of protein data according to an embodiment of the present disclosure improve the performance of the protein language model in different generation tasks and reduce the cost of protein language model-assisted protein design by combining reinforcement learning techniques and protein language models.

[0111] On the other hand, the training method and device of the RLAIF-based protein language model proposed by the present disclosure contribute to promoting research and development in the fields of protein design, drug discovery, etc.

[0112] On the other hand, the training method and device of the RLAIF-based protein language model and the prediction method and device of protein data according to the embodiments of the present disclosure can also capture long-range dependencies in sequences.

[0113] On the other hand, through the effective protein model alignment method proposed by the present disclosure, the efficiency and generation accuracy in protein tasks can be greatly improved.

[0114] Although some embodiments of the present disclosure have been disclosed and described, those skilled in the art should understand that these embodiments can be modified and varied without departing from the concept and spirit of the present disclosure defined by the claims and their equivalents.

Claims

1. A method for training a protein language model, characterized in that: The training method comprises: Obtaining a target task for predicting the amino acid sequence of a protein; Based on the target task, generating a first input data set corresponding to the target task, wherein the first input data set includes task description data related to the target task; Inputting the first input data set into a pre-trained protein language model to obtain a first output data set, wherein the first output data set includes a predicted amino acid sequence of the protein; Inputting the first output data set into a preset protein structure prediction model to obtain a second output data set for indicating a protein sequence folding result; According to the second output data set, determining an optimal evaluation model for quantifying the stability of the protein folding state as a scoring model; Inputting the first output data set into the scoring model to obtain corresponding reward value output data; According to the reward value output data, various parameters of the pre-trained protein language model are adjusted to obtain a trained protein language model.

2. The training method according to claim 1, characterized in that: The pre-trained protein language model is obtained by: Get the preset protein language model; The first input data set is used to adjust various parameters of the preset protein language model through a gradient back propagation algorithm to obtain an initial actor model as the pre-trained protein language model.

3. The training method according to claim 1, characterized in that: The step of adjusting various parameters of the pre-trained protein language model according to the reward value output data to obtain the trained protein language model comprises: According to the reward value output data, the optimization target value for adjusting the various parameters is calculated by the following formula: Wherein, opt represents the optimization target value, x represents the first input data set, V(x) represents the reward value of the prediction result of x by the trained protein language model evaluated by the scoring model, β represents the preset parameter value, and V ref (x) represents the reward value of the prediction result of the pre-trained protein language model for x evaluated by the scoring model, p θ represents the probability distribution of the trained protein language model, p ref represents the probability distribution of the pre-trained protein language model, t represents the number of predicted amino acid sequences, a t It represents the process of predicting the amino acid sequence of the next protein based on the amino acid sequence of the current protein. t A state indicating the number of amino acid sequences of the protein that currently exist; According to the optimization target value, the corresponding loss value is calculated by the following formula: loss = -opt, where loss represents the loss value; The various parameters are adjusted based on the loss value to obtain a trained protein language model.

4. The training method according to claim 1, characterized in that: After the step of acquiring the target task for predicting the amino acid sequence of the protein, the training method further comprises: determining the target task input format and the target task output format based on the target task, and The step of generating a first input data set corresponding to the target task based on the target task includes: Based on the target task, a first input data set having an input format of the target task is generated as a first input data set corresponding to the target task.

5. The training method according to claim 1, characterized in that: The step of determining the optimal evaluation model for quantifying the stability of the protein folding state as a scoring model according to the second output data set comprises: determining, based on the second output data set, a plurality of evaluation models corresponding to a plurality of performance indicators for quantifying the stability of a protein folding state; According to the importance of the multiple performance indicators, the multiple evaluation models are weighted and summed to obtain the optimal evaluation model as the scoring model.

6. A method for predicting protein data, characterized in that: The prediction method comprises: Obtaining a target task for predicting the amino acid sequence of a protein; The target task is input into the protein language model to obtain the predicted amino acid sequence of the protein as the prediction result of the protein data. Wherein, the protein language model is trained using the training method according to any one of claims 1 to 5.

7. A protein language model training device, characterized in that: The training device comprises: The task acquisition module is configured to: acquire a target task for predicting the amino acid sequence of a protein; A data construction module is configured to: generate a first input data set corresponding to the target task based on the target task, wherein the first input data set includes task description data related to the target task; A prediction module is configured to: input the first input data set into a pre-trained protein language model to obtain a first output data set, wherein the first output data set includes a predicted amino acid sequence of the protein; The reward value determination module is configured as follows: Inputting the first output data set into a preset protein structure prediction model to obtain a second output data set for indicating a protein sequence folding result, According to the second output data set, determining an optimal evaluation model for quantifying the stability of the protein folding state as a scoring model, Inputting the first output data set into the scoring model to obtain corresponding reward value output data; The optimization module is configured to: adjust various parameters of the pre-trained protein language model according to the reward value output data to obtain a trained protein language model.

8. A protein data prediction device, characterized in that: The prediction device comprises: An acquisition module is configured to: acquire a target task for predicting the amino acid sequence of a protein; The prediction module is configured to: input the target task into the protein language model to obtain the predicted amino acid sequence of the protein as the prediction result of the protein data, Wherein, the protein language model is trained using the training method according to any one of claims 1 to 5.

9. A computer program product, characterized in that The computer program product comprises a computer program / instruction, and when the computer program / instruction is executed by a processor, the method for training a protein language model according to any one of claims 1 to 5 and the method for predicting protein data according to claim 6 are implemented.

10. A computing device, characterized in that The computing device includes: at least one processor; and at least one memory storing computer executable instructions, wherein the computer executable instructions, when executed by the at least one processor, cause the at least one processor to execute the protein language model training method as described in any one of claims 1 to 5 and the protein data prediction method as described in claim 6.

Citation Information

Cited By

  • Reinforcement learning-based protein sequence generation model training method and equipment

    CN122117069A