Protein generation model optimization method based on deep learning
By using a pre-trained diffusion model and a reward threshold partitioning method, the protein generation model is optimized using single-point experimental feedback data. This solves the dependence of the protein generation model on high-quality paired data and achieves clear optimization and performance improvement in a sparse reward environment.
Patent Information
- Application Number
- CN202511820546.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-03-06
AI Technical Summary
Existing protein generation models rely on high-quality paired sample data, which makes the determination of quality complex and scarce, affecting optimization efficiency.
A pre-trained diffusion model is used to generate electron density distribution maps of candidate proteins, which are then converted into amino acid sequences. The feedback dataset is divided into positive and negative samples by setting a reward threshold. A utility function is constructed for iterative optimization to avoid the complexity of paired data. Function-oriented fine-tuning is performed using single-point experimental feedback data.
We have achieved clear and stable protein generation model optimization in sparse reward environments, reducing data acquisition costs and improving the quality and functionality of generated protein sequences.
Smart Images

Figure CN121617461A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of protein generation technology, and more specifically to a method for optimizing a protein generation model based on deep learning. Background Technology
[0002] The function of a protein is determined by its amino acid sequence. Currently, protein generation models can be pre-trained on large-scale protein sequence databases to generate protein sequences with feasible applications based on specific functional or structural requirements. However, due to the extremely vast protein sequence space, the vast majority of protein sequences generated by these models do not exist in nature and are foldable or lack the required function. To obtain protein sequences that meet specific application needs, function-oriented fine-tuning of the protein model is required to guide it in exploring reasonable and diverse protein sequences within the target functional region.
[0003] Therefore, a feasible fine-tuning method is to optimize the objective function of Direct Preference Optimization (DPO) based on reinforcement learning. Specifically, DPO can bypass the construction of an explicit reward model and use preference data to fine-tune the policy model to improve the protein generation model's ability to generate functional protein sequences. However, DPO and similar optimization methods are based on contrastive learning mechanisms, meaning that these optimization methods rely on high-quality pairwise sample data. Specifically, in high-quality pairwise data samples, the protein generation model needs to learn preference relationships from samples that clearly distinguish between "good" and "bad".
[0004] However, in the field of protein design, the sequence space is extremely vast. Determining whether a protein sequence is "good" or "bad" involves complex indicators across multiple dimensions, and the evaluation criteria are often ambiguous. The process of determining "good" or "bad" relies on wet experimental verification (such as measuring protein stability, activity, and solubility) or on complex computational simulations.
[0005] Specifically, when comparing the merits of paired candidate protein samples, multiple performance metrics must be considered comprehensively. Ideally, candidate protein sample A should significantly outperform candidate protein sample B in one or more key metrics, while also performing no worse than candidate protein sample B in other important metrics. However, in most real-world scenarios, there are conflicting trade-offs in evaluating the merits of candidate protein samples. For example, candidate protein sample A may outperform candidate protein sample B in some important metrics, but underperform in others. This makes high-quality paired sample data that clearly determines the merits extremely scarce, and this scarcity severely restricts the optimization efficiency of protein generation models. Summary of the Invention
[0006] This invention provides a deep learning-based protein generation model optimization method to address the problem that traditional optimization methods in the prior art rely on large-scale, high-quality protein sequence preference pairwise data.
[0007] To address the aforementioned technical problems, the present invention provides a deep learning-based protein generation model optimization method, which includes: A pre-trained diffusion model is used to generate electron density distribution maps of candidate proteins based on the structural features of the target site; The electron density distribution map of the candidate protein is converted into the corresponding amino acid sequence; The functional index values corresponding to the amino acid sequence are determined, and a feedback dataset is constructed. The feedback dataset consists of multiple independent data units. Each data unit contains an amino acid sequence, a functional index value and their corresponding relationship, and there is no pairing preference relationship between the data units. A reward threshold is set, and the feedback dataset is labeled as a positive sample dataset and a negative sample dataset according to the reward threshold, wherein the functional index value of any sample in the positive sample dataset is better than the reward threshold, and the functional index value of any sample in the negative sample dataset is worse than the reward threshold. Construct a utility function with the reward threshold as a reference point, and calculate the utility values of the positive sample dataset and the negative sample dataset respectively; The diffusion model is iteratively optimized based on the utility value.
[0008] This invention performs binary partitioning of each sample in the feedback dataset based on the relative relationship between the corresponding functional indicator value and the reward threshold. Compared with traditional paired preference data, this invention effectively avoids the complexity and high cost of data acquisition in traditional preference optimization methods (such as DPO). Furthermore, the explicit reward threshold makes the data labeling standard clearer and more objective.
[0009] Based on the utility value, positive sample datasets are further assigned encouraging properties, while negative sample datasets are assigned penalizing properties. This allows the positive sample dataset to guide the diffusion model to increase its tendency to generate such high-quality samples in subsequent optimizations, while the negative sample dataset can guide the model to suppress its tendency to generate such low-quality samples. This iterative optimization can achieve continuous improvement in the performance of the protein generation model.
[0010] Furthermore, by establishing a relative evaluation system with the reward threshold as a reference point, a clear and stable optimization guide can be given to the protein generation model in the sparse reward environment of protein generation sequences, effectively avoiding the protein generation model from being too dependent on a single reward mode.
[0011] According to another specific embodiment of the present invention, the step of generating a candidate protein electron density distribution map based on the structural features of the target point using a pre-trained diffusion model includes inputting the structural features of the target point into the pre-trained diffusion model, wherein the diffusion model is a U-Net architecture diffusion model trained in the electron density latent space, and the structural features of the target point include three-dimensional electron density distribution map data; performing iterative denoising sampling based on the three-dimensional electron density distribution map data to generate an electron density latent space representation of the candidate binding protein; and decoding the electron density latent space representation of the candidate binding protein to generate the candidate protein electron density distribution map.
[0012] According to another specific embodiment of the present invention, the determination of the functional index value corresponding to the amino acid sequence includes obtaining the functional index value corresponding to the amino acid sequence through computer simulation calculation and / or experimental determination.
[0013] According to another specific embodiment of the present invention, the step of constructing a utility function with the reward threshold as a reference point and calculating the utility values of the positive sample dataset and the negative sample dataset respectively includes constructing a utility function. The evaluation of the feedback dataset by the utility function is based on the relative difference between the functional index value of the feedback dataset and the reward threshold. The utility function is configured such that the utility function of samples with functional index values better than the reward threshold is concave in the reward domain, the utility function of samples with functional index values worse than the reward threshold is convex in the loss domain, and the utility function is more sensitive to loss than to an equal amount of gain.
[0014] According to another specific embodiment of the present invention, the step of constructing a utility function with the reward threshold as a reference point and calculating the utility values of the positive sample dataset and the negative sample dataset respectively includes: constructing the utility function based on prospect theory, denoted as: Where Q is the functional indicator value, Q ref The reward threshold is defined; the utility value of the positive sample dataset is calculated and denoted as: ; Calculate the utility value of the negative sample dataset, denoted as: .
[0015] According to another specific embodiment of the present invention, the iterative optimization of the diffusion model based on the utility value includes establishing an optimization objective function based on the utility function to maximize the expected value of the utility value, denoted as: ;in, This represents the diffusion model strategy to be optimized. This represents the mathematical expectation of the utility function U.
[0016] According to another specific embodiment of the present invention, the step of establishing an optimization objective function based on the utility function to maximize the expected utility value further includes: ; in, Let D represent the sample weight function of the feedback sample set, D represent the feedback dataset, t represent the diffusion time step of sampling from the uniform distribution U[0,1], β represent the hyperparameter of the deviation of the control policy, and Q represent the sample weight function of the feedback sample set. ref Indicates the reward threshold. π represents the diffusion model strategy to be optimized. ref This indicates the reference diffusion strategy.
[0017] According to another specific embodiment of the present invention, the step of establishing an optimization objective function based on the utility function to maximize the expected utility value further includes calculating the diffusion model strategy to be optimized. Compared with the reference diffusion strategy π ref The average behavioral difference value between the two is denoted as the reward threshold.
[0018] According to another specific embodiment of the present invention, the reward threshold is denoted as or Wherein, π(a|s) represents the probability that the diffusion model generates a specific candidate protein a based on the state s of the target site. This represents the diffusion model strategy to be optimized. To reference diffusion strategy π ref KL divergence between them.
[0019] According to another specific embodiment of the present invention, the iterative optimization of the diffusion model based on the utility value includes modeling the denoising process of the diffusion model as a Markov decision process, wherein each denoising step is defined as a state transition process, and the denoising direction and magnitude are defined as the action space; and iteratively updating the diffusion model based on the policy gradient method. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein: Figure 1This is a flowchart illustrating an embodiment of an optimization method provided in this application. Figure 1 ; Figure 2 This is a flowchart illustrating an embodiment of an optimization method provided in this application. Figure 2 ; Figure 3 This is a schematic diagram illustrating the change of the loss function during the training process of the diffusion model provided in this application; Figure 4 This describes the change in reward value after iterations during the training process of the diffusion model provided in this application. Detailed Implementation
[0021] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention is presented in conjunction with preferred embodiments, this does not mean that the features of the invention are limited to these embodiments. On the contrary, the purpose of describing the invention in conjunction with embodiments is to cover other options or modifications that may be derived based on the claims of the present invention. To provide a deep understanding of the invention, many specific details will be included in the following description. The invention may also be implemented without using these details. Furthermore, to avoid confusion or obscuring the focus of the invention, some specific details will be omitted in the description. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0022] It should be noted that in this specification, similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0024] Please see Figure 1 This application provides a deep learning-based protein generation model optimization method that can use unpaired single-point experimental feedback data to perform function-oriented fine-tuning of the protein generation model, thereby improving the quality and functionality of the protein sequences generated by the model.
[0025] The deep learning-based protein generation model optimization method includes: Step S1: Use a pre-trained diffusion model to generate an electron density distribution map of candidate proteins based on the structural features of the target site.
[0026] Using this pre-trained diffusion model as the initial model, which has learned general prior knowledge of protein structure, and by inputting the structural features of the target site as conditional information into the model, it is possible to utilize the generative capabilities of the model while ensuring that the structure of the protein sequence generated through the conditional mechanism has a specific binding potential with the target site.
[0027] Step S2: Convert the electron density distribution map of the candidate protein into the corresponding amino acid sequence.
[0028] The three-dimensional electron density distribution maps of candidate proteins are converted into one-dimensional amino acid sequences to ensure data standardization. Specifically, the protein electron density distribution map contains the spatial probability distribution of electrons in the protein.
[0029] This application does not limit the specific translation process; this step can be achieved through sequence decoding methods involving template matching and iterative refinement. For example, it can use current modeling software to trace the protein backbone, accurately match local density features with known chemical structures, and then reconstruct the corresponding amino acid sequence.
[0030] Step S3: Determine the functional index values corresponding to the amino acid sequence and construct a feedback dataset.
[0031] The feedback dataset constructed in step S3 is a "non-paired" single-point dataset. The feedback dataset consists of multiple independent data units. Each data unit contains an amino acid sequence, one or more functional index values corresponding to the amino acid sequence, and the correspondence between the amino acid sequence and the functional index values. There is no pairing preference relationship between the data units.
[0032] In other words, each amino acid sequence is sampled, measured, and evaluated as an independent sample, and there is no comparison of superiority or inferiority between different amino acid sequences based on pairing. The functional index value of an amino acid sequence is a self-contained and absolute recording property of that amino acid (such as a directly measurable melting temperature), which is independent of direct comparison of superiority or inferiority with other amino acid sequences.
[0033] It should be noted that the “paired” samples in this application specifically refer to samples that have a clear distinction in quality between a pair of annotations for the same prompt or condition, which are required in existing preference optimization methods (such as DPO).
[0034] Step S4: Set a reward threshold and label the feedback dataset as a positive sample dataset and a negative sample dataset according to the reward threshold.
[0035] Each sample in the feedback dataset is binary-coded based on the relative relationship between its corresponding functional indicator value and the reward threshold. Samples with functional indicator values better than the reward threshold are classified as positive samples, and samples with functional indicator values worse than the reward threshold are classified as negative samples. This ensures that there is no overlap between the positive and negative sample datasets.
[0036] Compared to traditional paired preference data, the binary partitioning method based on a single threshold relies solely on the independent functional index value of each sample. On the one hand, it effectively avoids the complexity and high cost of data acquisition in traditional preference optimization methods (such as DPO). On the other hand, the explicit reward threshold makes the data labeling standard clearer and more objective.
[0037] Step S5: Construct a utility function with the reward threshold as a reference point, and calculate the utility values of the positive sample dataset and the negative sample dataset respectively.
[0038] By quantifying the functional indicators of proteins into well-defined utility values through utility functions, a clear and explicit evaluation system for feedback datasets (positive sample datasets and negative sample datasets) can be established.
[0039] Step S6: Iteratively optimize the diffusion model based on the utility value.
[0040] Therefore, based on the calculated utility values, positive sample datasets can be further assigned incentive properties, allowing them to guide the diffusion model to increase its tendency to generate such high-quality samples in subsequent optimizations; and negative sample datasets can be assigned penalty properties, allowing them to guide the model to suppress its tendency to generate such low-quality samples in subsequent optimizations. This iterative optimization can achieve continuous improvement in the performance of the protein generation model (diffusion model).
[0041] Combining steps S1-S6 above, the reliance on paired data can be effectively eliminated, requiring only the evaluation of single-sample functional indicators, thus significantly reducing data acquisition costs. Furthermore, optimization based on a utility function, this explicit utility-oriented mechanism, ensures that the model evolves towards better performance in each iteration. For example, the optimized candidate protein sequence may be superior to the previously generated candidate protein sequence.
[0042] The inventors also recognized that traditional optimization methods (such as DPO) have some inherent optimization problems, which are amplified, especially in the complex, high-dimensional output space of protein generation models. Specifically, during the training of protein generation models, traditional optimization methods lead to a decrease in the "implicit reward" of the protein generation model on the preferred data.
[0043] The reason is that protein generation operates in a sparse reward environment. Since most generated protein sequences lack clear functional annotations, it's difficult to obtain accurate training signals, leading to biased estimations of the sequence's true value by the protein generation model. Consequently, the model over-utilizes known high-reward patterns, repeatedly generating "pseudo-high-performance" protein sequences. This typically means a decline in the quality of protein generation.
[0044] By employing steps S1-S6 above, a relative evaluation system with a reward threshold as a reference point is established, providing a clear and stable optimization guide for the protein generation model under sparse reward conditions. This strengthens the protein generation model's ability to suppress low-quality protein generation, thereby effectively preventing the model from becoming overly reliant on a single reward pattern.
[0045] It should be noted that any of the following embodiments in this application primarily pertain to a protein generation diffusion model (i.e., the pre-trained diffusion model described above). The diffusion model may include reversible diffusion processes, as detailed below: The forward diffusion process q can be represented as a fixed Markov chain, roughly described as starting from an initial sample x0 and continuously adding noise until standard Gaussian noise x is reached. T The process is represented as: .
[0046] The reverse diffusion process p starts from a random distribution of Gaussian random noise, gradually denoising it to finally reconstruct a sampled data that conforms to the original data distribution. This process is also modeled as a Markov chain: .
[0047] In summary, by minimizing the loss function, the diffusion model can learn to reconstruct data from noise. Once the diffusion model is trained, the reverse diffusion process p described above can be used to iteratively denoise the data, starting from pure noise, and finally sample and generate the required data sample x0.
[0048] In one possible specific implementation, step S1 above includes the following steps: Step S10: Input the structural features of the target point into the pre-trained diffusion model.
[0049] For example, the diffusion model is a U-Net architecture diffusion model trained in the electron density latent space, and the structural features of the target point include three-dimensional electron density distribution data. Of course, the target point, as a target protein, may also include, but is not limited to, the following data: atomic-level three-dimensional coordinates, molecular surface geometric features, electrostatic potential distribution, hydrophobic field distribution, and functional site information.
[0050] Step S11: Based on the three-dimensional electron density distribution map data, perform iterative denoising sampling to generate the electron density latent space representation of the candidate binding protein.
[0051] Step S12: Decode the latent space representation of the electron density of the candidate binding protein to generate the electron density distribution map of the candidate protein.
[0052] The electron density latent space is constructed by encoding the electron density data in the real space into the reciprocal space through Fourier transform. This representation method can effectively capture the global features and local details of protein structure.
[0053] The diffusion model using the U-Net architecture incorporates target structural condition information at each step of the denoising process through its encoder-decoder structure, ensuring that the electron density distribution of the generated candidate protein is effectively complementary to the target in terms of spatial conformation and physicochemical properties.
[0054] For example, step S3 can obtain functional index values corresponding to the amino acid sequence through computer simulation calculation and / or experimental determination.
[0055] The computer simulation calculations include, but are not limited to, calculations based on molecular dynamics simulations, binding free energy calculations, or sequence-structure feature predictions. The biological experimental assays include, but are not limited to, surface plasmon resonance, isothermal titration calorimetry, or enzyme-linked immunosorbent assays. The aforementioned biological experimental assays have superior accuracy compared to computer simulation calculations. For example, a specific biological experimental assay involves using a high-throughput biological experimental platform to simultaneously measure the functional index values of each binding protein, constructing the aforementioned feedback dataset. This feedback dataset uses amino acid sequences and functional index values as core units, but may still cover other indicators; this application does not limit its scope.
[0056] In one possible specific implementation, the reward threshold in step S4 (denoted as Q) ref It can be adaptively adjusted according to the actual situation.
[0057] The reward threshold can be based on the diffusion model strategy to be optimized. Compared with the reference diffusion strategy π ref The average behavioral difference value between the two is determined, and this average behavioral difference value is used as the reward threshold.
[0058] Among them, the reference diffusion strategy π ref This represents the diffusion model before fine-tuning begins or after the previous round of fine-tuning. Its parameters remain constant during fine-tuning. Therefore, it can be used as a benchmark to prevent the optimized diffusion model from deviating excessively from the learned, reasonable prior protein knowledge. Correspondingly, the behavioral difference value is calculated by evaluating the diffusion model strategy to be optimized. Compared with the reference diffusion strategy πref It is measured by the logarithm of the probability ratio of actions occurring under the same conditions.
[0059] In one scheme, the aforementioned reward threshold can be dynamically set, denoted as... Here, β is a hyperparameter controlling the degree of deviation from the strategy, and π(a|s) represents the probability that the diffusion model generates a specific candidate protein a based on the state s of the target site.
[0060] The reward threshold is theoretically proportional to the Kullback-Leibler divergence between the diffusion model policy to be optimized and the fixed reference diffusion model policy. However, since it is not feasible to accurately calculate the KL divergence in high-dimensional space, the theoretical KL divergence can be estimated unbiasedly using the Monte Carlo sampling method.
[0061] In the current optimization environment, sample M states (states s based on the target point) and actions (generating a specific candidate protein a) on the sample. ,calculate The average value is used to estimate the theoretical KL divergence.
[0062] The aforementioned reward threshold can then be denoted as: .
[0063] Thus, the adaptive determination of the reward threshold can effectively reflect the degree of deviation of the current policy from the reference policy (initial policy or previous policy), thereby enabling the construction of a dynamic utility function in step S5.
[0064] Of course, the aforementioned reward threshold can also be set based on static experience of domain knowledge, such as setting a fixed value according to the current optimization goal, historical data, etc. More specifically, a constant can be specified according to the expected performance improvement.
[0065] In one possible implementation, the utility function constructed in step S5 above evaluates the feedback dataset based on the relative difference between the functional index value of the feedback dataset and the reward threshold.
[0066] Specifically, the utility function is configured such that the utility function of samples with a functional index value better than the reward threshold is concave in the payoff domain (i.e., risk aversion), the utility function of samples with a functional index value worse than the reward threshold is convex in the loss domain (i.e., risk seeking), and the utility function is more sensitive to loss than to an equivalent amount of gain.
[0067] In other words, the utility function employs an asymmetric preference design to reflect the loss aversion characteristic of prospect theory. This ensures that during the optimization process, the penalty for generating poor-quality protein sequences in the protein generation model is greater than the reward for generating high-quality samples.
[0068] In the specific implementation of the utility function, this can be manifested by using a steeper function slope in the loss domain or assigning higher weights to the loss domain. This achieves a balance between strong inhibition of low-quality protein sequence generation and milder encouragement of high-quality protein sequence generation.
[0069] The utility function U constructed in step S5 above can be denoted as: Where Q is the functional indicator value.
[0070] In a specific application scenario, a utility function U with loss aversion characteristics is constructed using the hyperbolic tangent function. Thus, the utility value of the positive sample dataset is calculated. ; Calculate the utility value of the negative sample dataset. .
[0071] The hyperbolic tangent function can effectively satisfy the need for loss aversion. When the absolute value of the functional index value minus the reward threshold is the same, due to the characteristics of the hyperbolic tangent function, the rate of change of the function value in the loss domain is greater than the rate of change in the gain domain.
[0072] Step S6 above includes: establishing an optimization objective function based on the utility function, with the goal of maximizing the expected utility value. The constructed objective function can be expressed as follows: .in, This represents the diffusion model strategy to be optimized. This represents the mathematical expectation of the utility function U.
[0073] In a specific application scenario, maximizing the expected utility value as the objective can also be expressed as: in, This represents the sample weight function of the feedback sample set. For example, if a sample in the positive sample dataset is represented as "excellent", then... Conversely, if samples in a negative sample dataset are represented as "inferior", then... .
[0074] X0 represents the amino acid sequence sampled from the feedback dataset (denoted as D); t represents the diffusion time step sampled from the uniform distribution U[0,1]. This represents the diffusion strategy to be optimized. This indicates the reference diffusion strategy.
[0075] In this way, behavioral constraints are imposed on the diffusion model, ensuring that its generation behavior in each denoising step does not deviate excessively from the reference strategy. This guarantees that, driven by the aforementioned utility function U, the diffusion model can be guided to a higher performance platform, improving the generation of protein sequences with higher functional value.
[0076] In one specific implementation, step S6 includes: modeling the denoising process of the diffusion model as a Markov decision process, wherein each denoising step is defined as a state transition process, and the denoising direction and magnitude are defined as the action space; and iteratively updating the diffusion model based on the policy gradient method.
[0077] Since policy optimization is one of the core implementation methods of deep learning, it can be adapted to the diffusion model used in this application. Furthermore, the forward and backward processes of the diffusion model (the addition of noise at each step) can essentially be represented as Markov chains. Therefore, the forward and backward processes of the diffusion model can be incorporated into the modeling framework of deep learning. The policy probability in reinforcement learning can be equated to the probability distribution predicted at each step of the aforementioned diffusion model.
[0078] Thus, by transforming the diffusion process into a sequence decision problem, the diffusion model can continuously optimize its generation strategy based on functional feedback signals, ultimately achieving targeted enhancement of protein functional properties.
[0079] In a specific application scenario, virtual computational metrics are used to verify the feasibility of the entire closed-loop framework (steps S1~S6), and its loss reduction state is as follows: Figure 3 As shown.
[0080] See Figure 3 As the number of iterations increases, the loss value shows a monotonically decreasing trend. Furthermore, in this specific application scenario, the loss curve gradually stabilizes. This indicates that the diffusion model performs stably during training, without significant overfitting or underfitting. In other words, based on steps S1-S6, the diffusion model training exhibits superior convergence and effectiveness.
[0081] In the same application scenario, the increase in the model's reward value (a virtual indicator of generation quality) after each iteration is as follows: Figure 4 As shown in the figure, the blue curve represents the first fine-tuning, the red curve represents the second fine-tuning, and the gray curve represents the third fine-tuning.
[0082] See Figure 4All three curves show a monotonically increasing trend, indicating that the diffusion model can continuously improve with each fine-tuning. Furthermore, the starting and ending reward values for each round of fine-tuning are higher than the previous round; specifically, the third fine-tuning is more effective than the second, and the second fine-tuning is more effective than the first. This demonstrates that the learning efficiency and generation quality of the diffusion model improve after each fine-tuning.
[0083] Furthermore, the fact that the three curves gradually stabilize also indicates that the diffusion model of this application can stably reach a higher performance level after each round of fine-tuning.
[0084] based on Figure 3 , Figure 4 The experimental phenomena shown demonstrate that steps S1 to S6 can not only effectively improve the quality of the protein generation model, but also show a trend of continuous improvement.
[0085] In some implementations, this application also provides a computer program product, which includes a computer program / instructions that, when executed by a processor, implement the aforementioned deep learning-based protein generation model optimization method.
[0086] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. It is clear to those skilled in the art that various implementations can be implemented using software plus necessary general-purpose hardware platforms, and of course, also using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the prior art, can be embodied in the form of software products. These computer software products can be stored in computer-readable storage media, such as ROM / RAM, magnetic disks, optical disks, etc., and include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or certain parts of embodiments.
[0087] The detailed descriptions listed above are merely specific descriptions of feasible embodiments of the present invention, and are not intended to limit the scope of protection of the present invention. All equivalent embodiments or modifications made without departing from the spirit of the present invention should be included within the scope of protection of the present invention.
Claims
1. A deep learning-based protein generation model optimization method, characterized in that, The method comprises the following steps: generating a candidate protein electron density map based on the structural features of the target point by using a pre-trained diffusion model; converting the candidate protein electron density map into a corresponding amino acid sequence; determining the functional indicator value corresponding to the amino acid sequence, and constructing a feedback dataset, wherein the feedback dataset is composed of multiple independent data units, each data unit contains an amino acid sequence, a functional indicator value and their corresponding relationship, and there is no paired preference relationship between the data units; setting a reward threshold, and marking the feedback dataset as a positive sample dataset and a negative sample dataset according to the reward threshold, wherein the functional indicator value of any sample in the positive sample dataset is better than the reward threshold, and the functional indicator value of any sample in the negative sample dataset is worse than the reward threshold; constructing an utility function with the reward threshold as a reference point, and calculating the utility values of the positive sample dataset and the negative sample dataset respectively; iteratively optimizing the diffusion model according to the utility values. 2.The deep learning-based protein generation model optimization method of claim 1, wherein, The method of generating a candidate protein electron density map based on the structural features of the target point by using a pre-trained diffusion model comprises the following steps: inputting the structural features of the target point into the pre-trained diffusion model, wherein the diffusion model is a U-Net architecture diffusion model trained in an electron density hidden space, and the structural features of the target point contain three-dimensional electron density distribution data; performing iterative denoising sampling based on the three-dimensional electron density distribution data to generate an electron density hidden space representation of the candidate binding protein; decoding the electron density hidden space representation of the candidate binding protein to generate the candidate protein electron density map. 3.The deep learning-based protein generation model optimization method of claim 1, wherein, The method of determining the functional indicator value corresponding to the amino acid sequence comprises the following steps: obtaining the functional indicator value corresponding to the amino acid sequence through computer simulation calculation and / or experimental determination. 4.The deep learning-based protein generation model optimization method of claim 1, wherein, The method of constructing an utility function with the reward threshold as a reference point and calculating the utility values of the positive sample dataset and the negative sample dataset respectively comprises the following steps: constructing an utility function, wherein the evaluation of the feedback dataset based on the relative difference between the functional indicator value of the feedback dataset and the reward threshold, and the utility function is configured to: the utility function of the sample with a functional indicator value better than the reward threshold is concave in the gain domain, the utility function of the sample with a functional indicator value worse than the reward threshold is convex in the loss domain, and the sensitivity of the utility function to loss is higher than that to equivalent gain. 5.The deep learning-based protein generation model optimization method of claim 1 or 4, wherein, The method of constructing an utility function with the reward threshold as a reference point and calculating the utility values of the positive sample dataset and the negative sample dataset respectively comprises the following steps: constructing the utility function based on the prospect theory, denoted as: ; Wherein, Q is the functional index value, Q ref is the reward threshold value; calculating the utility value of the positive sample dataset, denoted as: ; calculating the utility value of the negative sample dataset, denoted as: 。 6.The deep learning-based protein generation model optimization method of claim 1 or 4, wherein, The method of iteratively optimizing the diffusion model according to the utility values comprises the following steps: establishing an optimization objective function based on the utility function to maximize the expected value of the utility value, denoted as: ; wherein, denotes the diffusion model strategy to be optimized, denotes the mathematical expectation of the utility function U.
7. The deep learning-based protein generation model optimization method of claim 6, wherein, The method of establishing an optimization objective function based on the utility function to maximize the expected value of the utility value further comprises the following steps: ; wherein, denotes a sample weight function of the feedback sample set, D denotes a feedback dataset, t denotes a diffusion time step sampled from a uniform distribution U[0, 1], β denotes a hyperparameter controlling the degree of policy deviation, Q ref denotes a reward threshold, denotes a diffusion model policy to be optimized, π ref denotes a reference diffusion policy. 8.The deep learning-based protein generation model optimization method of claim 7, wherein, further comprising: Computing a diffusion model policy to optimize The average behavior difference value between the reference diffusion policy π ref is denoted as the reward threshold. 9.The deep learning-based protein generation model optimization method of claim 8, wherein, The reward threshold is denoted as: or ; wherein π(a|s) represents a probability of generating a specific candidate protein a based on a state s of the target site by the diffusion model, represents the KL divergence between the diffusion model strategy to be optimized and the reference diffusion strategy π ref . 10.The deep learning based protein generation model optimization method of claim 1, wherein, The iterative optimization of the diffusion model according to the utility value comprises: Modeling the denoising process of the diffusion model as a Markov decision process, wherein each denoising step is defined as a state transition process, and the denoising direction and amplitude are defined as an action space; Iteratively updating the diffusion model based on a policy gradient method.
Citation Information
Cited By
Protein inverse folding method based on conditional information
CN121922190A
Reinforcement learning-based protein sequence generation model training method and equipment
CN122117069A
Training method and device of protein sequence generation model based on reinforcement learning
CN122117069B