Optimization of antibody generative model

By introducing an energy preference optimization goal into the antibody generation model and using a pre-trained model to generate and label energy labels, the antibody generation model is optimized, solving the problems of insufficient data and inaccurate learning objectives, and improving the energy performance and binding ability of antibodies.

WO2025194466A1PCT designated stage Publication Date: 2025-09-25BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/083132
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing antibody generation models suffer from insufficient data and inaccurate learning objectives during the design process, resulting in the generated antibodies having unreasonable energy performance and being unable to effectively bind to antigens.

Method used

By introducing the energy-biased optimization objective, the antibody generation model is fine-tuned, multiple antibodies are generated using the pre-trained model, energy labels are annotated to construct positive and negative samples, and the model is optimized using the energy-biased optimization algorithm to improve the quality of antibody generation.

Benefits of technology

It effectively avoids the problem of insufficient data in the antibody generation task, improves the energy performance of antibodies, and enhances the ability of antibodies to bind to antigens, so that the generated antibodies are more in line with the energy requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024083132_25092025_PF_FP_ABST
    Figure CN2024083132_25092025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure relate to the optimization of an antibody generative model. A method for model training comprises: on the basis of information of a target antigen, using a current antibody generative model to generate a plurality of first antibodies; respectively determining antibody energies of the plurality of first antibodies; on the basis of the antibody energies of the plurality of first antibodies, constructing a plurality of first samples for the antibody generative model, wherein the plurality of first samples comprise at least one positive sample and at least one negative sample, an energy label corresponding to the first antibody in the positive sample indicates that the antibody energy thereof satisfies an energy preference associated with the target antigen, and an energy label corresponding to the first antibody in the negative sample indicates that the antibody energy thereof does not satisfy the energy preference; and using the plurality of first samples to execute fine tuning on the antibody generative model on the basis of an optimization objective, wherein the optimization objective is based on the discrepancy between predicted energy labels of antibodies generated by the fine-tuned antibody generative model and the energy labels of the plurality of first samples.
Need to check novelty before this filing date? Find Prior Art

Description

Optimization of antibody production models Technical Field

[0001] Example embodiments of the present disclosure relate generally to the field of computers, and more particularly to optimization of antibody production models. Background Art

[0002] Antibodies are one of the important components of the immune system in animals. Specifically, antibodies are a type of "Y"-shaped protein composed of two chains, light and heavy, that can specifically bind to pathogens in animals. Their high specificity gives antibodies a promising application as a macromolecular drug. The specificity of antibodies comes from the complementarity determining region (CDR) of the antibody. CDR is usually also the part where the antibody binds to the antigen, and they play a key role in the recognition and binding of the antigen. In addition, the rest of the antibody is relatively conserved, so the design of antibodies can be regarded as the design of CDRs with binding ability to specific antigens. Understanding the importance of sequence analysis of the antibody CDR region is crucial to understanding the specific binding mechanism between antibodies and antigens.

[0003] Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for model training is provided. The method includes: generating multiple first antibodies based on information of a target antigen using a current antibody generation model; determining the antibody energy of each of the multiple first antibodies; constructing multiple first samples for the antibody generation model based on the antibody energy of the multiple first antibodies, the multiple first samples including at least one positive sample and at least one negative sample, the energy label corresponding to the first antibody in the positive sample indicating that its antibody energy satisfies the energy preference associated with the target antigen, and the energy label corresponding to the first antibody in the negative sample indicating that its antibody energy does not satisfy the energy preference; and using the multiple first samples, fine-tuning the antibody generation model according to an optimization target, the optimization target being based on the difference between the predicted energy label of the antibody generated by the fine-tuned antibody generation model and the energy label between the multiple first samples.

[0005] In a second aspect of the present disclosure, a device for model training is provided. The device includes: an antibody generation module, configured to generate multiple first antibodies based on information of the target antigen using the current antibody generation model; an energy determination module, configured to respectively determine the antibody energy of each of the multiple first antibodies; a sample construction module, configured to construct multiple first samples for the antibody generation model based on the antibody energy of the multiple first antibodies, the multiple first samples including at least one positive sample and at least one negative sample, the energy label corresponding to the first antibody in the positive sample indicates that its antibody energy satisfies the energy preference associated with the target antigen, and the energy label corresponding to the first antibody in the negative sample indicates that its antibody energy does not satisfy the energy preference; and a model fine-tuning module, configured to use the multiple first samples to perform fine-tuning on the antibody generation model according to an optimization target, the optimization target being based on the difference in energy labels between the predicted energy labels of the antibodies generated by the fine-tuned antibody generation model and the multiple first samples.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method of the first aspect of the present disclosure.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium and can be executed by a processor to perform the method according to the first aspect of the present disclosure.

[0008] In a fifth aspect of the present disclosure, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions, which, when executed by a device, cause the device to perform the method of the first aspect.

[0009] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG2 shows a schematic diagram of a model training system for an antibody generation model according to some embodiments of the present disclosure;

[0013] FIG3 shows a schematic diagram of an environment in which embodiments of the present disclosure can be implemented;

[0014] FIG4 shows a flowchart of a process for model training according to some embodiments of the present disclosure;

[0015] FIG5 shows a block diagram of an apparatus for model training according to some embodiments of the present disclosure; and

[0016] FIG6 illustrates a block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0017] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.

[0019] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0020] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0021] It is understandable that before using the technical solutions disclosed in the various embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the information involved in the present disclosure should be informed to relevant users and authorization should be obtained from relevant users in an appropriate manner in accordance with relevant laws and regulations. The relevant users may include any type of right holders, such as individuals, enterprises, and groups.

[0022] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly prompt the relevant user that the operation requested to be performed will require obtaining and using the information of the relevant user, so that the relevant user can independently choose whether to provide information to the software or hardware such as the electronic device, application, server or storage medium that executes the operation of the technical solution of the present disclosure based on the prompt message.

[0023] As an optional but non-limiting implementation, in response to receiving an active request from a relevant user, a prompt message may be sent to the relevant user in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide information to the electronic device.

[0024] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure. The activation of the digital assistant-related functions of the embodiment of the present disclosure, the data obtained, the processing and storage of the data, etc., shall all obtain the prior authorization of the user and other rights holders associated with the user, and shall comply with the provisions of relevant laws and regulations and the rules of agreement between rights holders.

[0025] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0026] Reinforcement learning, also known as reinforcement learning, evaluation learning, or enhanced learning, is a machine learning technique used to describe and solve the problem of an intelligent agent learning strategies to maximize rewards or achieve specific goals during its interaction with the environment. Reinforcement learning focuses on the interaction between the agent and the environment, and its goal is generally to maximize rewards. In other words, reinforcement learning is a learning mechanism that learns how to map states to behaviors in order to maximize rewards. Such an intelligent agent needs to continuously experiment in the environment, continuously optimizing the state-behavior relationship through feedback (rewards) provided by the environment.

[0027] Reinforcement learning systems generally involve four elements: policy, reward, value, and environment or model. The following sections introduce these four elements separately.

[0028] A policy defines the actions a model should take in a given state, essentially mapping states to actions. A state refers to the state perceived by the model. Typically, the policy is the core of a reinforcement learning system, as it determines the actions to take in each state. Depending on the configuration, the policy itself can be a specific mapping or a random distribution.

[0029] Rewards define the objective of a reinforcement learning problem. At each time step, the environment sends a scalar value to the reinforcement learning system. Rewards determine how well a model performs. Therefore, reward signals are the primary factor influencing the policy. The model's task is to maximize the total reward accumulated over a period of time.

[0030] Value, or the value function, is a crucial concept in reinforcement learning. Unlike immediate rewards, a value function measures long-term benefits. It evaluates the benefits of a current action from a long-term perspective, rather than focusing solely on the immediate reward. Calculating the value function requires analyzing transitions between states.

[0031] The environment, also known as the model, is used to predict the next state and corresponding reward after a state and action are given.

[0032] Based on reinforcement learning, model training through reinforcement learning (RL) and feedback is also proposed. Such a training scheme is also called reinforcement learning from human feedback (RLHF). In the model training system of RLHF, the target model to be trained is also called the action model (actor model), which is used to map the state s in the reinforcement learning environment to the action a. The state s in reinforcement learning corresponds to the model input, and the action a corresponds to the model output. The output of the target model is the action logic, which includes the score determined by the target model for each potential action based on the input, and the action with the highest score is determined as the final selection of the target model. The reward model is configured to determine the reward score of the output estimated by the target model for the input. The reward model is used to evaluate the action a and state s of the target model, and determine the reward score based on the quality of the model output. In RLHF, it is expected to maximize the reward score.

[0033] FIG1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In environment 100, a data processing system 110 can utilize an antibody generation model 120 to perform an antibody generation task. In some implementations, processing system 110 can utilize antibody generation model 120 to generate antibodies 112 based on input information 102. Antibodies 112 are sometimes also referred to as "synthetic antibodies."

[0034] In Figure 1, the processing system 110 can be any type of device with computing capabilities, including a terminal device or a server device. The terminal device can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. The server device can, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like.

[0035] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0036] There are generally two modeling schemes for designing antibodies against specific antigens: (1) adding information about the antigen to the model input so that the model generates specific antibodies based on the antigen; (2) collecting a large number of antibodies with known activity against the target antigen, directly learning from these antibodies, and generating antibodies similar to existing antibodies. In the former scheme, the input information 102 of the antibody generation model includes information about the specific antigen, and sometimes also includes conditional constraints on the antibodies to be generated. For the latter scheme, the input information of the antibody generation model can be arbitrary, because the model has been designed for the target antigen, and the antibodies it outputs are considered to be antibodies with known activity against the target antigen.

[0037] Since the first solution explicitly models the dependency between antibodies and antigens (e.g., the binding of antibodies to targets), it has the ability to design antibodies for unseen antigens. However, model training for this solution is more difficult. The use case of the second solution is relatively simple and of limited significance. This is because in reality, when faced with a certain antigen, the antibody for it is often unknown, and the hope is to use the model to generate antibodies that can bind to the antigen.

[0038] In addition to the differences in task definitions, existing solutions can also be divided into the following categories based on the information used to generate antibodies: (1) sequence design; (2) backbone structure design; and (3) sequence-structure co-design. Because of the close relationship between protein functionality and structure, solutions targeting the third option (i.e., sequence-structure co-design) can utilize complete structural information, and the results are generally more reliable than those of the first two options.

[0039] Several antibody CDR sequence-structure co-design schemes based on diffusion probability models and graph neural networks have been proposed. These schemes learn antibody design by maximizing the probability of the model generating a realistic antibody based on a training set. However, these schemes ignore the interaction between antibodies and antigens. Furthermore, the limited amount of antibody data also limits the model's likelihood learning, resulting in antibodies generated by existing schemes exhibiting strong structural conflicts and lacking antigen binding ability.

[0040] Antibody design schemes are generally limited in two ways. First, for sequence and structure learning objectives, the accuracy of sequence identity and average structural deviation is low, and the interaction between antibody and antigen is also ignored. Furthermore, due to the extremely limited amount of antibody-antigen data, these models are not adequately trained. These limitations make current antibody generation schemes unable to design satisfactory antibodies, especially showing a high degree of irrationality in terms of energy.

[0041] According to an embodiment of the present disclosure, a model optimization scheme based on energy preference is proposed for antibody generation. According to this scheme, for a pre-trained antibody generation model, an energy preference associated with a specific antigen is introduced as an optimization target. Multiple antibodies are generated using the pre-trained antibody generation model, and energy data for each of the multiple antibodies is calculated. Based on the energy preference associated with the specific antigen, energy labels are annotated for the generated multiple antibodies, thereby constructing positive and negative samples. Then, an energy preference-based optimization algorithm is used to fine-tune the antibody generation model based on the constructed samples.

[0042] The antibody generation task has been restructured, dividing it into model pre-training and targeted model optimization phases. Energy is introduced as a generation objective in the second phase of optimization to address deficiencies in the learning objective. Furthermore, energy can be leveraged to combine the computational properties of the software with the model generated in the first phase for large-scale sampling, effectively circumventing the data shortage issue in the antibody generation task. This optimization process effectively circumvents the data shortage issue in the antibody generation task. Leveraging a large amount of energy-labeled data and using a direct preference optimization algorithm allows for more comprehensive model learning, effectively improving the quality of the designed antibodies and ensuring that the generated antibodies are more energy-relevant.

[0043] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0044] FIG2 shows a schematic diagram of a model training system 200 for an antibody generation model 120 according to some embodiments of the present disclosure.

[0045] In the architecture 200, the antibody generation model 120 can be constructed based on various generative model structures. In some embodiments, the antibody generation model 120 can be a diffusion model, a graph neural network, or other model suitable for processing data modalities similar to antibodies and antigens. Of course, any other generative model can be applied to construct the antibody generation model 120. The embodiments of the present disclosure do not specifically limit the model structure. In the following, for the purpose of explanation, the antibody generation model 120 is mainly described based on the diffusion probability model.

[0046] The antibody generation model 120 is configured to generate antibodies that match the antigen based on at least the information of the antigen. As the model input of the antibody generation model 120, the information of the antigen can include the antibody structure. The antibody belongs to a protein, which can be composed of one or more amino acids. The antibody generation model 120 can generate the protein sequence and structure of the antibody, for example, the sequence and structure corresponding to the CDR (particularly CDR-H3) of the antibody can be generated, and CDR-H3 contributes the most to the diversity and specificity of the antibody. In certain embodiments, the model input of the antibody generation model 120 can also include the information of the rest of the antibody except the CDR or CDR-H3, including antibody architecture, other CDRs, etc.

[0047] Since antibodies are composed of amino acids, each amino acid can be represented by the following: i ∈{ACDEFGHIKLMNPQRSTVWY}, C α coordinate and frame orientation O i ∈SO(3), where i=1,...N, N is the number of amino acids in the antibody. If the CDR-H3 to be generated by the antibody generation model 120 has m amino acids, this can be expressed as: The rest of the antigen-antibody pair can be represented as Thus, the antibody generation task of the antibody generation model 120 can be Generated under conditions is represented by the conditional distribution Modeling.

[0048] In some embodiments, the antibody generation model 120 can be based on a diffusion probability model in a generative model, although other types of generative models are also feasible. For better understanding, the diffusion probability model will be briefly introduced below.

[0049] Diffusion probability models are a type of generative model, but their data generation process is based on a pair of Markov processes: a forward diffusion process and a backward denoising process. The forward diffusion process gradually perturbs the input data, generating a static noise distribution through T steps of incremental noise addition. Through model training, the learned backward denoising process performs the opposite process, gradually denoising the samples toward the data distribution, resulting in the desired data. Therefore, the backward denoising process can correspond to the desired antibody generation process, ultimately producing the antibody.

[0050] For antibody production, the forward diffusion process of the diffusion probability model can be expressed as follows:

[0051] in is the noise-free j-th amino acid at step 0, and is the noise-added amino acid in step t. is the noise imposed during the diffusion process, and defines and K is the number of amino acid types. They are category distribution, Gaussian distribution on , and isotropic Gaussian distribution on SO(3). ScaleRot scales the rotation angle about a fixed rotation axis to modify the rotation matrix.

[0052] The backward denoising process (or backward generation process) of the diffusion probability model is to recover the antibody by iterative denoising. The denoising process from step t to step t-1 is: is defined as:

[0053] in is the noise sequence and structure of CDR-H3 of the antibody to be generated in step t, are the parameters of the SE(3) equivariant neural network, and f(·)[j] represents the output corresponding to the jth amino acid. The training objective of the reverse generation process is to minimize the Kullback–Leibler (KL) divergence between the variable distribution p and the posterior distribution q, which is expressed as follows:

[0054] After some algebraic transformations, the above training objective can be simplified, and the reconstruction loss at step t can be obtained as follows:

[0055] in And||·|| F is the matrix Frobenius norm.

[0056] In the above equation (3), the first and third equations can be empirical perturbation denoising. In this way, the total loss function of the antibody generation model 120 can be expressed as By optimizing this loss function, one can start from the noise of the prior distribution and eventually generate antibodies by applying the reverse denoising process.

[0057] In some embodiments, model optimization can also be performed based on RLHF. For the pre-trained antibody generation model 120, the optimization goal is to maximize the reward score of the reward model, as follows:

[0058] where p θ (resp.p ref ) is the distribution of the fixed pre-trained model introduced after the model is fine-tuned, and β is a hyperparameter used to control the KL divergence regularization. The optimal solution to the above optimization objective is:

[0059] In the RLHF-based optimization scheme for the diffusion model, model optimization can be achieved by referring to direct preference optimization (DPO), and the optimization objective can be expressed as:

[0060] where σ(·) is the sigmoid function, and is a pair of “winning” and “losing” data samples (i.e. ), which is determined by the reward score r(·) of the reward model, that is, "Winning" Data Sample Corresponding to the positive samples in the optimization of the antibody generation model 120, the “failed” data samples Corresponds to negative samples in the optimization of the antibody generation model 120 .

[0061] It should be understood that while the above provides a modeling scheme for antibody production based on a diffusion probability model and an example optimization scheme based on RLHF, other modeling schemes and optimization schemes based on diffusion probability models may exist. The modeling schemes for antibody production may also be different in non-diffusion probability models, such as those based on graph neural networks.

[0062] In an embodiment of the present disclosure, the training process of the antibody generation model 120 is divided into two stages. The first stage is a pre-training stage, and the second stage is a model optimization stage based on energy preference.

[0063] In the pre-training stage, the antibody generation model 120 is pre-trained based on a pre-training data set, which includes multiple pairs of mutually matching sample antibodies and sample antigens. In the pre-training stage, the samples in the pre-training data set used can include known antigens and antibodies with binding ability to the antigens. In other words, mutually matching antibodies and antigens refer to the ability of the CDR portion of the antibody to bind to the corresponding antibody. In some embodiments, the pre-training data set can include a SAbDab antigen-antibody data set. Through pre-training, the antibody generation model 120 can have antibody generation capabilities.

[0064] During the model optimization phase, the pre-trained antibody production model 120 can be further optimized. In the disclosed embodiments, antibody energy is used as the antibody production target and introduced into the second-stage optimization to address deficiencies in the model learning objective. Furthermore, the energy calculation characteristics of the software can be combined with the model obtained in the first stage for large-scale sampling, thus circumventing the problem of insufficient antibody data. This approach can effectively improve the quality of designed antibodies, making their energy performance more suitable for demanding applications.

[0065] As shown in Figure 2, based on at least the information of the target antigen, a plurality of antibodies (e.g., antibody 202 and antibody 204) are generated using the current antibody generation model 120. In certain embodiments, the antibody generation model 120 can be an optimization performed for a specific target antigen. In this way, the information of the target antigen is utilized to generate antibodies for constructing samples in each round of fine-tuning. In certain embodiments, the antibody generation model 120 can be optimized for a plurality of antigens, so the information of the antigen used in the fine-tuning of different rounds will be different. In addition to the information of the target antigen, as mentioned above, when generating antibodies, the model input can also include other constraint information for the generated antibodies, such as antibody architecture, other CDRs, etc. The model output of the antibody generation model 120 can include amino acid type, coordinates, frame orientation, etc. In this way, based on the model output, the sequence and / or structure of each antibody can be determined.

[0066] During the optimization phase, energy preference optimization is performed using a large amount of antigen-antibody energy data sampled from the pre-trained antibody generation model 120. Then, during the energy calculation phase 230, the antibody energies of each of the multiple antibodies are determined, and based on the antibody energies of the multiple antibodies, multiple samples for the antibody generation model 120 are constructed.

[0067] In some embodiments, the antibody energy of each antibody can be calculated by a predetermined energy function, or an energy calculation algorithm. In some embodiments, the energy at the CDR level in the antibody can be calculated. Antibody energy can be calculated using various known or future antibody energy calculation methods. Then, based on the antibody energy, energy labels are marked 210 for each antibody (e.g., antibodies 202 and 204). Specifically, the plurality of samples include at least one positive sample and at least one negative sample, and the energy label corresponding to the antibody in the positive sample indicates that its antibody energy meets the energy preference associated with the target antigen, and the energy label corresponding to the antibody in the negative sample indicates that its antibody energy does not meet the energy preference.

[0068] The energy bias associated with a target antigen refers to the energy preference for the antibody to meet the expected energy after binding to the target antigen. In antibody design, it is desirable to generate antibodies with lower total energy. Thus, antibodies with lower total energy have energy labels corresponding to positive samples, while antibodies with higher total energy have energy labels corresponding to negative samples.

[0069] In some embodiments, in an RLHF-based optimization scheme, the reward model r(·) can be defined as where ε(·) is the energy function, is a temperature parameter (which is a pre-settable hyperparameter). In some embodiments, for the diffusion model, the total energy of the antibody can be calculated as: That is, the sum of the energy of the amino acids in the antibody.

[0070] The antibody production model 120 is fine-tuned using the plurality of samples according to an optimization target based on the difference between the predicted energy signatures of the antibodies generated by the fine-tuned antibody production model 120 and the energy signatures between the plurality of samples.

[0071] In the field of model training, although there are similar two-stage training models of pre-training and optimization, most require the use of high-quality training data in the second stage to optimize the model through supervised fine-tuning schemes. This approach can also achieve the purpose of further optimizing the model, but it cannot utilize a large amount of low-quality training data. According to an embodiment of the present disclosure, an energy preference-based optimization method is used in the second stage. This method can use a model that has not yet completed training to construct sample pairs, thereby being able to utilize high-quality and low-quality training data at the same time, and can achieve better performance than traditional supervised fine-tuning schemes.

[0072] In some embodiments, during the optimization (fine-tuning) stage, the optimization objective of the antibody generation model 120 is configured so that the difference between the predicted energy labels of the antibodies generated using the fine-tuned antibody generation model 120 and the energy labels corresponding to the positive samples is reduced, and the difference between the predicted energy labels and the energy labels corresponding to the negative samples is increased.

[0073] In some embodiments, if the antibody generation model 120 is based on a diffusion model and the optimization is performed using a reinforcement learning scheme of RLHF, the optimization objective of equation (7) above may be used to perform model fine-tuning.

[0074] In some implementations, under the diffusion model, the DPO optimization scheme can be transformed into optimizing an upper bound through Jensen's inequality. This involves performing direct preference optimization at discrete time steps, where the gradient is equivalent to optimizing the likelihood of positive samples and reducing the likelihood of negative samples. Positive samples are reasonable antibody samples with lower energy (e.g., antibody samples with lower CDR energy), while negative samples are antibody samples with higher energy (e.g., antibody samples with higher CDR energy). Thus, after optimization, the model possesses the ability to generate reasonable antibody CDRs with lower energy.

[0075] In some implementations, due to It is difficult to solve, so we can introduce hidden variables. And use evidence lower bound optimization (ELBO). In particular, L DPO Can be modified to:

[0076] in and

[0077] In some implementations, under the diffusion model, the DPO optimization scheme can be optimized by exploiting Jensen's inequality and the convexity of the -logσ function to optimize L DPO-Diffusion The upper bound of , that is, performing direct preference optimization at discrete time steps, its gradient is equivalent to optimizing the likelihood of positive samples and reducing the likelihood of negative samples. This can be expressed as follows:

[0078] in and and From the data and The reverse generation process of as well as

[0079] In some embodiments, energy preference-based optimization can be directly applied at the level of the entire sample, that is, the model is optimized by calculating the overall energy of the antibody. Sometimes this approach may obscure amino acids with excellent energy performance in antibodies that are slightly inferior overall. To address this, in some embodiments, a more fine-grained energy preference-based optimization is proposed, that is, energy preference optimization at the amino acid level. This can greatly improve the learning efficiency of the second stage. Energy preference usually defines rewards at the level of the entire antibody sample, while the energy of the antibody CDR region is essentially the sum of the energy at the amino acid level.

[0080] When calculating energy, the antibody energy of each of the multiple antibodies includes the energy at the amino acid residue (residue) level of the antibody. Accordingly, the energy preference can be measured at the amino acid level, and the energy label annotation at the amino acid residue level can be performed, such as the energy label annotations 210-1, 210-2, 210-3, ... 210-N (collectively referred to as energy label annotation 210) at each amino acid level as shown in Figure 2. In this way, for each of the multiple amino acids of the antibody, it can be determined based on the energy of the amino acid that the amino acid is a positive sample that satisfies the energy preference at the amino acid level, or a negative sample that does not satisfy the energy preference at the amino acid residue level. The optimization target for the antibody generation model 120 includes an optimization target at the amino acid residue level, which is configured to make a difference between the energy labels at the amino acid residue level of the antibody generated by the fine-tuned antibody generation model 120 and the amino acid residue level of the multiple samples.

[0081] In order to perform energy preference optimization at the amino acid level, in some embodiments, Jensen's inequality is used to scale the energy preference optimization loss under the original antibody generation model to an upper bound. This upper bound is the energy preference optimization of the diffusion model at both the discrete time step and amino acid level. Specifically, by using Jensen's inequality and the convexity of the -logσ function, the optimization target at the amino acid level (which is Upper bound of ):

[0082] The fine-tunable parameter set θ for the antibody generation model 120 of Lresidue-DPO-Diffusion can be calculated as follows:

[0083] as well as

[0084] in It can be regarded as the current policy p θ The estimated reward score of It is simply expressed as follows:

[0085] It can be seen that the gradient is reweighted using the estimated reward score of the entire antibody The gradient The reweighting is performed using the estimated reward scores of the amino acids themselves. It can increase the likelihood of all amino acids in the positive sample (winning sample) as a whole, and reduce the likelihood of all amino acids in the negative sample (failure sample). We can fully exploit residue-level signals from the estimated reward scores to efficiently optimize antibodies. Thanks to the energy information at the amino acid level, the model gradients are more reasonable and the optimization converges faster for the new upper bound loss.

[0086] According to experience, directly optimizing the total energy of the antibody will lead to some unwanted "shortcuts". Specifically, in some cases, repulsive force dominates the energy of the antibody, so the antibody generation model will be optimized to push the antibody to the farthest place from the antigen so that the repulsive force can be reduced during the optimization process, and may eventually fall into a local minimum. This effectively reduces the repulsive force, but it may also completely eliminate the attraction between the antibody and the antigen, which will seriously damage the function of the antibody. In some embodiments, when antibody energy is used as the optimization target, the antibody energy can be decomposed into two aspects: the rationality of its own structure and the functionality of antigen binding. Specifically, the antibody energy corresponding to each of the multiple antibodies includes the antibody energy corresponding to multiple energy types. The first energy type includes the total energy of the antibody, which can measure the rationality of the antibody's own structure. The second energy type can include the attraction between the antibody and the target antigen, and the third energy type can include the repulsive force between the antibody and the antigen. The attractive force and the repulsive force can be used to measure the ability of the antibody to bind to the antigen. It is expected that the attraction (or attractive energy) between the antibody and the target antigen is stronger and the repulsive force (repulsive energy) is weaker.

[0087] In some embodiments, for an antibody, in the energy function used (e.g., Rosetta energy function), the energy on each amino acid can be roughly divided into attractive force and repulsive force. In order to avoid the model from falling into a local optimum by focusing on optimizing the numerically dominant repulsive force in a shortcut manner, the energy of the antibody is split into the total antibody energy, attractive force and repulsive force. As shown in FIG2 , the total energy 212 (expressed as CDR E ) of each of the antibodies 202 and 204 can be calculated for each amino acid. Total ), with antigen attraction 214 (denoted as CDR-Ag E Attraction ), and the repulsive force 216 with the antigen (expressed as CDR-Ag E Repulsion ).

[0088] Although three example energy types are given here, in fact, antibody energy can be decomposed into more other types of energy. In the case of non-generality, the energy calculation formula can be defined as where V is the number of energy types, and is the weight of the v-th energy type.

[0089] For multiple energy types of antibodies, corresponding energy labels can be labeled based on the energies corresponding to different energy types (indicating whether the antibody is a positive sample or a negative sample under the energy type). In this way, the multiple samples constructed include positive samples and negative samples corresponding to multiple energy types. The energy labels corresponding to the antibodies in the positive samples corresponding to each energy type indicate that the antibody energy meets the energy preference associated with the target antigen under the energy type, and the energy labels corresponding to the antibodies in the negative samples indicate that the antibody energy does not meet the energy preference under the energy type.

[0090] When the antibody energy is decomposed into multiple energy types, the optimization objective includes multiple sub-optimization objectives corresponding to each of the multiple energy types. The sub-optimization objective corresponding to each energy type is configured to reduce the difference between the predicted energy signature corresponding to the predicted antibody energy of the antibody generated by the fine-tuned antibody generation model 120 and the energy signature corresponding to the positive sample corresponding to that energy type, and to increase the difference between the predicted energy signature and the energy signature corresponding to the negative sample corresponding to that energy type. By optimizing the three energy types separately, the model can balance the rationality from a comprehensive perspective, while generating CDRs that can reasonably interact with the antigen.

[0091] For each energy type ε v (·), the energy preference gradient corresponding to the energy type can be calculated in a similar way as in formula (12) above For example, for a total energy 212 (denoted as CDR E Total ), with antigen attraction 214 (denoted as CDR-Ag E Attraction ), and the repulsive force with the antigen 216, can determine the gradient at the amino acid residue level. and The parameter set of the antibody generation model 120 can then be adjusted based on the gradients of the individual amino acid residues.

[0092] Certain types of energy and properties in antibody-antigen binding are often mutually exclusive, and optimizing one property may lead to the degradation of another. The challenge in antibody design is generating CDRs that interact with the antigen in a rational manner, i.e., balancing attractive and repulsive forces. For example, it is generally desirable to have stronger attractive forces and weaker repulsive forces between the antibody and the target antigen, so that the total energy after binding is reduced, and the designed antibody is more effective against the antigen.

[0093] How to make the optimized antibody improve in multiple properties at the same time, or optimize another property while ensuring that a certain property remains unchanged, is a problem to be noted in the model optimization process. In some embodiments of the present disclosure, in order to avoid the conflict caused by optimizing attraction and repulsion at the same time, by introducing the gradient conflict solution into the training of the model, by eliminating the conflict between the gradients corresponding to each property, the above-mentioned mutually exclusive problem is solved to a certain extent. In some embodiments, the energy preferences associated with the target antigen under two or more energy types in the calculated multiple energy types conflict with each other. When performing fine-tuning for antibody generation model 120, the gradient conflict problem is solved by gradient trimming 220. In gradient trimming 220, if the energy preferences corresponding to the first energy type and the second energy type conflict with each other (for example, the conflict between attraction and repulsion), for the first sub-optimization target corresponding to the first energy type and the second sub-optimization target corresponding to the second energy type, the first sub-optimization target and the first sub-optimization target are calculated respectively with respect to the first gradient and the second gradient in the fine-tunable parameter set of antibody generation model 120. By gradient projection, the first gradient is projected to the normal plane corresponding to the second gradient, and the second gradient is projected to the normal plane corresponding to the first gradient. Then, fine-tuning of the antibody generation model 120 may be performed based on at least the first gradient and the second gradient after being projected on the normal plane.

[0094] In some embodiments, for each energy type ε v (·), calculate the gradient After that, the gradient can be projected (in random order) into the normal plane of other gradients (if these gradients conflict). This process can be expressed as follows:

[0095] Where v∈{1, ..., V}, and u=Shuffle(1, ..., V), V is the number of energy types, and Shuffle means random selection.

[0096] The above describes a round of fine-tuning of the antibody generation model 120, starting from a pre-trained antibody generation model 120. Fine-tuning the antibody generation model 120 can include multiple rounds of iterations. In the first round, antibodies are generated using the pre-trained antibody generation model 120, and samples are constructed using the process described above to fine-tune the antibody generation model 120. In the next round, the fine-tuned antibody generation model 120 is continued to be used to generate new antibodies, and samples are constructed using the process described above to further fine-tune the antibody generation model 120. If the antibody generation model 120 is to be optimized for a specific target antigen, then in each round of iteration, new antibodies are generated using information about the target antigen, and the energy labels of the antibodies are also labeled based on the energy preference corresponding to the target antigen. Through the energy optimization-based fine-tuning process, the antibody generation model 120 can continue to be optimized on more training data with energy labels, achieving more comprehensive learning. Through multiple rounds of iterative fine-tuning, the antibody generation model 120 can reach the final optimization goal and produce antibodies that are more energy-relevant to actual application requirements.

[0097] After the optimization is completed, the antibody generation model 120 can be used to generate one or more antibodies against the target antigen based on the target antigen information (and possibly other input information). In the antibody generation stage, the antibody framework region and target antigen information (e.g., amino acid type and coordinates of the target antigen) can be input to the antibody generation model 120 to generate antibodies that match the target antigen, such as outputting the corresponding CDR sequence and structure of the antibody.

[0098] Figure 3 shows a schematic diagram of an environment 300 in which embodiments of the present disclosure can be implemented. In the environment 300 of Figure 3, the model is generally shown to involve different stages, including a training stage 302 and an application stage 306. After the training stage is completed, there may also be a testing stage, which is not shown in the figure.

[0099] In the training phase 302, the model training system 310 is configured to perform training of the model 305 using the training data set 332. The model 305 can be, for example, the antibody production model 120 in Figures 1 and 2. At the beginning of training, the model can have initial parameter values. The training process is to update the parameter values ​​of the model 305 to the desired values ​​based on the training data. The model training system 310 can be configured to implement the model training system 200 of Figure 2.

[0100] In the application phase 306, the obtained model 305 has trained parameter values ​​and can be provided to the model application system 330 for use. In the application phase 306, the model 305 can be used to process the corresponding target input 332 in the actual scenario and provide the corresponding target output 334. The model application system 330 can be configured to implement the data processing system 110 of FIG. 1, which is configured to apply the antibody production model 120 trained by the model training system 200.

[0101] In Figure 3, the model training system 310 and the model application system 330 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices may include any type of mobile, fixed, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0102] It should be understood that the components and arrangements in the environment 300 shown in FIG3 are merely examples, and a computing system suitable for implementing the exemplary implementations described in the present disclosure may include one or more different components, other components, and / or different arrangements. For example, although shown as separate, the model training system 310 and the model application system 330 may be integrated into the same system or device. Implementations of the present disclosure are not limited in this respect.

[0103] FIG4 shows a flowchart of a process 400 for model training according to some embodiments of the present disclosure. Process 400 can be implemented at model training system 200 of FIG2 or model training system 310 of FIG3. For ease of discussion, process 400 will be described with reference to FIG2.

[0104] In block 410 , the model training system 200 generates a plurality of first antibodies based on information of the target antigen using the current antibody generation model.

[0105] At block 420 , the model training system 200 determines the antibody energy of each of the plurality of first antibodies.

[0106] In box 430, the model training system 200 constructs multiple first samples for the antibody generation model based on the antibody energy of the multiple first antibodies, and the multiple first samples include at least one positive sample and at least one negative sample. The energy label corresponding to the first antibody in the positive sample indicates that its antibody energy satisfies the energy preference associated with the target antigen, and the energy label corresponding to the first antibody in the negative sample indicates that its antibody energy does not satisfy the energy preference.

[0107] In block 440 , the model training system 200 uses the plurality of first samples to fine-tune the antibody production model according to an optimization objective, where the optimization objective is based on the difference between the predicted energy signatures of the antibodies generated by the fine-tuned antibody production model and the energy signatures of the plurality of first samples.

[0108] In some embodiments, the optimization objective is configured such that the difference between the predicted energy labels of antibodies generated using the fine-tuned antibody generation model and the energy labels corresponding to the positive samples is reduced, and the difference between the predicted energy labels and the energy labels corresponding to the negative samples is increased.

[0109] In some embodiments, the antibody energies corresponding to each of the plurality of first antibodies include antibody energies corresponding to a plurality of energy types. In some embodiments, the plurality of first samples include positive samples and negative samples corresponding to a plurality of energy types, respectively. The energy label corresponding to the first antibody in the positive sample corresponding to each energy type indicates that the antibody energy satisfies the energy preference associated with the target antigen under the energy type, and the energy label corresponding to the first antibody in the negative sample indicates that the antibody energy does not satisfy the energy preference under the energy type.

[0110] In some embodiments, the optimization objective includes multiple sub-optimization objectives corresponding to multiple energy types, and the sub-optimization objective corresponding to each energy type is configured to reduce the difference between the predicted energy label corresponding to the predicted antibody energy of the antibody generated by the fine-tuned antibody generation model and the energy label corresponding to the positive sample corresponding to the energy type, and increase the difference between the predicted energy label and the energy label corresponding to the negative sample corresponding to the energy type.

[0111] In some embodiments, the energy preferences associated with the target antigen under at least a first energy type and a second energy type in the plurality of energy types conflict with each other. In some embodiments, fine-tuning the antibody production model according to the optimization objective comprises: calculating, for a first sub-optimization objective corresponding to the first energy type and a second sub-optimization objective corresponding to the second energy type, a first gradient and a second gradient of the first sub-optimization objective in a set of fine-tunable parameters of the antibody production model, respectively; projecting the first gradient onto a normal plane corresponding to the second gradient and projecting the second gradient onto the normal plane corresponding to the first gradient by gradient projection; and fine-tuning the antibody production model based on at least the first gradient and the second gradient after being projected onto the normal plane.

[0112] In some embodiments, the antibody energy of each of the plurality of first antibodies comprises an amino acid residue-level energy of the antibody. In some embodiments, the optimization target comprises an amino acid residue-level optimization target, wherein the amino acid residue-level optimization target is configured to achieve a difference between the amino acid residue-level predicted energy signature of the antibody generated using the fine-tuned antibody generation model and the amino acid residue-level energy signatures of the plurality of first samples.

[0113] In some embodiments, the antibody energy corresponding to the multiple energy types includes at least one of the following: the total energy of the antibody, the attraction between the antibody and the target antigen, and the repulsion between the antibody and the antigen.

[0114] In some embodiments, process 400 also includes: generating multiple second antibodies based on information of the target antigen using the fine-tuned antibody generation model; determining the antibody energy of each of the multiple second antibodies; constructing multiple second samples for the antibody generation model based on the antibody energy of the multiple second antibodies, the multiple second samples including at least one positive sample and at least one negative sample, the energy label corresponding to the second antibody in the positive sample indicates that its antibody energy meets the energy preference associated with the target antigen, and the energy label corresponding to the second antibody in the negative sample indicates that its antibody energy does not meet the energy preference; and using the multiple second samples, further fine-tuning the antibody generation model according to the optimization target.

[0115] In some embodiments, the current antibody generation model is pre-trained based on a pre-training dataset, which includes multiple pairs of mutually matched sample antibodies and sample antigens.

[0116] In some embodiments, the antibody production model is based on a diffusion probability model.

[0117] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG5 shows a block diagram of an apparatus 500 for model training according to certain embodiments of the present disclosure. The apparatus 500 can be implemented as or included in the data processing system 110. The various modules / components in the apparatus 500 can be implemented by hardware, software, firmware, or any combination thereof.

[0118] As shown in FIG5 , the apparatus 500 includes an antibody generation module 510 configured to generate a plurality of first antibodies based on information of the target antigen using a current antibody generation model; and an energy determination module 520 configured to respectively determine the antibody energy of each of the plurality of first antibodies.

[0119] The apparatus 500 further includes a sample construction module 530 configured to construct a plurality of first samples for the antibody generation model based on the antibody energies of the plurality of first antibodies. The plurality of first samples include at least one positive sample and at least one negative sample, wherein the energy label corresponding to the first antibody in the positive sample indicates that the antibody energy satisfies the energy preference associated with the target antigen, and the energy label corresponding to the first antibody in the negative sample indicates that the antibody energy does not satisfy the energy preference.

[0120] The apparatus 500 further includes a model fine-tuning module 540 configured to fine-tune the antibody production model using the multiple first samples according to an optimization objective, wherein the optimization objective is based on the difference between the predicted energy signatures of the antibodies generated by the fine-tuned antibody production model and the energy signatures between the multiple first samples.

[0121] In some embodiments, the optimization objective is configured such that the difference between the predicted energy labels of antibodies generated using the fine-tuned antibody generation model and the energy labels corresponding to the positive samples is reduced, and the difference between the predicted energy labels and the energy labels corresponding to the negative samples is increased.

[0122] In some embodiments, the antibody energies corresponding to each of the plurality of first antibodies include antibody energies corresponding to a plurality of energy types. In some embodiments, the plurality of first samples include positive samples and negative samples corresponding to a plurality of energy types, respectively. The energy label corresponding to the first antibody in the positive sample corresponding to each energy type indicates that the antibody energy satisfies the energy preference associated with the target antigen under the energy type, and the energy label corresponding to the first antibody in the negative sample indicates that the antibody energy does not satisfy the energy preference under the energy type.

[0123] In some embodiments, the optimization objective includes multiple sub-optimization objectives corresponding to multiple energy types, and the sub-optimization objective corresponding to each energy type is configured to reduce the difference between the predicted energy label corresponding to the predicted antibody energy of the antibody generated by the fine-tuned antibody generation model and the energy label corresponding to the positive sample corresponding to the energy type, and increase the difference between the predicted energy label and the energy label corresponding to the negative sample corresponding to the energy type.

[0124] In some embodiments, the energy preferences associated with the target antigen under at least the first energy type and the second energy type in the plurality of energy types conflict with each other. In some embodiments, the model fine-tuning module 540 includes: a gradient calculation module configured to calculate, for a first sub-optimization objective corresponding to the first energy type and a second sub-optimization objective corresponding to the second energy type, a first gradient and a second gradient of the first sub-optimization objective in a set of fine-tunable parameters of the antibody generation model, respectively; a gradient projection module configured to project the first gradient onto a normal plane corresponding to the second gradient and project the second gradient onto the normal plane corresponding to the first gradient through gradient projection; and a projection-based fine-tuning module configured to perform fine-tuning on the antibody generation model based on at least the first gradient and the second gradient after being projected onto the normal plane.

[0125] In some embodiments, the antibody energy of each of the plurality of first antibodies comprises an amino acid residue-level energy of the antibody. In some embodiments, the optimization target comprises an amino acid residue-level optimization target, wherein the amino acid residue-level optimization target is configured to achieve a difference between the amino acid residue-level predicted energy signature of the antibody generated using the fine-tuned antibody generation model and the amino acid residue-level energy signatures of the plurality of first samples.

[0126] In some embodiments, the antibody energy corresponding to the multiple energy types includes at least one of the following: the total energy of the antibody, the attraction between the antibody and the target antigen, and the repulsion between the antibody and the antigen.

[0127] In some embodiments, the device 500 also includes: a second antibody generation module, configured to generate multiple second antibodies based on information of the target antigen using the fine-tuned antibody generation model; a second energy determination module, configured to respectively determine the antibody energy of each of the multiple second antibodies; a second sample construction module, configured to construct multiple second samples for the antibody generation model based on the antibody energy of the multiple second antibodies, the multiple second samples including at least one positive sample and at least one negative sample, the energy label corresponding to the second antibody in the positive sample indicates that its antibody energy meets the energy preference associated with the target antigen, and the energy label corresponding to the second antibody in the negative sample indicates that its antibody energy does not meet the energy preference; and a second model fine-tuning module, configured to use the multiple second samples to perform further fine-tuning on the antibody generation model according to the optimization objective.

[0128] In some embodiments, the current antibody generation model is pre-trained based on a pre-training dataset, which includes multiple pairs of mutually matched sample antibodies and sample antigens.

[0129] In some embodiments, the antibody production model is based on a diffusion probability model.

[0130] The units and / or modules included in the device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 500 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0131] FIG6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in FIG6 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 600 shown in FIG6 can be used to implement the data processing system 110 of FIG1 , the model training system 200 of FIG2 , the model training system 310 or the model application system 320 of FIG3 , or the apparatus 500 of FIG5 .

[0132] As shown in FIG6 , electronic device 600 is a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 600.

[0133] The electronic device 600 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 600.

[0134] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG6 , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0135] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 600 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0136] The input device 650 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 660 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 600, or with any device that allows the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0137] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0138] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0139] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0140] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0141] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0142] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for model training, comprising: Based on the information of the target antigen, multiple first antibodies are generated using the current antibody generation model; determining the antibody energy of each of the plurality of first antibodies; Based on the antibody energies of the multiple first antibodies, constructing multiple first samples for the antibody generation model, the multiple first samples including at least one positive sample and at least one negative sample, the energy label corresponding to the first antibody in the positive sample indicates that its antibody energy satisfies the energy preference associated with the target antigen, and the energy label corresponding to the first antibody in the negative sample indicates that its antibody energy does not satisfy the energy preference; as well as The antibody production model is fine-tuned using the plurality of first samples according to an optimization target based on differences in predicted energy signatures of antibodies generated by the fine-tuned antibody production model and energy signatures between the plurality of first samples.

2. According to claim 1, wherein the optimization objective is configured so that the difference between the predicted energy label of the antibody generated using the fine-tuned antibody generation model and the energy label corresponding to the positive sample is reduced, and the difference between the predicted energy label and the energy label corresponding to the negative sample is increased.

3. The method according to claim 1, wherein the antibody energies corresponding to each of the plurality of first antibodies include antibody energies corresponding to a plurality of energy types, and The multiple first samples include positive samples and negative samples corresponding to the multiple energy types respectively, the energy label corresponding to the first antibody in the positive sample corresponding to each energy type indicates that its antibody energy satisfies the energy preference associated with the target antigen under the energy type, and the energy label corresponding to the first antibody in the negative sample indicates that its antibody energy does not satisfy the energy preference under the energy type.

4. The method according to claim 3, wherein the optimization objective includes multiple sub-optimization objectives corresponding to the multiple energy types, and the sub-optimization objective corresponding to each energy type is configured to reduce the difference between the predicted energy label corresponding to the predicted antibody energy of the antibody generated by the fine-tuned antibody generation model and the energy label corresponding to the positive sample corresponding to the energy type, and increase the difference between the predicted energy label and the energy label corresponding to the negative sample corresponding to the energy type.

5. The method of claim 4, wherein the energy preferences associated with the target antigen under at least a first energy type and a second energy type in the plurality of energy types conflict with each other, and wherein fine-tuning the antibody production model according to an optimization objective comprises: For a first sub-optimization objective corresponding to the first energy type and a second sub-optimization objective corresponding to the second energy type, respectively calculating a first gradient and a second gradient of the first sub-optimization objective with respect to a set of fine-tunable parameters of the antibody production model; Projecting the first gradient onto a normal plane corresponding to the second gradient, and projecting the second gradient onto a normal plane corresponding to the first gradient, through gradient projection; as well as Fine-tuning the antibody production model is performed based on at least the first gradient and the second gradient after being projected onto a normal plane.

6. The method according to claim 1, wherein the antibody energy of each of the plurality of first antibodies comprises energy at the level of amino acid residues of the antibody, and The optimization target includes an optimization target at the amino acid residue level, and the optimization target at the amino acid residue level is configured to make the difference between the predicted energy signature at the amino acid residue level of the antibody generated based on the fine-tuned antibody production model and the energy signature at the amino acid residue level of the multiple first samples.

7. The method according to claim 3, wherein the antibody energies corresponding to the multiple energy types include at least one of the following: The total energy of the antibody, The attraction between the antibody and the target antigen, The repulsive force between the antibody and the antigen.

8. The method according to claim 1, further comprising: Based on the information of the target antigen, a plurality of second antibodies are generated using the fine-tuned antibody production model; respectively determining the antibody energy of each of the plurality of second antibodies; Based on the antibody energy of the multiple second antibodies, a model for the antibody production is constructed. a plurality of second samples, the plurality of second samples including at least one positive sample and at least one negative sample, the energy label corresponding to the second antibody in the positive sample indicates that the antibody energy satisfies the energy preference associated with the target antigen, and the energy label corresponding to the second antibody in the negative sample indicates that the antibody energy does not satisfy the energy preference; as well as The antibody production model is further fine-tuned using the plurality of second samples according to the optimization objective. 9 . The method according to claim 1 , wherein the current antibody generation model is pre-trained based on a pre-training dataset, wherein the pre-training dataset includes a plurality of pairs of mutually matched sample antibodies and sample antigens.

10. The method according to any one of claims 1 to 9, wherein the antibody production model is based on a diffusion probability model.

11. A device for model training, comprising: an antibody generation module, configured to generate a plurality of first antibodies based on information of the target antigen using a current antibody generation model; an energy determination module, configured to respectively determine the antibody energy of each of the plurality of first antibodies; a sample construction module configured to construct a plurality of first samples for the antibody generation model based on the antibody energies of the plurality of first antibodies, the plurality of first samples including at least one positive sample and at least one negative sample, the energy label corresponding to the first antibody in the positive sample indicating that the antibody energy thereof satisfies the energy preference associated with the target antigen, and the energy label corresponding to the first antibody in the negative sample indicating that the antibody energy thereof does not satisfy the energy preference; as well as The model fine-tuning module is configured to use the multiple first samples to fine-tune the antibody production model according to an optimization target, wherein the optimization target is based on the difference between the predicted energy signatures of the antibodies generated by the fine-tuned antibody production model and the energy signatures of the multiple first samples.

12. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processor The electronic device comprises a processing unit and stores instructions for execution by the at least one processing unit, wherein the instructions, when executed by the at least one processing unit, cause the electronic device to execute the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 10.

14. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Cellular based fret assay for the determination of simultaneous binding

    CN108139394A

  • Attention mechanism-based antibody non-sequencing prediction method and device

    CN114822696A

  • Antibody affinity prediction method and system based on heterogeneous protein characteristic hybridization

    CN116030879A

  • Data processing method and related device

    CN116959125A

  • Machine learning techniques for predicting thermostability

    US20230368861A1