Method, device and computer program product for optimization of ligand generation model

By optimizing ligand generation models using Direct Preference Optimization and Diffusion-DPO with multi-granularity control, the method addresses the limitations of existing SBDD models, enhancing molecule diversity and quality in structure-based drug design.

WO2026006986A1PCT designated stage Publication Date: 2026-01-08BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/103125
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-02
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

The challenge in structure-based drug design (SBDD) lies in the lack of high-quality protein-ligand pair data, which restricts the effectiveness of generative models, and existing optimization methods lack flexibility and diversity in molecule design due to fixed model parameters.

Method used

A method is introduced to optimize a ligand generation model by generating ligands, determining a training objective based on superiority comparisons, and updating the model to favor superior ligands, using Direct Preference Optimization (DPO) and Diffusion-DPO, which incorporates multi-granularity control through global and local preference alignments.

Benefits of technology

This approach enhances the flexibility and effectiveness of ligand generation models by improving the diversity and quality of designed molecules, aligning with desired properties through dynamic model updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024103125_08012026_PF_FP_ABST
    Figure CN2024103125_08012026_PF_FP_ABST
Patent Text Reader

Abstract

A method is proposed for optimizing a ligand generation model. In the method, a first ligand and a second ligand for binding to a protein target are generated by using the ligand generation model. In accordance with a determination that the first ligand is superior to the second ligand in terms of at least one property, a training objective is determined. The training objective causes the ligand generation model to have a propensity to generate the first ligand as compared to the second ligand. The ligand generation model is updated according to the training objective.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD, DEVICE AND COMPUTER PROGRAM PRODUCT FOR OPTIMIZATION OF LIGAND GENERATION MODELFIELD

[0001] The present disclosure generally relates to the field of computer, and more specifically, to method, device, and computer program product for optimization of a ligand generation model.BACKGROUND

[0002] Structure-based drug design (SBDD) represents a strategic approach within medicinal chemistry and pharmaceutical research, leveraging the three-dimensional structures of biomolecules to guide the creation and optimization of novel therapeutic agents. The primary aim of SBDD is to engineer molecules that effectively bind to specific protein targets. Recently, this challenge has been approached as a conditional generative task, adopting a data-driven methodology and incorporating advanced generative models that utilize geometric deep learning.SUMMARY

[0003] In a first aspect of the present disclosure, there is provided a method of optimizing a ligand generation model. The method includes generating, by using the ligand generation model, a first ligand and a second ligand for binding to a protein target; in accordance with a determination that the first ligand is superior to the second ligand in terms of at least one property, determining a training objective which causes the ligand generation model to have a propensity to generate the first ligand as compared to the second ligand; and updating the ligand generation model according to the training objective.

[0004] In a second aspect of the present disclosure, there is provided an electronic device. The electronic device comprises: a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements a method according to the first aspect of the present disclosure.

[0005] In a third aspect of the present disclosure, there is provided a computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform a method according to the first aspect of the present disclosure.

[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to  identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Through the more detailed description of some embodiments of the present disclosure in the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, wherein the same reference generally refers to the same components in the embodiments of the present disclosure.

[0008] FIG. 1 illustrates an example environment in which example embodiments of the present disclosure can be implemented;

[0009] FIG. 2 illustrates a schematic diagram of an example architecture for optimizing a ligand generation model according to some embodiments of the present disclosure;

[0010] FIG. 3 illustrates a schematic diagram of an example architecture for optimizing a ligand generation model according to some embodiments of the present disclosure;

[0011] FIG. 4 illustrates an example flowchart of a method of optimizing a ligand generation model according to some embodiments of the present disclosure; and

[0012] FIG. 5 illustrates a block diagram of an electronic device in which various embodiments of the present disclosure can be implemented.DETAILED DESCRIPTION

[0013] Principle of the present disclosure will now be described with reference to some embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below.

[0014] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0015] References in the present disclosure to “one embodiment, ” “an embodiment, ” “an example embodiment, ” and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that it is  within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0016] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.

[0018] Principle of the present disclosure will now be described with reference to some embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below. In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0019] It may be understood that data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with requirements of corresponding laws and regulations and relevant rules.

[0020] It may be understood that, before using the technical solutions disclosed in various embodiment of the present disclosure, the user should be informed of the type, scope of use, and use scenario of the information involved in the present disclosure in an appropriate manner in accordance with relevant laws and regulations, and the user’s authorization should be obtained.

[0021] For example, in response to receiving an active request from the user, prompt information is sent to the user to explicitly inform the user that the requested operation will need to acquire and use the user’s information. Therefore, the user may independently choose, according to the prompt information, whether to provide the information to software or hardware  such as electronic devices, applications, servers, or storage media that perform operations of the technical solutions of the present disclosure.

[0022] As an optional but non-limiting implementation, in response to receiving an active request from the user, the way of sending prompt information to the user, for example, may include a pop-up window, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose “agree” or “disagree” to provide the information to the electronic device.

[0023] It may be understood that the above process of notifying and obtaining the user authorization is only illustrative and does not limit the implementation of the present disclosure. Other methods that satisfy relevant laws and regulations are also applicable to the implementation of the present disclosure.

[0024] As used herein, the term “model” is referred to as an association between an input and an output learned from training data, and thus a corresponding output may be generated for a given input after the training. The generation of the model may be based on a machine learning technique. In general, a machine learning model may be built, which receives input information and makes predictions based on the input information. For example, a classification model may predict a class of the input information among a predetermined set of classes. As used herein, “model” may also be referred to as “machine learning model” , “learning model” , “machine learning network” , or “learning network, ” which are used interchangeably herein.

[0025] As used herein, the expression “generating a ligand” or the like does not mean that the ligand is produced or fabricated physically. Rather, it means that the structure of the ligand is determined or the like.

[0026] Diffusion models have been employed to simulate the distribution of ligand atoms by type and position. However, the lack of high-quality protein-ligand pair data presents a major challenge for advancing generative models in the realm of SBDD. Deep learning’s effectiveness typically depends on extensive datasets. However, assembling datasets of protein-ligand interactions is difficult and restricted, given the complexity and the resource-heavy nature of the experimental processes involved.

[0027] Notably, a widely-used dataset for SBDD consists of ligands that are docked into multiple similar binding pockets across the Protein Data Bank using docking software. This may be regarded as a form of data augmentation; while it expands the dataset’s size, it may unavoidably introduce some low-quality data. Besides, the number of unique ligands remain the same before and after this data augmentation. Further, the ligands in the dataset have moderate binding affinities, which do not meet the stringent demands of drug design. To address the aforementioned  challenge, another solution provides a straightforward method for searching molecules with desired properties in the extensive chemical space. However, pure searching or optimization methods lack generative capabilities and fall short in the diversity of the designed molecules. A further solution proposes to integrate conditional diffusion models with iterative optimization by providing molecular substructures as conditions of conditional diffusion models and iteratively replacing the substructures with better ones. This solution achieves better properties and maintains certain diversity. Nonetheless, the performance of this solution is still limited due to fixed model parameters during the optimization process.

[0028] Embodiments of the present disclosure propose solutions for optimization of a ligand generation model. According to embodiments of the present disclosure, a first ligand and a second ligand for binding to a protein target are generated by using the ligand generation model. If the first ligand is superior to the second ligand in terms of at least one property, a training objective is determined. The training objective causes the ligand generation model to have a propensity to generate the first ligand as compared to the second ligand. The ligand generation model is updated according to the training objective. In this way, flexibility and effectiveness of the optimization of the ligand generation model can be improved.

[0029] Theoretical foundation of the present disclosure and example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0030] FIG. 1 illustrates a block diagram of an example environment 100 in which various embodiments of the present disclosure may be implemented. In the environment 100 of FIG. 1, an electronic device 120 receives a training dataset 101 and train a ligand generation model 110, for example optimizing the ligand generation model 110. The training dataset 101 may include information concerning ligands for binding to a protein target.

[0031] In FIG. 1, the electronic device 120 may include any computing system with computing capability, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices may include any type of mobile terminals, fixed terminals, or portable terminals, including mobile phones, desktop computers, laptops, netbooks, tablets, media computers, multimedia tablets, or any combination of the aforementioned, including accessories and peripherals of these devices or any combination thereof. Servers include but are not limited to mainframe, edge computing nodes, computing devices in cloud environment, etc.

[0032] To better understand the embodiments of the present disclosure, some preliminaries are first described. First, a SBDD task may be defined and a decomposed diffusion model for this task may be introduced.

[0033] In the context of SBDD, generative models are conditioned on the protein binding site, represented as to generate ligands  that bind to this site. Herein,  and are the number of atoms in the protein and ligand, respectively. For both protein and ligand,  represents the coordinates of atoms, the types of atoms, and the bonds between atoms. Herein, h types of atoms (i.e., H, C, N, O, S, Se) and 5 types of bonds (i.e., non-bond, single, double, triple, aromatic) are considered.

[0034] Following the decomposed diffusion model, each ligand is decomposed into fragments  comprising several arms connected by at most one scaffold  Based on the decomposed substructures, informative data-dependent priors  are estimated from atom positions by maximum likelihood estimation. This data-dependent prior enhances the training efficacy of the diffusion model, where is gradually diffused with a fixed schedule{λt}t=1, …, T. Further, αt=1-λt and are denoted. The i-th atom position is shifted to its corresponding prior center:  The noisy data distribution at time t derived from the distribution at time t -1 is computed as follows:

[0035] where Ka and Kb represent the number of atom types and bond types used for featurization. The perturbed structure is then fed into the prediction model, then the reconstruction loss at the time t can be derived from the KL divergence as follows:

[0036] where (x0, v0, b0) , (xt, vt, bt) ,  represent true atoms positions, types, and bond types at time 0, time t, predicted atoms positions, types, and bonds types at time t;  denotes mixed categorical distribution with weight and  The overall loss is with γv,γb as weights of reconstruction loss of atom and bond type. To better illustrate decomposition, a decomposed molecule with the arms highlighted is shown in FIG. 2.

[0037] Reference is now made to FIG. 2, which illustrates an example architecture 200 for optimizing a ligand generation model according to some embodiments of the present disclosure. As shown in FIG. 2, the ligand generation model 110 generates a plurality of ligands for binding to a protein target, for example, a pocket. A ligand pair of the plurality of ligands may be compared with each other in terms of at least one property. The at least one property may include any suitable property for evaluating the ligands. For example, the at least one property may be related to target binding affinity and molecular properties, and molecular conformation. To assess target binding affinity, Vina score, Vina Min and Vina Dock may be used. The vina Score quantifies the direct binding affinity between a molecule and the target protein, Vina Min measures the affinity after local structural optimization via force field, Vina Dock assesses the affinity after re-docking the ligand into the target protein, and High Affinity measures the proportion of generated molecules with a Vina Dock score higher than that of the reference ligand. Regarding molecular properties, drug-likeness (QED) , synthetic accessibility (SA) and diversity may be used.

[0038] In the following, the ligand pair is described by taking the first ligand 210 and the second ligand 220 as an example. By comparing the first ligand 210 and the second ligand 220 in terms of the at least one property, it may be determined that the first ligand 210 is superior to the second ligand in terms of at least one property. In other words, the first ligand 210 is preferred over the second ligand 220. As shown in FIG. 2, the first ligand 210 is shown with “WIN” and the second ligand 220 is shown with “LOSE” . The first ligand 210 and the second ligand 220 may be referred to as “win ligand” and “lose ligand” . Any suitable approach can be used to identify the win and lose ligands and protection scope is not limited in this regard.

[0039] Then, a training objective for the ligand generation model 110 is determined. The training objective is configured to cause the ligand generation model 110 to have a propensity to generate the first ligand compared to the second ligand. In other words, the training objective is used to cause the ligand generation model 110 to have a probability for generating the first ligand 110 higher than a probability for generating the second ligand 120.

[0040] Next, as shown by the dash line, the ligand generation model 110 is updated according to the training objective 201. For example, the pre-trained ligand generation model 110 is fine-tuned or optimized according to the training objective 201.

[0041] In the example architecture, Direct Preference Optimization (DPO) is used for the ligand generation model 110. To better understand the embodiments of the present disclosure, example principles are now described.

[0042] In some embodiments, the ligand generation model 110 may include a diffusion model. In the following, examples may be described with respect to the diffusion model. However, it is to be understood that other models are possible and the concept described with reference to the diffusion model may be extended to other models.

[0043] Despite the diffusion models achieving promising results in modeling existing molecules, it may fail to meet the practical goal of SBDD is to synthesize molecules with the desired properties. By maximizing the reward function, the pre-learned model distribution may be optimized towards regions of higher quality. Drawing on Reinforcement Learning from human / AI feedback (RLHF) , this optimization problem is formulated in SBDD as follows:

[0044] where is the distribution of the training set,  is the reward function that evaluates properties of generated molecules, and β>0 is a hyperparameter for the Kullback-Leibler (KL) divergence regularization, which controls the derivation from the reference model pref for ligand generation. For example, the reference model may a pre-trained version of the ligand generation model 110.

[0045] Following Direct Preference Optimization (DPO) , by leveraging preference data pair  with superior and inferior properties, the DPO training loss is expressed as follows:

[0046] where represents the win ligand, for example, the first ligand 110, and represents the lose ligand, for example, the second ligand 220.

[0047] An idea of DPO in LLM is borrowed, proposing Diffusion-DPO, which can directly optimize the distribution learned by diffusion models with the preference defined over the entire diffusion process. This preference in SBDD may be turned as  Herein,  denotes the diffusion trajectories from the reverse process The training loss for Diffusion-DPO is defined to capture the differences in rewards across these trajectories:

[0048] where and represent trajectories of molecules with superior and inferior properties, respectively. By leveraging inequality to externalize the expectation and replacing  by q, the training loss for Diffusion-DPO is then reformulated as:

[0049] where the training loss as shown in the equation (7) may be an example of the training objective 201.

[0050] In some embodiments, decomposable optimization may be used and accordingly, the training objective may be decomposed to objectives for a global ligand structure and local ligand structure (which is also referred to as a substructure) , respectively. For example, a first objective for a global ligand structure may be determined based on a first propensity difference in generating an entire structure of the first ligand and an entire structure of the second ligand. The first objective is also referred to as a global objective or GLOBALDPO.

[0051] A second objective for a local ligand structure may be determined based on a second propensity difference in generating a first substructure of the first ligand and a second substructure of the second ligand. The first and second substructures corresponds to a same ligand fragment. The second objective is also referred to as a local objective or local DPO loss.

[0052] Then, the training objective may be derived based on the first and second objectives.

[0053] Next, the alignment of pre-trained diffusion models with preferences defined within the decomposed space may be described. Then, it will be highlighted how multi-granularity control enhances the effectiveness and flexibility of the optimization according to the embodiments of the present disclosure.

[0054] Reference is now made to FIG. 3. The architecture 300 shown in FIG. 3 may be considered as an example of the architecture 200. As shown, the ligand generation model 110 is a diffusion model which implements a forward process and a reverse process. In the reverse process, a plurality of ligands for binding to the protein target 305 (for example, a pocket) may be generated. Then, the first ligand 210 and the second ligand 220 may be identified as win and lose, respectively, in terms of at least one property. In the global preference alignment 301, the first objective, for example the global DPO loss, may be calculated. In the local preference alignment 302, the second objective, for example the local DPO loss, may be calculated.

[0055] As shown in FIG. 3, a ligand for binding to the protein target may include a plurality of ligand fragments. As a result, the first ligand 210 may be decomposed into a plurality of substructures 310-1, 310-2, 310-3, 310-4, which is also referred to as a first substructure 310 and first substructures 310. Similarly, the second ligand 220 may be decomposed into a plurality of substructures 320-1, 320-2, 320-3, 320-4, which is also referred to as a second substructure 320  and second substructures 320. A first substructure and a second substructure correspond to a same fragment, as shown by the shading of each substructure. Specifically, in the example of FIG. 3, the first substructure 310-1 and the second substructure 320-1 correspond to the same fragment; the first substructure 310-2 and the second substructure 320-2 correspond to the same fragment; the first substructure 310-3 and the second substructure 320-3 correspond to the same fragment; and the first substructure 310-4 and the second substructure 320-4 correspond to the same fragment.

[0056] In some embodiments, if the ligand for binding to the protein target comprises a plurality of ligand fragments, respective second propensity differences may be determined for the plurality of ligand fragments. Accordingly, the second objective may be determined based on the respective second propensity differences. For example, the local DPO loss may be determined for each fragment based on the first substructure 310 and the second substructure 320 corresponding to this fragment.

[0057] In some embodiments, the first objective may be determined with respect to at least one first property which is non-decomposable at a ligand fragment level, and the second objective may be determined with respect to at least one second property which is decomposable at the ligand fragment level. Examples of non-decomposable property may include the QED and SA. Examples of decomposable property may include the Vina score, Vina Min, and Vina Dock.

[0058] The introduction of de-composition not only induces better evidence lower bound for diffusion model, but also shows potential in molecular optimization with a controllable diffusion model. It is observed that different molecular substructures, especially decomposed arms, may independently contribute to specific properties such as binding affinities and clash scores.

[0059] Leveraging the above observation, some embodiments of the present disclosure propose to extend the decomposition concept to preference optimization with the ligand generation model. Specifically, preference pairs are constructed at the substructure level for objectives that can be decomposed and perform local preference alignment with local DPO loss, which will be introduced in detail below. Such substructure-level preference is more precise and may provide more flexibility, thus potentially improving the optimization efficacy. For properties that are inherently non-decomposable, the molecule-level preference is used, and global preference alignment is performed with Diffusion-DPO, which is termed as global DPO below.

[0060] The following will describe the global DPO and local DPO, respectively.

[0061] Since certain optimization objectives, like Vina Minimize Score, can be calculated as summations of contributions from different substructures, we integrate decomposition into Diffusion-DPO and define the training loss of global DPO as follows:

[0062] where represents the i-th decomposed substructure. In GLOBALDPO, the preferences assigned to decomposed substructures are derived holistically from the preferences of the molecules. In other words, + and - in the substructure indicate that it is decomposed from the winning or losing molecule, regardless of its own properties.

[0063] In some embodiments, to determine the first objective for the global ligand structure, a first propensity metric for the ligand generation model may be determined based on a probability of the ligand generation model to generate the first ligand and a probability of the ligand generation model to generate the second ligand. Then, the first objective may be derived based on the first propensity metric. For example, the term  in equation (7) and the term  in equation (8) may be examples of the first propensity metric. Then, the global DPO loss may be determined according to the equation (7) or the equation (8) .

[0064] Some embodiments of the present disclosure further introduce to construct preference pairs directly with substructures’ properties and define the training loss of local DPO as:

[0065] where

[0066] In Equation (9) ,  represents the reward of the decomposed substructure evaluated at time 0. A (i) defined in Equation (9) is referred as the preference learned by the model.

[0067] In some embodiments, to determine the second objective for the local ligand structure, a second propensity metric for the ligand generation model may be determined based on a probability of the ligand generation model to generate the first substructure and a probability of the ligand generation model to generate the second substructure. A first reward score of the first substructure and a second reward score of the second substructure may be compared. Then, the second objective may be determined based on the second propensity metric and a comparing result of the first reward score and the second reward score. For example, the term A (i) in the equation (9) may be considered as an example of the second propensity metric, and  may be an example of comparing the first reward score and the second reward score. Then, the local DPO loss may be derived according to the equation (9) .

[0068] The first and second reward score may be compared in any suitable manner. In some embodiments, if the first reward score exceeds the second reward score, the second objective may be derived by applying a positive factor (for example, a positive value, such as 1) to the second propensity metric. If the first reward score is below the second reward score, the second objective may be derived by applying a negative factor (for example, a negative value, such as -1) to the second propensity metric. Continuing with the example of the equation (9) , sign () function is used to apply the positive factor or the negative factor to the term A (i) according to the comparison of the reward and the reward

[0069] The advantage of the local DPO becomes evident when considering the consistency of substructure preferences with overall molecular preferences. When there exists a conflict, that is, if the local DPO adjusts accordingly: if the  preference learned by the model is consistent with the preference defined with substructures (A (i) <0, sign (·) <0) , this results in a larger summation within -logσ. It further leads to a smaller loss because -logσ is monotonically decreasing.

[0070] Conversely, if the preference learned by the model is inconsistent with the preference defined with substructures (A (i)>0, sign (·) <0) , it will result in a larger loss. In extreme cases, where the preference learned by the model invariably opposes the preferences of substructures, the local DPO loss sets an upper bound for Diffusion-DPO.

[0071] Thanks to the global DPO loss and local DPO loss introduced above, it is now possible to integrate preferences across different granularities for effective multi-objective optimization.

[0072] In some embodiments, a plurality of properties may be considered, and the training objective may be determined with respect to the plurality of properties. For example, for a decomposable property, the local DPO loss as described above may be determined. For a non-decomposable property, the global DPO loss as described above may be determined. In this case, the final training loss for optimizing the ligand generation model may be a sum of at least one local DPO loss and at least one global DPO loss.

[0073] In an example, the final objective may be expressed as follows:

[0074] where and represent the set of decomposable and non-decomposable properties, separately. The local DPO loss is used to optimize decomposable properties and the global DPO loss is used to optimize non-decomposable properties. Such a dual granularity allows for more precise control over the optimization process and provides greater flexibility in selecting preferences to meet the diverse needs of molecular design.

[0075] The above equations introduce a parameter β, which is a schedule factor.

[0076] In some embodiments, a linear beta schedule may be employed. A sampled schedule factor at a sampled time step may be determined based on a linear relation with respect to a  reference schedule factor at a reference time step. In this case, the training objective is determined based on the sampled schedule factor, for example, as described above with respect to the equations (7) - (9) . In some embodiments, the reference time step is the maximum time step.

[0077] As an example, in diffusion models, the final steps of the reverse process play a crucial role, as they determine the types and positions of the atoms, thereby directly influencing the properties of generated molecules. These steps are also important for optimization, as achieving improved property distributions requires effective alignment with the desired properties. These embodiments of the present disclosure propose a linear beta schedule to improve optimization efficiency and better balance the exploration of novel molecular structures with the adherence to pre-learned prior distribution.

[0078] At each time step t, the schedule factor βt is defined as where βT is the initial beta parameter defined at the last time step T. This progressive scaling reduces the regularization impact of the reference model during the crucial steps of the diffusion process, allowing the model to better align with preferences.

[0079] Last but not least, in the present disclosure, DPO may be used to optimize non-decomposable properties of ligand molecules generated by diffusion models, which is referred to as global DPO. Inspired by the decomposition nature of ligand molecules when they bind to proteins, for the properties that are decomposable (e.g, binding affinity, which is measured by Vina) , decomposed DPO loss, termed as local DPO loss, where preferences are defined over substructures instead of the whole molecules. Additionally, in some embodiments, a linear beta schedule which also improves the performance remarkedly. Embodiments of the present disclosure may be used to various scenarios, for example but not limited to: structure-based molecular generation and structure-based molecular optimization. Under both setting, the proposed optimization of the ligand generation model can significantly outperform baselines, demonstrating its flexibility and effectiveness.

[0080] Example process and device

[0081] FIG. 4 illustrates a flowchart of a method 400 of optimizing a ligand generation model in accordance with some example implementations of the present disclosure. The method 400 may be implemented at the electronic device 120 as illustrated in FIG. 1.

[0082] At a block 410, the electronic device 120 generates, by using the ligand generation model, a first ligand and a second ligand for binding to a protein target. At a block 420, in accordance with a determination that the first ligand is superior to the second ligand in terms of at least one property, the electronic device 120 determines a training objective which causes the ligand generation model to have a propensity to generate the first ligand compared to the second ligand. At a block 430, the electronic device 120 updates the ligand generation model according to the training objective.

[0083] In some embodiments, determining the training objective comprises: determining a first objective for a global ligand structure based on a first propensity difference in generating an entire structure of the first ligand and an entire structure of the second ligand; determining a second objective for a local ligand structure based on a second propensity difference in generating a first substructure of the first ligand and a second substructure of the second ligand, the first and second substructures corresponding to a same ligand fragment; and deriving the training objective based on the first and second objectives.

[0084] In some embodiments, determining the first objective for the global ligand structure comprises: determining a first propensity metric for the ligand generation model based on a probability of the ligand generation model to generate the first ligand and a probability of the ligand generation model to generate the second ligand; and deriving the first objective based on the first propensity metric.

[0085] In some embodiments, determining the second objective for the local ligand structure comprises: determining a second propensity metric for the ligand generation model based on a probability of the ligand generation model to generate the first substructure and a probability of the ligand generation model to generate the second substructure; comparing a first reward score of the first substructure and a second reward score of the second substructure; and determining the second objective based on the second propensity metric and a comparing result.

[0086] In some embodiments, determining the second objective based on the second propensity metric and the comparing result comprises: in response to that the first reward score exceeds the second reward score, deriving the second objective by applying a positive factor to the second propensity metric; and in response to that the first reward score is below the second  reward score, deriving the second objective by applying a negative factor to the second propensity metric.

[0087] In some embodiments, the first objective is determined with respect to at least one first property which is non-decomposable at a ligand fragment level, and the second objective is determined with respect to at least one second property which is decomposable at the ligand fragment level.

[0088] In some embodiments, a ligand for binding to the protein target comprises a plurality of ligand fragments, and respective second propensity differences are determined for the plurality of ligand fragments, and the second objective is determined based on the respective second propensity differences.

[0089] In some embodiments, the ligand generation model comprises a diffusion model, and the method further comprises: determining a sampled schedule factor at a sampled time step based on a linear relation with respect to a reference schedule factor at a reference time step, and wherein the training objective is determined based on the sampled schedule factor.

[0090] In some embodiments, the reference time step is the maximum time step.

[0091] In some embodiments of the present disclosure, there is provided a non-transitory computer program product, the non-transitory computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform a method of optimizing a ligand generation model. The method comprises: generating, by using the ligand generation model, a first ligand and a second ligand for binding to a protein target; in accordance with a determination that the first ligand is superior to the second ligand in terms of at least one property, determining a training objective which causes the ligand generation model to have a propensity to generate the first ligand compared to the second ligand; and updating the ligand generation model according to the training objective.

[0092] FIG. 5 illustrates a block diagram of an electronic device 500 in which various embodiments of the present disclosure can be implemented. It would be appreciated that the electronic device 500 shown in FIG. 5 is merely for purpose of illustration, without suggesting any limitation to the functions and scopes of the present disclosure in any manner. The electronic device 500 may be used to implement the above method 500. As shown in FIG. 5,  the electronic device 500 may be a general-purpose electronic device. The electronic device 500 may at least comprise one or more processors or processing units 510, a memory 520, a storage unit 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560.

[0093] The processing unit 510 may be a physical or virtual processor and can implement various processes based on programs 525 stored in the memory 520. In a multi-processor system, multiple processing units execute computer executable instructions in parallel so as to improve the parallel processing capability of the electronic device 500. The processing unit 510 may also be referred to as a central processing unit (CPU) , a microprocessor, a controller, or a microcontroller.

[0094] The electronic device 500 typically includes various computer storage medium. Such medium can be any medium accessible by the electronic device 500, including, but not limited to, volatile and non-volatile medium, or detachable and non-detachable medium. The memory 520 can be a volatile memory (for example, a register, cache, Random Access Memory (RAM) ) , a non-volatile memory (such as a Read-Only Memory (ROM) , Electrically Erasable Programmable Read-Only Memory (EEPROM) , or a flash memory) , or any combination thereof. The storage unit 530 may be any detachable or non-detachable medium and may include a machine-readable medium such as a memory, flash memory drive, magnetic disk, or another other media, which can be used for storing information and / or data and can be accessed in the electronic device 500.

[0095] The electronic device 500 may further include additional detachable / non-detachable, volatile / non-volatile memory medium. Although not shown in FIG. 5, it is possible to provide a magnetic disk drive for reading from and / or writing into a detachable and non-volatile magnetic disk and an optical disk drive for reading from and / or writing into a detachable non-volatile optical disk. In such cases, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0096] The communication unit 540 communicates with a further electronic device via the communication medium. In addition, the functions of the components in the electronic device 500 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the electronic device 500 can operate  in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs) or further general network nodes.

[0097] The input device 550 may be one or more of a variety of input devices, such as a mouse, keyboard, tracking ball, voice-input device, and the like. The output device 560 may be one or more of a variety of output devices, such as a display, loudspeaker, printer, and the like. By means of the communication unit 540, the electronic device 500 can further communicate with one or more external devices (not shown) such as the storage devices and display device, with one or more devices enabling the user to interact with the electronic device 500, or any devices (such as a network card, a modem, and the like) enabling the electronic device 500 to communicate with one or more other electronic devices, if required. Such communication can be performed via input / output (I / O) interfaces (not shown) .

[0098] In some embodiments, instead of being integrated in a single device, some, or all components of the electronic device 500 may also be arranged in cloud computing architecture. In the cloud computing architecture, the components may be provided remotely and work together to implement the functionalities described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage service, which will not require end users to be aware of the physical locations or configurations of the systems or hardware providing these services. In various embodiments, the cloud computing provides the services via a wide area network (such as Internet) using suitable protocols. For example, a cloud computing provider provides applications over the wide area network, which can be accessed through a web browser or any other computing components. The software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote position. The computing resources in the cloud computing environment may be merged or distributed at locations in a remote data center. Cloud computing infrastructures may provide the services through a shared data center, though they behave as a single access point for the users. Therefore, the cloud computing architectures may be used to provide the components and functionalities described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.

[0099] The functionalities described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware  logic components that can be used include Field-Programmable Gate Arrays (FPGAs) , Application-specific Integrated Circuits (ASICs) , Application-specific Standard Products (ASSPs) , System-on-a-chip systems (SOCs) , Complex Programmable Logic Devices (CPLDs) , and the like.

[0100] Program code for carrying out the methods of the subject matter described herein may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, special purpose computer, or other programmable data processing apparatus such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely or partly on a machine, executed as a stand-alone software package partly on the machine, partly on a remote machine, or entirely on the remote machine or server.

[0101] In the context of this disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0102] Further, while operations are illustrated in a particular order, this should not be understood as requiring that such operations are performed in the particular order shown or in sequential order, or that all illustrated operations are performed to achieve the desired results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the subject matter described herein, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in the context of separate embodiments may also be implemented in combination in a  single implementation. Rather, various features described in a single implementation may also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0103] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter specified in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0104] From the foregoing, it will be appreciated that specific embodiments of the presently disclosed technology have been described herein for purposes of illustration, but that various modifications may be made without deviating from the scope of the disclosure. Accordingly, the presently disclosed technology is not limited except as by the appended claims.

[0105] Embodiments of the subject matter and the functional operations described in the present disclosure can be implemented in various systems, digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing unit” or “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0106] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document) , in a single file dedicated to the program in question, or in multiple  coordinated files (e.g., files that store one or more modules, sub programs, or portions of code) . A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0107] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0108] It is intended that the specification, together with the drawings, be considered exemplary only, where exemplary means an example. As used herein, the use of “or” is intended to include “and / or” , unless the context clearly indicates otherwise.

[0109] While the present disclosure contains many specifics, these should not be construed as limitations on the scope of any disclosure or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular disclosures. Certain features that are described in the present disclosure in the context of separate embodiments can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0110] Similarly, while operations are illustrated in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in  the present disclosure should not be understood as requiring such separation in all embodiments. Only a few embodiments and examples are described, and other embodiments, enhancements and variations can be made based on what is described and illustrated in the present disclosure.

Claims

1.A method of optimizing a ligand generation model, comprising:generating, by using the ligand generation model, a first ligand and a second ligand for binding to a protein target;in accordance with a determination that the first ligand is superior to the second ligand in terms of at least one property, determining a training objective which causes the ligand generation model to have a propensity to generate the first ligand compared to the second ligand; andupdating the ligand generation model according to the training objective.2.The method of claim 1, wherein determining the training objective comprises:determining a first objective for a global ligand structure based on a first propensity difference in generating an entire structure of the first ligand and an entire structure of the second ligand;determining a second objective for a local ligand structure based on a second propensity difference in generating a first substructure of the first ligand and a second substructure of the second ligand, the first and second substructures corresponding to a same ligand fragment; andderiving the training objective based on the first and second objectives.3.The method of claim 2, wherein determining the first objective for the global ligand structure comprises:determining a first propensity metric for the ligand generation model based on a probability of the ligand generation model to generate the first ligand and a probability of the ligand generation model to generate the second ligand; andderiving the first objective based on the first propensity metric.4.The method of claim 2, wherein determining the second objective for the local ligand structure comprises:determining a second propensity metric for the ligand generation model based on a probability of the ligand generation model to generate the first substructure and a probability of the ligand generation model to generate the second substructure;comparing a first reward score of the first substructure and a second reward score of the second substructure; anddetermining the second objective based on the second propensity metric and a comparing result.5.The method of claim 4, wherein determining the second objective based on the second propensity metric and the comparing result comprises:in response to that the first reward score exceeds the second reward score, deriving the second objective by applying a positive factor to the second propensity metric; andin response to that the first reward score is below the second reward score, deriving the second objective by applying a negative factor to the second propensity metric.6.The method of claim 2, wherein the first objective is determined with respect to at least one first property which is non-decomposable at a ligand fragment level, andthe second objective is determined with respect to at least one second property which is decomposable at the ligand fragment level.7.The method of claim 2, wherein a ligand for binding to the protein target comprises a plurality of ligand fragments, and respective second propensity differences are determined for the plurality of ligand fragments, and the second objective is determined based on the respective second propensity differences.8.The method of claim 1, wherein the ligand generation model comprises a diffusion model, and the method further comprises:determining a sampled schedule factor at a sampled time step based on a linear relation with respect to a reference schedule factor at a reference time step, andwherein the training objective is determined based on the sampled schedule factor.9.The method of claim 8, wherein the reference time step is the maximum time step.10.An electronic device, comprising a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements acts of optimizing a ligand generation model, the acts comprising:generating, by using the ligand generation model, a first ligand and a second ligand for binding to a protein target;in accordance with a determination that the first ligand is superior to the second ligand in terms of at least one property, determining a training objective which causes the ligand generation model to have a propensity to generate the first ligand compared to the second ligand; andupdating the ligand generation model according to the training objective.11.The device of claim 10, wherein determining the training objective comprises:determining a first objective for a global ligand structure based on a first propensity difference in generating an entire structure of the first ligand and an entire structure of the second ligand;determining a second objective for a local ligand structure based on a second propensity difference in generating a first substructure of the first ligand and a second substructure of the second ligand, the first and second substructures corresponding to a same ligand fragment; andderiving the training objective based on the first and second objectives.12.The device of claim 11, wherein determining the first objective for the global ligand structure comprises:determining a first propensity metric for the ligand generation model based on a probability of the ligand generation model to generate the first ligand and a probability of the ligand generation model to generate the second ligand; andderiving the first objective based on the first propensity metric.13.The device of claim 11, wherein determining the second objective for the local ligand structure comprises:determining a second propensity metric for the ligand generation model based on a probability of the ligand generation model to generate the first substructure and a probability of the ligand generation model to generate the second substructure;comparing a first reward score of the first substructure and a second reward score of the second substructure; anddetermining the second objective based on the second propensity metric and a comparing result.14.The device of claim 13, wherein determining the second objective based on the second propensity metric and the comparing result comprises:in response to that the first reward score exceeds the second reward score, deriving the second objective by applying a positive factor to the second propensity metric; andin response to that the first reward score is below the second reward score, deriving the second objective by applying a negative factor to the second propensity metric.15.The device of claim 11, wherein the first objective is determined with respect to at least one first property which is non-decomposable at a ligand fragment level, andthe second objective is determined with respect to at least one second property which is decomposable at the ligand fragment level.16.The device of claim 11, wherein a ligand for binding to the protein target comprises a plurality of ligand fragments, and respective second propensity differences are determined for the plurality of ligand fragments, and the second objective is determined based on the respective second propensity differences.17.The device of claim 10, wherein the ligand generation model comprises a diffusion model, and the method further comprises:determining a sampled schedule factor at a sampled time step based on a linear relation with respect to a reference schedule factor at a reference time step, andwherein the training objective is determined based on the sampled schedule factor.18.The device of claim 17, wherein the reference time step is the maximum time step.19.A computer program product, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform acts of optimizing a ligand generation model, the acts comprising:generating, by using the ligand generation model, a first ligand and a second ligand for binding to a protein target;in accordance with a determination that the first ligand is superior to the second ligand in terms of at least one property, determining a training objective which causes the ligand generation model to have a propensity to generate the first ligand compared to the second ligand; andupdating the ligand generation model according to the training objective.20.The computer program product of claim 19, wherein determining the training objective comprises:determining a first objective for a global ligand structure based on a first propensity difference in generating an entire structure of the first ligand and an entire structure of the second ligand;determining a second objective for a local ligand structure based on a second propensity difference in generating a first substructure of the first ligand and a second substructure of the second ligand, the first and second substructures corresponding to a same ligand fragment; andderiving the training objective based on the first and second objectives.

Citation Information

Patent Citations

  • Medicine screening method and device and electronic equipment

    CN111816252A

  • Method and device for designing ligand molecules

    CN113838541A

  • Ligand generation method and electronic equipment

    CN116364172A

  • Method and device for generating ligand molecules, electronic equipment and storage medium

    CN117316266A

  • Method and device for training and optimizing analysis model, electronic equipment and storage medium

    CN117316316A