Method, device and computer program product for protein structure prediction

The method uses physics-based force guidance in a diffusion process to improve protein structure prediction, addressing inefficiencies in conventional simulations by aligning predicted conformations with the Boltzmann distribution.

WO2025189315A1PCT designated stage Publication Date: 2025-09-18BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/080994
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-11
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Conventional molecular dynamics simulations face limitations in sampling protein conformations due to long timescales and rare event sampling, while existing diffusion models lack physics-based guidance, leading to inefficient and inaccurate protein structure prediction.

Method used

A method involving a diffusion process guided by physics-based force fields to predict protein structures, combining sequence-conditional and unconditional score models to ensure low-energy conformations align with the Boltzmann distribution.

Benefits of technology

Enhances the prediction quality and fidelity of protein structures by ensuring they comply with the equilibrium distribution, balancing sampling diversity and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024080994_18092025_PF_FP_ABST
    Figure CN2024080994_18092025_PF_FP_ABST
Patent Text Reader

Abstract

A method is proposed for protein structure prediction. In the method, for a diffusion process, an initial structure of a target protein is obtained based on composition information concerning amino acids of the target protein. The diffusion process on the initial structure is performed based on a force field prediction for the target protein and the composition information. A target structure of the target protein is determined based on a result of the diffusion process.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD, DEVICE AND COMPUTER PROGRAM PRODUCT FOR PROTEIN STRUCTURE PREDICTIONFIELD

[0001] The present disclosure generally relates to the field of computer, and more specifically, to method, device, and computer program product for protein structure prediction.BACKGROUND

[0002] Proteins are dynamic macromolecules that play pivotal roles in various biological processes. Their functionality is realized primarily through conformational changes-structural alterations that enable proteins to interact with other molecules. Depicting the protein conformational landscape provides vital insights for identifying potential druggable sites hidden beneath the protein surface and revealing transition pathways between multiple metastable states. A comprehensive understanding of protein conformations facilitates the elucidation of biological reaction mechanisms, thereby empowering researchers to design targeted inhibitors and therapeutic agents with improved specificity and efficacy.SUMMARY

[0003] In a first aspect of the present disclosure, there is provided a method of protein structure prediction. The method includes obtaining, for a diffusion process, an initial structure of a target protein based on composition information concerning amino acids of the target protein; performing the diffusion process on the initial structure based on a force field prediction for the target protein and the composition information; and determining a target structure of the target protein based on a result of the diffusion process.

[0004] In a second aspect of the present disclosure, there is provided an electronic device. The electronic device comprises: a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements a method according to the first aspect of the present disclosure.

[0005] In a third aspect of the present disclosure, there is provided a computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform a method according to the first aspect of the present disclosure.

[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Through the more detailed description of some embodiments of the present disclosure in the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, wherein the same reference generally refers to the same components in the embodiments of the present disclosure.

[0008] FIG. 1 illustrates an example environment in which example embodiments of the present disclosure can be implemented;

[0009] FIG. 2 illustrates a schematic diagram of an example diffusion process for protein structure prediction according to some embodiments of the present disclosure;

[0010] FIG. 3 illustrates a schematic diagram of an example generation process of protein structure according to some embodiments of the present disclosure;

[0011] FIG. 4 illustrates an example flowchart of a method of protein structure prediction according to some embodiments of the present disclosure; and

[0012] FIG. 5 illustrates a block diagram of an electronic device in which various embodiments of the present disclosure can be implemented.DETAILED DESCRIPTION

[0013] Principle of the present disclosure will now be described with reference to some embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below.

[0014] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0015] References in the present disclosure to “one embodiment, ” “an embodiment, ” “an example embodiment, ” and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not  necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0016] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.

[0018] Principle of the present disclosure will now be described with reference to some embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below. In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0019] It may be understood that data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with requirements of corresponding laws and regulations and relevant rules.

[0020] It may be understood that, before using the technical solutions disclosed in various embodiment of the present disclosure, the user should be informed of the type, scope of use, and use scenario of the information involved in the present disclosure in an appropriate manner in accordance with relevant laws and regulations, and the user’s authorization should be obtained.

[0021] For example, in response to receiving an active request from the user, prompt information is sent to the user to explicitly inform the user that the requested operation will need  to acquire and use the user’s information. Therefore, the user may independently choose, according to the prompt information, whether to provide the information to software or hardware such as electronic devices, applications, servers, or storage media that perform operations of the technical solutions of the present disclosure.

[0022] As an optional but non-limiting implementation, in response to receiving an active request from the user, the way of sending prompt information to the user, for example, may include a pop-up window, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose “agree” or “disagree” to provide the information to the electronic device.

[0023] It may be understood that the above process of notifying and obtaining the user authorization is only illustrative and does not limit the implementation of the present disclosure. Other methods that satisfy relevant laws and regulations are also applicable to the implementation of the present disclosure.

[0024] As used herein, the term “model” is referred to as an association between an input and an output learned from training data, and thus a corresponding output may be generated for a given input after the training. The generation of the model may be based on a machine learning technique. In general, a machine learning model may be built, which receives input information and makes predictions based on the input information. For example, a classification model may predict a class of the input information among a predetermined set of classes. As used herein, “model” may also be referred to as “machine learning model” , “learning model” , “machine learning network” , or “learning network, ” which are used interchangeably herein.

[0025] Traditional physics-based simulation methods such as molecular dynamics (MD) simulations have been extensively studied for protein conformation sampling. With a well-designed empirical force field and numerical integrator, the model propagates the three-dimensional (3D) structure of a protein system over time following Newtonian mechanics. MD simulations converge towards the equilibrium distribution (i.e., the Boltzmann distribution) given sufficient time, which facilitates estimation of significant thermodynamic properties, e.g., binding free energy change.

[0026] However, to preserve energy conservation and ensure numerical stability, the time step of MD simulations is typically only a few femtoseconds. This poses a challenge as certain biological processes of interest, such as protein folding, span much longer timescales, ranging from microseconds to seconds. This results in limited sampling efficiency within conventional MD simulations, further compounded by the rare event sampling problem, impeding the research community to widely adopt MD for high throughput studies.

[0027] With the development of machine learning, machine learning models such as deep neural networks have been used for protein structure prediction. Building upon the cornerstone of powerful folding models, several attempts have been made to tailor these deep neural networks for protein conformation sampling. By perturbing the model input, such as multiple sequence alignment (MSA) masking or clustering, the folding model provides a more diverse set of possible folded structures, i.e., alternative conformations. However, this heuristic approach cannot guarantee the predicted structure to be a low energy state of the target sequence.

[0028] Some solutions have incorporated diffusion models for protein conformation generation. By pretraining on a large amount of known protein structures and efficient sampling through a predefined stochastic process, these models have shown promise in exploring diverse protein conformational states. Nevertheless, existing diffusion models fall short in utilizing important physical prior information, such as the MD force field, to guide the diffusion processes, hampering the capability to faithfully sample diverse protein conformations complying with the Boltzmann distribution.

[0029] Embodiments of the present disclosure propose solutions for protein structure prediction. According to embodiments of the present disclosure, for a diffusion process, an initial structure of a target protein is obtained based on composition information concerning amino acids of the target protein. The diffusion process on the initial structure is performed based on a force field prediction for the target protein and the composition information. A target structure of the target protein is determined based on a result of the diffusion process.

[0030] In the embodiments of the present disclosure, physics-based force guidance is used for protein structure prediction. Such novel physics-based guidance strategies with theoretical guarantee effectively guide the diffusion sampler to generate low energy conformations better complying with the underlying Boltzmann distribution. In this way, prediction quality and fidelity can be ensured.

[0031] Theoretical foundation of the present disclosure and example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0032] FIG. 1 illustrates a block diagram of an example environment 100 in which various embodiments of the present disclosure may be implemented. In the environment 100 of FIG. 1, an electronic device 120 receives composition information 101 and output a target structure 102. The composition information 101 concerns amino acids of a target protein, for example, an amino acid sequence of the target protein. The target structure 102 is a structure of the target protein. Herein, the structure is also referred to as conformation.

[0033] The electronic device 120 includes or is deployed with a structure prediction model 110. To predict the structure of the target protein, the structure prediction model 110 generates an initial structure of the target protein based on the composition information. The initial structure may be a structure with noises, for example, sampling from a Gaussian distribution.

[0034] The diffusion process is performed on the initial structure based on a force field prediction for the target protein and the composition information 101. The diffusion process and the force field prediction will be described below. Then, the target structure 102 of the target protein may be determined based on a result of the diffusion process. For example, an output of the diffusion process may be used as the predicted structure for the target protein.

[0035] To better understand the embodiments of the present disclosure, some preliminaries are first described. The protein backbone may be represented as below: for a protein with N amino acid residues, its backbone atomic coordinates can be parameterized by a collection of N orientation preserving rigid transformations (i.e., frames) to the local [N, C, C, O] backbone atoms in each residue. The positions of all N frames are denoted collectively by x0= [T0, R0] ∈SE (3) N, where and R0∈SO (3) N denote the corresponding translation and rotation operations, respectively. With an additional backbone torsion angle ψ describing the rotation of the oxygen atom around the C-Cα bond within each residue, the protein backbone structure can be reconstructed from the frame representations.

[0036] Diffusion modeling on manifold SE (3) N is employed for protein backbone generation. In some embodiments, two independent diffusion processes may be defined for the translation and rotation subspaces, respectively:

[0037] where subscript t denotes the diffusion time variable in [0, 1] , βt, and σt, are predefined time-dependent noise schedules, P is a projection operator removing the center of mass, and  represents the standard Wiener process in The transition kernel of T satisfies  where The rotational transition kernel satisfies pt (Rr|R0) =IGSO3 (Rt; R0, t) , where IGSO3 is the isotropic Gaussian distribution on SO (3) .

[0038] The associated reverse-time stochastic differential equation (SDE) is as follows:

[0039] where denotes another standard Wiener process in reverse time.

[0040] Guided sampling has emerged as a critical strategy in developing diffusion models capable of generating samples complying with human instructions. Consider existing paired data x0~p0 (x0|c) with a conditioning variable c (subscript denotes diffusion time, where t=0 corresponds to the original data) , the conditional probability density is typically perceived at time t as t aa pt (xt|c) . Applying Bayes’ rule,  can be obtained. Therefore, a classifier may be trained to predict the conditioning probability pt (c|xt) with given noisy data xt, and use the score of the classifier output as guidance. Instead of training a separate classifier to estimate  utilizing an implicit classifier is proposed. Given that, in some embodiments, a linear combination of an unconditional score estimator sθ (xt) and a conditional score estimator sθ (xt, c) may be used to jointly estimate a target score function:

[0041] where γ is a hyperparameter controlling the guidance strength. When γ=0, it reduces to an unconditional model, while at γ=1, it becomes a pure conditional model. These two models can be simultaneously trained under the same hood, where the unconditional model receives masked conditioning variable c during training.

[0042] Some example embodiments and the corresponding theoretical foundations are now described to better understand the benefits of these embodiments. In some embodiments, the diffusion process for protection structure prediction may comprise a plurality of denoising steps. In any given denoising step of the plurality of denoising steps, a target score function may be determined based on a first intermediate structure (which is the input structure for the given denoising step) and the composition information. The first intermediate structure may comprise the initial structure or an output structure from a denoising step previous to the given denoising step. Specifically, for the first denoising step, the first intermediate structure is the initial structure. For any denoising step after the first denoising step, the first intermediate structure is the output structure from a previous denoising step.

[0043] In some embodiments, the target score function may comprise a translational component and a rotational component. In other words, the structure of the target protein may be represented by a translational component and a rotational component.

[0044] For example, an example structure is represented as xN with two components, e.g., a translational component TN and a rotational component RN. Then, xN= [TN, RN] . As mentioned above, the initial structure is a structure with noises which may be represented as The structure prediction model 110 generates such initial structure and performs the diffusion process starting from the initial structure. In the time step (i.e., given denoising step) i=N, ..., 1, the input structure may be represented by the structure xi, and a target score function represented as may be determined. Finally, the structure prediction model 110 determines the target structure through the plurality of denoising steps.

[0045] In some embodiments, the composition information may comprise an amino acid sequence of the target protein. A score based model conditioned on the amino acid sequence and a score based model unconditioned on the amino acid sequence may be combined. For example, the target score function may include a first score function and a second score function. The first score function may be generated based on the first intermediate structure by using a first diffusion model conditioned on the amino acid sequence. The second score function may be generated based on the first intermediate structure by using a second diffusion model without being conditioned on the amino acid sequence. Then, the target score function may be derived based on the first score function, the second score function, and a first predetermined weight. In this way, diffusion models may be introduced for protein structure generation.

[0046] Reference is now made to FIG. 2, which illustrates an example diffusion process 200 of protein structure according to some embodiments of the present disclosure. The first intermediate structure is represented as xi. A diffusion model 210 may be the first diffusion model, which may be also referred to as a conditional score model. The diffusion model 210 may receive xi and output a score function 211 (i.e., the first score function) . A diffusion model 220 may be the second diffusion model, which is also referred to as an unconditional score model. The diffusion model 220 may receive xi and output a score function 221 (i.e., the second score function) . A target score function 202 may be obtained based at least on the score function 211 and the score function 221. Assuming the first predetermined weight as γ, the target score function 202 may be derived from γ×score function 211 and (1-γ) ×score function 221.

[0047] The conditional score model and the unconditional score model may be collectively referred to as a baseline model, and will be described in detail with examples below.

[0048] As an example, the structure prediction model 110 may include an unconditional score model  (i.e., the second diffusion model) and a sequence-conditional one  The unconditional model is trained on protein structures without any sequence information, effectively capturing the conformation distribution of general proteins. On the other hand, the sequence-conditional score model has access to both protein sequence information (denoted as seq) and the corresponding structure xt, at time t.

[0049] Any suitable network architecture may be adopted to parameterize the corresponding score functions. For example, the unconditional score model may take sinusoidal embedding of the residue index and the diffusion time t as its single ( {si} ) and pair ( {zij} ) embeddings, and the conditional score model (i.e., the first diffusion model) additionally concatenates precomputed representations of the amino acid sequence (for example, from a folding model) to its single embedding. Note that the choice of sequence representations for the conditional score model is flexible. Using pretrained representations from folding models helps diffusion models generate reasonable protein structures, while the unconditional score model can effectively improve sampling diversity. Both the score models may be trained with the denoising score matching (DSM) loss function:

[0050] where λ (t) is a reweighting function inversely proportional to the score norm. During the reverse sampling process, a hyperparameterγ (i.e., the first predetermined weight) is used to  control the classifier-free guidance strength from the conditional score model, so that the score function may be estimated by:

[0051] where sθ (xt, t|seq) is the target score function, the term is the first score function and the term is the second score function. For notation simplicity, hereafter the sequence conditional term is omitted in the baseline score model, i.e., sθ (xt, t) =sθ (xt, t|seq) .

[0052] Continue with the given denoising step. In the given denoising step, after determining the target score function, the target score function is updated based on a predicted force field for the first intermediate structure. A second intermediate structure as an output structure of the given step is then determined based on the updated target score function and the first intermediate structure. If the given denoising step is the last denoising step of the diffusion process, the second intermediate structure is the output of the diffusion process, which may be then determined as the target structure or used to determine the target structure. If the given denoising step is not the last denoising step, the second intermediate structure is an input structure of a next denoising step.

[0053] Referring back to FIG. 2, a diffusion model 230 may receive xi and output a predicted force field 231, which is used to update the target function 202. The second intermediate structure xi-1 is determined based on the updated target score function and the first intermediate structure xi.

[0054] From the predicted force field, the diffusion model may successfully reweight the generated structure to ensure that they adhere better to the equilibrium distribution. A visual depiction is shown in FIG. 3, which illustrates a diagram of an example generation process of protein structure according to some embodiments of the present disclosure. In this way, as a force-guided model, the structure prediction model 110 of the present disclosure targets multi-conformation generation for proteins. Employing a sequence-based conditional score network to guide an unconditional score model, the structure prediction model 110 achieves reasonable conformation diversity while ensuring sample quality. Building upon the energy guidance foundations, a novel force-guided sampling strategy is used to estimate the intermediate force function, which is then embedded within a reverse time sampling process. As shown in the upper branch, with a mixture of sequence-conditional and unconditional score models, diverse conformations can be sampled with reasonable quality. As shown in the lower sample, by  incorporating force guidance, structures with lower energy can be generated to better comply with the Boltzmann distribution.

[0055] Some example embodiments regarding the force field prediction are now described below.

[0056] In some embodiments, the force field prediction is performed by determining a predicted energy for an intermediate structure generated in the diffusion process by using a diffusion model for energy prediction (also referred to as diffusion energy model) . A predicted force field is derived for the intermediate structure based on the predicted energy and the intermediate structure.

[0057] Despite the diffusion model’s capability to generate diverse structures, these conformations are not always reasonable in the sense that they may reside in the high energy region of the potential energy surface. Due to the limited availability of multi-conformation data within existing protein structure databases, the training data distribution does not comply with the equilibrium distribution, but rather only containing a few data points near the potential energy minima for each protein sequence. This necessitates the development of a generative model propelled by data and steered by physics-based guidance towards generating samples according to the Boltzmann distribution. To that end, energy guidance is introduced, where the predicted energy function attributes a reward to guide the conformation generation process.

[0058] Given an existing (baseline) diffusion model which generates samples x0~q0 (x0) , the goal is to sample protein conformations from the equilibrium distribution,

[0059] where is the intractable normalizing constant, ε0 is the energy function (for example, generated by any suitable molecular simulation tool) which evaluates the potential energy of each generated conformation x0. k is the inverse temperature factor. Given any test function F (x0) , p0 (x0) gives a more accurate estimate through importance sampling, defined as

[0060] In the forward diffusion process, the original samples x0~q0 (x0) are perturbed according to the SDE described in Eq. (1) , and it is asserted that pt (xt|x0) : =qt (xt|x0) .Whereas it is the marginal distribution pt (xt) that truly captures the attention. There are the following properties for the energy function.

[0061] In the first proposition, it is supposed that and for t∈ (0, 1] , pt (xt|x0) : =qt (xt|x0) . Then, the marginal distribution satisfies  where qt (xt) is the data-based marginal distribution, and εt (xt) satisfies:

[0062] where εtis referred to as the intermediate energy function, capturing the dynamic changes in energy at various stages during the diffusion process.

[0063] With the exact formula in Eq. (7) , a neural network fφ (xt, t) parameterized by φis utilized to approximate the intermediate energy function εt (xt) by the optimization problem,

[0064] Given that the optimal satisfies  sampling can be carried out according to the reverse-time SDE with score function:

[0065] where η is the hyperparameter controlling the guidance strength, and sθ (xt, t) refers to the baseline score model. Note that sθ (xt, t) and the intermediate energy networkfφ (xt, t) are trained independently.

[0066] If the score model only samples protein backbone structures, full-atom structures may be generated by any suitable manner for energy evaluation using a suitable molecular simulation tool. However, there are inherent challenges with normalization intricacies in the energy computed by the molecular simulation tool. Within limited batch size, significant numerical volatility is witnessed in estimating. In contrast, the numerical value of atomic force will be much more stable. Consequently, the design of a force-guided policy tailored for protein conformation generation proves to be incredibly significant. Compared with energy guidance, force guidance does not suffer from the energy fluctuation problem and can be directly applied as guidance in the reverse sampling process. How to realize intermediate force guidance will be illustrated below.

[0067] To this end, in some embodiments, the force field prediction may be directed performed by using a diffusion model for force prediction (also referred to as diffusion force model) . In other words, a predicted force field for an intermediate structure generated in the diffusion process may be generated by the diffusion force model.

[0068] In the context of protein conformation modeling, both a physics-based energy function ε0 (x0) (i.e., interatomic potential energy) as well as its gradient  (i.e., force on each atom) have been accessed. In contrast to the unnormalized potential energy function, atomic force is more local and exhibits better numerical stability, which also aligns better with the score matching objective. Following the marginal distribution of energy in Eq. (7) , there are the following properties for the force function.

[0069] In the second proposition, the presumption of pt (xt|x0) : =qt (xt|x0) is given, where0<t≤1, the energy function of intermediate state follows Eq. (7) . The corresponding intermediate force has the following formula:

[0070] where Denoting implicit distribution the score function follows

[0071] where η is a hyperparameter controlling the strength of force guidance.

[0072] The second proposition unveils the precise force at time t, thereby advancing the understanding of the intermediate energy function. The intermediate force formula consists of the ground truth potential energy, the diffusion transition kernel, and a marginal score function (estimated by the score model) . As t approaches 0, qt (x0|xt) converges to the delta function and the intermediate force converges to the ground truth force. On the other hand, when t approaches 1, qt (xt|x0) reduces to qt (xt) , and the force vanishes.

[0073] In some embodiments, as mentioned above, the target score function may include the translational component and the rotational component. Further, in some embodiments, the translational component may be updated based on the predicted force field and a second predetermined weight, while the rotational component may not be updated. In this way, structural stability can be maintained.

[0074] Continue with the given denoising step. In some embodiments, to determine a second intermediate structure as an output structure of the given step, a transition to the first intermediate structure may be determined based on the updated target score function and the first intermediate structure. As such, the second intermediate structure may be derived by applying the transition to the first intermediate structure.

[0075] Starting from the prior distribution in SE (3) N, an example entire inference process is illustrated in Algorithm 1 (as shown in Table 1) . Score network  (i.e., target score function) is the baseline model mixing the sequence conditional score and unconditional score. Force guidance term hψ (xt, t) is incorporated to the score function term at every time step during the reverse sampling process, whose guidance strength is controlled by hyperparameter η. To maintain structural stability, force guidance is only applied to the translational components (i.e., α-carbons) in

[0076] Table 1 example inference process

[0077] As shown in Table 1, in line 1, the initial structure xN is obtained. In lines 2 to 9 correspond to the diffusion process starting from the initial structure. Specifically, in line 4, the target score is obtained based on the conditional score function sθ (xi, i, seq) , the unconditional score function sθ (xi, i) and the first predetermined weight γ. In line 5, the predicted force field H is obtained. In lines 6 and 7, the transitions  and are obtained. Note that the target score function is updated with the predicted force field H. In this case, the transitional component of the target score function is updated with the predicted force field H. As mentioned above, this may maintain structural stability. In line 8, the second intermediate structure xi-1 is obtained. After the plurality of denoising steps, the target structure x0 is determined.

[0078] In the embodiments described above, the diffusion process is performed by using at least one diffusion model for score estimation (for example, the conditional score model and the unconditional score model) and a further diffusion model for energy prediction or force prediction (for example, the force prediction model) . In some embodiments, the one or more score models may be trained separately from the force or energy prediction model.

[0079] In some embodiments, the force or energy prediction model may be trained after the one or more score models. For example, in the training process of the force or energy prediction model, a score function may be generated based on an intermediate structure for a sample protein by using the trained one or more score models. A predicted force field may be generated based on the intermediate structure by using the force or energy prediction model. An estimated force field  may be derived based on the score function, a ground truth structure and a ground truth energy of the sample protein. Then, parameters of the force or energy prediction model may be updated based on a difference between the predicted force field and the estimated force field. For example, the force or energy prediction model is trained by minimizing a loss function based on the difference between the predicted force field and the estimated force field.

[0080] An example training process is now described. A baseline score model is employed to generate protein conformations from q0 (x0) , which are then used to train an independent intermediate force network the tangent space of  at xt. Let K be a positive integer denoting the training batch size. To ensure that data in a batch adheres to the same Boltzmann distribution, a protein sequence is first randomly chosen, then subsequently draw K samples from this sequence. By adding noise to the sampled data x0 following the SDE in Eq. (1) , perturbed data is obtained at time t as  such that  The intermediate force loss function is defined as

[0081] where The latter component of Eq. (11b) signifies the precise value of the intermediate force at time t, where is the intractable score function estimated by sθ (xt, t) in the baseline and qt (xt|x0) is the tractable Gaussian distribution. Instead of computing in the entire generated data collection, the expectation with the K samples is estimated in a batch. The force-guided training process is summarized in Algorithm 2 (as shown in Table 2) .

[0082] Table 2 example training process

[0083] As shown in Table 2, in line 3, an intermediate structure xt for time step t is obtained. Then, a score function sθ (xt, t) is generated based on the intermediate structure xt. In lines 5 to 7, a predicted force field hψ (xt, t) at time t is determined using the diffusion model for force prediction. An estimated force field is derived based on the score function sθ (xt, t) , a ground truth structure x0 and a ground truth energy ε0 (x0) . In line 8, parameters of the further diffusion model may be updated based on a difference In this way, the force prediction model is trained.

[0084] Taking into account that the exact intermediate force formula in Eq. (10) equals to the MD force-field at t=0, and turns 0 when t=0, the network hψ (xt, t) is constructed as the interpolation form following:

[0085] where gψ (xt, t) is a neural network estimating the intermediate term within the interpolation construction. It ensures the boundary conditions at t∈ {0, 1} and it is empirically found that this construction reduces the variance during training.

[0086] In some embodiments, the diffusion process described above may be one of a plurality of diffusion processes used for structure prediction on the target protein. These diffusion processes may have different initial structures at the start of denoising steps. A distribution of protein energies with respect to protein structures of the target protein may be constructed based on respective target structures generated by the plurality of diffusion processes. In this way, the protein confirmation landscape may be constructed, and one or more stable protein structures can be determined.

[0087] Last but not least, the embodiments of the present disclosure introduce diffusion models for protein structure prediction. By mixing a sequence-conditional score network with an unconditional model, a balance is achieved between sampling quality and diversity through classifier-free guidance. Upon such a baseline model, novel physics-based energy and force guidance strategies with theoretical guarantee are proposed, which effectively guide the diffusion sampler to generate low energy conformations better complying with the underlying Boltzmann distribution.

[0088] Example process and device

[0089] FIG. 4 illustrates a flowchart of a method 400 of protein structure prediction in accordance with some example implementations of the present disclosure. The method 400 may be implemented at the electronic device 120 as illustrated in FIG. 1. At a block 410, for a diffusion process, an initial structure of a target protein is obtained based on composition information concerning amino acids of the target protein. At a block 420, the diffusion process on the initial structure is performed based on a force field prediction for the target protein and the composition information. At a block 430, a target structure of the target protein is determined based on a result of the diffusion process.

[0090] In some embodiments, the diffusion process comprises a plurality of denoising steps, and a given denoising step of the plurality of denoising steps comprises: determining a target score function based on a first intermediate structure and the composition information, wherein the first  intermediate structure comprises the initial structure or an output structure from a denoising step previous to the given denoising step; updating the target score function based on a predicted force field for the first intermediate structure; and determining a second intermediate structure as an output structure of the given step based on the updated target score function and the first intermediate structure.

[0091] In some embodiments, the composition information comprises an amino acid sequence of the target protein, and determining a target score function based on a first intermediate structure and the composition information comprises: generating a first score function based on the first intermediate structure by using a first diffusion model conditioned on the amino acid sequence; generating a second score function based on the first intermediate structure by using a second diffusion model without being conditioned on the amino acid sequence; and deriving the target score function based on the first score function, the second score function, and a first predetermined weight.

[0092] In some embodiments, the target score function comprises a translational component and a rotational component, and updating the target score function comprises: updating the translational component based on the predicted force field and a second predetermined weight.

[0093] In some embodiments, determining a second intermediate structure as an output structure of the given step based on the updated target score function and the first intermediate structure comprises: determining a transition to the first intermediate structure based on the updated target score function and the first intermediate structure; and deriving the second intermediate structure by applying the transition to the first intermediate structure.

[0094] In some embodiments, the force field prediction is performed by: determining a predicted energy for an intermediate structure generated in the diffusion process by using a diffusion model for energy prediction; and deriving a predicted force field for the intermediate structure based on the predicted energy and the intermediate structure.

[0095] In some embodiments, the force field prediction is performed by: determining a predicted force field for an intermediate structure generated in the diffusion process by using a diffusion model for force prediction.

[0096] In some embodiments, the diffusion process is one of a plurality of diffusion processes used for structure prediction on the target protein, and the method further comprises: constructing  a distribution of protein energies with respect to protein structures of the target protein based on respective target structures generated by the plurality of diffusion processes.

[0097] In some embodiments, the diffusion process is performed by using at least one diffusion model for score estimation and a further diffusion model for energy prediction or force prediction, and the further diffusion model is trained by: generating a score function based on an intermediate structure for a sample protein by using the trained at least one diffusion model; determining a predicted force field based on the intermediate structure by using the further diffusion model; deriving an estimated force field based on the score function, a ground truth structure and a ground truth energy of the sample protein; and updating parameters of the further diffusion model based on a difference between the predicted force field and the estimated force field.

[0098] In some embodiments of the present disclosure, there is provided a non-transitory computer program product, the non-transitory computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform a method of protein structure prediction. The method comprises: obtaining, for a diffusion process, an initial structure of a target protein based on composition information concerning amino acids of the target protein; performing the diffusion process on the initial structure based on a force field prediction for the target protein and the composition information; and determining a target structure of the target protein based on a result of the diffusion process. In some embodiments of the present disclosure, the method further comprises other steps as described in the present disclosure.

[0099] FIG. 5 illustrates a block diagram of an electronic device 500 in which various embodiments of the present disclosure can be implemented. It would be appreciated that the electronic device 500 shown in FIG. 5 is merely for purpose of illustration, without suggesting any limitation to the functions and scopes of the present disclosure in any manner. The electronic device 500 may be used to implement the above method 500. As shown in FIG. 5, the electronic device 500 may be a general-purpose electronic device. The electronic device 500 may at least comprise one or more processors or processing units 510, a memory 520, a storage  unit 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560.

[0100] The processing unit 510 may be a physical or virtual processor and can implement various processes based on programs 525 stored in the memory 520. In a multi-processor system, multiple processing units execute computer executable instructions in parallel so as to improve the parallel processing capability of the electronic device 500. The processing unit 510 may also be referred to as a central processing unit (CPU) , a microprocessor, a controller, or a microcontroller.

[0101] The electronic device 500 typically includes various computer storage medium. Such medium can be any medium accessible by the electronic device 500, including, but not limited to, volatile and non-volatile medium, or detachable and non-detachable medium. The memory 520 can be a volatile memory (for example, a register, cache, Random Access Memory (RAM)) , a non-volatile memory (such as a Read-Only Memory (ROM) , Electrically Erasable Programmable Read-Only Memory (EEPROM) , or a flash memory) , or any combination thereof. The storage unit 530 may be any detachable or non-detachable medium and may include a machine-readable medium such as a memory, flash memory drive, magnetic disk, or another other media, which can be used for storing information and / or data and can be accessed in the electronic device 500.

[0102] The electronic device 500 may further include additional detachable / non-detachable, volatile / non-volatile memory medium. Although not shown in FIG. 5, it is possible to provide a magnetic disk drive for reading from and / or writing into a detachable and non-volatile magnetic disk and an optical disk drive for reading from and / or writing into a detachable non-volatile optical disk. In such cases, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0103] The communication unit 540 communicates with a further electronic device via the communication medium. In addition, the functions of the components in the electronic device 500 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs) or further general network nodes.

[0104] The input device 550 may be one or more of a variety of input devices, such as a mouse, keyboard, tracking ball, voice-input device, and the like. The output device 560 may be one or more of a variety of output devices, such as a display, loudspeaker, printer, and the like. By means of the communication unit 540, the electronic device 500 can further communicate with one or more external devices (not shown) such as the storage devices and display device, with one or more devices enabling the user to interact with the electronic device 500, or any devices (such as a network card, a modem, and the like) enabling the electronic device 500 to communicate with one or more other electronic devices, if required. Such communication can be performed via input / output (I / O) interfaces (not shown) .

[0105] In some embodiments, instead of being integrated in a single device, some, or all components of the electronic device 500 may also be arranged in cloud computing architecture. In the cloud computing architecture, the components may be provided remotely and work together to implement the functionalities described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage service, which will not require end users to be aware of the physical locations or configurations of the systems or hardware providing these services. In various embodiments, the cloud computing provides the services via a wide area network (such as Internet) using suitable protocols. For example, a cloud computing provider provides applications over the wide area network, which can be accessed through a web browser or any other computing components. The software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote position. The computing resources in the cloud computing environment may be merged or distributed at locations in a remote data center. Cloud computing infrastructures may provide the services through a shared data center, though they behave as a single access point for the users. Therefore, the cloud computing architectures may be used to provide the components and functionalities described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.

[0106] The functionalities described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-Programmable Gate Arrays (FPGAs) , Application-specific Integrated Circuits (ASICs) , Application-specific Standard Products  (ASSPs) , System-on-a-chip systems (SOCs) , Complex Programmable Logic Devices (CPLDs) , and the like.

[0107] Program code for carrying out the methods of the subject matter described herein may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, special purpose computer, or other programmable data processing apparatus such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely or partly on a machine, executed as a stand-alone software package partly on the machine, partly on a remote machine, or entirely on the remote machine or server.

[0108] In the context of this disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0109] Further, while operations are illustrated in a particular order, this should not be understood as requiring that such operations are performed in the particular order shown or in sequential order, or that all illustrated operations are performed to achieve the desired results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the subject matter described herein, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single implementation. Rather, various features described in a single implementation may also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0110] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter specified in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0111] From the foregoing, it will be appreciated that specific embodiments of the presently disclosed technology have been described herein for purposes of illustration, but that various modifications may be made without deviating from the scope of the disclosure. Accordingly, the presently disclosed technology is not limited except as by the appended claims.

[0112] Embodiments of the subject matter and the functional operations described in the present disclosure can be implemented in various systems, digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing unit” or “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0113] Acomputer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document) , in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code) . A computer program can be deployed to be executed on one computer or on multiple computers  that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0114] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0115] It is intended that the specification, together with the drawings, be considered exemplary only, where exemplary means an example. As used herein, the use of “or” is intended to include “and / or” , unless the context clearly indicates otherwise.

[0116] While the present disclosure contains many specifics, these should not be construed as limitations on the scope of any disclosure or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular disclosures. Certain features that are described in the present disclosure in the context of separate embodiments can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0117] Similarly, while operations are illustrated in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in the present disclosure should not be understood as requiring such separation in all embodiments.  Only a few embodiments and examples are described, and other embodiments, enhancements and variations can be made based on what is described and illustrated in the present disclosure.

Claims

1.A method of protein structure prediction, comprising:obtaining, for a diffusion process, an initial structure of a target protein based on composition information concerning amino acids of the target protein;performing the diffusion process on the initial structure based on a force field prediction for the target protein and the composition information; anddetermining a target structure of the target protein based on a result of the diffusion process.2.The method of claim 1, wherein the diffusion process comprises a plurality of denoising steps, and performing a given denoising step of the plurality of denoising steps comprises:determining a target score function based on a first intermediate structure and the composition information, wherein the first intermediate structure comprises the initial structure or an output structure from a denoising step previous to the given denoising step;updating the target score function based on a predicted force field for the first intermediate structure; anddetermining a second intermediate structure as an output structure of the given step based on the updated target score function and the first intermediate structure.3.The method of claim 2, wherein the composition information comprises an amino acid sequence of the target protein, and determining a target score function based on a first intermediate structure and the composition information comprises:generating a first score function based on the first intermediate structure by using a first diffusion model conditioned on the amino acid sequence;generating a second score function based on the first intermediate structure by using a second diffusion model without being conditioned on the amino acid sequence; andderiving the target score function based on the first score function, the second score function, and a first predetermined weight.4.The method of claim 2, wherein the target score function comprises a translational component and a rotational component, and updating the target score function comprises:updating the translational component based on the predicted force field and a second predetermined weight.5.The method of claim 2, wherein determining a second intermediate structure as an output structure of the given step based on the updated target score function and the first intermediate structure comprises:determining a transition to the first intermediate structure based on the updated target score function and the first intermediate structure; andderiving the second intermediate structure by applying the transition to the first intermediate structure.6.The method of claim 1, wherein the force field prediction is performed by:determining a predicted energy for an intermediate structure generated in the diffusion process by using a diffusion model for energy prediction; andderiving a predicted force field for the intermediate structure based on the predicted energy and the intermediate structure.7.The method of claim 1, wherein the force field prediction is performed by:determining a predicted force field for an intermediate structure generated in the diffusion process by using a diffusion model for force prediction.8.The method of claim 1, wherein the diffusion process is one of a plurality of diffusion processes used for structure prediction on the target protein, and the method further comprises:constructing a distribution of protein energies with respect to protein structures of the target protein based on respective target structures generated by the plurality of diffusion processes.9.The method of claim 1, wherein the diffusion process is performed by using at least one diffusion model for score estimation and a further diffusion model for energy prediction or force prediction, and the further diffusion model is trained by:generating a score function based on an intermediate structure for a sample protein by using the trained at least one diffusion model;determining a predicted force field based on the intermediate structure by using the further diffusion model;deriving an estimated force field based on the score function, a ground truth structure and a ground truth energy of the sample protein; andupdating parameters of the further diffusion model based on a difference between the predicted force field and the estimated force field.10.An electronic device, comprising a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements acts of protein structure prediction, the acts comprising:obtaining, for a diffusion process, an initial structure of a target protein based on composition information concerning amino acids of the target protein;performing the diffusion process on the initial structure based on a force field prediction for the target protein and the composition information; anddetermining a target structure of the target protein based on a result of the diffusion process.11.The device of claim 10, wherein the diffusion process comprises a plurality of denoising steps, and performing a given denoising step of the plurality of denoising steps comprises:determining a target score function based on a first intermediate structure and the composition information, wherein the first intermediate structure comprises the initial structure or an output structure from a denoising step previous to the given denoising step;updating the target score function based on a predicted force field for the first intermediate structure; anddetermining a second intermediate structure as an output structure of the given step based on the updated target score function and the first intermediate structure.12.The device of claim 11, wherein the composition information comprises an amino acid sequence of the target protein, and determining a target score function based on a first intermediate structure and the composition information comprises:generating a first score function based on the first intermediate structure by using a first diffusion model conditioned on the amino acid sequence;generating a second score function based on the first intermediate structure by using a second diffusion model without being conditioned on the amino acid sequence; andderiving the target score function based on the first score function, the second score function, and a first predetermined weight.13.The device of claim 11, wherein the target score function comprises a translational component and a rotational component, and updating the target score function comprises:updating the translational component based on the predicted force field and a second predetermined weight.14.The device of claim 11, wherein determining a second intermediate structure as an output structure of the given step based on the updated target score function and the first intermediate structure comprises:determining a transition to the first intermediate structure based on the updated target score function and the first intermediate structure; andderiving the second intermediate structure by applying the transition to the first intermediate structure.15.The device of claim 10, wherein the force field prediction is performed by:determining a predicted energy for an intermediate structure generated in the diffusion process by using a diffusion model for energy prediction; andderiving a predicted force field for the intermediate structure based on the predicted energy and the intermediate structure.16.The device of claim 10, wherein the force field prediction is performed by:determining a predicted force field for an intermediate structure generated in the diffusion process by using a diffusion model for force prediction.17.The device of claim 10, wherein the diffusion process is one of a plurality of diffusion processes used for structure prediction on the target protein, and the acts further comprises:constructing a distribution of protein energies with respect to protein structures of the target protein based on respective target structures generated by the plurality of diffusion processes.18.The device of claim 10, wherein the diffusion process is performed by using at least one diffusion model for score estimation and a further diffusion model for energy prediction or force prediction, and the further diffusion model is trained by:generating a score function based on an intermediate structure for a sample protein by using the trained at least one diffusion model;determining a predicted force field based on the intermediate structure by using the further diffusion model;deriving an estimated force field based on the score function, a ground truth structure and a ground truth energy of the sample protein; andupdating parameters of the further diffusion model based on a difference between the predicted force field and the estimated force field.19.A computer program product, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform acts of protein structure prediction, the acts comprising:obtaining, for a diffusion process, an initial structure of a target protein based on composition information concerning amino acids of the target protein;performing the diffusion process on the initial structure based on a force field prediction for the target protein and the composition information; anddetermining a target structure of the target protein based on a result of the diffusion process.20.The computer program product of claim 19, wherein the diffusion process comprises a plurality of denoising steps, and performing a given denoising step of the plurality of denoising steps comprises:determining a target score function based on a first intermediate structure and the composition information, wherein the first intermediate structure comprises the initial structure or an output structure from a denoising step previous to the given denoising step;updating the target score function based on a predicted force field for the first intermediate structure; anddetermining a second intermediate structure as an output structure of the given step based on the updated target score function and the first intermediate structure.