Combining conformation prediction method, model training method, device and storage medium

By generating a set of binding conformation samples and adjusting the parameters of the binding conformation prediction model, and using deep learning models and Gaussian mixture function optimization, the problems of lead compound optimization and candidate drug failure in new drug research and development are solved, efficient and accurate drug virtual screening is achieved, and the time and cost of new drug research and development are reduced.

CN118506855BActive Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410643026.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-22
Publication Date
2025-09-23
Estimated Expiration
2044-05-22

AI Technical Summary

Technical Problem

In the development of new drugs with existing technologies, the optimization and modification of lead compounds and the failure of candidate drugs lead to high R&D costs and long cycles, making it difficult to quickly screen out high-quality potential molecules that interact with targets.

Method used

By generating a set of binding conformation samples, adjusting the parameters of the binding conformation prediction model, using a deep learning model to dock the target protein and the molecular library, generating candidate binding conformations, and optimizing the model based on the loss function of the Gaussian mixture function, the prediction accuracy is improved, and finally the target binding conformation of the target protein is determined.

Benefits of technology

It improves the accuracy of the binding conformation prediction model, enhances the efficiency and accuracy of drug virtual screening, and can more quickly screen out drug molecules with higher affinity to the target protein, reducing the time and cost of new drug research and development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118506855B_ABST
    Figure CN118506855B_ABST
Patent Text Reader

Abstract

The present disclosure provides a binding conformation prediction method, a model training method, an apparatus and a storage medium, which relate to the field of computer technology, in particular to the field of artificial intelligence and biological computing. The specific implementation scheme is: based on the binding conformation sample set generated by the target protein, the parameters of the first binding conformation prediction model are adjusted to obtain a second binding conformation prediction model; a plurality of candidate binding conformations corresponding to the target protein are respectively input into the second binding conformation prediction model to obtain a prediction result for each candidate binding conformation; based on the prediction result of each candidate binding conformation, the target binding conformation of the target protein is obtained. According to the embodiment of the present disclosure, the first binding conformation prediction model can be adjusted based on the binding conformation sample set generated by the target protein to obtain a second binding conformation prediction model, thereby improving the accuracy of the second binding conformation prediction model's prediction result for the binding conformation of the target protein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to the fields of artificial intelligence and biological computing. Background Art

[0002] Repeated optimization and modification of lead compounds and the failure of later-stage drug candidates are major reasons for the time-consuming and costly nature of new drug development. High-quality lead compounds can minimize the drug development cycle, enabling new drugs to enter and pass clinical trials more quickly. Virtual screening of large compound spaces is an important approach for discovering high-quality lead compounds. Virtual drug screening technology, through computer-aided drug design, can rapidly identify lead compounds, such as potential molecules, that interact with specific targets from vast databases, shortening the drug development cycle. Summary of the Invention

[0003] The present disclosure provides a method for combining conformation prediction, a model training method, an apparatus, a device, and a storage medium.

[0004] According to one aspect of the present disclosure, a method for predicting binding conformations is provided, comprising:

[0005] Generate a collection of binding conformation samples based on the target protein;

[0006] Adjusting the parameters of the first binding conformation prediction model based on the binding conformation sample set to obtain a second binding conformation prediction model;

[0007] Inputting multiple candidate binding conformations corresponding to the target protein into the second binding conformation prediction model respectively to obtain a prediction result for each of the multiple candidate binding conformations;

[0008] Based on the prediction results of each candidate binding conformation, the target binding conformation of the target protein is obtained.

[0009] According to another aspect of the present disclosure, a method for training a binding conformation prediction model is provided, comprising:

[0010] Inputting the binding conformation sample into the first binding conformation prediction model to obtain prediction parameters of the binding conformation sample;

[0011] Determine the loss function based on the predicted parameters of the bound conformational samples;

[0012] adjusting the parameters of the first binding conformation prediction model based on the loss function;

[0013] When the loss function converges, a second binding conformation prediction model is obtained.

[0014] According to another aspect of the present disclosure, there is provided a binding conformation prediction apparatus, comprising:

[0015] A sample generation module, used to generate a binding conformation sample set based on the target protein;

[0016] a model adjustment module, configured to adjust parameters of the first binding conformation prediction model based on the binding conformation sample set to obtain a second binding conformation prediction model;

[0017] a candidate binding conformation prediction module, configured to input a plurality of candidate binding conformations corresponding to the target protein into a second binding conformation prediction model to obtain a prediction result for each of the plurality of candidate binding conformations;

[0018] The target binding conformation determination module is used to obtain the target binding conformation of the target protein based on the prediction results of each candidate binding conformation.

[0019] According to another aspect of the present disclosure, there is provided a combined conformation prediction model training device, comprising:

[0020] a sample prediction module, configured to input the binding conformation sample into the first binding conformation prediction model to obtain prediction parameters of the binding conformation sample;

[0021] A loss function determination module, for determining a loss function based on prediction parameters of the binding conformation samples;

[0022] a parameter adjustment module, configured to adjust parameters of the first binding conformation prediction model based on a loss function;

[0023] The model determination module is used to obtain a second binding conformation prediction model when the loss function converges.

[0024] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0025] at least one processor; and

[0026] a memory communicatively connected to the at least one processor; wherein,

[0027] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method of any embodiment of the present disclosure.

[0028] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method according to any embodiment of the present disclosure.

[0029] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method according to any embodiment of the present disclosure when executed by a processor.

[0030] In the disclosed embodiment, the parameters of the first binding conformation prediction model are first adjusted based on the binding conformation sample set generated for the target protein to obtain a second binding conformation prediction model, which can improve the accuracy of the prediction results of the second binding conformation prediction model for the binding conformation of the target protein.

[0031] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0033] Figure 1 is a schematic flow chart of a binding conformation prediction method according to an embodiment of the present disclosure;

[0034] Figure 2 is a schematic flow chart of a binding conformation prediction method according to another embodiment of the present disclosure;

[0035] Figure 3 is a schematic flow chart of a binding conformation prediction method according to another embodiment of the present disclosure;

[0036] Figure 4 is a schematic flow chart of a binding conformation prediction method according to another embodiment of the present disclosure;

[0037] Figure 5 is a schematic diagram of an application scenario of a method for combining conformation prediction according to an embodiment of the present disclosure;

[0038] Figure 6 is a schematic flow chart of a binding conformation prediction method according to another embodiment of the present disclosure;

[0039] Figure 7 is a schematic flow chart of a binding conformation prediction method according to another embodiment of the present disclosure;

[0040] Figure 8 is a schematic flow chart of a binding conformation prediction method according to another embodiment of the present disclosure;

[0041] Figure 9 is a schematic flow chart of a binding conformation prediction method according to another embodiment of the present disclosure;

[0042] Figure 10is a schematic flow chart of a binding conformation prediction method according to another embodiment of the present disclosure;

[0043] Figure 11a and Figure 11b is a schematic diagram of an application scenario of a method for combining conformation prediction according to an embodiment of the present disclosure;

[0044] Figure 12 is a flowchart of a method for training a combined conformation prediction model according to an embodiment of the present disclosure;

[0045] Figure 13 is a flowchart of a method for training a combined conformation prediction model according to another embodiment of the present disclosure;

[0046] Figure 14 is a flowchart of a method for training a combined conformation prediction model according to another embodiment of the present disclosure;

[0047] Figure 15 is a schematic structural diagram of a binding conformation prediction device according to an embodiment of the present disclosure;

[0048] Figure 16 is a schematic structural diagram of a binding conformation prediction device according to another embodiment of the present disclosure;

[0049] Figure 17 is a schematic structural diagram of a binding conformation prediction device according to another embodiment of the present disclosure;

[0050] Figure 18 This is a structural diagram of a combined conformation prediction model training device according to an embodiment of the present disclosure.

[0051] Figure 19 is a block diagram of an electronic device for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0052] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0053] Figure 1 FIG. 1 is a flow chart of a method for predicting conformations according to an embodiment of the present disclosure. The method may include:

[0054] S110, generating a binding conformation sample set based on the target protein.

[0055] S120 , adjusting parameters of the first binding conformation prediction model based on the binding conformation sample set to obtain a second binding conformation prediction model.

[0056] S130, inputting the multiple candidate binding conformations corresponding to the target protein into the second binding conformation prediction model respectively to obtain a prediction result for each of the multiple candidate binding conformations.

[0057] S140, based on the prediction results of each candidate binding conformation, a target binding conformation of the target protein is obtained.

[0058] The binding conformation prediction method of the embodiment of the present disclosure can be executed by an electronic device. The electronic device can be a terminal or a server. Exemplarily, the electronic device can be a terminal device with computing capabilities. Exemplarily, the electronic device can be a server in the cloud; the server can be a single server, or can be one or more servers in a server cluster, or can be a server in a distributed system (or can be called a computing node). It should be understood that the above is only an exemplary description of an electronic device, and actual processing may not be limited to the devices mentioned in the above examples. As long as the electronic device can execute the binding conformation prediction method provided by this embodiment, it is within the scope of protection of this embodiment.

[0059] In the disclosed embodiments, the binding conformation may include a binding conformation diagram of a target protein and a molecule. The target protein may also be referred to as a target, a target protein, a target protein, etc., or simply as a protein. The molecule may also be referred to as a drug molecule, a small molecule, etc. The binding conformation diagram of the target protein and the molecule may also be referred to as protein-small molecule complex data, protein-small molecule complex co-crystallization data, drug-protein complex co-crystallization data, etc. The molecules included in different binding conformation samples in the binding conformation sample set are different. A pre-set set of molecules and a target protein may be docked to generate a plurality of binding conformation samples to form a binding conformation sample set. The samples in the binding conformation sample set corresponding to different target proteins may be different. The first binding conformation prediction model is trained using the binding conformation sample set corresponding to the target protein, and the parameters of the first binding conformation prediction model may be fine-tuned so that the adjusted second binding conformation prediction model is more suitable for predicting the binding conformation corresponding to the target protein.

[0060] In an embodiment of the present disclosure, the first binding conformation prediction model and the second binding conformation prediction model may include a deep learning model, such as a neural network model.

[0061] In the disclosed embodiments, multiple candidate binding conformations can be obtained by docking a target protein with a molecular library. The molecular library may be different from a pre-set molecular set, and the number of molecules in the molecular library may be much larger than the number of molecules in the pre-set molecular set. The candidate binding conformations are input into the fine-tuned second binding conformation prediction model, and a prediction result of the candidate binding conformation can be output. The prediction result may include an affinity score for the target protein and the molecule in the candidate binding conformation. The specific form of the affinity score is not limited, for example, it can be 0-10 points, or 0-100 points, etc.

[0062] In the disclosed embodiments, after obtaining the prediction results for each candidate binding conformation, the target binding conformation can be selected based on the prediction results. For example, the one with the highest score among multiple candidate binding conformations can be used as the target binding conformation for the target protein. Multiple candidate binding conformations can also be ranked according to the prediction results. For example, they can be ranked in descending order of scores. Then, the target binding conformation for the target protein can be screened out based on the ranking results. For example, the top 20% of the candidate binding conformations can be used as the target binding conformation, or the top 10 candidate binding conformations in the prediction results can be used as the target binding conformation. It should be noted that this is only an exemplary description, and the specific screening method can be set according to actual conditions and is not limited here. Furthermore, based on the target binding conformation, the target molecule corresponding to the target protein can be determined. The target molecule has a high affinity for the target protein. Affinity can also be referred to as binding affinity, interaction force, etc.

[0063] In the disclosed embodiments, the parameters of a first binding conformation prediction model are first adjusted based on a set of binding conformation samples generated for a target protein to obtain a second binding conformation prediction model. This can improve the accuracy of the second binding conformation prediction model's prediction results for the target protein's binding conformation, thereby improving the accuracy of the target binding conformation. For example, during drug development, this solution can be used to perform a virtual drug screening process, which can screen out more accurate binding conformations among the binding conformations of the target protein and drug molecules, thereby identifying drug molecules with higher affinity for the target protein.

[0064] FIG2 is a schematic flow diagram of a method for predicting binding conformations according to another embodiment of the present disclosure, which may include one or more features of the aforementioned methods. In one embodiment, S110 generates a sample set of binding conformations based on a target protein, including:

[0065] S 210, generating a set of binding conformation samples based on the target protein and the molecules in the first molecule set by docking.

[0066] In the disclosed embodiments, the first molecule set may be a pre-set set of molecules used to generate training samples required for model fine-tuning. The first molecule set may be a small-scale molecule set (also referred to as a small-scale molecule library or a small-scale small molecule library, etc.), which can be set based on actual circumstances. For example, the small-scale molecule set may include 1,000 small molecules, which can be docked with the target protein to generate 1,000 binding conformation samples. The number of small molecules in this example is merely illustrative; in actual processing, the small-scale molecule set may include more or fewer molecules, and this is not a limitation here.

[0067] In the disclosed embodiments, docking can be implemented using one of the following docking algorithms: the Qvina2 algorithm, the Vina algorithm, or the Glide algorithm. This docking algorithm can be executed by a docking module. The user inputs information such as the target protein and the docking pocket (docking location) into the docking module. The docking module can then use the Qvina2 algorithm, the Vina algorithm, or the Glide algorithm, based on the target protein and docking pocket information, to dock with molecules in the first molecule set, generating binding conformation samples.

[0068] In some examples, based on the map of the target protein and the maps of multiple molecules in the first set of molecules, any of the above-described docking algorithms can generate multiple binding conformation samples, which can constitute a binding conformation sample set. A binding conformation sample can include a binding conformation map. For example, a map of a molecule can include a molecular atomic covalent map, etc. A map of the target protein can include an amino acid-level non-covalent map of the target protein.

[0069] In the disclosed embodiment, based on the target protein to be predicted and the molecules in the first molecule set, a binding conformation sample set is generated by docking, which can make the binding conformation sample set more relevant to the target protein, improve the accuracy of the second binding conformation prediction model in predicting the binding conformation of the target protein, and thus improve the accuracy of the virtual screening results.

[0070] FIG3 is a schematic flow diagram of a method for predicting binding conformations according to another embodiment of the present disclosure. This method may include one or more features of the aforementioned method for predicting binding conformations. In one embodiment, S120 adjusts the parameters of a first binding conformation prediction model based on a set of binding conformation samples to obtain a second binding conformation prediction model, including:

[0071] S310 , inputting the binding conformation samples in the binding conformation sample set into a first binding conformation prediction model to obtain prediction parameters of the binding conformation samples.

[0072] S320 , determining a loss function based on the predicted parameters of the bound conformational samples.

[0073] S330, adjusting parameters of the first binding conformation prediction model based on the loss function.

[0074] S340: When the loss function converges, a second binding conformation prediction model is obtained.

[0075] In the embodiment of the present disclosure, the prediction parameters output by the first binding conformation prediction model are related to the loss function formula. For example, if the loss function is determined based on the Gaussian formula, the prediction parameters may be relevant parameters in the Gaussian formula. For example, one binding conformation sample may correspond to a set of prediction parameters, and the loss function may be determined based on multiple sets of prediction parameters corresponding to multiple binding conformation samples. In the case that the value of the loss function calculated based on the prediction parameters does not meet the convergence condition, the parameters of the first binding conformation prediction model may be adjusted. After the adjustment, the model may be continuously trained using the samples until the value of the loss function meets the convergence condition, thereby obtaining a second binding conformation prediction model. For example, the convergence condition of the loss function may be: the value of the loss function is less than or equal to a threshold value.

[0076] In the disclosed embodiment, the above steps S310, S320, and S330 may be performed multiple times. The first binding conformation prediction model may be trained using a loss function determined by prediction parameters of the binding conformation samples, thereby improving the prediction accuracy of the trained second binding conformation prediction model.

[0077] FIG4 is a flow diagram of a method for predicting binding conformations according to another embodiment of the present disclosure. This method may include one or more features of the aforementioned method for predicting binding conformations. In one embodiment, S310 inputs a binding conformation sample from a binding conformation sample set into a first binding conformation prediction model to obtain prediction parameters for the binding conformation sample, including:

[0078] S410, inputting a covalent graph of molecular nodes included in the binding conformation sample into a first network of a first binding conformation prediction model to obtain features of the molecules included in the binding conformation sample;

[0079] S420, inputting the amino acid-level non-covalent map of the target protein included in the binding conformation sample into the second network of the first binding conformation prediction model to obtain features of the target protein included in the binding conformation sample;

[0080] S430, obtaining a splicing feature of the bound conformation sample based on features of the molecules contained in the bound conformation sample and features of the target protein;

[0081] S440 , inputting the splicing features of the binding conformation sample into the third network of the first binding conformation prediction model to obtain prediction parameters of the binding conformation sample.

[0082] In the disclosed embodiment, the binding conformation sample may include a binding conformation map, which may include a non-covalent map of the target protein at the amino acid level and a covalent map of the molecular nodes docked with the target protein.

[0083] like Figure 5 As shown, a molecular node covalent map (i.e., a molecular atomic covalent map) in a binding conformation map is input into a first network 501, which outputs the molecular features of the binding conformation map. The amino acid-level non-covalent map of the target protein in the binding conformation map is input into a second network 502, which obtains the target protein features of the binding conformation map. For example, the first network of the first binding conformation prediction model and / or the second network of the first binding conformation prediction model can be graph neural networks (GNNs). For another example, the first network of the first binding conformation prediction model and / or the second network of the first binding conformation prediction model can be composed of N graph neural networks (GNNs), where N is a positive integer greater than or equal to 1. The first binding conformation prediction model can be referred to as a molecular GNN×N, and the second binding conformation prediction model can be referred to as a protein GNN×N.

[0084] like Figure 5 As shown, the features of the molecule and the target protein in the bound conformational image are processed by a random dropout layer 503, a pooling layer 504, and a splicing layer 506 to obtain a spliced ​​feature of the bound conformational image. Specifically, the features of the molecule can be input into a first random dropout layer, and the output of the first random dropout layer can be input into a first pooling layer; the features of the target protein can be input into a second random dropout layer, and the output of the second random dropout layer can be input into a second pooling layer; the outputs of the first and second pooling layers are spliced ​​together by a splicing layer 506 to obtain a spliced ​​feature of the bound conformational image. For example, the first and / or second random dropout layers can be implemented using a dropout algorithm.

[0085] like Figure 5 As shown, the splicing features of the binding conformation map are input into the third network 507 of the first binding conformation prediction model to obtain prediction parameters for the binding conformation sample. The prediction parameters may include the Gaussian mean vector μ, standard deviation vector σ, and Gaussian coefficient π of the target protein node and the molecule node in the binding conformation map. For example, the third network may be a feed-forward neural network (FFN).

[0086] In the disclosed embodiment, the prediction parameters of the binding conformation sample are obtained more accurately based on the splicing features of the molecule features output by the first network and the features of the target protein output by the second network.

[0087] In another embodiment of the present disclosure, the first network includes a first node embedding layer, a first edge embedding layer, a first attention neural network, a first node hidden embedding layer and a first edge hidden embedding layer; the second network includes a second node embedding layer, a second edge embedding layer, a second attention neural network, a second node hidden embedding layer and a second edge hidden embedding layer.

[0088] In some examples, such as Figure 5 As shown, the molecular nodes in the molecular node covalent graph (i.e., the molecular atomic covalent graph) can be input into the first node embedding layer 5011 (i.e., the atomic embedding layer), and the edges between different molecular nodes can be input into the first edge embedding layer 5012. The outputs of the first node embedding layer and the first edge embedding layer are processed by the first attention neural network 5013, the first node hidden embedding layer 5014 (i.e., the atomic hidden embedding layer), and the first edge hidden embedding layer 5015 to obtain the features of each molecule node. The target protein nodes in the amino acid-level non-covalent graph of the target protein are input into the second node embedding layer 5021 (i.e., the atomic embedding layer), and the edges between different target protein nodes are input into the second edge embedding layer 5022. The outputs of the second node embedding layer and the second edge embedding layer are processed by the second attention neural network 5023, the second node hidden embedding layer 5024 (i.e., the atomic hidden embedding layer), and the second edge hidden embedding layer 5025 to obtain the features of each target protein node. The first attention neural network and / or the second attention neural network can be an attention graph neural network (Attention GGN).

[0089] In the embodiment of the present disclosure, through the node embedding layer, edge embedding layer, attention neural network, node hidden embedding layer and edge hidden embedding layer, etc., the node features of the molecules and the node features of the target protein extracted by the first network and the second network can be more accurate.

[0090] FIG6 is a flow chart of a method for predicting binding conformations according to another embodiment of the present disclosure, which may include one or more features of the aforementioned method for predicting binding conformations. In one embodiment, the prediction parameters include the mean vector, standard deviation vector, and Gaussian coefficient of the target protein node and the molecule node. S320 determines a loss function based on the prediction parameters of the binding conformation sample, including:

[0091] S 610, determining a Gaussian mixture function based on a mean vector, a standard deviation vector, and a Gaussian coefficient of the target protein node and the molecule node of the bound conformational sample, and a Euclidean distance between the target protein node and the molecule node;

[0092] S 620 , determining a likelihood value of a distance distribution between a target protein node and a molecule node based on a Gaussian mixture function, wherein the loss function is determined based on the likelihood values ​​of one or more binding conformation samples.

[0093] In the disclosed embodiment, the target protein node may include a target protein atom, and the molecular node may include a molecular atom. In some examples, the prediction parameters may be calculated based on the embedded feature vectors and linear transformation weights of the target protein node and the molecular node. For example: referring to Formula 1, the target protein node and the molecular node are embedded with the feature vector, and the mean vector linear transformation weight is substituted into the first nonlinear activation function to obtain the mean vector of the target protein node and the molecular node. For another example, referring to Formula 2, the target protein node and the molecular node are embedded with the feature vector, and the standard deviation vector linear transformation weight is substituted into the first nonlinear activation function to obtain the standard deviation vector of the target protein node and the molecular node combined with the conformation sample. For another example, referring to Formula 3, the target protein node and the molecular node are embedded with the feature vector, and the Gaussian coefficient linear transformation weight is substituted into the second nonlinear activation function to obtain the Gaussian coefficient of the target protein node and the molecular node combined with the conformation sample. The specific calculation formula is as follows:

[0094] Formula 1

[0095] in is the atomic number of the target protein, is the atomic number of the molecule, is the mean vector of target protein nodes and molecule nodes, is the first nonlinear activation function, is the linear transformation weight of the mean vector, Embed feature vectors for target protein nodes and molecule nodes.

[0096] Formula 2

[0097] in, is the standard deviation vector of target protein nodes and molecule nodes, is the linear transformation weight of the standard deviation vector.

[0098] Formula 3

[0099] in, is the Gaussian coefficient of the target protein node and the molecule node, is the second nonlinear activation function, is the Gaussian coefficient linear transformation weight.

[0100] In some examples, the first non-linear activation function is as follows:

[0101] Formula 4

[0102] Among them, when calculating the mean vector Can be , in the case of calculating the standard deviation vector Can be .

[0103] In some examples, the second non-linear activation function is as follows:

[0104] Formula 5

[0105] in, Can be used to transform the feature vector Normalization. . Can be .

[0106] In the embodiment of the present disclosure, a Gaussian mixture function can be determined based on the probability of the Euclidean distance under the mean vector condition of the target protein node and the molecular node, the normal distribution of the standard deviation vector of the target protein node and the molecular node, and the Gaussian coefficient of the target protein node and the molecular node.

[0107] A Gaussian mixture function An example of the calculation formula is as follows:

[0108] Formula 6

[0109] in, is the number of Gaussian distributions, is the index of the Gaussian distribution (indicating the n-th Gaussian distribution of p and l), is the atomic number of the target protein, is the atomic number of the molecule, is the Gaussian coefficient of the target protein node and the molecule node, is the mean vector of target protein nodes and molecule nodes, is the standard deviation vector of target protein nodes and molecule nodes, is a normal distribution, is the Euclidean distance between the target protein node and the molecule node, ∣ represents the conditional probability, The distance between the target protein node and the small molecule atom node is the mean vector , and the standard deviation is The probability under the distribution.

[0110] In the embodiment of the present disclosure, the likelihood value of the distance distribution between the target protein node and the molecular node is determined based on the Gaussian mixture function. An example of the calculation formula is as follows:

[0111] Formula 7

[0112] In the disclosed embodiment, the loss function may be calculated based on the likelihood value of a binding conformation sample in one binding conformation sample, or may be calculated based on the likelihood values ​​of binding conformation samples in multiple binding conformation samples by performing operations such as consecutive addition.

[0113] In the disclosed embodiments, the first binding conformation prediction model and the second structural conformation prediction model can be conformation rescoring models based on a mixed probability density function (Gaussian mixture function). The Gaussian mixture function is obtained by combining the mean vector, standard deviation vector, and Gaussian coefficient of the target protein node and the molecular node of the conformation sample, as well as the Euclidean distance between the target protein node and the molecular node, to determine the likelihood value of one or more binding conformation samples to obtain the loss function of the model. In this way, the use of the Gaussian mixture function can make the loss function of the model converge quickly, thereby improving the accuracy of the prediction of the second structural conformation model after training.

[0114] Figure 7 is a schematic flow diagram of a binding conformation prediction method according to another embodiment of the present disclosure, which may include one or more features of the aforementioned binding conformation prediction method. In one embodiment, S130 inputs multiple candidate binding conformations corresponding to the target protein into a second binding conformation prediction model, obtaining a prediction result for each of the multiple candidate binding conformations, including:

[0115] S710, inputting the candidate binding conformation into a second binding conformation prediction model to obtain prediction parameters of the binding conformation;

[0116] S720, based on the prediction parameters of the candidate binding conformations and the Gaussian mixture function, obtain a prediction result for each candidate binding conformation in multiple candidate binding conformations; wherein the Gaussian mixture function is determined based on the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node and the Euclidean distance between the target protein node and the molecular node.

[0117] In the disclosed embodiments, candidate binding conformations generated by molecular docking between the target protein and the second molecule set are input into a second binding conformation prediction model to obtain prediction parameters for the candidate binding conformations. The working principle of the second binding conformation prediction model is similar to that of the first binding conformation prediction model. For details, please refer to the above-mentioned process of using the first binding conformation prediction model to process the binding conformation samples to obtain prediction parameters. For the sake of brevity, this will not be repeated here.

[0118] In the embodiment of the present disclosure, a Gaussian mixture function The calculation formula of can be found in the above formula 6; the prediction results of the candidate binding conformations can be further calculated according to the Gaussian mixture function. The example of the formula of the prediction result is as follows:

[0119]

[0120] in, is the prediction result of the candidate binding conformation, is the total number of target protein nodes in a candidate binding conformation, is the index of the target protein node, is the total number of molecular nodes of a candidate binding conformation, The index of the molecule node for the candidate binding conformation.

[0121] In the disclosed embodiment, a prediction result for each of the plurality of candidate binding conformations is obtained based on the prediction parameters of the binding conformation obtained by the second binding conformation prediction model and a Gaussian mixture function. This can make the obtained prediction result more accurate.

[0122] Figure 8 is a flow chart of a binding conformation prediction method according to another embodiment of the present disclosure, which may include one or more features of the above-mentioned binding conformation prediction method. In one embodiment, S710 inputs a candidate binding conformation into a second binding conformation prediction model to obtain prediction parameters for the binding conformation, and obtains a prediction result for each candidate binding conformation in a plurality of candidate binding conformations, including:

[0123] S810, inputting the molecular node covalent graph included in the candidate binding conformation into the first network of the second binding conformation prediction model to obtain the characteristics of the molecules included in the candidate binding conformation.

[0124] S820, inputting the amino acid-level non-covalent map of the target protein included in the candidate binding conformation into the second network of the second binding conformation prediction model to obtain the characteristics of the target protein included in the candidate binding conformation.

[0125] S830: Based on the characteristics of the molecules contained in the candidate binding conformation and the characteristics of the target protein, obtain the splicing characteristics of the candidate binding conformation.

[0126] S840, inputting the splicing features of the candidate binding conformation into the third network of the second binding conformation prediction model to obtain prediction parameters of the candidate binding conformation.

[0127] In the disclosed embodiments, the candidate binding conformations may include a binding conformation map, which may include a non-covalent map of the target protein at the amino acid level and a covalent map of the molecular nodes docked with the target protein.

[0128] In the embodiment of the present disclosure, the model architecture of the second binding conformation prediction model is basically the same as that of the first network, the second network and the third network of the first binding conformation prediction model, and the parameters may be different. The covalent map of the molecular nodes in the binding conformation map of a candidate binding conformation is input into the first network, and the characteristics of the molecules in the binding conformation map can be output. The amino acid-level non-covalent map of the target protein in the binding conformation map is input into the second network to obtain the characteristics of the target protein in the binding conformation map. After the characteristics of the molecules in the binding conformation map and the characteristics of the target protein are processed by the random discard layer, the pooling layer and the splicing layer, the splicing characteristics of the binding conformation sample are obtained. The splicing characteristics of the binding conformation map are input into the third network of the first binding conformation prediction model to obtain the prediction parameters of the candidate binding conformation. The specific structure and working principle of the first network and the second network can be found in Figure 5 and its related descriptions.

[0129] In the disclosed embodiment, based on the splicing features of the molecule output by the first network and the target protein output by the second network, the prediction parameters of the candidate binding conformations are more accurate.

[0130] Figure 9 FIG2 is a flow chart of a method for combining conformation prediction according to another embodiment of the present disclosure. The method may include one or more features of the above-mentioned method for combining conformation prediction. In one embodiment, the method further includes:

[0131] S910, based on the target protein and the molecules in the second molecule set, generates multiple candidate binding conformations through docking.

[0132] In the embodiment of the present disclosure, the second molecule set can be obtained by using molecules in a molecule library input by a user, wherein the molecule library input by the user can be at least one group library among a plurality of molecule libraries.

[0133] In the embodiment of the present disclosure, the docking method can be implemented by using one of the following docking algorithms: Qvina2 algorithm, Vina algorithm, Glide algorithm. The docking algorithm can be executed by the docking module.

[0134] Users can input information such as target protein, docking pocket (docking position), etc. into the docking module. The docking module then uses the Qvina2 algorithm, Vina algorithm, or Glide algorithm to dock the target protein with each molecule in the second molecule set based on the target protein and docking pocket information, generating multiple candidate binding conformations.

[0135] In some examples, based on a map of the target protein and maps of multiple molecules in the second set of molecules, as well as parameters such as the binding pocket, multiple candidate binding conformations can be generated using any of the above-described docking algorithms. A candidate binding conformation can include a binding conformation map. For example, a map of a molecule can include a molecular atomic covalent map, etc. A map of the target protein can include an amino acid-level non-covalent map of the target protein.

[0136] In the disclosed embodiments, based on the target protein to be predicted and the molecules in the second molecule set, multiple binding conformations can be quickly generated through docking, thereby improving the efficiency of virtual screening results.

[0137] Figure 10 FIG2 is a flow chart of a method for predicting binding conformations according to another embodiment of the present disclosure, which may include one or more features of the aforementioned method for predicting binding conformations. In one embodiment, S140 obtains a target binding conformation of the target protein based on the prediction results for each candidate binding conformation, including:

[0138] S1010, based on the prediction results of each candidate binding conformation in the multiple candidate binding conformations and the binding mode between the target protein and the molecule, screen and obtain a target binding conformation of the target protein from the multiple candidate binding conformations; wherein the binding mode includes at least one of hydrogen bonding properties, hydrophobic properties, and binding distance properties.

[0139] In the disclosed embodiments, the target binding conformation can be determined directly based on the prediction results of each candidate binding conformation. For example, the candidate binding conformation ranked first according to the prediction results can be used as the target binding conformation. For another example, the top 20% of the candidate binding conformations can be used as the target binding conformation, or the top 10 candidate binding conformations can be used as the target binding conformation.

[0140] In the disclosed embodiments, target binding conformations can be screened according to the prediction results and binding modes of each candidate binding conformation. The binding mode can be input by the user or preset. After obtaining the prediction results for each candidate binding conformation in a plurality of candidate binding conformations, the target binding conformation is screened from the sorted plurality of candidate binding conformations according to the binding mode of the target protein and the molecule. For example, the candidate binding objects are sorted in descending order according to their affinity scores. Then, the target binding conformation of the target protein is screened according to the sorting results and the binding mode. For example, the binding conformation that meets the binding mode among the top 20% of the candidate binding conformations can be used as the target binding conformation, or the binding conformation that meets the binding mode among the top 10 candidate binding conformations in the prediction results can be used as the target binding conformation. It should be noted that this is only an exemplary description, and the specific screening method can be set according to the actual situation and is not limited here. Furthermore, the target molecule corresponding to the target protein can be determined based on the target binding conformation. The target molecule has a high affinity for the target protein.

[0141] In the disclosed embodiments, the process of screening a plurality of candidate binding conformations for a target protein and a molecule's binding patterns to identify a candidate binding conformation that matches the binding pattern can be performed by a binding pattern screening module. A user can input screening parameters such as a binding pattern into the binding pattern screening module. The binding pattern can include at least one of a hydrogen bonding property, a hydrophobicity property, and a binding distance property.

[0142] In the disclosed embodiments, candidate binding conformations can be further screened based on the binding mode between the target protein and the molecule to obtain the target binding conformation, thereby improving the accuracy of screening the target binding conformation.

[0143] Combine Figure 11a and Figure 11b The application scenarios of the above binding conformation prediction method are exemplified. Figure 11a and Figure 11b As shown, based on the above method, a high-performance virtual screening platform can be constructed. This platform can integrate molecular library preprocessing, molecular docking, rescoring, and post-screening functions in a single click. This platform is a small molecule drug virtual screening system that can sample multiple binding conformations and integrate high-precision rescoring models. Specifically, examples of the platform's integrated functions are as follows:

[0144] The functional modules of the user interface include a user input module and a user reception and output module. The user input module receives user input information 1101 through the user interface. This user input information includes at least one of a protein file (such as a user-specified target protein and docking pocket), a small molecule virtual screening library (i.e., a second set of molecules), and user-set screening parameters. The user-set screening parameters include at least one of the number of screenings and the binding mode (also referred to as the target binding mode). The user reception and output module receives the final structure.

[0145] The functional modules of the virtual screening process include:

[0146] Data preprocessing module: Receives user-provided protein files and small molecule virtual screening libraries for preprocessing. For example, the data preprocessing module includes a protein preprocessing module 1102 that can standardize the target protein and docking pocket in the protein file to obtain preprocessed target protein and docking pocket.

[0147] Docking modules, such as the Qvina docking module 1103, integrate the faster Qvina2 to perform high-precision conformational search and generate multiple conformations. After obtaining the pre-processed target protein and docking pocket from the pre-processing module, the Qvina docking module 1103 uses the Qvina2 algorithm to dock the specified target protein with molecules in the small molecule virtual screening library based on the pre-processed target protein and docking pocket information, and can obtain multiple candidate binding conformations, such as Figure 11a Optionally, the designated target protein is first docked with the molecules in the small-scale molecular library (first molecular set) 1104 to generate a set of binding conformation samples for fine-tuning the model, as shown in FIG. Figure 11b shown.

[0148] The rescoring module, such as the AI ​​rescoring ranking module 1105, adopts a conformational AI rescoring model based on a mixed probability density function. The model is pre-trained with large-scale protein-small molecule complex data, and uses the protein file input by the user to perform fine-tuning for a certain number of training steps through a built-in high-quality small-scale molecule library (or called a small-scale small molecule library, docking small molecule library, etc.). The AI ​​rescoring model scores the multiple binding conformations generated by the aforementioned docking module based on the small molecule virtual screening library and selects the conformation with the highest score. Figure 11b As shown, after generating the binding conformation sample set, the parameters of the AI ​​rescoring ranking model 1105 (i.e., the first binding conformation prediction model) can be adjusted based on the binding conformation sample set to obtain the adjusted AI rescoring ranking model 1105 (i.e., the second binding conformation prediction model). This method is also called fine-tuning the AI ​​rescoring ranking model 1105 through self-distillation technology. Figure 11a and Figure 11bAs shown, after generating multiple candidate binding conformations, the multiple candidate binding conformations are input into the adjusted AI rescoring ranking model 1105 to obtain prediction results of the multiple candidate binding conformations, and the multiple candidate binding conformations are ranked according to the prediction results.

[0149] The binding mode (interaction) screening module 1107 calculates interactions such as hydrogen bonds and hydrophobicity based on the results of the AI ​​model screening in the previous step, supports user-defined interaction sites, and thus screens out binding modes that meet the user's expectations. After sorting multiple candidate binding conformations, it is determined whether to perform binding mode screening. If the screening parameters entered by the user do not include a binding mode, no binding mode screening is performed and the final result is output directly. If binding mode screening is performed, the sorted multiple candidate binding conformations are input into the binding mode screening module 1107, and the final result 1108 (i.e., target binding conformation or target molecule) of the target protein is screened based on the screening parameters entered by the user, such as the binding mode and screening number, and the final result 1108 is sent to the user. For example, if the user enters a screening number of 10, this can determine that the candidate binding conformations ranked in the top 10 of the predicted results are used as the target binding conformations.

[0150] If the user input information obtained through the user interface includes screening parameters such as binding mode and number of screenings, multiple candidate binding conformations are input into the binding mode screening module 1107. Based on the binding mode, the mode screening module 1107 identifies candidate binding conformations that match the binding mode from the multiple candidate binding conformations. For example, interactions (binding modes) such as hydrogen bonds and / or hydrophobicity are calculated, and user-defined interaction sites are supported. This allows the final result 1108 (target binding conformation or target molecule) to be screened for binding modes that meet the user's expectations. Final result 1108 is then sent to the user interface.

[0151] In the disclosed embodiment, by performing a small-scale docking process on the target (target protein) (for example, using Qvina2 to dock the specified target protein with the molecules in the first molecule set 1104 to generate a set of binding conformation samples), the binding conformation samples are used to fine-tune the re-scoring model, thereby improving the virtual screening accuracy of the binding conformation, and by adopting the Qvina2 docking step, the speed can be greatly improved without affecting the accuracy.

[0152] This solution enables the construction of a user-friendly, high-performance virtual screening platform for drug development. This platform utilizes an improved deep learning model, integrated with a deep learning-based scoring function, and trained on large-scale protein-small molecule complex data. Self-distillation technology further enhances the accuracy of virtual screening. Computational speed is increased by adapting a faster and more accurate molecular docking algorithm, achieving simultaneous improvements in both accuracy and speed. This solution has the following features:

[0153] The model was upgraded and adopted a conformational rescoring model based on a mixed probability density function. The training data included all known drug-protein complex co-crystallization data. The effect was significantly better than the affinity model and the baseline model.

[0154] Self-distillation technology is introduced to conduct a small-scale docking process on the target, and the obtained conformation is used to fine-tune the re-scoring model to further improve the virtual screening accuracy of the target.

[0155] Multiple docking conformations are generated, and the optimal conformational structure (target binding conformation) is selected through the re-scoring model, which brings a significant improvement in the accuracy of the virtual screening task.

[0156] The docking algorithm has been upgraded, and the docking speed has been greatly improved without affecting the accuracy.

[0157] This solution can provide a virtual screening platform that is portable, robust, user-friendly, and has advantages in speed and accuracy. Users only need to prepare customized protein files and small molecule virtual screening libraries for screening, set the ratios of the above modules as needed, and then they can run it with one-click convenience.

[0158] This solution can use faster docking software to achieve conformation generation and search, which greatly improves the overall screening efficiency. At the same time, the AI ​​model has been upgraded. The model is trained with large-scale protein-small molecule complex data and based on the interaction information between protein and small molecule, multiple conformation scores are predicted through mixed probability density functions. The optimal conformation is selected according to the score. The overall model structure diagram is shown in Figure 5 . Combining multi-conformational input and AI model scoring, the accuracy of virtual screening is greatly improved. The main application areas of this solution include virtual screening of small molecule drugs for a certain protein target. Due to its high-performance computing characteristics, it can screen large-scale (hundreds of millions) molecular virtual screening libraries to achieve large-scale virtual screening. At the same time, it integrates deep learning models obtained through large-scale training data, and its virtual screening accuracy reaches a high level. It also has user-friendly features and is suitable for pharmaceutical companies, pharmaceutical chemistry researchers, etc. in the field of drug research and development. Compared with traditional solutions, this platform has accuracy advantages in the virtual screening tasks of multiple protein targets in the kinase family.

[0159] The present disclosure provides a method for training a conformation prediction model. Figure 12 As shown, including:

[0160] S1210 , inputting the binding conformation sample into a first binding conformation prediction model to obtain prediction parameters of the binding conformation sample.

[0161] S1220, determining a loss function based on the predicted parameters of the bound conformational samples.

[0162] S1230, adjusting parameters of the first binding conformation prediction model based on the loss function.

[0163] S1240: When the loss function converges, a second binding conformation prediction model is obtained.

[0164] In the embodiments of the present disclosure, the combined conformation prediction model training method can be executed by an electronic device. The electronic device can be a terminal or a server. Exemplarily, the electronic device can be a terminal device with computing capabilities. Exemplarily, the electronic device can be a server in the cloud; the server can be a single server, or can be one or more servers in a server cluster, or can be a server in a distributed system (or can be called a computing node). It should be understood that the above is only an exemplary description of an electronic device, and the actual processing may not be limited to the devices mentioned in the above examples. As long as the electronic device can execute the combined conformation prediction model training method provided in this embodiment, it is within the protection scope of this embodiment.

[0165] In the disclosed embodiments, the binding conformation may include a binding conformation diagram of the target protein and the molecule. The binding conformation diagram is the same as the binding conformation diagram in the above-mentioned binding conformation prediction method embodiment, and is not described here for brevity.

[0166] In the embodiments of the present disclosure, the binding conformation samples used for initial training may be samples from a large-scale binding conformation sample set. The binding conformation samples used for fine-tuning may be samples from a small-scale binding conformation sample set. Among them, the large-scale binding conformation sample set may include all known drug-protein complex co-crystallization data, such as binding conformation maps of all known proteins and molecules. The number of samples in the large-scale binding conformation sample set is much larger than the number of samples used for fine-tuning the model in the above embodiments. For example: a large-scale binding conformation sample set may include 1 million binding conformation samples. The number of samples in the large-scale binding conformation sample set in this example is only for illustrative purposes. In actual processing, the large-scale binding conformation sample set may include more or fewer samples, which is not limited here.

[0167] In the disclosed embodiment, each binding conformation sample in a large-scale binding conformation sample set can be input into the first binding conformation prediction model to obtain prediction parameters for each binding conformation sample; and a loss function is determined based on the prediction parameters for each binding conformation sample.

[0168] In the embodiment of the present disclosure, S1210~S1240 and Figure 3 The processing methods for S310~S340 are basically the same and will not be described here.

[0169] In the embodiment of the present disclosure, the above steps S1210 to S1230 can be performed multiple times. The first binding conformation prediction model can be trained using a loss function determined by the prediction parameters of the binding conformation sample, thereby improving the accuracy of the prediction of the second binding conformation prediction model after training. If the initial training data includes all known drug-protein complex co-crystallization data, such as the binding conformation maps of all known drug molecules and proteins, the accuracy of the prediction of the binding conformation prediction model can be improved. If the training data used for fine-tuning includes a small-scale training sample corresponding to the target protein, the accuracy of the trained prediction model can be fine-tuned so that the prediction accuracy of the prediction model for the binding conformation corresponding to the target protein can be further improved.

[0170] Figure 13 2 is a flow chart illustrating a method for training a bound conformation prediction model according to another embodiment of the present disclosure. This method may include one or more features of the aforementioned method for training a bound conformation prediction model. In one embodiment, S1210 inputs a bound conformation sample into a first bound conformation prediction model to obtain prediction parameters for the bound conformation sample, including:

[0171] S1310, inputting the molecular node covalent graph included in the binding conformation sample into the first network of the first binding conformation prediction model to obtain features of the molecules included in the binding conformation sample;

[0172] S1320, inputting the amino acid-level non-covalent map of the target protein included in the binding conformation sample into the second network of the first binding conformation prediction model to obtain features of the target protein included in the binding conformation sample;

[0173] S1330, obtaining a splicing feature of the bound conformation sample based on features of the molecules contained in the bound conformation sample and features of the target protein;

[0174] S1340: Input the splicing features of the binding conformation sample into the third network of the first binding conformation prediction model to obtain prediction parameters of the binding conformation sample.

[0175] Among them, S1310~S1340 and Figure 4 The processing methods for S410~S440 are basically the same. For the specific processing process, see Figure 5 And the related descriptions will not be repeated here.

[0176] In another embodiment of the present disclosure, the first network includes a first node embedding layer, a first edge embedding layer, a first attention neural network, a first node hidden embedding layer, and a first edge hidden embedding layer; the second network includes a second node embedding layer, a second edge embedding layer, a second attention neural network, a second node hidden embedding layer, and a second edge hidden embedding layer.

[0177] In the embodiment of the present disclosure, the first network includes the first node embedding layer, the first edge embedding layer, the first attention neural network, the first node hidden embedding layer and the first edge hidden embedding layer. See Figure 5 and its related description, as well as the processing method of the second network including the second node embedding layer, the second edge embedding layer, the second attention neural network, the second node hidden embedding layer and the second edge hidden embedding layer on the amino acid level non-covalent graph of the target protein, see Figure 5 And its related description. No further details will be given here.

[0178] In the embodiment of the present disclosure, the node features of the molecules and the node features of the target protein extracted using the first network and the second network are more accurate.

[0179] Figure 14 This is a flow chart of a method for training a bound conformation prediction model according to another embodiment of the present disclosure. This method may include one or more features of the aforementioned method for training a bound conformation prediction model. In one embodiment, the prediction parameters include the mean vector, standard deviation vector, and Gaussian coefficient of the target protein node and the molecule node. S1220 determines a loss function based on the prediction parameters of the bound conformation sample, including:

[0180] S 1410, determining a Gaussian mixture function based on a mean vector, a standard deviation vector, and a Gaussian coefficient of the target protein node and the molecule node of the bound conformational sample, and a Euclidean distance between the target protein node and the molecule node;

[0181] S1420, based on the Gaussian mixture function, determining the likelihood value of the distance distribution between the target protein node and the molecule node, and the loss function is determined based on the likelihood value of one or more binding conformation samples.

[0182] Among them, S1410~S1430 and Figure 6 The processing methods for S610~S630 are basically the same and will not be described here.

[0183] In the embodiment of the present disclosure, the Gaussian mixture function is shown in the above formula 6, and the likelihood value of the distance distribution between the target protein node and the molecule node is determined based on the Gaussian mixture function. Detailed description is omitted here.

[0184] In the disclosed embodiments, a loss function is obtained by combining the likelihood of one or more bound conformational samples using a Gaussian mixture function derived from the mean vector, standard deviation vector, and Gaussian coefficient of the target protein node and molecule node of the conformational sample, as well as the Euclidean distance between the target protein node and the molecule node. This can make the resulting loss function more accurate, thereby improving the accuracy of the prediction of the trained second structural conformational model.

[0185] Figure 15: is a schematic structural diagram of a binding conformation prediction device according to an embodiment of the present disclosure, comprising:

[0186] A sample generation module 1501 is used to generate a binding conformation sample set based on the target protein;

[0187] A model adjustment module 1502 is configured to adjust parameters of the first binding conformation prediction model based on the binding conformation sample set to obtain a second binding conformation prediction model;

[0188] The candidate binding conformation prediction module 1503 is used to input the multiple candidate binding conformations corresponding to the target protein into the second binding conformation prediction model to obtain a prediction result for each of the multiple candidate binding conformations;

[0189] The target binding conformation determination module 1504 is used to obtain the target binding conformation of the target protein based on the prediction results of each candidate binding conformation.

[0190] In one embodiment, the sample generation module 1501 is used to generate a binding conformation sample set based on the target protein and the molecules in the first molecule set by docking.

[0191] In one embodiment, Figure 16 As shown, the model adjustment module 1502 includes:

[0192] The sample prediction module 1601 is used to input the binding conformation samples in the binding conformation sample set into the first binding conformation prediction model to obtain prediction parameters of the binding conformation samples;

[0193] a loss function determination module 1602 for determining a loss function based on the prediction parameters of the bound conformational samples;

[0194] a parameter adjustment module 1603, configured to adjust the parameters of the first binding conformation prediction model based on the loss function;

[0195] The model determination module 1604 is used to obtain a second binding conformation prediction model when the loss function converges.

[0196] In one embodiment, the sample prediction module 1601 is used to input the covalent map of the molecular nodes contained in the binding conformation sample into the first network of the first binding conformation prediction model to obtain the characteristics of the molecules contained in the binding conformation sample; input the amino acid-level non-covalent map of the target protein contained in the binding conformation sample into the second network of the first binding conformation prediction model to obtain the characteristics of the target protein contained in the binding conformation sample; based on the characteristics of the molecules contained in the binding conformation sample and the characteristics of the target protein, obtain the splicing characteristics of the binding conformation sample; input the splicing characteristics of the binding conformation sample into the third network of the first binding conformation prediction model to obtain the prediction parameters of the binding conformation sample.

[0197] In one embodiment, the prediction parameters include the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node. The loss function determination module 1602 is used to determine the Gaussian mixture function based on the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node of the binding conformation sample, and the Euclidean distance between the target protein node and the molecular node; based on the Gaussian mixture function, determine the likelihood value of the distance distribution between the target protein node and the molecular node, and the loss function is determined based on the likelihood value of one or more binding conformation samples.

[0198] In one embodiment, the candidate binding conformation prediction module 1503 is used to input the candidate binding conformation into the second binding conformation prediction model to obtain prediction parameters of the binding conformation; based on the prediction parameters of the candidate binding conformation and the Gaussian mixture function, obtain the prediction results of each candidate binding conformation in multiple candidate binding conformations; wherein the Gaussian mixture function is determined based on the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node and the Euclidean distance between the target protein node and the molecular node.

[0199] In one embodiment, the candidate binding conformation prediction module 1503 is used to input the covalent map of the molecular nodes contained in the candidate binding conformation into the first network of the second binding conformation prediction model to obtain the characteristics of the molecules contained in the candidate binding conformation; input the amino acid-level non-covalent map of the target protein contained in the candidate binding conformation into the second network of the second binding conformation prediction model to obtain the characteristics of the target protein contained in the candidate binding conformation; based on the characteristics of the molecules contained in the candidate binding conformation and the characteristics of the target protein, obtain the splicing characteristics of the candidate binding conformation; input the splicing characteristics of the candidate binding conformation into the third network of the second binding conformation prediction model to obtain the prediction parameters of the candidate binding conformation.

[0200] In one embodiment, the first network includes a first node embedding layer, a first edge embedding layer, a first attention neural network, a first node hidden embedding layer, and a first edge hidden embedding layer; the second network includes a second node embedding layer, a second edge embedding layer, a second attention neural network, a second node hidden embedding layer, and a second edge hidden embedding layer.

[0201] like Figure 17 As shown, the binding conformation prediction device also includes:

[0202] The candidate binding conformation generation module 1701 is used to generate multiple candidate binding conformations based on the target protein and the molecules in the second molecule set through docking.

[0203] In one embodiment, the target binding conformation determination module 1504 is used to screen the target binding conformation of the target protein from multiple candidate binding conformations based on the prediction results of each candidate binding conformation in the multiple candidate binding conformations and the binding mode between the target protein and the molecule; wherein the binding mode includes at least one of hydrogen bonding properties, hydrophobic properties, and binding distance properties.

[0204] Figure 18 : is a structural diagram of a combined conformation prediction model training device according to an embodiment of the present disclosure, comprising:

[0205] The sample prediction module 1801 is used to input the binding conformation sample into the first binding conformation prediction model to obtain prediction parameters of the binding conformation sample;

[0206] a loss function determination module 1802 for determining a loss function based on the prediction parameters of the bound conformational samples;

[0207] a parameter adjustment module 1803, configured to adjust the parameters of the first binding conformation prediction model based on the loss function;

[0208] The model determination module 1804 is used to obtain a second binding conformation prediction model when the loss function converges.

[0209] In one embodiment, the sample prediction module 1801 is used to input the covalent map of the molecular nodes contained in the binding conformation sample into the first network of the first binding conformation prediction model to obtain the characteristics of the molecules contained in the binding conformation sample; input the amino acid-level non-covalent map of the target protein contained in the binding conformation sample into the second network of the first binding conformation prediction model to obtain the characteristics of the target protein contained in the binding conformation sample; based on the characteristics of the molecules contained in the binding conformation sample and the characteristics of the target protein, obtain the splicing characteristics of the binding conformation sample; input the splicing characteristics of the binding conformation sample into the third network of the first binding conformation prediction model to obtain the prediction parameters of the binding conformation sample.

[0210] In one embodiment, the prediction parameters include the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node. The loss function determination module 1802 is used to determine the Gaussian mixture function based on the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node of the binding conformation sample, and the Euclidean distance between the target protein node and the molecular node; based on the Gaussian mixture function, the likelihood value of the distance distribution between the target protein node and the molecular node is determined, and the loss function is determined based on the likelihood value of one or more binding conformation samples.

[0211] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0212] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0213] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0214] Figure 19 A schematic block diagram of an example electronic device 1900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0215] like Figure 19 As shown, device 1900 includes a computing unit 1901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1902 or a computer program loaded from a storage unit 1908 into a random access memory (RAM) 1903. RAM 1903 may also store various programs and data required for the operation of device 1900. Computing unit 1901, ROM 1902, and RAM 1903 are connected to each other via a bus 1904. An input / output (I / O) interface 1905 is also connected to bus 1904.

[0216] Various components in device 1900 are connected to I / O interface 1905, including: an input unit 1906, such as a keyboard, mouse, etc.; an output unit 1907, such as various types of displays, speakers, etc.; a storage unit 1908, such as a magnetic disk, optical disk, etc.; and a communication unit 1909, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1909 allows device 1900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0217] Computing unit 1901 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1901 performs the various methods and processes described above. For example, in some embodiments, the above-described methods may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1908. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 1900 via ROM 1902 and / or communication unit 1909. When the computer program is loaded into RAM 1903 and executed by computing unit 1901, one or more steps of the above-described methods may be performed. Alternatively, in other embodiments, computing unit 1901 may be configured to perform the above-described methods in any other suitable manner (e.g., via firmware).

[0218] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0219] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0220] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0221] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0222] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0223] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0224] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0225] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for predicting binding conformation, comprising: Generate a collection of binding conformation samples based on the target protein; Adjusting parameters of the first binding conformation prediction model based on the binding conformation sample set to obtain a second binding conformation prediction model; Inputting the plurality of candidate binding conformations corresponding to the target protein into the second binding conformation prediction model respectively to obtain a prediction result for each of the plurality of candidate binding conformations; Based on the prediction results of each candidate binding conformation, obtaining the target binding conformation of the target protein; Wherein, the multiple candidate binding conformations corresponding to the target protein are respectively input into the second binding conformation prediction model to obtain the prediction results of each candidate binding conformation in the multiple candidate binding conformations, including: inputting the candidate binding conformation into the second binding conformation prediction model to obtain the prediction parameters of the binding conformation; based on the prediction parameters of the candidate binding conformation and a Gaussian mixture function, obtaining the prediction results of each candidate binding conformation in the multiple candidate binding conformations; wherein, the Gaussian mixture function is determined based on the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node and the Euclidean distance between the target protein node and the molecular node.

2. The method according to claim 1, wherein The method of generating a binding conformation sample set based on the target protein includes: Based on the target protein and the molecules in the first molecule set, the binding conformation sample set is generated by docking.

3. The method according to claim 1, wherein The adjusting the parameters of the first binding conformation prediction model based on the binding conformation sample set to obtain the second binding conformation prediction model includes: Inputting the binding conformation samples in the binding conformation sample set into the first binding conformation prediction model to obtain prediction parameters of the binding conformation samples; determining a loss function based on the predicted parameters of the bound conformation samples; adjusting parameters of the first binding conformation prediction model based on the loss function; When the loss function converges, the second binding conformation prediction model is obtained.

4. The method according to claim 3, wherein: The step of inputting the binding conformation samples in the binding conformation sample set into the first binding conformation prediction model to obtain prediction parameters of the binding conformation samples comprises: Inputting the molecular node covalent graph included in the binding conformation sample into the first network of the first binding conformation prediction model to obtain characteristics of the molecules included in the binding conformation sample; inputting the amino acid-level non-covalent map of the target protein contained in the binding conformation sample into the second network of the first binding conformation prediction model to obtain the characteristics of the target protein contained in the binding conformation sample; Obtaining a splicing feature of the binding conformation sample based on features of the molecules contained in the binding conformation sample and features of the target protein; The splicing features of the binding conformation sample are input into the third network of the first binding conformation prediction model to obtain the prediction parameters of the binding conformation sample.

5. The method according to claim 3, wherein: The prediction parameters include the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecule node, and the loss function is determined based on the prediction parameters of the binding conformation sample, including: Determining a Gaussian mixture function based on the mean vector, standard deviation vector, and Gaussian coefficient of the target protein node and the molecular node of the binding conformation sample, as well as the Euclidean distance between the target protein node and the molecular node; Based on the Gaussian mixture function, the likelihood value of the distance distribution between the target protein node and the molecule node is determined, and the loss function is determined based on the likelihood value of one or more binding conformation samples.

6. The method according to any one of claims 1 to 5, wherein Inputting the candidate binding conformation into a second binding conformation prediction model to obtain prediction parameters of the binding conformation comprises: Inputting the molecular node covalent graph included in the candidate binding conformation into the first network of the second binding conformation prediction model to obtain the characteristics of the molecules included in the candidate binding conformation; Inputting the amino acid-level non-covalent map of the target protein included in the candidate binding conformation into the second network of the second binding conformation prediction model to obtain the characteristics of the target protein included in the candidate binding conformation; Based on the characteristics of the molecules included in the candidate binding conformation and the characteristics of the target protein, obtaining the splicing characteristics of the candidate binding conformation; The splicing features of the candidate binding conformation are input into the third network of the second binding conformation prediction model to obtain the prediction parameters of the candidate binding conformation.

7. The method according to claim 4, wherein: The first network includes a first node embedding layer, a first edge embedding layer, a first attention neural network, a first node hidden embedding layer, and a first edge hidden embedding layer; The second network includes a second node embedding layer, a second edge embedding layer, a second attention neural network, a second node hidden embedding layer, and a second edge hidden embedding layer.

8. The method according to any one of claims 1 to 5, wherein The method further comprises: Based on the target protein and the molecules in the second molecule set, a plurality of candidate binding conformations are generated by docking.

9. The method according to any one of claims 1 to 5, wherein The step of obtaining the target binding conformation of the target protein based on the prediction results of each candidate binding conformation comprises: Based on the prediction results of each candidate binding conformation in the multiple candidate binding conformations and the binding mode between the target protein and the molecule, a target binding conformation of the target protein is screened from the multiple candidate binding conformations; wherein the binding mode includes at least one of hydrogen bonding properties, hydrophobic properties, and binding distance properties.

10. A method for training a bound conformation prediction model, comprising: Inputting the binding conformation sample into the first binding conformation prediction model to obtain prediction parameters of the binding conformation sample; determining a loss function based on the predicted parameters of the bound conformation samples; adjusting parameters of the first binding conformation prediction model based on the loss function; When the loss function converges, a second binding conformation prediction model is obtained; Wherein, the prediction parameters include the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node, and the prediction parameters based on the binding conformation sample are used to determine the loss function, including: determining a Gaussian mixture function based on the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node of the binding conformation sample, and the Euclidean distance between the target protein node and the molecular node; determining the likelihood value of the distance distribution between the target protein node and the molecular node based on the Gaussian mixture function, and the loss function is determined based on the likelihood value of one or more binding conformation samples.

11. The method according to claim 10, wherein: The step of inputting the binding conformation sample into the first binding conformation prediction model to obtain prediction parameters of the binding conformation sample comprises: Inputting the molecular node covalent graph included in the binding conformation sample into the first network of the first binding conformation prediction model to obtain characteristics of the molecules included in the binding conformation sample; inputting the amino acid-level non-covalent map of the target protein contained in the binding conformation sample into the second network of the first binding conformation prediction model to obtain the characteristics of the target protein contained in the binding conformation sample; Obtaining a splicing feature of the binding conformation sample based on features of the molecules contained in the binding conformation sample and features of the target protein; The splicing features of the binding conformation sample are input into the third network of the first binding conformation prediction model to obtain the prediction parameters of the binding conformation sample.

12. A binding conformation prediction device comprising: A sample generation module, used to generate a binding conformation sample set based on the target protein; a model adjustment module, configured to adjust the parameters of the first binding conformation prediction model based on the binding conformation sample set to obtain a second binding conformation prediction model; a candidate binding conformation prediction module, configured to input the plurality of candidate binding conformations corresponding to the target protein into the second binding conformation prediction model, respectively, to obtain a prediction result for each of the plurality of candidate binding conformations; a target binding conformation determination module, configured to obtain a target binding conformation of the target protein based on the prediction results of each candidate binding conformation; The candidate binding conformation prediction module is used to input the candidate binding conformation into the second binding conformation prediction model to obtain the prediction parameters of the binding conformation; based on the prediction parameters of the candidate binding conformation and the Gaussian mixture function, obtain the prediction results of each candidate binding conformation in the multiple candidate binding conformations; wherein, the Gaussian mixture function is determined based on the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node, and the Euclidean distance between the target protein node and the molecular node.

13. The device according to claim 12, wherein The sample generation module is used to generate the binding conformation sample set by docking based on the target protein and the molecules in the first molecule set.

14. The device according to claim 12, wherein The model adjustment module includes: a sample prediction module, configured to input the binding conformation samples in the binding conformation sample set into the first binding conformation prediction model to obtain prediction parameters of the binding conformation samples; a loss function determination module, configured to determine a loss function based on the predicted parameters of the binding conformation sample; a parameter adjustment module, configured to adjust the parameters of the first binding conformation prediction model based on the loss function; A model determination module is used to obtain the second binding conformation prediction model when the loss function converges.

15. The device according to claim 14, wherein The sample prediction module is used to input the molecular node covalent graph contained in the binding conformation sample into the first network of the first binding conformation prediction model to obtain the characteristics of the molecules contained in the binding conformation sample; The amino acid-level non-covalent map of the target protein contained in the binding conformation sample is input into the second network of the first binding conformation prediction model to obtain the characteristics of the target protein contained in the binding conformation sample; based on the characteristics of the molecules contained in the binding conformation sample and the characteristics of the target protein, the splicing characteristics of the binding conformation sample are obtained; the splicing characteristics of the binding conformation sample are input into the third network of the first binding conformation prediction model to obtain the prediction parameters of the binding conformation sample.

16. The device according to claim 14, wherein The prediction parameters include the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node. The loss function determination module is used to determine the Gaussian mixture function based on the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node of the binding conformation sample, and the Euclidean distance between the target protein node and the molecular node; based on the Gaussian mixture function, determine the likelihood value of the distance distribution between the target protein node and the molecular node, and the loss function is determined based on the likelihood value of one or more binding conformation samples.

17. The device according to any one of claims 12 to 16, wherein The candidate binding conformation prediction module is used to input the molecular node covalent map contained in the candidate binding conformation into the first network of the second binding conformation prediction model to obtain the characteristics of the molecules contained in the candidate binding conformation; input the amino acid-level non-covalent map of the target protein contained in the candidate binding conformation into the second network of the second binding conformation prediction model to obtain the characteristics of the target protein contained in the candidate binding conformation; based on the characteristics of the molecules contained in the candidate binding conformation and the characteristics of the target protein, obtain the splicing characteristics of the candidate binding conformation; input the splicing characteristics of the candidate binding conformation into the third network of the second binding conformation prediction model to obtain the prediction parameters of the candidate binding conformation.

18. The device according to claim 15, wherein The first network includes a first node embedding layer, a first edge embedding layer, a first attention neural network, a first node hidden embedding layer, and a first edge hidden embedding layer; The second network includes a second node embedding layer, a second edge embedding layer, a second attention neural network, a second node hidden embedding layer, and a second edge hidden embedding layer.

19. The device according to any one of claims 12 to 16, wherein Also includes: The candidate binding conformation generation module is used to generate multiple candidate binding conformations through docking based on the target protein and the molecules in the second molecule set.

20. The device according to any one of claims 12 to 16, wherein The target binding conformation determination module is used to screen the target binding conformation of the target protein from the multiple candidate binding conformations based on the prediction results of each candidate binding conformation in the multiple candidate binding conformations and the binding mode between the target protein and the molecule; wherein the binding mode includes at least one of hydrogen bonding properties, hydrophobic properties, and binding distance properties.

21. A combined conformation prediction model training device, comprising: a sample prediction module, configured to input a binding conformation sample into a first binding conformation prediction model to obtain prediction parameters of the binding conformation sample; a loss function determination module, configured to determine a loss function based on the predicted parameters of the binding conformation sample; a parameter adjustment module, configured to adjust the parameters of the first binding conformation prediction model based on the loss function; A model determination module, configured to obtain a second binding conformation prediction model when the loss function converges; Among them, the prediction parameters include the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node, and the loss function determination module is used to determine the Gaussian mixture function based on the mean vector, standard deviation vector and Gaussian coefficient of the target protein node and the molecular node of the binding conformation sample, and the Euclidean distance between the target protein node and the molecular node; based on the Gaussian mixture function, the likelihood value of the distance distribution between the target protein node and the molecular node is determined, and the loss function is determined based on the likelihood value of one or more binding conformation samples.

22. The device according to claim 21, wherein The sample prediction module is used to input the molecular node covalent graph contained in the binding conformation sample into the first network of the first binding conformation prediction model to obtain the characteristics of the molecules contained in the binding conformation sample; The amino acid-level non-covalent map of the target protein contained in the binding conformation sample is input into the second network of the first binding conformation prediction model to obtain the characteristics of the target protein contained in the binding conformation sample; based on the characteristics of the molecules contained in the binding conformation sample and the characteristics of the target protein, the splicing characteristics of the binding conformation sample are obtained; the splicing characteristics of the binding conformation sample are input into the third network of the first binding conformation prediction model to obtain the prediction parameters of the binding conformation sample.

23. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 9 or claims 10 to 11.

24. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9 or claims 10-11.

25. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9 or claims 10 to 11.

Citation Information

Patent Citations

  • Transaction two-party relationship information identification method and device

    CN112215604A

  • Protein ligand affinity prediction method, related device and equipment

    CN115116538A

  • Generation method and device of drug target affinity prediction model, computer equipment and storage medium

    CN116525028A