Molecular docking information prediction model training method and device, equipment and medium

CN116959550BActive Publication Date: 2026-09-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310361613.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2026-09-29
Estimated Expiration
2043-03-31

AI Technical Summary

Technical Problem

[0004]然而,蛋白质具有结构柔性,蛋白质中的氨基酸位置会发生移动,蛋白质的结构柔性会造成蛋白质对接的预测过程准确性下降

Benefits of technology

[0059]通过对第一样本结构和第二样本结构分别进行结构扰动处理,改变了样本分子中至少两个分子部分之间的相对位置;模拟了样本分子进行结构柔性变化得到第一扰动结构和第二扰动结构,调用预测模型对第一扰动结构和第二扰动结构进行处理,并结合样本分子的第一对接关键点和第二对接关键点对预测模型进行训练,实现了预测模型具有对不同结构状态的样本分子预测对接关键点的能力,避免了样本分子具有结构柔性对预测造成的影响。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116959550B_ABST
    Figure CN116959550B_ABST
Patent Text Reader

Abstract

The application discloses a molecular docking information prediction model training method and device, equipment and medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring sample information pairs, wherein the sample information pairs comprise a first sample structure, a second sample structure, a first docking key point and a second docking key point; performing structure disturbance processing on the first sample structure and the second sample structure respectively to obtain a first disturbance structure corresponding to the first sample structure and a second disturbance structure corresponding to the second sample structure; calling a prediction model to perform prediction processing on the first disturbance structure and the second disturbance structure respectively to obtain a first prediction key point corresponding to the first disturbance structure and a second prediction key point corresponding to the second disturbance structure; and training the prediction model based on the difference between the first prediction key point and the first docking key point and the difference between the second prediction key point and the second docking key point to obtain a trained prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and medium for training a predictive model of molecular docking information. Background Technology

[0002] Protein-protein complexes are complexes formed by the docking of two proteins. The structural information of protein-protein complexes is helpful for the design of organic compounds.

[0003] In related technologies, the structure of a protein-protein complex formed when two proteins dock is predicted by using the structural information of the two proteins.

[0004] However, proteins are structurally flexible, and the positions of amino acids within proteins can shift. This structural flexibility can reduce the accuracy of protein docking prediction. Summary of the Invention

[0005] This application provides a method, apparatus, device, and medium for training a prediction model of molecular docking information, the technical solution of which is as follows:

[0006] According to one aspect of this application, a method for training a prediction model for molecular docking information is provided, the method comprising:

[0007] A sample information pair is obtained, the sample information pair including a first sample structure, a second sample structure, a first docking key point corresponding to the first sample structure, and a second docking key point corresponding to the second sample structure; the first sample structure is used to indicate the positional features of at least two molecular parts in the first sample molecule, and the second sample structure is used to indicate the positional features of at least two molecular parts in the second sample molecule.

[0008] The first sample structure and the second sample structure are subjected to structural perturbation processing to obtain a first perturbation structure corresponding to the first sample structure and a second perturbation structure corresponding to the second sample structure. The structural perturbation processing is used to modify the relative positions between at least two molecular parts.

[0009] The prediction model is invoked to perform prediction processing on the first disturbance structure and the second disturbance structure respectively, to obtain the first prediction key point corresponding to the first disturbance structure and the second prediction key point corresponding to the second disturbance structure;

[0010] Based on the difference between the first prediction key point and the first docking key point, and the difference between the second prediction key point and the second docking key point, the prediction model is trained to obtain the trained prediction model. The first docking key point and the second docking key point are used to indicate the key molecular parts of the first sample molecule and the second sample molecule when they dock.

[0011] According to another aspect of this application, a predictive model training apparatus for molecular docking information is provided, the apparatus comprising:

[0012] The acquisition module is used to acquire sample information pairs, the sample information pairs including a first sample structure, a second sample structure, a first docking key point corresponding to the first sample structure, and a second docking key point corresponding to the second sample structure; the first sample structure is used to indicate the positional features of at least two molecular parts in the first sample molecule, and the second sample structure is used to indicate the positional features of at least two molecular parts in the second sample molecule.

[0013] The processing module is used to perform structural perturbation processing on the first sample structure and the second sample structure respectively to obtain a first perturbation structure corresponding to the first sample structure and a second perturbation structure corresponding to the second sample structure. The structural perturbation processing is used to modify the relative positions between at least two molecular parts.

[0014] The prediction module is used to call the prediction model to perform prediction processing on the first disturbance structure and the second disturbance structure respectively, so as to obtain the first prediction key point corresponding to the first disturbance structure and the second prediction key point corresponding to the second disturbance structure.

[0015] The training module is used to train the prediction model based on the difference between the first prediction key point and the first docking key point, and the difference between the second prediction key point and the second docking key point, to obtain the trained prediction model. The first docking key point and the second docking key point are used to indicate the key molecular parts of the first sample molecule and the second sample molecule when they dock.

[0016] In an optional design of this application, the processing module is further configured to:

[0017] Based on the first elastic distance threshold between the molecular parts of the first sample molecule, the first sample structure is subjected to the structural perturbation process to obtain the first perturbation structure;

[0018] Based on the second elastic distance threshold between the molecular parts of the second sample molecule, the second sample structure is subjected to the structural perturbation process to obtain the second perturbation structure.

[0019] In an optional design of this application, the processing module is further configured to:

[0020] Based on the first sample structure and the first elastic distance threshold, a first distance matrix is ​​constructed;

[0021] Based on the inverse of the first distance matrix, determine the first correction constraint for at least two molecular parts of the first sample molecule;

[0022] Using the first modified constraint as a constraint condition, the first sample structure is subjected to the structural perturbation process to obtain the first perturbation structure;

[0023] Based on the second sample structure and the second elastic distance threshold, a second distance matrix is ​​constructed;

[0024] Based on the inverse of the second distance matrix, determine the second correction constraint for at least two molecular parts of the second sample molecule;

[0025] Using the second modified constraint as a constraint condition, the second sample structure is subjected to the structural perturbation process to obtain the second perturbation structure.

[0026] In one alternative design of this application,

[0027] The a-th molecular part of the first sample molecule corresponds to the a-th molecular part constraint in the first correction constraint. There is a correlation between the a-th molecular part constraint and the a-th matrix value on the diagonal of the inverse matrix of the first distance matrix. a is a positive integer not exceeding n. The first sample molecule includes n molecular parts.

[0028] And / or,

[0029] The b-th molecular part of the second sample molecule corresponds to the b-th molecular part constraint in the second correction constraint. There is a correlation between the b-th molecular part constraint and the b-th matrix value on the diagonal of the inverse matrix of the second distance matrix. b is a positive integer not exceeding m. The second sample molecule includes m molecular parts.

[0030] In an optional design of this application, the prediction model includes a dimensionality reduction network and an attention network; the prediction module is further configured to:

[0031] The dimensionality reduction network is invoked to perform dimensionality reduction processing on the first perturbation structure to obtain the first structural feature corresponding to the first perturbation structure;

[0032] The dimensionality reduction network is invoked to perform dimensionality reduction processing on the second perturbation structure to obtain the second structural features corresponding to the second perturbation structure;

[0033] The attention network is invoked to perform prediction processing on the first structural feature to obtain the first predicted key point;

[0034] The attention network is invoked to perform prediction processing on the second structural features to obtain the second predicted key points.

[0035] In an optional design of this application, the prediction module is further configured to:

[0036] The attention network is invoked to predict the i-th and j-th sub-features in the first structural feature to obtain the first importance information of the i-th molecular part relative to the j-th molecular part in the first sample molecule.

[0037] Based on the first importance information, the molecular parts in the first sample molecule are ordered sequentially, and the first prediction key point is determined in the sequential order of the first sample molecule.

[0038] The attention network is invoked to predict the k-th and l-th sub-features in the second structural feature to obtain the second importance information of the k-th molecular part relative to the l-th molecular part in the second sample molecule.

[0039] Based on the second importance information, the molecular parts in the second sample molecule are ordered sequentially, and the second prediction key point is determined in the sequential ordering of the second sample molecule.

[0040] In one alternative design of this application,

[0041] The acquisition module is further configured to acquire docking structure information of docking molecules, wherein the docking molecules are molecules obtained by docking the first sample molecule and the second sample molecule, and the docking structure information is used to indicate the position information of at least four molecular parts in the docking molecule.

[0042] The prediction module is also used to call the attention network to perform docking prediction on the first structural feature to obtain the first docking position corresponding to the first sample molecule.

[0043] The prediction module is also used to call the attention network to perform docking prediction on the second structural feature to obtain the second docking position corresponding to the second sample molecule;

[0044] The processing module is further configured to determine predicted docking information based on the first docking position and the second docking position;

[0045] The training module is further used to train the prediction model based on the differences between the first predicted key point and the first docking key point, the differences between the second predicted key point and the second docking key point, and the differences between the predicted docking information and the docking structure information, so as to obtain the trained prediction model.

[0046] In an optional design of this application, the acquisition module is further configured to:

[0047] Obtain the first position information of at least two molecular parts in the first sample molecule and the second position information of at least two molecular parts in the second sample molecule;

[0048] Based on the first proximity distance threshold between the molecular parts of the first sample molecule, and according to the first position information, the first proximity features of at least two molecular parts in the first sample molecule are determined.

[0049] Based on the second proximity distance threshold between the molecular parts of the second sample molecule, and according to the second position information, the second proximity features of at least two molecular parts in the second sample molecule are determined.

[0050] The first sample structure is determined based on the first proximity feature and the first location information, and the second sample structure is determined based on the second proximity feature and the second location information;

[0051] Obtain the first docking key point and the second docking key point.

[0052] In one alternative design of this application,

[0053] The acquisition module is also used to acquire the first molecular structure of the first predicted molecule and the second molecular structure of the second predicted molecule;

[0054] The prediction module is further configured to call the trained prediction model to perform prediction processing on the first molecular structure and the second molecular structure respectively, to obtain the first key point corresponding to the first predicted molecule and the second key point corresponding to the second predicted molecule.

[0055] According to another aspect of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the molecular docking information prediction model training method as described above.

[0056] According to another aspect of this application, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the method for training a predictive model of molecular docking information as described above.

[0057] According to another aspect of this application, a computer program product is provided, the computer program product including computer instructions stored in a computer-readable storage medium, wherein a processor reads from the computer-readable storage medium and executes the computer instructions to implement the molecular docking information prediction model training method described above.

[0058] The beneficial effects of the technical solution provided in this application include at least the following:

[0059] By perturbing the first and second sample structures respectively, the relative positions between at least two molecular parts in the sample molecule are changed. The first and second perturbed structures are obtained by simulating the structural flexibility changes of the sample molecule. The prediction model is called to process the first and second perturbed structures, and the prediction model is trained by combining the first and second docking key points of the sample molecule. This realizes that the prediction model has the ability to predict the docking key points of sample molecules with different structural states, and avoids the influence of the structural flexibility of the sample molecule on the prediction. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 This is a schematic diagram of a computer system provided in an exemplary embodiment of this application;

[0062] Figure 2 This is a schematic diagram illustrating the training of a prediction model for molecular docking information provided in an exemplary embodiment of this application;

[0063] Figure 3 This is a flowchart of a method for training a prediction model of molecular docking information provided in an exemplary embodiment of this application;

[0064] Figure 4 This is a flowchart of a method for training a prediction model of molecular docking information provided in an exemplary embodiment of this application;

[0065] Figure 5 This is a flowchart of a method for training a prediction model of molecular docking information provided in an exemplary embodiment of this application;

[0066] Figure 6 This is a flowchart of a method for training a prediction model of molecular docking information provided in an exemplary embodiment of this application;

[0067] Figure 7 This is a flowchart of a method for training a prediction model of molecular docking information provided in an exemplary embodiment of this application;

[0068] Figure 8 This is a flowchart of a method for training a prediction model of molecular docking information provided in an exemplary embodiment of this application;

[0069] Figure 9 This is a flowchart of a method for training a prediction model of molecular docking information provided in an exemplary embodiment of this application;

[0070] Figure 10 This is a block diagram of a molecular docking information prediction model training device provided in an exemplary embodiment of this application;

[0071] Figure 11 This is a structural block diagram of a server provided in an exemplary embodiment of this application.

[0072] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0074] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0075] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0076] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions. This information was obtained under full authorization.

[0077] It should be understood that although the terms first, second, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, a first parameter may also be referred to as a second parameter without departing from the scope of this disclosure, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0078] Figure 1 A schematic diagram of a computer system provided in one embodiment of this application is shown. This computer system can implement a system architecture for a predictive model training method for molecular docking information. The computer system may include: a terminal 100 and a server 200.

[0079] Terminal 100 can be an electronic device such as a mobile phone, tablet computer, vehicle terminal (vehicle system), wearable device, or PC (Personal Computer). A client application for the target application can be installed and run on terminal 100. This target application can be a predictive model training application for molecular docking information, or other applications that provide predictive model training functions for molecular docking information; this application does not limit the specific application. Furthermore, this application does not limit the form of the target application, including but not limited to apps, mini-programs, etc., installed on terminal 100, and can also be in web page form.

[0080] Server 200 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. Server 200 can be the backend server for the aforementioned target application, used to provide backend services to the clients of the target application.

[0081] The molecular docking information prediction model training method provided in this application embodiment can be executed by a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. Figure 1 Taking the implementation environment of the scheme shown as an example, the prediction model training method for molecular docking information can be executed by the terminal 100 (such as the client of the target application installed and running in the terminal 100 executing the prediction model training method for molecular docking information), or by the server 200 executing the prediction model training method for molecular docking information, or by the terminal 100 and the server 200 interacting and cooperating to execute it. This application does not limit this.

[0082] Furthermore, the technical solution of this application can be combined with blockchain technology. For example, in the molecular docking information prediction model training method disclosed in this application, some data involved (physiological images, distance transformation images, closed regions, etc.) can be stored on the blockchain. The terminal 100 and the server 200 can communicate through a network, such as a wired or wireless network.

[0083] Next, we will introduce the process of training the prediction model for molecular docking information:

[0084] Figure 2 This illustration shows a schematic diagram of training a prediction model for molecular docking information according to an embodiment of this application.

[0085] Obtain sample information pair 300; sample information pair 300 includes first sample structure 302 of first sample protein, second sample structure 304 of second sample protein, first docking key point 312 corresponding to first sample structure 302, and second docking key point 314 corresponding to second sample structure 304.

[0086] For example, the first sample structure 302 is used to indicate the positional features of amino acids in the first sample protein, and the second sample structure 304 is used to indicate the positional features of amino acids in the second sample protein. For example, the positional features include the positional information of the amino acids and the proximity features between amino acids. Specifically, the first sample structure 302 includes the positional information of n amino acids in the first sample protein in a three-dimensional coordinate system, such as three-dimensional coordinates; it also includes the proximity features between each amino acid in the first sample protein and other amino acids, for example: the proximity feature is a feature matrix of dimension n multiplied by n, where the feature value in the x-th row and y-th column of the feature matrix is ​​1 if the distance between the x-th amino acid and the y-th amino acid is greater than a proximity distance threshold, and 0 otherwise. x and y are both positive integers not exceeding n. Similarly, the second sample structure 304 includes the positional information of m amino acids in the second sample protein in a three-dimensional coordinate system, and the proximity features between each amino acid in the second sample protein and other amino acids.

[0087] For example, the first docking key point 312 is an amino acid in the first sample protein that generates an interaction force between the second sample protein and the first sample protein when the first sample protein and the second sample protein are docked. Similarly, the second docking key point 314 is an amino acid in the second sample protein that generates an interaction force between the second sample protein and the first protein when the first sample protein and the second sample protein are docked.

[0088] The docking structure information 306 is obtained from the docking of the first and second sample proteins. The docking protein is also called a protein dimer, protein-protein complex, or simply any of these protein complexes. The docking structure information 306 is used to indicate the positional information of each amino acid in the docking protein.

[0089] Based on the elastic distance threshold between amino acids, the first sample structure 302 and the second sample structure 304 are subjected to structural perturbation processing to obtain the first perturbation structure 302a and the second perturbation structure 304a. An exemplary elastic distance threshold between amino acids is 6 EGStrands. Structural perturbation is used to modify the relative positions of at least two amino acids in a protein. For example, structural perturbation, such as through random sampling, can be used to augment data on a first and second sample protein to simulate protein flexibility. This allows the model to exhibit protein flexibility invariance during training, meaning that different morphologies of the same protein do not affect the model's predictions.

[0090] The prediction model 320 includes a dimensionality reduction network 322 and an attention network 324.

[0091] The dimensionality reduction network 322 is invoked to perform dimensionality reduction processing on the first perturbation structure 302a and the second perturbation structure 304a respectively, to obtain the first structural feature 332 corresponding to the first perturbation structure 302a and the second structural feature 334 corresponding to the second perturbation structure 304a.

[0092] The attention network 324 is invoked to perform prediction processing on the first structural feature 332 and the second structural feature 334 respectively, to obtain the first predicted key point 332a and the first docking position 332b corresponding to the first structural feature 332, and the second predicted key point 334a and the second docking position 334b ​​corresponding to the second structural feature 334.

[0093] For example, the first docking position 332b is the position of the first sample molecule in the docking protein predicted by the prediction model 320 when the first sample protein and the second sample protein are docked; similarly, the second docking position 334b ​​is the position of the second sample molecule in the docking protein. Further, it refers to the positions of individual amino acids in the first sample protein and the second sample protein.

[0094] For example, the first prediction key point 332a is the amino acid in the first sample protein that generates an interaction force between the second sample protein and the first sample protein when the first sample protein and the second sample protein perform protein docking, as predicted by the prediction model 320; similarly, the second prediction key point 334a is the amino acid in the second sample protein that generates an interaction force between the second sample protein and the first protein.

[0095] Based on the first docking position 332b and the second docking position 334b, the predicted docking information is obtained by splicing together;

[0096] The prediction model 320 is trained based on the first difference 342 between the first prediction key point 332a and the first docking key point 312, the second difference 344 between the second prediction key point 334a and the second docking key point 314, and the third difference 346 between the prediction docking information and the docking structure information 306, to obtain the trained prediction model.

[0097] To improve the prediction accuracy of the molecular docking information prediction model, it is necessary to train the molecular docking information prediction model. The following examples will introduce the training method of the molecular docking information prediction model.

[0098] Figure 3 A flowchart illustrating a method for training a predictive model of molecular docking information according to an exemplary embodiment of this application is shown. This method can be executed by a computer device. The method includes:

[0099] Step 510: Obtain sample information pairs;

[0100] For example, the sample information pair includes relevant information about the first sample molecule and the second sample molecule; specifically, the sample information pair includes the first sample structure, the second sample structure, the first docking key point corresponding to the first sample structure, and the second docking key point corresponding to the second sample structure.

[0101] The first sample structure indicates the positional features of at least two molecular parts in the first sample molecule, and the second sample structure indicates the positional features of at least two molecular parts in the second sample molecule. For example, the first and second sample structures can directly carry positional information of the molecular parts, or they can be positional features determined based on the positional information of the molecular parts. For example, the first and second docking key points indicate the key molecular parts of the first and second sample molecules during molecular docking.

[0102] For example, the first sample molecule and / or the second sample molecule can be inorganic or organic molecules; further, they can be any one of nucleic acids, proteins, carbohydrates, and lipids. The first and second sample molecules are usually of the same type, but differences are not excluded. Furthermore, the molecular components in the sample molecule are used to compose the sample molecule; for example, if the first sample molecule is a protein, its molecular component is the amino acids in the protein; similarly, the molecular component of nucleic acids is nucleotides; and the molecular component of inorganic molecules is atoms. It should be noted that the above is merely an example, and molecules can be broken down into other molecular components according to different granularities.

[0103] Step 520: Perform structural perturbation processing on the first sample structure and the second sample structure respectively to obtain the first perturbation structure corresponding to the first sample structure and the second perturbation structure corresponding to the second sample structure.

[0104] For example, structural perturbation processing is used to modify the relative positions between at least two molecular parts; by performing structural perturbation processing on the first sample structure and the second sample structure respectively, the relative positions between at least two molecular parts in the first sample structure and the second sample structure are changed. Similar to the first sample structure and the second sample structure mentioned above, structural perturbation processing can directly change the positional information of the molecular parts, or it can change the positional features determined based on the positional information of the molecular parts.

[0105] In one example, the structural perturbation processes performed on the first sample structure and the second sample structure are independent of each other. For instance, structural perturbation is performed on the first sample structure to obtain the first perturbation structure corresponding to the first sample structure; structural perturbation is performed on the second sample structure to obtain the second perturbation structure corresponding to the second sample structure.

[0106] Step 530: Call the prediction model to perform prediction processing on the first disturbance structure and the second disturbance structure respectively, to obtain the first prediction key point corresponding to the first disturbance structure and the second prediction key point corresponding to the second disturbance structure.

[0107] For example, the prediction model is used to perform prediction processing to obtain the predicted keypoints corresponding to the perturbation structure input to the prediction model. In one example, the prediction processing of the first perturbation structure and the second perturbation structure is independent of each other. For example, the prediction model is invoked to perform prediction processing on the first perturbation structure to obtain the first predicted keypoints corresponding to the first perturbation structure; the prediction model is invoked to perform prediction processing on the second perturbation structure to obtain the second predicted keypoints corresponding to the second perturbation structure.

[0108] For example, the first predicted key point is the key molecular part of the first perturbation structure predicted by the prediction model, and it is the key molecular part of the first perturbation structure when the molecules corresponding to the first perturbation structure and the molecules corresponding to the second perturbation structure dock. Similarly, the second predicted key point is the key molecular part of the second perturbation structure when the molecules dock, as predicted by the prediction model. It should be noted that, due to the structural perturbation processing performed in step 520, the molecular structure corresponding to the first perturbation structure is different from the molecular structure of the first sample molecule.

[0109] Step 540: Based on the difference between the first predicted key point and the first docking key point, and the difference between the second predicted key point and the second docking key point, train the prediction model to obtain the trained prediction model.

[0110] For example, the first docking key point and the second docking key point are used to indicate the respective key molecular portions when the first sample molecule and the second sample molecule perform molecular docking. The first docking key point is the molecular portion within the first sample molecule that satisfies the docking key conditions between the first sample molecule and the second sample molecule when they dock. The docking key conditions include, but are not limited to, at least one of generating an interaction force, a distance less than an interaction threshold, and property matching of the local surface of the molecules. Similarly, the second docking key point is the molecular portion within the second sample molecule that satisfies the docking key conditions between the second sample molecule and the first molecule when they dock.

[0111] For example, the purpose of training the prediction model is to minimize the difference between the first predicted key point and the first docking key point, and the difference between the second predicted key point and the second docking key point. Training the prediction model is achieved by adjusting the parameters of the prediction model; specifically, some or all of the parameters in the prediction model can be used.

[0112] In summary, the method provided in this embodiment changes the relative positions between at least two molecular parts in a sample molecule by performing structural perturbation processing on the first and second sample structures respectively; it simulates the structural flexibility changes of the sample molecule to obtain the first and second perturbation structures, calls the prediction model to process the first and second perturbation structures, and trains the prediction model in combination with the first and second docking key points of the sample molecule, thereby realizing the prediction model's ability to predict docking key points for sample molecules with different structural states, and avoiding the influence of the structural flexibility of the sample molecule on the prediction.

[0113] Figure 4 A flowchart illustrating a method for training a predictive model of molecular docking information according to an exemplary embodiment of this application is shown. This method can be executed by a computer device. That is, in Figure 3 In the illustrated embodiment, step 520 can be implemented as steps 522 and 524:

[0114] Step 522: Based on the first elastic distance threshold between the molecular parts of the first sample molecule, perform structural perturbation processing on the first sample structure to obtain the first perturbation structure;

[0115] For example, the first elastic distance threshold can be preset or determined based on the first sample structure of the first sample molecule; this embodiment does not limit the method of determining the first elastic distance threshold.

[0116] It should be noted that the first elastic distance threshold can be directly used to determine the perturbation mode of the molecular part of the first sample molecule during structural perturbation processing, and can also be used to determine the perturbation constraint conditions of the molecular part of the first sample molecule during structural perturbation processing. Furthermore, the above-mentioned perturbation mode and perturbation constraint conditions can be determined directly or indirectly based on the first elastic distance threshold.

[0117] Step 524: Based on the second elastic distance threshold between the molecular parts of the second sample molecule, perform structural perturbation processing on the second sample structure to obtain the second perturbation structure;

[0118] For example, the second elastic distance threshold is similar to the first elastic distance threshold, and this embodiment does not limit the method of determining the second elastic distance threshold. The second elastic distance threshold and the first elastic distance threshold are usually the same, but differences are not excluded. In a preferred embodiment, both the first elastic distance threshold and the second elastic distance threshold are 6 EGStrons.

[0119] Similarly, the second elastic distance threshold can directly or indirectly determine the perturbation mode or perturbation constraint condition of the molecular part of the second sample molecule during structural perturbation processing.

[0120] In summary, the method provided in this embodiment perturbs the structure of the first sample structure using a first elastic distance threshold and the structure of the second sample structure using a second elastic distance threshold, thereby changing the relative positions between at least two molecular parts in the sample molecule. It simulates the flexible structural changes of the sample molecule to obtain the first and second perturbated structures, calls a prediction model to process the first and second perturbated structures, and trains the prediction model by combining the first and second docking key points of the sample molecule. This enables the prediction model to predict docking key points for sample molecules in different structural states, avoiding the influence of the structural flexibility of the sample molecule on the prediction.

[0121] Figure 5 A flowchart illustrating a method for training a predictive model of molecular docking information according to an exemplary embodiment of this application is shown. This method can be executed by a computer device. That is, in Figure 3 In the illustrated embodiment, step 522 can be implemented as steps 522a, 522b, and 522c; step 524 can be implemented as steps 524a, 524b, and 524c.

[0122] Step 522a: Construct a first distance matrix based on the first sample structure and the first elastic distance threshold;

[0123] For example, the first distance matrix is ​​used to indicate whether the distance between two molecular parts in the first sample molecule exceeds a first elasticity threshold. For example, when the first sample molecule comprises n molecular parts, the first distance matrix is ​​a feature matrix of dimension n by n. Specifically, the first distance matrix includes:

[0124]

[0125] Among them, P ij Let s represent the eigenvalue in the i-th row and j-th column of the first distance matrix. ij r represents the distance between the i-th and j-th molecular parts in the first sample molecule. c This represents the first elastic distance threshold. The first distance matrix is ​​formed by P. ij Composed of.

[0126] Step 522b: Determine the first correction constraint of at least two molecular parts of the first sample molecule based on the inverse of the first distance matrix;

[0127] For example, the inverse of the first distance matrix carries, directly or indirectly, a first correction constraint in at least two numerator parts. The first correction constraint is used to indicate the constraint conditions for structural perturbation of the first sample structure, so that structural perturbation of the first sample structure will not exceed the first correction constraint.

[0128] In one optional implementation of this embodiment, the a-th molecular part of the first sample molecule corresponds to the a-th molecular part constraint in the first modified constraint, and there is a correlation between the a-th molecular part constraint and the a-th matrix value on the diagonal of the inverse of the first distance matrix. For example, the diagonal is the main diagonal of the inverse of the first distance matrix, but it is not excluded that it is the secondary diagonal. For example, there is a positive correlation between the a-th molecular part constraint and the a-th matrix value on the diagonal of the inverse of the first distance matrix. For example, taking a first sample molecule comprising n molecular parts as an example, a is a positive integer not exceeding n.

[0129] Step 522c: Using the first modified constraint as a constraint condition, perform structural perturbation processing on the first sample structure to obtain the first perturbation structure;

[0130] For example, the first correction constraint indicates the perturbation boundary for the structural perturbation process. The specific method of structural perturbation can be randomly determined or determined based on the first sample molecule; this embodiment does not impose any restrictions on this. It should be noted that the structural perturbation process changes the relative positions between at least two molecular parts in the first sample molecule. Taking a protein as an example, after structural perturbation, the distance between at least two amino acids in the protein changes. It can be understood that the structural perturbation process performs positional changes at the molecular site level to achieve structural changes between the molecular parts of the first sample molecule.

[0131] Step 524a: Construct a second distance matrix based on the second sample structure and the second elastic distance threshold;

[0132] For example, the second distance matrix is ​​used to indicate whether the distance between two molecular parts in the second sample molecule exceeds the first elastic threshold. For example, when the second sample molecule includes m molecular parts, the first distance matrix is ​​a feature matrix of dimension m by m. The calculation method of the second distance matrix is ​​described in step 522a above, and will not be repeated here.

[0133] Step 524b: Determine the second correction constraint for at least two molecular parts of the second sample molecule based on the inverse of the second distance matrix;

[0134] For example, the inverse of the second distance matrix carries, directly or indirectly, a second correction constraint on at least two molecular parts.

[0135] In one optional implementation of this embodiment, the b-th molecular part of the second sample molecule corresponds to the b-th molecular part constraint in the second correction constraint, and there is a correlation between the b-th molecular part constraint and the b-th matrix value on the diagonal of the inverse of the second distance matrix. For example, the diagonal is the main diagonal of the inverse of the second distance matrix, but it is not excluded that it is the secondary diagonal. For example, the b-th molecular part constraint and the b-th matrix value on the diagonal of the inverse of the second distance matrix are positively correlated. For example, taking a second sample molecule comprising m molecular parts as an example, b is a positive integer not exceeding m.

[0136] Step 524c: Using the second modified constraint as a constraint condition, perform structural perturbation processing on the second sample structure to obtain the second perturbation structure;

[0137] For example, the second correction constraint indicates the perturbation boundary for the structural perturbation process. The specific method of structural perturbation process can be determined randomly or based on the second sample molecule; this embodiment does not impose any restrictions on this. It is understood that the structural perturbation process involves positional changes at the molecular level to achieve structural changes between the molecular parts of the second sample molecule.

[0138] In summary, the method provided in this embodiment determines a first correction constraint through a first elastic distance threshold as a constraint condition for structural perturbation processing of the first sample structure, and determines a second correction constraint through a second elastic distance threshold as a constraint condition for structural perturbation processing of the second sample structure; it changes the relative positions between at least two molecular parts in the sample molecule; it simulates the structural flexibility changes of the sample molecule to obtain a first perturbation structure and a second perturbation structure, calls the prediction model to process the first perturbation structure and the second perturbation structure, and trains the prediction model in combination with the first docking key point and the second docking key point of the sample molecule, thereby realizing the prediction model's ability to predict docking key points of sample molecules with different structural states, and avoiding the influence of the structural flexibility of the sample molecule on the prediction.

[0139] Next, we will further introduce the prediction model for molecular docking information.

[0140] Figure 6 A flowchart illustrating a method for training a predictive model of molecular docking information according to an exemplary embodiment of this application is shown. This method can be executed by a computer device. That is, in Figure 3 In the illustrated embodiment, step 530 can be implemented as steps 532, 534, 536, and 538:

[0141] Step 532: Call the dimensionality reduction network to perform dimensionality reduction processing on the first perturbation structure to obtain the first structural features corresponding to the first perturbation structure;

[0142] For example, the dimensionality reduction network is used to perform dimensionality reduction processing on the first perturbation structure to obtain the first structural feature. The first structural feature can be a part of the first perturbation structure, or it can be the hidden layer feature obtained by feature extraction of the first perturbation structure. This embodiment does not limit the method of dimensionality reduction processing.

[0143] For example, the dimensionality reduction network includes, but is not limited to, at least one of graph attention networks and graph neural networks (GNNs). Furthermore, when the dimensionality reduction network is a GNN, the number of layers in the GNN does not exceed 10, in order to avoid oversmoothing and / or overfitting during model training.

[0144] Step 534: Use a dimensionality reduction network to perform dimensionality reduction on the second perturbation structure to obtain the second structural features corresponding to the second perturbation structure;

[0145] For example, the dimensionality reduction network is also used to reduce the dimensionality of the second perturbation structure. In one example, the processes of dimensionality reduction of the first perturbation structure and the dimensionality reduction of the second perturbation structure are independent of each other. For example, the dimensionality reduction network uses the same network parameters for the above dimensionality reduction process, and the network parameters are updated during model training.

[0146] Step 536: Call the attention network to predict the first structural features and obtain the first predicted key points;

[0147] For example, the attention network is used to perform predictive processing on the first structural features. By predicting the first structural features, a first predicted key point is obtained, and the key molecular part of the docking between the molecule corresponding to the first perturbation structure and the molecule corresponding to the second perturbation structure is predicted.

[0148] In one optional implementation of this embodiment, step 536 can be implemented as follows:

[0149] An attention network is invoked to predict the i-th and j-th sub-features in the first structural feature to obtain the first importance information of the i-th molecular part relative to the j-th molecular part in the first sample molecule; based on the first importance information, the molecular parts in the first sample molecule are ordered in order, and the first prediction key point is determined in the order of the first sample molecule.

[0150] For example, during the prediction process of the first structural feature using an attention network, the attention network predicts two sub-features within the first structural feature, and these two sub-features correspond one-to-one with two molecular parts in the first perturbation structure. The first importance information between the two molecular parts is obtained; this first importance information represents the degree of importance between the two molecular parts in the first perturbation structure. For example, the importance degree is the degree of correlation between the two molecular parts during the prediction of the first keypoint. For instance, the degree of correlation and the first importance information are positively correlated. For example, the prediction process is related to the information corresponding to the molecular parts in the first perturbation structure, and the importance information is used to indicate the degree of correlation between different molecular parts.

[0151] In one example, the first importance information is calculated as follows:

[0152] e ij =a(Wh i Wh j )

[0153]

[0154] Among them, e ij This indicates the importance of the i-th molecule relative to the j-th molecule in the first importance information. a(.) represents the attention network, W represents the parameter matrix of the attention network, and h... i h represents the i-th sub-feature corresponding to the i-th molecular part in the first structural feature. j E represents the j-th sub-feature corresponding to the j-th molecular part in the first structural feature. i This represents the sum of the importance of the i-th molecular part relative to all the molecular parts corresponding to the first perturbation structure, and is also represented as the first importance information of the i-th molecular part. N represents the number of each molecular part corresponding to the first perturbation structure.

[0155] Based on the first importance information, the molecular parts corresponding to the first perturbation structure are sorted. For example, they are sorted from largest to smallest according to the first importance information of each molecular part. The first prediction key point is determined in the sequential sorting of the first sample molecules. For example, the first 10 molecular parts in the sequential sorting are determined as the first prediction key points. It is understood that in different embodiments, more or fewer molecular parts in the sequential sorting can be used, and different selection methods can be adopted, such as selecting the later or middle molecular parts. This embodiment does not impose any limiting provisions.

[0156] Step 538: Call the attention network to predict the second structural features and obtain the second predicted key points;

[0157] For example, the attention network is used to perform prediction processing on the second structural features. By predicting the second structural features, a second prediction key point is obtained, and the key molecular part when the molecule corresponding to the second perturbation structure and the molecule corresponding to the first perturbation structure dock is predicted.

[0158] In one optional implementation of this embodiment, step 538 can be implemented as follows:

[0159] An attention network is invoked to predict the k-th and 1-th sub-features in the second structural features, thereby obtaining the second importance information of the k-th molecular part relative to the 1-th molecular part in the second sample molecule; based on the second importance information, the molecular parts in the second sample molecule are ordered, and the second prediction key point is determined in the order of the second sample molecules.

[0160] For example, the prediction process of the second prediction key point is similar to that of the first prediction key point. For a description of the above steps, please refer to step 536 above, which will not be repeated here.

[0161] In one example, the attention network's prediction processing of the first structural feature and the prediction processing of the second structural feature are independent of each other. For instance, the attention network uses the same network parameters for the above prediction processing, and these parameters are updated during model training.

[0162] In summary, the method provided in this embodiment changes the relative positions between at least two molecular parts in a sample molecule by performing structural perturbation processing on the first and second sample structures respectively; it simulates the flexible structural changes of the sample molecule to obtain the first and second perturbation structures; it calls the dimensionality reduction network and attention network in the prediction model to process the first and second perturbation structures; and it trains the prediction model by combining the first and second docking key points of the sample molecule, thus realizing the prediction model's ability to predict docking key points for sample molecules with different structural states, avoiding the influence of the structural flexibility of the sample molecule on the prediction.

[0163] Figure 7 A flowchart illustrating a method for training a predictive model of molecular docking information according to an exemplary embodiment of this application is shown. This method can be executed by a computer device. That is, in Figure 6 In the illustrated embodiment, steps 510a, 539a, 539b, and 539c are also included; step 540 can be implemented as step 542:

[0164] Step 510a: Obtain docking structure information of docking molecules;

[0165] For example, the docking molecule is a molecule obtained by docking a first sample molecule and a second sample molecule. The docking structure information is used to indicate the positional information of at least four molecular parts in the docking molecule. In one example, the number of molecular parts in the docking molecule is equal to the sum of the number of molecular parts in the first sample molecule and the second sample molecule. The number of molecular parts in the first sample molecule and the second sample molecule can be the same or different.

[0166] Step 539a: Call the attention network to perform docking prediction on the first structural features to obtain the first docking position corresponding to the first sample molecule;

[0167] For example, docking prediction and the prediction processing in step 536 can be independent or interconnected; for instance, docking prediction can be performed based on the first prediction key point to obtain the first docking position. For example, the first docking position includes the positional information of at least two molecular parts corresponding to the first perturbation structure after docking of the molecules corresponding to the first and second perturbation structures. It should be noted that the first docking position only includes the positional information of at least two molecular parts corresponding to the first perturbation structure, which can be some or all of the molecular parts corresponding to the first perturbation structure; the relative positions between the molecular parts in the first docking position can be the same as or different from the relative positions between the molecular parts in the first perturbation structure.

[0168] It should be noted that in this embodiment, step 539a is executed after step 536, but it is not excluded that step 539a is executed before or at the same time as step 536. This embodiment does not impose any restrictions.

[0169] Step 539b: Call the attention network to predict the docking of the second structural features and obtain the second docking position corresponding to the second sample molecule;

[0170] For example, similar to step 539a, the prediction processes in this step and step 538 can be independent or related. In one example, the processes of the attention network performing docking prediction for the first structural feature and for the second structural feature are independent. For example, the attention network uses the same network parameters for the above docking prediction, and the network parameters are updated during model training. For a description of the second docking position and the docking prediction in this step, please refer to step 539a above; it will not be repeated here.

[0171] It should be noted that in this embodiment, step 539b is executed after step 538, but it is not excluded that step 539b is executed before or at the same time as step 538. This embodiment does not impose any restrictions.

[0172] Step 539c: Determine the predicted docking information based on the first docking position and the second docking position;

[0173] For example, the first docking position and the second docking position are spliced ​​together to obtain the predicted docking information. In one example, the first docking position and the second docking position usually correspond to the same reference frame, but it is not impossible for the reference frames to be different. If the reference frames corresponding to the first docking position and the second docking position are different, a reference frame conversion is required before splicing the first docking position and the second docking position to unify the two pieces of information in the same reference frame before splicing.

[0174] Step 542: Train the prediction model based on the differences between the first predicted key point and the first docking key point, the differences between the second predicted key point and the second docking key point, and the differences between the predicted docking information and the docking structure information, to obtain the trained prediction model;

[0175] In a specific implementation, the above differences are calculated as follows:

[0176] L ce =CE(E1, E1) * )+CE(E2,E2 * )

[0177]

[0178] L = L ce +L mse

[0179] Wherein, CE(E1, E1) * E1 represents the difference between the first predicted keypoint and the first docking keypoint. * Indicates the first key docking point; CE(E2, E2) * E1 represents the difference between the second predicted keypoint and the second docking keypoint, and E2 represents the predicted second predicted keypoint. * This represents the second docking critical point; CE(.,.) represents the cross-entropy loss function. ce This represents the sum of the differences between the first prediction key point and the first docking key point, and between the second prediction key point and the second docking key point. This indicates the position information of the k-th molecule in the docking structure information. Let represent the position information of the k-th molecule in the predicted docking information, and p represent the sum of the molecule portions in the first sample molecule and the second sample molecule. L mseThe mean squared error (MSE) is represented by the difference between predicted docking information and docking structure information. L represents L0. ce and L mse The sum of these factors, with L as the loss function, is used to train the prediction model, resulting in the trained prediction model.

[0180] In summary, the method provided in this embodiment alters the relative positions of at least two molecular parts in a sample molecule by performing structural perturbation on the first and second sample structures respectively. It simulates the flexible structural changes of the sample molecule to obtain the first and second perturbation structures. The method then uses a dimensionality reduction network and an attention network in the prediction model to process these structures, with the attention network also predicting the first and second docking positions. By combining the first and second docking key points of the sample molecule with the docking structure information of the docking molecule, the prediction model is trained, enabling it to predict docking key points for sample molecules in different structural states, thus avoiding the influence of the structural flexibility of the sample molecule on the prediction.

[0181] Figure 8 A flowchart illustrating a method for training a predictive model of molecular docking information according to an exemplary embodiment of this application is shown. This method can be executed by a computer device. That is, in Figure 3 In the illustrated embodiment, step 510 can be implemented as steps 512, 514a, 514b, 516, and 518:

[0182] Step 512: Obtain the first position information of at least two molecular parts in the first sample molecule and the second position information of at least two molecular parts in the second sample molecule;

[0183] For example, the first positional information is the coordinate information of at least two molecular parts in the first sample molecule in three-dimensional space. Taking the first sample molecule as having n molecular parts as an example, the first positional information can be represented by a matrix of dimension n×3, denoted as A1. Similarly, taking the second sample molecule as having m molecular parts as an example, the second positional information can be represented by a matrix of dimension m×3, denoted as A2.

[0184] Step 514a: Based on the first proximity distance threshold between the molecular parts of the first sample molecule, determine the first proximity features of at least two molecular parts in the first sample molecule according to the first position information;

[0185] For example, the first proximity distance threshold can be preset or determined based on the first sample structure of the first sample molecule; this embodiment does not limit the method of determining the first proximity distance threshold. Furthermore, the first proximity threshold and the first elastic distance threshold in step 522 can be the same or different; this embodiment does not limit the correlation between the two.

[0186] For example, the first neighbor feature is represented by an n×n matrix, denoted as X1. If the distance between the x-th and y-th molecular parts is greater than the neighbor distance threshold, the feature value of the x-th row and y-th column of the first neighbor feature is 1; otherwise, it is 0. x and y are both positive integers not exceeding n.

[0187] Step 514b: Based on the second proximity distance threshold between the molecular parts of the second sample molecule, determine the second proximity features of at least two molecular parts in the second sample molecule according to the second position information;

[0188] For example, the second proximity distance threshold and the first proximity distance threshold are similar, and this embodiment does not limit the method of determining the second proximity distance threshold. The second proximity distance threshold and the first proximity distance threshold can be the same or different, and this embodiment does not limit the correlation between the second proximity distance threshold and the first proximity distance threshold, nor the correlation between the second elastic distance threshold and the second proximity distance threshold. For example, the first proximity feature is represented by a matrix of dimension m×m, denoted as X2.

[0189] Step 516: Determine the first sample structure based on the first proximity feature and the first location information, and determine the second sample structure based on the second proximity feature and the second location information;

[0190] For example, the first neighbor feature and the first location information are concatenated to obtain the first sample structure, which is denoted as G1 and has a dimension of n×(3+n); for example, G1=(A1,X1).

[0191] Similarly, the second sample structure is denoted as G2, with dimensions m×(3+m); for example, G2=(A2,X2).

[0192] Step 518: Obtain the first and second docking key points.

[0193] It should be noted that this embodiment does not restrict the execution order between step 518 and the four steps mentioned above. Step 518 can be executed before, after, or simultaneously with any one of the four steps mentioned above.

[0194] In summary, the method provided in this embodiment obtains first and second position information, processes the two pieces of information to obtain a first sample structure and a second sample structure, and mines and transforms the position information; by performing structural perturbation processing on the first and second sample structures respectively, the relative positions between at least two molecular parts in the sample molecule are changed; the first and second perturbation structures are obtained by simulating the structural flexibility changes of the sample molecule; the prediction model is called to process the first and second perturbation structures; and the prediction model is trained in combination with the first and second docking key points of the sample molecule, thereby realizing the prediction model's ability to predict docking key points for sample molecules in different structural states, avoiding the influence of the structural flexibility of the sample molecule on the prediction.

[0195] Figure 9 A flowchart illustrating a method for training a predictive model of molecular docking information according to an exemplary embodiment of this application is shown. This method can be executed by a computer device. That is, in Figure 3 The illustrated embodiment further includes steps 552 and 554:

[0196] Step 552: Obtain the first molecular structure of the first predicted molecule and the second molecular structure of the second predicted molecule;

[0197] For example, the first and second predicted molecules are molecules to be predicted. The structures of the first and second molecules are similar to the first and second sample structures in step 510, and can directly or indirectly carry positional information of the molecular parts.

[0198] Step 554: Call the trained prediction model to predict the first molecular structure and the second molecular structure respectively, and obtain the first key point corresponding to the first predicted molecule and the second key point corresponding to the second predicted molecule.

[0199] For example, the trained prediction model is obtained by training the prediction model through any of the above embodiments. This embodiment does not limit the specific training method. This embodiment only illustrates the training of the prediction model through steps 510 to 540 and does not impose any restrictions.

[0200] The first critical point and the second critical point are the critical molecular parts of the first and second predicted molecules respectively when they dock. The molecular parts corresponding to the first critical point and the second critical point satisfy the docking critical conditions, which include at least one of the following: generating an interaction force, the distance being less than the interaction threshold, and the property matching of the local surface of the molecule.

[0201] In one example, during the prediction of molecular docking information using the xx dataset, the trained prediction model showed good training performance, as detailed in the table below:

[0202] Patch Dock 16.17 10.03 6780 HDOCK 13.45 9.64 993 Equi Dock 11.26 8.50 17 The trained prediction model 9.92 5.03 12

[0203] As can be seen, the trained prediction model provided in this embodiment has smaller errors in terms of root-mean-square deviation standard deviation (RMSD_dev) and root-mean-square deviation mean (RMSD_mean) compared to the patch dock, highly integrated suite dock, and equivariant dock, and it takes less time.

[0204] In a specific example, taking proteins as the first and second predicted molecules, the prediction model is trained to possess protein flexibility invariance, meaning that different forms of the same protein do not affect the model's prediction results. The trained prediction model achieves docking with flexible proteins, increasing the amount of structural information available for understanding protein function, and can be further used for cell signaling pathway research and drug design of protein-protein complexes.

[0205] In summary, the method provided in this embodiment changes the relative positions between at least two molecular parts in a sample molecule by performing structural perturbation processing on the first and second sample structures respectively; it simulates the structural flexibility changes of the sample molecule to obtain the first and second perturbation structures; it calls the prediction model to process the first and second perturbation structures; and it trains the prediction model by combining the first and second docking key points of the sample molecule, thus realizing the prediction model's ability to predict docking key points for sample molecules with different structural states. Calling the trained prediction model avoids the influence of the structural flexibility of the sample molecule on the prediction.

[0206] Those skilled in the art will understand that the above embodiments can be implemented independently, or the above embodiments can be freely combined to create new embodiments to realize the molecular docking information prediction model training method of this application.

[0207] Figure 10 A block diagram of a molecular docking information prediction model training apparatus provided in an exemplary embodiment of this application is shown. The apparatus includes:

[0208] The acquisition module 810 is used to acquire sample information pairs, the sample information pairs including a first sample structure, a second sample structure, a first docking key point corresponding to the first sample structure and a second docking key point corresponding to the second sample structure; the first sample structure is used to indicate the positional features of at least two molecular parts in the first sample molecule, and the second sample structure is used to indicate the positional features of at least two molecular parts in the second sample molecule.

[0209] The processing module 820 is used to perform structural perturbation processing on the first sample structure and the second sample structure respectively to obtain a first perturbation structure corresponding to the first sample structure and a second perturbation structure corresponding to the second sample structure. The structural perturbation processing is used to modify the relative positions between at least two molecular parts.

[0210] The prediction module 830 is used to call the prediction model to perform prediction processing on the first disturbance structure and the second disturbance structure respectively, so as to obtain the first prediction key point corresponding to the first disturbance structure and the second prediction key point corresponding to the second disturbance structure.

[0211] The training module 840 is used to train the prediction model based on the difference between the first prediction key point and the first docking key point, and the difference between the second prediction key point and the second docking key point, to obtain the trained prediction model. The first docking key point and the second docking key point are used to indicate the key molecular parts of the first sample molecule and the second sample molecule when they dock.

[0212] In an optional design of this embodiment, the processing module 820 is further configured to:

[0213] Based on the first elastic distance threshold between the molecular parts of the first sample molecule, the first sample structure is subjected to the structural perturbation process to obtain the first perturbation structure;

[0214] Based on the second elastic distance threshold between the molecular parts of the second sample molecule, the second sample structure is subjected to the structural perturbation process to obtain the second perturbation structure.

[0215] In an optional design of this embodiment, the processing module 820 is further configured to:

[0216] Based on the first sample structure and the first elastic distance threshold, a first distance matrix is ​​constructed;

[0217] Based on the inverse of the first distance matrix, determine the first correction constraint for at least two molecular parts of the first sample molecule;

[0218] Using the first modified constraint as a constraint condition, the first sample structure is subjected to the structural perturbation process to obtain the first perturbation structure;

[0219] Based on the second sample structure and the second elastic distance threshold, a second distance matrix is ​​constructed;

[0220] Based on the inverse of the second distance matrix, determine the second correction constraint for at least two molecular parts of the second sample molecule;

[0221] Using the second modified constraint as a constraint condition, the second sample structure is subjected to the structural perturbation process to obtain the second perturbation structure.

[0222] In an optional design of this embodiment,

[0223] The a-th molecular part of the first sample molecule corresponds to the a-th molecular part constraint in the first correction constraint. There is a correlation between the a-th molecular part constraint and the a-th matrix value on the diagonal of the inverse matrix of the first distance matrix. a is a positive integer not exceeding n. The first sample molecule includes n molecular parts.

[0224] And / or, the b-th molecular part of the second sample molecule corresponds to the b-th molecular part constraint in the second correction constraint, and there is a correlation between the b-th molecular part constraint and the b-th matrix value on the diagonal of the inverse matrix of the second distance matrix, where b is a positive integer not exceeding m, and the second sample molecule includes m molecular parts.

[0225] In an optional design of this embodiment, the prediction model includes a dimensionality reduction network and an attention network; the prediction module 830 is further configured to:

[0226] The dimensionality reduction network is invoked to perform dimensionality reduction processing on the first perturbation structure to obtain the first structural feature corresponding to the first perturbation structure;

[0227] The dimensionality reduction network is invoked to perform dimensionality reduction processing on the second perturbation structure to obtain the second structural features corresponding to the second perturbation structure;

[0228] The attention network is invoked to perform prediction processing on the first structural feature to obtain the first predicted key point;

[0229] The attention network is invoked to perform prediction processing on the second structural features to obtain the second predicted key points.

[0230] In an optional design of this embodiment, the prediction module 830 is further configured to:

[0231] The attention network is invoked to predict the i-th and j-th sub-features in the first structural feature to obtain the first importance information of the i-th molecular part relative to the j-th molecular part in the first sample molecule.

[0232] Based on the first importance information, the molecular parts in the first sample molecule are ordered sequentially, and the first prediction key point is determined in the sequential order of the first sample molecule.

[0233] The attention network is invoked to predict the k-th and l-th sub-features in the second structural feature to obtain the second importance information of the k-th molecular part relative to the l-th molecular part in the second sample molecule.

[0234] Based on the second importance information, the molecular parts in the second sample molecule are ordered sequentially, and the second prediction key point is determined in the sequential ordering of the second sample molecule.

[0235] In an optional design of this embodiment,

[0236] The acquisition module 810 is further configured to acquire docking structure information of docking molecules, wherein the docking molecules are molecules obtained by docking the first sample molecule and the second sample molecule, and the docking structure information is used to indicate the position information of at least four molecular parts in the docking molecule.

[0237] The prediction module 830 is further configured to call the attention network to perform docking prediction on the first structural feature to obtain the first docking position corresponding to the first sample molecule.

[0238] The prediction module 830 is also used to call the attention network to perform docking prediction on the second structural feature to obtain the second docking position corresponding to the second sample molecule;

[0239] The processing module 820 is further configured to determine predicted docking information based on the first docking position and the second docking position;

[0240] The training module 840 is further configured to train the prediction model based on the differences between the first predicted key point and the first docking key point, the differences between the second predicted key point and the second docking key point, and the differences between the predicted docking information and the docking structure information, so as to obtain the trained prediction model.

[0241] In an optional design of this embodiment, the acquisition module 810 is further configured to:

[0242] Obtain the first position information of at least two molecular parts in the first sample molecule and the second position information of at least two molecular parts in the second sample molecule;

[0243] Based on the first proximity distance threshold between the molecular parts of the first sample molecule, and according to the first position information, the first proximity features of at least two molecular parts in the first sample molecule are determined.

[0244] Based on the second proximity distance threshold between the molecular parts of the second sample molecule, and according to the second position information, the second proximity features of at least two molecular parts in the second sample molecule are determined.

[0245] The first sample structure is determined based on the first proximity feature and the first location information, and the second sample structure is determined based on the second proximity feature and the second location information;

[0246] Obtain the first docking key point and the second docking key point.

[0247] In an optional design of this embodiment, the acquisition module 810 is further configured to acquire the first molecular structure of the first predicted molecule and the second molecular structure of the second predicted molecule;

[0248] The prediction module 830 is further configured to call the trained prediction model to perform prediction processing on the first molecular structure and the second molecular structure respectively, so as to obtain the first key point corresponding to the first predicted molecule and the second key point corresponding to the second predicted molecule.

[0249] It should be noted that the device provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules according to actual needs, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0250] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments of the relevant method; the technical effects achieved by each module performing its operation are the same as the technical effects in the embodiments of the relevant method, and will not be elaborated here.

[0251] This application also provides a computer device, which includes a processor and a memory, wherein the memory stores a computer program; the processor is used to execute the computer program in the memory to implement the molecular docking information prediction model training method provided in the above method embodiments.

[0252] Alternatively, the computer device is a server. For example, Figure 11 This is a structural block diagram of a server provided in an exemplary embodiment of this application.

[0253] Typically, server 2300 includes a processor 2301 and memory 2302.

[0254] Processor 2301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 2301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 2301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake-up state, also known as a Central Processing Unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, processor 2301 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 2301 may also include an Artificial Intelligence (AI) processor for handling computational operations related to machine learning. Memory 2302 may include one or more computer-readable storage media, which may be non-transitory. The memory 2302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 2302 is used to store at least one instruction, which is executed by the processor 2301 to implement the molecular docking information prediction model training method provided in the method embodiments of this application.

[0255] In some embodiments, the server 2300 may optionally include an input interface 2303 and an output interface 2304. The processor 2301, memory 2302, and input interfaces 2303 and 2304 can be connected via a bus or signal lines. Various peripheral devices can be connected to the input interfaces 2303 and 2304 via a bus, signal lines, or a circuit board. The input interfaces 2303 and 2304 can be used to connect at least one input / output (I / O) related peripheral device to the processor 2301 and memory 2302. In some embodiments, the processor 2301, memory 2302, and input interfaces 2303 and 2304 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 2301, memory 2302, and input interfaces 2303 and 2304 can be implemented on separate chips or circuit boards, and this application embodiment does not limit this.

[0256] Those skilled in the art will understand that the structure shown above does not constitute a limitation on server 2300, and may include more or fewer components than shown, or combine certain components, or employ different component arrangements.

[0257] In an exemplary embodiment, a chip is also provided, the chip including programmable logic circuits and / or program instructions, which, when the chip is run on a computer device, are used to implement the method for training a predictive model of molecular docking information as described above.

[0258] In an exemplary embodiment, a computer program product is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions to implement the molecular docking information prediction model training method provided in the above-described method embodiments.

[0259] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores a computer program that is loaded and executed by a processor to implement the method for training a predictive model of molecular docking information provided in the above-described method embodiments.

[0260] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk. Those skilled in the art should recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented by hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of computer programs from one place to another. Storage media can be any available medium accessible to general-purpose or special-purpose computers.

[0261] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for training a predictive model of molecular docking information, characterized in that, The method includes: A sample information pair is obtained, the sample information pair including a first sample structure, a second sample structure, a first docking key point corresponding to the first sample structure, and a second docking key point corresponding to the second sample structure; the first sample structure is used to indicate the positional features of at least two molecular parts in the first sample molecule, and the second sample structure is used to indicate the positional features of at least two molecular parts in the second sample molecule. Based on a first elastic distance threshold between molecular parts of the first sample molecule, the first sample structure is subjected to structural perturbation processing to obtain a first perturbation structure; and based on a second elastic distance threshold between molecular parts of the second sample molecule, the second sample structure is subjected to the same structural perturbation processing to obtain a second perturbation structure; the structural perturbation processing is used to modify the relative positions between at least two molecular parts; The prediction model is invoked to perform prediction processing on the first disturbance structure and the second disturbance structure respectively, to obtain the first prediction key point corresponding to the first disturbance structure and the second prediction key point corresponding to the second disturbance structure; Based on the differences between the first prediction key point and the first docking key point, and the differences between the second prediction key point and the second docking key point, the prediction model is trained to obtain the trained prediction model. The first docking key point and the second docking key point are used to indicate the key molecular parts of the first sample molecule and the second sample molecule when they are docked.

2. The method according to claim 1, characterized in that, The first sample structure is perturbed based on a first elastic distance threshold between molecular parts to obtain a first perturbed structure, including: Based on the first sample structure and the first elastic distance threshold, a first distance matrix is ​​constructed; Based on the inverse of the first distance matrix, determine the first correction constraint for at least two molecular parts of the first sample molecule; Using the first modified constraint as a constraint condition, the first sample structure is subjected to the structural perturbation process to obtain the first perturbation structure; The second sample structure is subjected to structural perturbation processing based on a second elastic distance threshold between molecular parts of the second sample molecule to obtain a second perturbation structure, including: Based on the second sample structure and the second elastic distance threshold, a second distance matrix is ​​constructed; Based on the inverse of the second distance matrix, determine the second correction constraint for at least two molecular parts of the second sample molecule; Using the second modified constraint as a constraint condition, the second sample structure is subjected to the structural perturbation process to obtain the second perturbation structure.

3. The method according to claim 2, characterized in that, The a-th molecular part of the first sample molecule corresponds to the a-th molecular part constraint in the first correction constraint. There is a correlation between the a-th molecular part constraint and the a-th matrix value on the diagonal of the inverse matrix of the first distance matrix. a is a positive integer not exceeding n. The first sample molecule includes n molecular parts. And / or, The b-th molecular part of the second sample molecule corresponds to the b-th molecular part constraint in the second correction constraint. There is a correlation between the b-th molecular part constraint and the b-th matrix value on the diagonal of the inverse matrix of the second distance matrix. b is a positive integer not exceeding m. The second sample molecule includes m molecular parts.

4. The method according to any one of claims 1 to 3, characterized in that, The prediction model includes a dimensionality reduction network and an attention network; The step of calling the prediction model to perform prediction processing on the first perturbation structure and the second perturbation structure respectively, to obtain the first prediction key point corresponding to the first perturbation structure and the second prediction key point corresponding to the second perturbation structure, includes: The dimensionality reduction network is invoked to perform dimensionality reduction processing on the first perturbation structure to obtain the first structural feature corresponding to the first perturbation structure; and the dimensionality reduction network is invoked to perform dimensionality reduction processing on the second perturbation structure to obtain the second structural feature corresponding to the second perturbation structure. The attention network is invoked to perform prediction processing on the first structural feature to obtain the first predicted key point; and the attention network is invoked to perform prediction processing on the second structural feature to obtain the second predicted key point.

5. The method according to claim 4, characterized in that, The step of calling the attention network to predict the first structural features and obtain the first predicted key points includes: The attention network is invoked to predict the i-th and j-th sub-features in the first structural feature to obtain the first importance information of the i-th molecular part relative to the j-th molecular part in the first sample molecule. Based on the first importance information, the molecular parts in the first sample molecule are ordered sequentially, and the first prediction key point is determined in the sequential order of the first sample molecule. The step of calling the attention network to predict the second structural features and obtain the second predicted key points includes: The attention network is invoked to predict the k-th and l-th sub-features in the second structural feature to obtain the second importance information of the k-th molecular part relative to the l-th molecular part in the second sample molecule. Based on the second importance information, the molecular parts in the second sample molecule are ordered sequentially, and the second prediction key point is determined in the sequential ordering of the second sample molecule.

6. The method according to claim 4, characterized in that, The method further includes: Obtain docking structure information of docking molecules, wherein the docking molecules are molecules obtained by molecular docking of the first sample molecule and the second sample molecule, and the docking structure information is used to indicate the position information of at least four molecular parts in the docking molecule; The attention network is invoked to perform docking prediction on the first structural feature to obtain the first docking position corresponding to the first sample molecule; and the attention network is invoked to perform docking prediction on the second structural feature to obtain the second docking position corresponding to the second sample molecule. Based on the first docking position and the second docking position, the predicted docking information is determined; The step of training the prediction model based on the difference between the first predicted key point and the first docking key point, and the difference between the second predicted key point and the second docking key point, to obtain the trained prediction model, includes: The prediction model is trained based on the differences between the first predicted key point and the first docking key point, the differences between the second predicted key point and the second docking key point, and the differences between the predicted docking information and the docking structure information, to obtain the trained prediction model.

7. The method according to any one of claims 1 to 3, characterized in that, The acquisition of sample information pairs includes: Obtain the first position information of at least two molecular parts in the first sample molecule and the second position information of at least two molecular parts in the second sample molecule; Based on a first proximity distance threshold between molecular parts of the first sample molecule, a first proximity feature of at least two molecular parts in the first sample molecule is determined according to the first location information; and based on a second proximity distance threshold between molecular parts of the second sample molecule, a second proximity feature of at least two molecular parts in the second sample molecule is determined according to the second location information. The first sample structure is determined based on the first proximity feature and the first location information, and the second sample structure is determined based on the second proximity feature and the second location information; Obtain the first docking key point and the second docking key point.

8. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain the first molecular structure of the first predicted molecule and the second molecular structure of the second predicted molecule; The trained prediction model is invoked to predict the first molecular structure and the second molecular structure respectively, to obtain the first key point corresponding to the first predicted molecule and the second key point corresponding to the second predicted molecule.

9. A predictive model training device for molecular docking information, characterized in that, The device includes: The acquisition module is used to acquire sample information pairs, the sample information pairs including a first sample structure, a second sample structure, a first docking key point corresponding to the first sample structure, and a second docking key point corresponding to the second sample structure; the first sample structure is used to indicate the positional features of at least two molecular parts in the first sample molecule, and the second sample structure is used to indicate the positional features of at least two molecular parts in the second sample molecule. The processing module is configured to perform structural perturbation processing on the first sample structure based on a first elastic distance threshold between molecular parts of the first sample molecule to obtain a first perturbation structure; and to perform the structural perturbation processing on the second sample structure based on a second elastic distance threshold between molecular parts of the second sample molecule to obtain a second perturbation structure; the structural perturbation processing is used to modify the relative positions between at least two molecular parts; The prediction module is used to call the prediction model to perform prediction processing on the first disturbance structure and the second disturbance structure respectively, so as to obtain the first prediction key point corresponding to the first disturbance structure and the second prediction key point corresponding to the second disturbance structure. The training module is used to train the prediction model based on the difference between the first prediction key point and the first docking key point, and the difference between the second prediction key point and the second docking key point, to obtain the trained prediction model. The first docking key point and the second docking key point are used to indicate the key molecular parts of the first sample molecule and the second sample molecule when they dock.

10. A computer device, characterized in that, The computer device includes: a processor and a memory, wherein the memory stores at least one program; the processor is configured to execute the at least one program in the memory to implement the molecular docking information prediction model training method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The readable storage medium stores executable instructions, which are loaded and executed by a processor to implement the molecular docking information prediction model training method as described in any one of claims 1 to 8.

12. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium. The processor reads and executes the computer instructions from the computer-readable storage medium to implement the molecular docking information prediction model training method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for training key point detection model and method for detecting key points of target object

    CN113095336A

  • Computational method for predicting functional sites of biological molecules

    US20140279758A1