Method and device for acquiring composite characterization information, storage medium, and equipment
By training a model to obtain the spatial conformational changes of specified nodes before and after mutation of antibody-antigen complex samples, this method solves the problem of inaccurate characterization information caused by the failure to consider three-dimensional spatial structure in existing technologies, and improves the accuracy of affinity change prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-12
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies do not consider the three-dimensional spatial structure of the complex when characterizing the nodes of the antibody-antigen complex after mutation, resulting in inaccurate characterization information and affecting the prediction results of changes in affinity.
By obtaining the first difference corresponding to the spatial conformational changes of a specified node before and after mutation of a complex sample, the initial information acquisition model is trained to obtain the target information acquisition model, and the characterization information of the complex is obtained using this model.
This improves the accuracy of predictions regarding changes in affinity, resulting in more reasonable and accurate characterization information.
Smart Images

Figure CN116682490B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital medical treatment, and in particular to a compound representation information acquisition method and device, a storage medium and a computer device. BACKGROUND
[0002] In the field of digital medical treatment, the interaction between proteins and proteins is of great significance for the treatment of diseases, for example, the interaction between antibodies and antigens. If the interaction between antibodies and antigens is good, the disease corresponding to the antigen can be treated by using the antibody.
[0003] In order to find a better antibody for treating a disease, a mutation is usually made based on an existing antibody, and then it is predicted whether the interaction between the mutated antibody and the antigen is good, that is, the change in affinity between the mutated antibody and the antigen is predicted. This method usually needs to obtain the node representation corresponding to the mutated antibody-antigen complex, so as to predict the change in affinity according to the node representation.
[0004] In the prior art, the node representation corresponding to the mutated antibody-antigen complex is usually obtained by using deep learning technology. However, when obtaining the node representation corresponding to the mutated antibody-antigen complex, the three-dimensional spatial structure of the complex is usually not considered, resulting in that the finally obtained node representation information is not accurate, and thus directly affecting the prediction result of the subsequent change in affinity. SUMMARY
[0005] Therefore, the present application provides a compound representation information acquisition method and device, a storage medium and a computer device. The first difference corresponding to the spatial conformation change of the specified node of the compound sample before and after mutation is used to train an initial information acquisition model to obtain a target information acquisition model. The model training process can learn more differences in spatial structure, so that more reasonable and accurate representation information can be obtained when the target information acquisition model is used to acquire the representation information of the compound, and the accuracy of the prediction result of the subsequent change in affinity is greatly improved.
[0006] According to one aspect of the present application, a compound representation information acquisition method is provided, comprising:
[0007] obtaining a first graph structure of each first compound sample in a first compound sample set and a second graph structure of each second compound sample in a second compound sample set, the first compound sample and the second compound sample corresponding one-to-one, the first compound sample being a compound before mutation, and the second compound sample being a compound after mutation;
[0008] calculate a first difference corresponding to each of the plurality of specified nodes of each of the first complex sample and the corresponding second complex sample, the first difference being determined based on a spatial difference of the specified node;
[0009] train an initial information acquisition model according to the first graph structure, the second graph structure, and the first differences of the plurality of specified nodes to obtain a target information acquisition model;
[0010] obtain a complex to be analyzed, input a target graph structure of the complex to be analyzed into the target information acquisition model, and obtain representation information of the complex to be analyzed, the representation information including node representation of each target node in the complex to be analyzed.
[0011] According to another aspect of the present application, a device for obtaining complex representation information is provided, comprising:
[0012] a graph structure obtaining module configured to obtain a first graph structure of each first complex sample in a first complex sample set and a second graph structure of each second complex sample in a second complex sample set, the first complex sample and the second complex sample corresponding to each other, the first complex sample being a pre-mutation complex, and the second complex sample being a post-mutation complex;
[0013] a difference calculating module configured to calculate a first difference corresponding to each of the plurality of specified nodes of each of the first complex sample and the corresponding second complex sample, the first difference being determined based on a spatial difference of the specified node;
[0014] a model training module configured to train an initial information acquisition model according to the first graph structure, the second graph structure, and the first differences of the plurality of specified nodes to obtain a target information acquisition model;
[0015] a representation information determining module configured to obtain a complex to be analyzed, input a target graph structure of the complex to be analyzed into the target information acquisition model, and obtain representation information of the complex to be analyzed, the representation information including node representation of each target node in the complex to be analyzed.
[0016] According to still another aspect of the present application, a storage medium having a computer program stored thereon is provided, the program being executed by a processor to implement the above method for obtaining complex representation information.
[0017] According to yet another aspect of the present application, a computer device is provided, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, the processor implementing the above method for obtaining complex representation information when executing the program.
[0018] By means of the technical scheme, the method and device for acquiring compound characterization information, the storage medium and the computer device provided by the application can first acquire a first graph structure of each first compound sample in a first compound sample set and a second graph structure of each second compound sample in a second compound sample set. Then, a plurality of specified nodes are determined in the first compound sample, and the three-dimensional spatial positions of the specified nodes in the first graph structure of the first compound sample and the three-dimensional spatial positions of the specified nodes in the second graph structure of the second compound sample are determined. After that, the first difference corresponding to each specified node can be obtained by performing difference calculation on the basis of the two three-dimensional spatial positions of each specified node. The initial information acquisition model is trained by using the first graph structure corresponding to each first compound sample, the second graph structure corresponding to each second compound sample, and the first difference corresponding to the specified nodes in each first graph structure and second graph structure. Finally, the target information acquisition model can be obtained. After obtaining the target information acquisition model, the target information acquisition model can be used to analyze the compound to be analyzed to obtain the characterization information corresponding to the compound to be analyzed. The first difference corresponding to the spatial conformation change of the specified nodes before and after the mutation of the compound sample is used to train the initial information acquisition model to obtain the target information acquisition model, so that more spatial structure differences can be learned in the model training process. Therefore, when the characterization information of the compound is acquired by using the target information acquisition model, more reasonable and accurate characterization information can be obtained, and the accuracy of the subsequent affinity change prediction result can be greatly improved.
[0019] The above description is only a summary of the technical scheme of the application. In order to more clearly understand the technical means of the application, the application can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS
[0020] The drawings described herein are used to provide further understanding of the application, and form a part of the application. The schematic embodiments of the application and the description thereof are used to explain the application, and do not constitute an improper limitation on the application. In the drawings:
[0021] Figure 1 A flowchart of a method for acquiring compound characterization information provided by an embodiment of the application is shown;
[0022] Figure 2 A flowchart of another method for acquiring compound characterization information provided by an embodiment of the application is shown;
[0023] Figure 3A structural schematic diagram of a device for acquiring complex characterization information is shown. DETAILED DESCRIPTION
[0024] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0025] In the present embodiment, a method for acquiring complex characterization information is provided, as shown in the figure, the method comprises: Figure 1
[0026] Step 101, acquiring a first graph structure of each first complex sample in a first complex sample set and a second graph structure of each second complex sample in a second complex sample set, the first complex sample and the second complex sample correspond to each other, the first complex sample is a pre-mutation complex, and the second complex sample is a post-mutation complex;
[0027] The method for acquiring complex characterization information provided in the present embodiment can be applied in the field of digital medical technology, and specifically can be applied in the process of acquiring characterization information of protein-protein complexes, wherein the protein-protein complex can be an antigen-antibody complex, etc. First, the graph structure corresponding to each complex sample in two complex sample sets can be acquired, wherein the two complex sample sets are a first complex sample set and a second complex sample set, the first complex sample set can include a plurality of first complex samples, and the second complex sample set can include a plurality of second complex samples. The number of first complex samples in the first complex sample set is the same as the number of second complex samples in the second complex sample set, and any first complex sample in the first complex sample set has a corresponding second complex sample, that is, the first complex sample and the second complex sample correspond to each other. The graph structure of the complex sample can be a three-dimensional space graph structure of the complex sample. The first complex sample can be a pre-mutation complex, and the second complex sample can be a post-mutation complex. For example, in the field of digital medical technology, the first complex sample can be an original antibody-antigen complex, and the second complex sample can be a mutated antibody-antigen complex.
[0028] Step 102, calculating a first difference corresponding to each specified node in a plurality of specified nodes of each of the first complex sample and the corresponding second complex sample, the first difference being determined based on the spatial difference of the specified node;
[0029] In this embodiment, in order to study the changes of the node representations of the first complex sample and the second complex sample before and after mutation, the specified nodes can be determined from the first complex sample, and only the node representations of the specified nodes are used to determine the changes of the first complex sample and the second complex sample before and after mutation. Specifically, the specified nodes can be nodes of interest, for example, when the first complex sample and the second complex sample are protein-protein complex samples, the specified nodes can be amino acids of interest, which can be determined artificially. After the specified nodes are determined in the first complex sample, these nodes are also taken as specified nodes in the second complex sample, because the nodes themselves do not change before and after mutation of the complex, only the spatial positions of the nodes in the complex change. Therefore, after a plurality of specified nodes are determined in the first complex sample, the three-dimensional spatial positions of the specified nodes in the first graph structure of the first complex sample can be determined, and the specified nodes can also be found in the second complex sample, and the three-dimensional spatial positions of the specified nodes in the second graph structure of the second complex sample are determined. Finally, the difference is calculated based on the two three-dimensional spatial positions of each specified node, and the first difference corresponding to the specified node is obtained.
[0030] In step 103, the initial information acquisition model is trained based on the first graph structure, the second graph structure, and the first differences of the plurality of specified nodes, and a target information acquisition model is obtained.
[0031] In this embodiment, an initial information acquisition model can be constructed. Specifically, the initial information acquisition model can be a graph attention network model. Then, the first graph structure corresponding to each first complex sample, the second graph structure corresponding to each second complex sample, and the first difference corresponding to each specified node in the first graph structure and the second graph structure can be used to train the initial information acquisition model, and finally the target information acquisition model can be obtained.
[0032] In step 104, a complex to be analyzed is obtained, and the target graph structure of the complex to be analyzed is input into the target information acquisition model to obtain the representation information of the complex to be analyzed, wherein the representation information includes the node representation of each target node in the complex to be analyzed.
[0033] In this embodiment, after obtaining the target information acquisition model, the target information acquisition model can be used to analyze the compound to be analyzed to obtain the characterization information corresponding to the compound to be analyzed. Specifically, the target graph structure of the compound to be analyzed can be obtained, and the target graph structure can be a three-dimensional spatial graph structure of the compound to be analyzed. Then, the target graph structure of the compound to be analyzed can be input into the target information acquisition model, and the target information acquisition model can correspondingly output the node characterization corresponding to each target node in the compound to be analyzed. Here, the target graph structure can only include target nodes, that is, the target graph structure is a graph structure composed of the spatial positions of the target nodes and the connection relationship therebetween, so that the node characterization corresponding to the target nodes can be directly output from the target information acquisition model; the target graph structure can also include nodes other than the target nodes, that is, the target graph structure is a graph structure composed of the spatial positions of all nodes and the connection relationship therebetween, but the target nodes can be marked, and the node characterization corresponding to the target nodes can be directly extracted from the node characterization output from the target information acquisition model to form the characterization information of the compound to be analyzed.
[0034] By applying the technical solution of this embodiment, first, the first graph structure of each first compound sample in the first compound sample set and the second graph structure of each second compound sample in the second compound sample set can be obtained. Then, a plurality of specified nodes are determined in the first compound sample, and the three-dimensional spatial positions of the specified nodes in the first graph structure of the first compound sample and the three-dimensional spatial positions of the specified nodes in the second graph structure of the second compound sample are determined, and then the first difference corresponding to each specified node is obtained by performing difference calculation based on the two three-dimensional spatial positions of each specified node. By using the first graph structure corresponding to each first compound sample, the second graph structure corresponding to each second compound sample, and the first difference corresponding to the specified nodes in each first graph structure and second graph structure, the initial information acquisition model is trained, and finally the target information acquisition model can be obtained. After obtaining the target information acquisition model, the target information acquisition model can be used to analyze the compound to be analyzed to obtain the characterization information corresponding to the compound to be analyzed. The present application trains the initial information acquisition model by using the first difference corresponding to the spatial conformation change of the specified nodes before and after the mutation of the compound sample, obtains the target information acquisition model, can make the model learn more differences in spatial structure during the model training process, and thus when the characterization information of the compound is obtained by using the target information acquisition model subsequently, more reasonable and accurate characterization information can be obtained, and the accuracy of the subsequent prediction result of the affinity change condition is greatly improved.
[0035] Further, as a refinement and extension of the above embodiment, in order to fully describe the specific implementation process of the embodiment, another method for acquiring compound characterization information is provided, as shown in Figure 2 The method comprises the following steps:
[0036] Step 201, determining the specified nodes in each of the first compound samples, taking the three-dimensional coordinates of the target atoms corresponding to the specified nodes as the three-dimensional coordinates of the specified nodes, and determining the specified nodes in each of the second compound samples, taking the three-dimensional coordinates of the target atoms corresponding to the specified nodes as the three-dimensional coordinates of the specified nodes.
[0037] In this embodiment, the specified nodes in the first compound sample can be determined first, and the specified nodes are the nodes of interest. Since the nodes in the first compound sample before mutation and the nodes in the second compound sample after mutation are the same, the difference lies in the change of the spatial position of the nodes, therefore, the specified nodes in the first compound sample can also appear in the second compound sample. After determining the specified nodes in the first compound sample, the three-dimensional coordinates of the target atoms in the specified nodes can be directly taken as the three-dimensional coordinates of the specified nodes. By using the same method, the three-dimensional coordinates of the specified nodes in the second compound sample can be determined. Here, the target atoms in the specified nodes can be pre-specified atoms, for example, the first compound sample and the second compound sample are both protein-protein compound samples, and the specified nodes can be specified amino acids therein, and the target atoms can be the middle carbon atoms in the amino acids, that is, the three-dimensional coordinates of the middle carbon atoms can be directly taken as the three-dimensional coordinates of the specified amino acids.
[0038] Step 202, constructing the first graph structure based on the three-dimensional coordinates corresponding to each of the specified nodes in the first compound sample, and constructing the second graph structure based on the three-dimensional coordinates corresponding to each of the specified nodes in the second compound sample.
[0039] In this embodiment, after determining the three-dimensional coordinates of each specified node in the first compound sample, the first graph structure can be constructed based on the three-dimensional coordinates of these specified nodes and the mutual connection relationship between each two specified nodes, that is, the first graph structure can be a graph structure composed of the specified nodes. Similarly, the second graph structure can be constructed.
[0040] Step 203, acquiring the first graph structure of each first compound sample in the first compound sample set and the second graph structure of each second compound sample in the second compound sample set, the first compound sample and the second compound sample corresponding to each other, the first compound sample being a compound before mutation, and the second compound sample being a compound after mutation.
[0041] Step 204, determining a first coordinate system corresponding to each specified node in the first composite sample, and a second coordinate system corresponding to each specified node in the second composite sample;
[0042] In this embodiment, each first composite sample can include a plurality of specified nodes, and a first coordinate system corresponding to each specified node can be determined, that is, each specified node corresponds to its own coordinate system, which can be a three-dimensional coordinate system, which can include an origin and three coordinate axes. Similarly, a second coordinate system corresponding to each specified node in each second composite sample can be determined.
[0043] Step 205, sequentially taking each specified node as a reference node;
[0044] In this embodiment, the first composite sample can include a plurality of specified nodes, and when calculating the first difference of each specified node, each specified node can be sequentially taken as a reference node.
[0045] Step 206, taking the first coordinate system corresponding to the reference node in the first composite sample as a reference coordinate system, and performing coordinate transformation on the first coordinate system of the remaining specified nodes to obtain a transformed coordinate corresponding to each remaining specified node;
[0046] In this embodiment, after the reference node is determined, the first coordinate system corresponding to the reference node can be marked as the reference coordinate system, and then the first coordinate system corresponding to the other remaining specified nodes in the specified nodes can be respectively transformed, and each first coordinate system can be transformed into the reference coordinate system, and a transformed coordinate corresponding to each remaining specified node (which can be a transformed coordinate of the origin in the first coordinate system of each remaining specified node) can be obtained. For example, the first composite sample includes five specified nodes, namely specified node 1 to specified node 5, and each of the five specified nodes corresponds to a first coordinate system. The five specified nodes can be taken as reference nodes, for example, specified node 1 is taken as a reference node, then the first coordinate system corresponding to specified node 1 can be taken as a reference coordinate system, and then the first coordinate systems of the remaining four specified nodes, namely specified node 2 to specified node 5, can be transformed into the reference coordinate system, and the transformed coordinates of specified node 2 to specified node 5 in the reference coordinate system can be obtained. Similarly, when specified node 2 to specified node 5 are sequentially taken as reference nodes, the transformed coordinates of the other four remaining specified nodes in the reference coordinate system can be obtained each time.
[0047] Step 207, taking the second coordinate system corresponding to the reference node in the second composite sample as a reference coordinate system, and performing coordinate transformation on the second coordinate system of the remaining specified nodes to obtain a transformed coordinate corresponding to each remaining specified node;
[0048] In this embodiment, each specified node in the second composite sample can also be taken as a reference node, and then the transformation coordinates of each remaining specified node corresponding to the reference node are obtained.
[0049] In step 208, the reference coordinate system of the reference node in the first composite sample is overlapped with the reference coordinate system of the reference node in the second composite sample, and the node sub-differences of the remaining specified nodes are calculated.
[0050] In this embodiment, for any specified node in the first composite sample, the specified node can also be found in the second composite sample. Therefore, when any specified node is taken as a reference node, the reference coordinate system of the reference node in the first composite sample can be overlapped with the reference coordinate system of the reference node in the second composite sample, specifically, the coordinate origin and the three coordinate axes of the two reference coordinate systems can be overlapped in turn, and then the node sub-differences corresponding to the remaining specified nodes are calculated. For example, the first composite sample includes five specified nodes, namely specified node 1 to specified node 5. After the specified node 1 is taken as a reference node, four transformation coordinates of the specified node 2 to the specified node 5 in the reference coordinate system of the first composite sample and four transformation coordinates of the specified node 2 to the specified node 5 in the reference coordinate system of the second composite sample are obtained. Then, the transformation coordinates of the specified node 2 in the two composite samples are subtracted to obtain the node sub-difference corresponding to the specified node 2. Similarly, the node sub-differences corresponding to the specified node 3 to the specified node 5 are obtained.
[0051] In step 209, the node sub-differences corresponding to each specified node under each reference node are added to obtain the first difference of each specified node.
[0052] In this embodiment, the node sub-differences of each specified node under each reference node can be added together, and then the first difference of each specified node can be obtained. For example, the first composite sample includes five specified nodes, namely, specified node 1 to specified node 5. When the specified node 1 is the reference node, the corresponding node sub-differences of the specified node 2 to the specified node 5 can be obtained; when the specified node 2 is the reference node, the corresponding node sub-differences of the specified node 1, the specified node 3 to the specified node 5 can be obtained; when the specified node 3 is the reference node, the corresponding node sub-differences of the specified node 1 to the specified node 2, the specified node 4 to the specified node 5 can be obtained; when the specified node 4 is the reference node, the corresponding node sub-differences of the specified node 1 to the specified node 3, the specified node 5 can be obtained; and when the specified node 5 is the reference node, the corresponding node sub-differences of the specified node 1 to the specified node 4 can be obtained. After the five specified nodes are taken as the reference nodes respectively, each specified node actually corresponds to four node sub-differences. Adding the four node sub-differences together, the first difference of each specified node can be obtained.
[0053] In step 210, the first graph structure is input into the initial information acquisition model, and the first representation corresponding to each specified node in the first graph structure is obtained.
[0054] In this embodiment, the input of the initial information acquisition model can be a graph structure, and the output can be a node representation corresponding to each specified node in the graph structure. Specifically, the first graph structure can be input into the initial information acquisition model, and the first representation of each specified node in the first graph structure is output.
[0055] In step 211, the second graph structure is input into the initial information acquisition model, and the second representation corresponding to each specified node in the second graph structure is obtained.
[0056] In this embodiment, the second representation of each specified node in the second graph structure can also be obtained.
[0057] In step 212, the node difference value of each specified node is calculated according to the first representation and the second representation.
[0058] In this embodiment, the first representation and the second representation can be vectors, specifically one-dimensional vectors. After the first representation and the second representation are obtained, the node difference value between the first representation and the second representation can be calculated, wherein the node difference value is also a one-dimensional vector.
[0059] In step 213, the node difference value is input into a preset multi-layer perception machine, and the second difference of each specified node is obtained.
[0060] In this embodiment, after obtaining the node difference value of each specified node, the node difference value can be input into the preset multi-layer perception, so that the preset multi-layer perception can correspondingly output the second difference corresponding to each specified node. For example, the node difference values corresponding to each specified node in the first complex sample can be input into the preset multi-layer perception together, and finally the preset multi-layer perception can correspondingly output the second difference corresponding to each specified node in the specified nodes. The second difference can indicate the difference degree between the first representation and the second representation of each specified node.
[0061] Step 214, calculating the difference difference value between the first difference and the second difference of each specified node, taking the sum of the difference difference values corresponding to each specified node as a target loss function, training the initial information acquisition model to obtain the target information acquisition model;
[0062] In this embodiment, the first difference corresponding to each specified node is actually the real spatial position difference of the specified node in the first complex sample and in the corresponding second complex sample; the second difference is actually the predicted difference of the specified node in the first complex sample and in the corresponding second complex sample, which is obtained based on the node representation. The difference difference value between the first difference and the second difference can obtain the gap between the prediction and the reality. When the second difference is closer to the first difference, it means that the accuracy of the target information acquisition model is better. Therefore, the sum of the difference difference values of each specified node can be taken as a target loss function to train the initial information acquisition model, and finally the target information acquisition model can be obtained. It should be noted that the first complex sample includes multiple, and the specified nodes in each first complex sample also include multiple. Assuming that the first complex sample is M, and the specified nodes in each first complex sample include N (actually the number of specified nodes in each first complex sample can be different), then the target loss function can be the sum of the difference difference values of M*N specified nodes.
[0063] Step 215, obtaining a complex to be analyzed, inputting the target graph structure of the complex to be analyzed into the target information acquisition model to obtain the representation information of the complex to be analyzed, wherein the representation information includes the node representation of each target node in the complex to be analyzed.
[0064] In this embodiment, after obtaining the target information acquisition model, the target information acquisition model can be used to analyze the to-be-analyzed complex to obtain the characterization information corresponding to the to-be-analyzed complex. Specifically, the target graph structure of the to-be-analyzed complex can be obtained, and the target graph structure can be a three-dimensional spatial graph structure of the to-be-analyzed complex. Then, the target graph structure of the to-be-analyzed complex can be input into the target information acquisition model, and the target information acquisition model can correspondingly output the node characterization corresponding to each target node in the to-be-analyzed complex.
[0065] In the embodiments of the present application, the "determining the first coordinate system corresponding to each specified node in the first complex sample" in step 204 comprises: determining the target structure corresponding to each specified node from the first complex sample, and determining the first coordinate system corresponding to the specified node based on the first sub-chain, the second sub-chain and the third sub-chain of the target structure; and the "determining the second coordinate system corresponding to each specified node in the second complex sample" in step 204 comprises: determining the target structure corresponding to each specified node from the second complex sample, and determining the second coordinate system corresponding to the specified node based on the first sub-chain, the second sub-chain and the third sub-chain of the internal structure.
[0066] In this embodiment, the first coordinate system of each specified node can be determined by the following method. First, the target structure corresponding to each specified node can be determined from the first complex sample. For example, when the first complex sample is an antibody-antigen complex sample and the specified node is a specified amino acid, the target structure can be a specified component in the specified amino acid, and the specified component can be the intermediate carbon atom, the amino group, the carboxyl group and the R group. Then, the first sub-chain, the second sub-chain and the third sub-chain can be determined from the target structure, and the first coordinate system of the specified node can be determined according to the three sub-chains. For example, taking the intermediate carbon atom as the coordinate origin, the amino group, the carboxyl group and the R group are determined as the three sub-chains respectively, each sub-chain corresponds to a coordinate axis, and then the first coordinate system of the specified node can be determined. Similarly, the second coordinate system corresponding to each specified node in the second complex sample can be determined.
[0067] Optionally, in this embodiment, step 214, "using the sum of the difference values corresponding to each of the specified nodes as the target loss function to train the initial information acquisition model to obtain the target information acquisition model," includes: when the sum is greater than a preset loss threshold, adjusting the model parameters in the initial information acquisition model and the preset multilayer perceptron, and updating the second difference based on the adjusted initial information acquisition model and the preset multilayer perceptron, recalculating the sum of the difference values corresponding to each of the specified nodes based on the updated second difference, until the sum is less than or equal to the preset loss threshold, thereby obtaining the target information acquisition model.
[0068] In this embodiment, after initially calculating the sum of the difference values corresponding to each specified node, the sum can be compared with a preset loss threshold. If the sum is less than or equal to the preset loss threshold, it indicates that the accuracy of the initial information acquisition model is within an acceptable range, and the initial information acquisition model can be directly used as the target information acquisition model. If the sum is greater than the preset loss threshold, it indicates that the accuracy of the initial information acquisition model still has a large gap and further training is needed. In this case, the model parameters in the initial information acquisition model and the preset multilayer perceptron can be adjusted. After the model parameters are adjusted, the first graph structure and the second graph structure can be input into the initial information acquisition model after adjusting the model parameters to obtain new first representations and second representations corresponding to each specified node. The new node difference between the new first representation and the second representation is calculated, and the new node difference is input into the preset multilayer perceptron after adjusting the model parameters to obtain a new second difference. Then, based on the new second difference and the previous first difference, the sum of the difference values of each specified node is recalculated and compared with the preset loss threshold again... until the sum is less than or equal to the preset loss threshold, thus obtaining the target information acquisition model.
[0069] In this embodiment of the application, optionally, the first complex sample is the protein complex before mutation, the second complex sample is the protein complex after mutation, and the designated node includes a designated amino acid node.
[0070] Furthermore, as Figure 1 To specifically implement the method, embodiments of this application provide a device for acquiring complex characterization information, such as... Figure 3 As shown, the device includes:
[0071] The first graph structure acquisition module is configured to acquire a first graph structure of each first complex sample in the first complex sample set and a second graph structure of each second complex sample in the second complex sample set, the first complex sample and the second complex sample corresponding to each other, the first complex sample being a pre-mutation complex, and the second complex sample being a post-mutation complex.
[0072] The difference calculation module is configured to calculate a first difference corresponding to each specified node in each of the first complex sample and the corresponding second complex sample, the first difference being determined based on a spatial difference of the specified node.
[0073] The model training module is configured to train an initial information acquisition model to obtain a target information acquisition model according to the first graph structure, the second graph structure, and the first difference of the plurality of specified nodes.
[0074] The representation information determination module is configured to acquire a complex to be analyzed, input a target graph structure of the complex to be analyzed into the target information acquisition model, and obtain representation information of the complex to be analyzed, the representation information including node representations of each target node in the complex to be analyzed.
[0075] Optionally, the difference calculation module includes:
[0076] The coordinate system determination unit is configured to determine a first coordinate system corresponding to each specified node in the first complex sample and a second coordinate system corresponding to each specified node in the second complex sample.
[0077] The reference node determination unit is configured to sequentially take each specified node as a reference node.
[0078] The coordinate transformation unit is configured to take the first coordinate system corresponding to the reference node in the first complex sample as a reference coordinate system, perform coordinate transformation on the first coordinate system of the remaining specified nodes to obtain a transformed coordinate corresponding to each remaining specified node, and take the second coordinate system corresponding to the reference node in the second complex sample as a reference coordinate system, perform coordinate transformation on the second coordinate system of the remaining specified nodes to obtain a transformed coordinate corresponding to each remaining specified node.
[0079] The coincidence unit is configured to coincide the reference coordinate system corresponding to the reference node in the first complex sample with the reference coordinate system corresponding to the reference node in the second complex sample, and calculate a node sub-difference of the remaining specified nodes.
[0080] The difference calculation unit is configured to add the node sub-difference corresponding to each specified node under each reference node to obtain the first difference of each specified node.
[0081] Optionally, the coordinate system determining unit is further configured to:
[0082] determine a target structure corresponding to each designated node from the first complex sample, and determine a first coordinate system corresponding to the designated node based on a first sub-chain, a second sub-chain and a third sub-chain of the target structure;
[0083] The coordinate system determining unit is further configured to:
[0084] determine a target structure corresponding to each designated node from the second complex sample, and determine a second coordinate system corresponding to the designated node based on a first sub-chain, a second sub-chain and a third sub-chain of the internal structure.
[0085] Optionally, the apparatus further comprises:
[0086] a three-dimensional coordinate determining module configured to, before the first graph structure of each first complex sample in the first complex sample set is acquired, determine the designated node in each first complex sample, and determine the designated node in each second complex sample, taking the three-dimensional coordinate of a target atom corresponding to the designated node as the three-dimensional coordinate of the designated node;
[0087] a graph structure constructing module configured to construct the first graph structure based on the three-dimensional coordinates of each designated node in the first complex sample, and construct the second graph structure based on the three-dimensional coordinates of each designated node in the second complex sample.
[0088] Optionally, the model training module comprises:
[0089] a representation determining unit configured to input the first graph structure into the initial information acquisition model to obtain a first representation corresponding to each designated node in the first graph structure, and input the second graph structure into the initial information acquisition model to obtain a second representation corresponding to each designated node in the second graph structure;
[0090] a node difference determining unit configured to calculate a node difference of each designated node according to the first representation and the second representation;
[0091] a second difference determining unit configured to input the node difference into a preset multi-layer perception machine to obtain a second difference of each designated node;
[0092] The model determination unit is configured to calculate a difference difference value between the first difference and the second difference of each of the specified nodes, take a sum of the difference difference values corresponding to the specified nodes as a target loss function, train the initial information acquisition model based on the target loss function, and obtain the target information acquisition model.
[0093] Optionally, the model determination unit is further configured to:
[0094] When the sum is greater than a preset loss threshold, adjust model parameters in the initial information acquisition model and the preset multi-layer perception machine, update the second difference based on the initial information acquisition model and the preset multi-layer perception machine after the model parameters are adjusted, recalculate the sum of the difference difference values corresponding to the specified nodes according to the updated second difference, and end until the sum is less than or equal to the preset loss threshold, and obtain the target information acquisition model.
[0095] Optionally, the first complex sample is a protein complex before mutation, the second complex sample is a protein complex after mutation, and the specified nodes include specified amino acid nodes.
[0096] It should be noted that other corresponding descriptions of the functions of the device for obtaining complex characterization information provided in the embodiments of the present application can be referred to the corresponding descriptions in the method, which will not be repeated here. Figures 1 to 2 The method.
[0097] Based on the above method as shown in Figures 1 to 2 Accordingly, the embodiments of the present application also provide a storage medium having a computer program stored thereon, which, when executed by a processor, implements the above method for obtaining complex characterization information as shown in Figures 1 to 2 .
[0098] Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.
[0099] Based on the above method as shown in Figures 1 to 2 , and Figure 3 the virtual device embodiment, in order to achieve the above purpose, the embodiments of the present application also provide a computer device, which can be a personal computer, a server, a network device, etc. The computer device includes a storage medium and a processor; the storage medium is configured to store a computer program; the processor is configured to execute the computer program to implement the above method for obtaining complex characterization information as shown in Figures 1 to 2 .
[0100] Optionally, the computer device can further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, and the like. The user interface can include a display, an input unit such as a keyboard, and the like. Optionally, the user interface can further include a USB interface, a card reader interface, and the like. The network interface can optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), and the like.
[0101] Those skilled in the art can understand that the computer device structure provided by the embodiment does not constitute a limitation on the computer device, and can include more or fewer components, or combine certain components, or different component arrangements.
[0102] The storage medium can further include an operating system and a network communication module. The operating system is a program for managing and saving computer device hardware and software resources, and supports the running of information processing programs and other software and / or programs. The network communication module is used to realize communication between components in the storage medium, and communication with other hardware and software in the entity device.
[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software with a necessary general hardware platform, or by hardware. First, the first graph structure of each first complex sample in the first complex sample set and the second graph structure of each second complex sample in the second complex sample set can be obtained. Then, a plurality of specified nodes in the first complex sample are determined, and the three-dimensional spatial positions of the specified nodes in the first graph structure of the first complex sample and the three-dimensional spatial positions of the specified nodes in the second graph structure of the second complex sample are determined. Subsequently, the first difference corresponding to each specified node can be obtained by performing difference calculation based on the two three-dimensional spatial positions of each specified node. The initial information acquisition model is trained by using the first graph structure corresponding to each first complex sample, the second graph structure corresponding to each second complex sample, and the first difference corresponding to the specified nodes in each first graph structure and second graph structure. Finally, the target information acquisition model can be obtained. After obtaining the target information acquisition model, the target information acquisition model can be used to analyze the complex to be analyzed to obtain the representation information corresponding to the complex to be analyzed. The present application trains the initial information acquisition model by using the first difference corresponding to the spatial conformation change of the specified nodes before and after the mutation of the complex sample, and obtains the target information acquisition model. This can enable the model to learn more differences in spatial structure during the training process, so that when the target information acquisition model is used to obtain the representation information of the complex, more reasonable and accurate representation information can be obtained, greatly improving the accuracy of the subsequent prediction results of the affinity change.
[0104] Those skilled in the art can understand that the modules or processes in the drawings are not necessarily required for the implementation of the present application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be changed and located in one or more devices different from the implementation scenario. The modules of the above implementation scenario can be combined into one module, or can be further split into a plurality of sub-modules.
[0105]
[0106] The present application is not limited to the above specific implementation scenarios, and any changes that can be thought of by those skilled in the art should fall within the scope of the present application.
Claims
1. A method of acquiring composite characterization information, characterized by, The method comprises the following steps: obtaining a first graph structure of each first complex sample in a first complex sample set and a second graph structure of each second complex sample in a second complex sample set, the first complex sample corresponding to the second complex sample one by one, the first complex sample being a pre-mutation complex, and the second complex sample being a post-mutation complex; calculating a first difference corresponding to each specified node in a plurality of specified nodes of each of the first complex sample and the corresponding second complex sample, the first difference of each specified node being calculated based on the following manner: determining a three-dimensional spatial position of the specified node in the first graph structure of the first complex sample, and determining a three-dimensional spatial position of the specified node in the second graph structure of the second complex sample, and performing difference calculation based on the two three-dimensional spatial positions of the specified node to obtain the first difference corresponding to the specified node; training an initial information acquisition model according to the first graph structure, the second graph structure and the first differences of the plurality of specified nodes to obtain a target information acquisition model; obtaining a complex to be analyzed, inputting a target graph structure of the complex to be analyzed into the target information acquisition model to obtain representation information of the complex to be analyzed, and the representation information comprising node representations of each target node in the complex to be analyzed; the training of the initial information acquisition model according to the first graph structure, the second graph structure and the first differences of the plurality of specified nodes to obtain the target information acquisition model comprises: inputting the first graph structure into the initial information acquisition model to obtain a first representation corresponding to each of the specified nodes in the first graph structure; inputting the second graph structure into the initial information acquisition model to obtain a second representation corresponding to each of the specified nodes in the second graph structure; calculating a node difference value of each of the specified nodes according to the first representation and the second representation; inputting the node difference value into a preset multi-layer perception machine to obtain a second difference of each of the specified nodes; calculating a difference difference value between the first difference and the second difference of each of the specified nodes, taking a sum of the difference difference values corresponding to each of the specified nodes as a target loss function, training the initial information acquisition model to obtain the target information acquisition model.
2. The method of claim 1, wherein, the calculation of the first difference corresponding to each specified node in the plurality of specified nodes of each of the first complex sample and the corresponding second complex sample comprises: determining a first coordinate system corresponding to each specified node in the first complex sample and a second coordinate system corresponding to each of the specified nodes in the second complex sample; sequentially taking each of the specified nodes as a reference node; taking the first coordinate system corresponding to the reference node as a reference coordinate system in the first complex sample, and performing coordinate transformation on the first coordinate system of the remaining specified nodes to obtain a transformed coordinate corresponding to each of the remaining specified nodes; The second coordinate system corresponding to the reference node in the second complex sample is taken as a reference coordinate system, and a coordinate transformation is performed on the second coordinate system of the remaining specified node to obtain a transformed coordinate corresponding to each remaining specified node; The reference coordinate system corresponding to the reference node in the first complex sample is overlapped with the reference coordinate system corresponding to the reference node in the second complex sample, and a node sub-difference of the remaining specified node is calculated; The node sub-difference corresponding to each specified node under each reference node is added to obtain a first difference of each specified node.
3. The method of claim 2, wherein, The determination of the first coordinate system corresponding to each specified node in the first complex sample comprises: The target structure corresponding to each specified node in the first complex sample is determined, and the first coordinate system corresponding to the specified node is determined based on the first sub-chain, the second sub-chain and the third sub-chain of the target structure; The determination of the second coordinate system corresponding to each specified node in the second complex sample comprises: The target structure corresponding to each specified node in the second complex sample is determined, and the second coordinate system corresponding to the specified node is determined based on the first sub-chain, the second sub-chain and the third sub-chain of the internal structure.
4. The method of claim 1, wherein, Before the first graph structure of each first complex sample in the first complex sample set is obtained, the method further comprises: The specified node in each first complex sample is determined, and the three-dimensional coordinates of the target atom corresponding to the specified node are taken as the three-dimensional coordinates of the specified node. The specified node in each second complex sample is determined, and the three-dimensional coordinates of the target atom corresponding to the specified node are taken as the three-dimensional coordinates of the specified node; Based on the three-dimensional coordinates corresponding to each specified node in the first complex sample, the first graph structure is constructed, and based on the three-dimensional coordinates corresponding to each specified node in the second complex sample, the second graph structure is constructed.
5. The method of claim 1, wherein, The sum of the difference difference values corresponding to each specified node is taken as a target loss function, the initial information acquisition model is trained to obtain the target information acquisition model, comprising: When the sum is greater than a preset loss threshold, the model parameters in the initial information acquisition model and the preset multi-layer perception machine are adjusted, and the second difference is updated based on the initial information acquisition model and the preset multi-layer perception machine after adjusting the model parameters, and the sum of the difference difference values corresponding to each specified node is recalculated according to the updated second difference, until the sum is less than or equal to the preset loss threshold, and the target information acquisition model is obtained.
6. The method according to any one of claims 2 to 5, characterized in that, The first complex sample is a protein complex before mutation, the second complex sample is a protein complex after mutation, and the specified node includes a specified amino acid node.
7. A device for acquiring complex characterization information, characterized in that, It comprises: The first graph structure acquisition module is configured to acquire a first graph structure of each first complex sample in the first complex sample set and a second graph structure of each second complex sample in the second complex sample set, the first complex sample and the second complex sample corresponding to each other, the first complex sample being a pre-mutation complex, and the second complex sample being a post-mutation complex. The difference calculation module is configured to calculate a first difference corresponding to each specified node in each of the first complex sample and the corresponding second complex sample, the first difference of each specified node being calculated based on the following manner: determining a three-dimensional space position of the specified node in the first graph structure of the first complex sample, determining a three-dimensional space position of the specified node in the second graph structure of the second complex sample, and performing difference calculation based on the two three-dimensional space positions of the specified node to obtain the first difference corresponding to the specified node. The model training module is configured to train an initial information acquisition model based on the first graph structure, the second graph structure, and the first differences of the plurality of specified nodes to obtain a target information acquisition model. The representation information determination module is configured to acquire a complex to be analyzed, input a target graph structure of the complex to be analyzed into the target information acquisition model, and obtain representation information of the complex to be analyzed, the representation information including node representations of each target node in the complex to be analyzed. The model training module includes: The representation determination unit is configured to input the first graph structure into the initial information acquisition model to obtain a first representation corresponding to each of the specified nodes in the first graph structure, and input the second graph structure into the initial information acquisition model to obtain a second representation corresponding to each of the specified nodes in the second graph structure. The node difference value determination unit is configured to calculate a node difference value of each of the specified nodes based on the first representation and the second representation. The second difference determination unit is configured to input the node difference value into a preset multi-layer perception machine to obtain a second difference of each of the specified nodes. The model determination unit is configured to calculate a difference difference value between the first difference and the second difference of each of the specified nodes, take a sum of the difference difference values corresponding to each of the specified nodes as a target loss function, train the initial information acquisition model based on the target loss function, and obtain the target information acquisition model.
8. A storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the method in any one of claims 1 to 6.
9. A computer device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, The processor executes the computer program to implement the method in any one of claims 1 to 6.
Citation Information
Patent Citations
Disease risk assessment system, method and device and storage medium
CN112259161A
Prediction model training method and device, data prediction method and device and storage medium
CN112735535A