A model training method, and a molecular property information prediction method and device
Patent Information
- Application Number
- CN202310714271.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2043-06-15
AI Technical Summary
[0005]然而,目前并没有有效地方式来对分子性质进行准确的预测
[0055] As can be seen from the above method, this application can construct a three-dimensional molecular map of the specified protein degradation target chimera by obtaining data on the specified protein degradation target chimera molecule. This three-dimensional molecular map data can fully characterize various features of the molecular structure of the specified protein degradation target chimera molecule. Then, after inputting the three-dimensional molecular map data into the prediction model, the prediction model will predict the molecular property information of the specified protein degradation target chimera molecule based on the three-dimensional molecular map data. Then, based on the deviation between the predicted molecular property information and the actual molecular property information corresponding to the specified protein degradation target chimera molecule, the prediction model is trained, so that in the subsequent process of predicting molecular property information, the molecular property information can be obtained quickly and accurately through the prediction model, thereby improving the efficiency and accuracy of determining molecular property information.
Smart Images

Figure CN116524998B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the fields of artificial intelligence and bioengineering, and in particular to a method for model training and a method and apparatus for predicting molecular property information. Background Technology
[0002] The ubiquitination enzyme system is a classic pathway for degrading proteins in the body. Utilizing this pathway, protein degradation-targeting chimeras with dual-function fragments can be designed to clear pathogenic proteins from the body. Among these, protein degradation-targeting chimeras are a targeted protein degradation technology that uses small molecule compounds to regulate protein levels. Compared to traditional small molecule drugs, they have a different mode of action, delivering proteins to the proteasome via unique targets to directly mediate protein ubiquitination and degradation.
[0003] Currently, protein degradation targeting chimera technology is a hot topic in related pharmaceutical research. These molecules have been used to develop drugs to treat diseases including cancer, immune system diseases, and nervous system diseases. Especially in the field of anti-cancer, protein degradation targeting chimeras can target proteins that induce cancer cells to produce, thereby eliminating the side effects caused by chemotherapy drugs.
[0004] Molecular property prediction is crucial for applications such as drug discovery or protein design. By analyzing a molecular structure model to estimate its relevant physical and chemical properties and calculating precise molecular properties, the efficiency and accuracy of technology design can be greatly improved.
[0005] However, there is currently no effective way to accurately predict molecular properties. Summary of the Invention
[0006] This specification provides a method for model training and a method and apparatus for recommending molecular structure information, in order to partially solve the aforementioned problems existing in the prior art.
[0007] The following technical solution is adopted in this specification:
[0008] This manual provides a method for model training, including:
[0009] Obtain data on the degradation of a specified protein targeting a chimeric molecule;
[0010] Based on the data, construct three-dimensional molecular map data of the specified protein degradation-targeting chimeric molecule;
[0011] The three-dimensional molecular map data of the specified protein degradation target chimeric molecule is input into the prediction model to be trained, so that the prediction model can predict the molecular property information of the specified protein degradation target chimeric molecule.
[0012] The prediction model is trained based on the deviation between the predicted molecular property information and the actual molecular property information corresponding to the specified protein degradation target chimeric molecule.
[0013] Optionally, the molecular property information includes at least one of the following: the lipophilicity of the specified protein degradation targeting chimeric molecule, the pH of the specified protein degradation targeting chimeric molecule, the molecular weight of the specified protein degradation targeting chimeric molecule, the hydrogen bond donors and acceptors in the specified protein degradation targeting chimeric molecule, the solubility of the specified protein degradation targeting chimeric molecule, and the permeability of the specified protein degradation targeting chimeric molecule.
[0014] Optionally, based on the data, three-dimensional molecular map data of the specified protein degradation-targeting chimeric molecule is constructed, specifically including:
[0015] Using each atom in the specified protein degradation targeting chimeric molecule as an atom node, an adjacency matrix corresponding to each atom is constructed to represent the edges between each atom node, so as to obtain the three-dimensional molecular graph data of the specified protein degradation targeting chimeric molecule. The adjacency matrix is used to represent the bonding information between connected atoms.
[0016] Optionally, the node information corresponding to the atomic nodes of each atom contained in the three-dimensional molecular graph data includes at least one of atom type information, atom three-dimensional coordinate information, and atomic nuclear charge number.
[0017] Optionally, the three-dimensional molecular map data of the specified protein degradation-targeting chimeric molecule is input into the prediction model to be trained, so that the prediction model predicts the molecular property information of the specified protein degradation-targeting chimeric molecule, specifically including:
[0018] The three-dimensional molecular map data of the chimeric molecule targeted by the specified protein degradation is input into the embedding layer of the prediction model to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer.
[0019] The embedding vector is input into the information update layer of the prediction model to determine the invariant and equivariant features of the specified protein degradation-targeting chimeric molecule;
[0020] The invariant and isovariant features are input into the output layer of the prediction model, so that the molecular property information of the predicted protein degradation-targeting chimeric molecule is output through the output layer.
[0021] Optionally, the three-dimensional molecular map data of the chimeric molecule targeted by the specified protein degradation is input into the embedding layer of the prediction model to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer, specifically including:
[0022] For each atom in the specified protein degradation targeting chimeric molecule, the atomic feature information corresponding to the atom is determined based on the three-dimensional molecular map data. The atomic feature information includes at least one of the following: atom type information of the atom, spacing information between the atom and other atoms, and neighboring atom information.
[0023] The atomic feature information corresponding to each atom is input into the embedding layer in the prediction model, so as to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer.
[0024] Optionally, the embedding vector is input into the information update layer of the prediction model to determine the invariant and equivariant features of the specified protein degradation-targeting chimeric molecule, specifically including:
[0025] The embedding vector is input into the information update layer of the prediction model to determine the initial invariant features of the specified protein degradation-targeting chimeric molecule through the information update layer.
[0026] The information update layer determines the attention weights corresponding to the initial invariant features, and based on the attention weights, determines the attention features corresponding to the specified protein degradation-targeting chimeric molecule.
[0027] The information update layer determines the invariant and equivariant features of the specified protein degradation-targeting chimeric molecule based on the attention features.
[0028] Optionally, through the information update layer, based on the attention features, the invariant and isovariant features of the designated protein degradation-targeting chimeric molecule are determined, specifically including:
[0029] The initial isomorphic features are determined through the information update layer;
[0030] Through the information update layer, the initial invariant features are updated according to the initial equivariant features to obtain the updated initial invariant features, and the initial equivariant features are updated according to the initial invariant features to obtain the updated initial equivariant features.
[0031] Through the information update layer, the invariant feature is determined based on the updated initial invariant feature and the attention feature, and the equivariant feature is determined based on the updated initial equivariant feature, the updated initial invariant feature and the attention feature.
[0032] This specification provides a method for predicting molecular property information, including:
[0033] Obtain data on the degradation of the target protein and the targeting chimeric molecule;
[0034] Based on the data, construct three-dimensional molecular map data of the target protein degradation targeting chimeric molecule;
[0035] The three-dimensional molecular map data of the target protein degradation-targeting chimeric molecule is input into a pre-trained prediction model so that the prediction model can predict the molecular property information of the target protein degradation-targeting chimeric molecule. The prediction model is trained using the model training method described above.
[0036] This specification provides a model training apparatus, comprising:
[0037] The acquisition module is used to acquire data on chimeric molecules that target the degradation of a specified protein.
[0038] A construction module is used to construct three-dimensional molecular map data of the specified protein degradation-targeting chimeric molecule based on the data;
[0039] The prediction module is used to input the three-dimensional molecular map data of the specified protein degradation target chimeric molecule into the prediction model to be trained, so that the prediction model can predict the molecular property information of the specified protein degradation target chimeric molecule.
[0040] The training module is used to train the prediction model based on the deviation between the predicted molecular property information and the actual molecular property information corresponding to the specified protein degradation target chimeric molecule.
[0041] Optionally, the molecular property information includes at least one of the following: the lipophilicity of the specified protein degradation targeting chimeric molecule, the pH of the specified protein degradation targeting chimeric molecule, the molecular weight of the specified protein degradation targeting chimeric molecule, the hydrogen bond donors and acceptors in the specified protein degradation targeting chimeric molecule, the solubility of the specified protein degradation targeting chimeric molecule, and the permeability of the specified protein degradation targeting chimeric molecule.
[0042] Optionally, the construction module is specifically used to take each atom contained in the specified protein degradation targeting chimeric molecule as an atomic node, and construct an adjacency matrix corresponding to each atom to represent the edges between each atomic node, so as to obtain the three-dimensional molecular graph data of the specified protein degradation targeting chimeric molecule, wherein the adjacency matrix is used to represent the bonding information between connected atoms.
[0043] Optionally, the node information corresponding to the atomic nodes of each atom contained in the three-dimensional molecular graph data includes at least one of atom type information, atom three-dimensional coordinate information, and atomic nuclear charge number.
[0044] Optionally, the prediction module is specifically configured to: input the three-dimensional molecular map data of the specified protein degradation-targeting chimeric molecule into the embedding layer of the prediction model to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer; input the embedding vector into the information update layer of the prediction model to determine the invariant and isovariant features of the specified protein degradation-targeting chimeric molecule; and input the invariant and isovariant features into the output layer of the prediction model to output the predicted molecular property information of the specified protein degradation-targeting chimeric molecule through the output layer.
[0045] Optionally, the prediction module is specifically used to determine the atomic feature information corresponding to each atom in the specified protein degradation target chimeric molecule based on the three-dimensional molecular map data. The atomic feature information includes at least one of the following: atomic type information of the atom, spacing information between the atom and other atoms, and neighboring atom information of the atom. The atomic feature information corresponding to each atom is input into the embedding layer in the prediction model to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer.
[0046] Optionally, the prediction module is specifically configured to: input the embedding vector into the information update layer of the prediction model, so as to determine the initial invariant features of the specified protein degradation targeting chimeric molecule through the information update layer; determine the attention weights corresponding to the initial invariant features through the information update layer, and determine the attention features corresponding to the specified protein degradation targeting chimeric molecule based on the attention weights; and determine the invariant features and equivalent features of the specified protein degradation targeting chimeric molecule based on the attention features through the information update layer.
[0047] Optionally, the prediction module is specifically configured to: determine initial isovariant features through the information update layer; update the initial invariant features according to the initial isovariant features through the information update layer to obtain updated initial invariant features, and update the initial isovariant features according to the initial invariant features to obtain updated initial isovariant features; determine the invariant features according to the updated initial invariant features and the attention features through the information update layer, and determine the isovariant features according to the updated initial isovariant features, the updated initial invariant features and the attention features.
[0048] This specification provides a device for predicting molecular property information, including:
[0049] The acquisition module is used to acquire data on the degradation of the target protein and the targeted chimeric molecule.
[0050] A construction module is used to construct three-dimensional molecular map data of the target protein degradation targeting chimeric molecule based on the data;
[0051] The prediction module is used to input the three-dimensional molecular map data of the target protein degradation targeting chimeric molecule into a pre-trained prediction model, so that the prediction model can predict the molecular property information of the target protein degradation targeting chimeric molecule. The prediction model is trained by the above-mentioned model training method.
[0052] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for training the model or the method for predicting molecular property information.
[0053] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method for training a model or a method for predicting molecular property information.
[0054] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:
[0055] As can be seen from the above method, this application can construct a three-dimensional molecular map of the specified protein degradation target chimera by obtaining data on the specified protein degradation target chimera molecule. This three-dimensional molecular map data can fully characterize various features of the molecular structure of the specified protein degradation target chimera molecule. Then, after inputting the three-dimensional molecular map data into the prediction model, the prediction model will predict the molecular property information of the specified protein degradation target chimera molecule based on the three-dimensional molecular map data. Then, based on the deviation between the predicted molecular property information and the actual molecular property information corresponding to the specified protein degradation target chimera molecule, the prediction model is trained, so that in the subsequent process of predicting molecular property information, the molecular property information can be obtained quickly and accurately through the prediction model, thereby improving the efficiency and accuracy of determining molecular property information. Attached Figure Description
[0056] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0057] Figure 1 This is a flowchart illustrating a model training method provided in this specification;
[0058] Figure 2 A schematic diagram illustrating the process of a method for predicting molecular property information provided in this specification;
[0059] Figure 3 This is a schematic diagram of the architecture of a molecular property prediction system provided in this specification;
[0060] Figure 4 A schematic diagram of a model training apparatus provided in this specification;
[0061] Figure 5 A schematic diagram of a molecular property information prediction device provided in this specification;
[0062] Figure 6 The one provided in this specification corresponds to Figure 1 or Figure 2 A schematic diagram of the structure of an electronic device. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0064] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0065] Figure 1 This is a flowchart illustrating a model training method provided in this specification, including the following steps:
[0066] S101: Obtain data on chimeric molecules that target the degradation of a specified protein.
[0067] The model training method provided in this manual can be executed by terminal devices such as desktop computers and laptops, or by servers. For ease of explanation, this manual only uses terminal devices as the execution subject to describe the provided model training method.
[0068] In this specification, the terminal device can acquire raw structural data of various protein degradation-targeting chimeric molecules to construct a dataset for training the model. This data can be obtained by crawling from external networks or by file entry.
[0069] After obtaining the above dataset, the terminal device needs to search for suitable data to train the prediction model. In other words, not all data on various protein degradation targeting chimeric molecules included in the dataset are suitable as samples for training. Some data may not have good labeled molecular fragments, and some data may be considered "dirty data".
[0070] Therefore, the terminal device needs to search the dataset for suitable protein degradation-targeting chimeric molecules to serve as training samples; that is, to identify the specified protein degradation-targeting chimeric molecules. Specifically, this can be achieved by creating a scalable 3D molecular graph structure data generator to clean, reconstruct, and optimize the data of protein degradation-targeting chimeric molecules, thereby searching for data of the specified protein degradation-targeting chimeric molecules to serve as training samples, and subsequently determining the 3D molecular graph data of the specified protein degradation-targeting chimeric molecules.
[0071] S102: Based on the data, construct the three-dimensional molecular map data of the specified protein degradation-targeting chimera.
[0072] In this specification, after the terminal device determines the data of the specified protein degradation targeting chimeric molecule from the aforementioned dataset, it can further determine the three-dimensional molecular map data of the specified protein degradation targeting chimeric molecule. The three-dimensional molecular map data mentioned here is used to represent the structural features of the specified protein degradation targeting chimeric molecule, such as the type of each atom, the position of each atom, and the connection relationships between the atoms.
[0073] Specifically, the aforementioned three-dimensional molecular graph data can specify each atom contained in the protein degradation targeting chimeric molecule as an atomic node, and construct an adjacency matrix corresponding to each atom to represent the edges between each atomic node, so as to obtain the three-dimensional molecular graph data of the specified protein degradation targeting chimeric molecule.
[0074] The adjacency matrix is used to represent the bonding information between connected atoms. For example, when two atoms are connected by a single bond, it can be represented by "0", when they are connected by a double bond, it can be represented by "1", when they are connected by a triple bond, it can be represented by "2", when they are connected by a quadruple bond, it can be represented by "3", and so on.
[0075] Furthermore, for the aforementioned three-dimensional molecular graph data, each atom node corresponds to a specific node information. For any given atom, the node information of its atomic nodes can be used to represent some inherent characteristics of that atom. For example, the node information of the atom node may include at least one of the following: atom type information, atom three-dimensional coordinate information, and atomic nuclear charge number.
[0076] Atom type information indicates which type of atom the atom belongs to, such as carbon (C), oxygen (O), sulfur (S), nitrogen (N), fluorine (F), chlorine (Cl), etc. There are several ways to determine the atom type information, such as using one-hot encoding, or by using pre-defined feature vectors corresponding to each element's symbol.
[0077] For the three-dimensional coordinate information of each atom, each atom can be projected into a preset three-dimensional coordinate system to determine the three-dimensional coordinate information of each atom in the three-dimensional coordinate system. The preset three-dimensional coordinate system can be a three-dimensional coordinate system constructed with the molecular centroid of the specified protein degradation target chimeric molecule as the origin.
[0078] As can be seen from the above, the three-dimensional molecular map data of the specified protein degradation target chimeric molecule determined in this specification can consider the structural characteristics of the specified protein degradation target chimeric molecule from multiple perspectives, thereby ensuring the accuracy and rationality of the subsequent prediction model in the prediction results.
[0079] S103: Input the three-dimensional molecular map data of the specified protein degradation target chimeric molecule into the prediction model to be trained, so that the prediction model can predict the molecular property information of the specified protein degradation target chimeric molecule.
[0080] In this specification, the prediction model described above includes an embedding layer. The main function of this embedding layer is to obtain an embedding vector after the three-dimensional molecular map data is input into the prediction model. Then, the molecular property information is predicted in subsequent processes using the embedding vector.
[0081] In addition, the prediction model in this specification may also include an information update layer. The main function of this information update layer is to input the above-mentioned embedding vector into the information update layer, and then obtain the invariant and isovariant features of the protein degradation targeting chimeric molecule through multiple data update operations performed by the information update layer. Then, the invariant and isovariant features are input into the output layer of the prediction model, so that the molecular property information of the predicted specified protein degradation targeting chimeric molecule can be output through the output layer.
[0082] Specifically, for each atom in the chimeric molecule targeting protein degradation, the terminal device can first determine the atomic feature information corresponding to the atom based on the three-dimensional molecular map data. The atomic feature information mentioned here is used to represent the characteristics of the atom, and may specifically include the atom type information of the atom (i.e., the atom type information is the information mentioned above used to indicate which type of atom it belongs to), the spacing information between the atom and other atoms, and the information of the atom's neighboring atoms, etc.
[0083] In this specification, the spacing information between two atoms can be determined in several ways. For example, after determining the three-dimensional coordinates of each atom, the distance between the three-dimensional coordinates of two atoms can be determined, and the determined distance can be directly used as the spacing information between the two atoms. Another example is the determination of the spacing information between two atoms using radial basis functions (RBF). The specific method for determining the spacing information using radial basis functions can be found in the following formula:
[0084]
[0085] Where, d ij d is used to represent the distance between atoms i and j. c This indicates the distance cutoff point; the specific value can be determined according to requirements. RBF (d ij Then, it is used to represent the spacing information between atoms i and j as determined by the radial basis function.
[0086] The adjacent atom information mentioned above can be determined using the following embedding formula:
[0087]
[0088] embed1 is used to represent the Cauchy image embedding function. Used to represent a linear mapping function, A ij L1 is used to represent the molecular adjacency matrix containing bonding information, and e is used to represent the linear transformation function. ij This is used to represent the neighboring atom information of atom i and atom j.
[0089] After determining the atomic feature information corresponding to each atom, the atomic feature information corresponding to each atom can be input into the embedding layer in the prediction model to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer.
[0090] As mentioned above, in this specification, the information update layer is mainly used to obtain the invariant and isovariant features of the protein degradation-targeting chimeric molecule through multiple data update operations performed by the information update layer after the embedding vector is input into it. Therefore, the information update layer mainly uses attention and message passing mechanisms to continuously update the invariant and isovariant features, and after the feature update is completed, the updated features are used to predict molecular property information.
[0091] Specifically, after the embedding vector is input into the information update layer of the prediction model, the initial invariant features of the specified protein degradation-targeting chimeric molecule can be determined through the information update layer. Then, the attention weights corresponding to the initial invariant features are determined through the information update layer, and the attention features corresponding to the specified protein degradation-targeting chimeric molecule are determined based on the attention weights. Finally, the invariant features and equivalent features of the specified protein degradation-targeting chimeric molecule are determined through the information update layer based on the attention features. This process can be considered to be completed entirely within the information update layer.
[0092] When determining the initial invariant characteristics of a chimeric molecule that targets the degradation of a specific protein, the following formula can be used:
[0093]
[0094] Among them, c i Used to represent the nuclear charge number of atom i, x i The atom type information of atom i is used to represent atom i, and embed2 is used to represent the Laplacian feature embedding function. Used to represent linear mapping functions, while This is used to represent the initial invariant characteristics of atom i.
[0095] As for the initial isovariant characteristics, they can be determined by initializing the values, as shown in the following formula:
[0096]
[0097] As can be seen from the above formula, the initial isovariant features can be initialized as follows:
[0098] Once the initial isovariant features and initial invariant features are determined, these two features can be updated to obtain the updated initial isovariant features and updated initial invariant features.
[0099] The prediction model can first determine the attention weights corresponding to the initial invariant features through the attention mechanism in the information update layer, which can be specifically determined by the following formula:
[0100] Q = W Q S+b Q
[0101] K = W K S+b k
[0102] Among them, W Q The weight coefficients, b, are used to represent the attention weights Q corresponding to unequal features. QW is used to represent the bias coefficient of the attention weight Q corresponding to unequal features. K The weight coefficients, b, are used to represent the attention weights K corresponding to unequal features. k The bias coefficient used to represent the attention weight K corresponding to unequal features.
[0103] Then, through the attention mechanism in the information update layer, the attention features corresponding to the specified protein degradation targeting chimeric molecule can be further determined based on the determined attention weights. See the following formula for details:
[0104]
[0105]
[0106] Among them, W a Let Q be the weight matrix. i The attention weights Q and K are used to represent the unequal features corresponding to the determined atom i. i The attention weights K, e, are used to represent the unequal features corresponding to the determined atom i. i Used to represent the neighboring atom information of atom i, y1, y2, and y3 have no particularly practical meaning and can be regarded as A output The output matrix was split into three matrices: y1, y2, and y3. This representation is primarily used to determine the invariant and isovariant characteristics of atoms in the target chimeric molecule for the specified protein degradation in subsequent processes.
[0107] Ω is used to represent a linear transformation function. The main purpose of processing y1 through this linear transformation function is to transform the obtained attention features y1 into a single feature vector.
[0108] For the initial equivariant features, this specification suggests using the VN-MLP equivariant multilayer perceptron mechanism for updating, as detailed in the following formula:
[0109]
[0110]
[0111] Where W represents the attention weight Q corresponding to the isovariant feature. V The weight coefficients, b1, are used to represent the attention weights Q corresponding to the equivariant features. V The bias coefficient, U, is used to represent the attention weight K corresponding to the equivariant feature. V The weight coefficients, b2, are used to represent the attention weights K corresponding to the equivariant features. V The bias coefficient.
[0112] Then, the attention features for the initial isovariant features can be determined by using the attention weights corresponding to the determined isovariant features. The specific formula is as follows:
[0113]
[0114] After obtaining the attention features for the initial invariant features and the initial isovariant features through the above method, the initial invariant features and the initial isovariant features can be updated through the information update layer.
[0115] Specifically, through the information update layer, the initial invariant features can be updated based on the initial isovariant features to obtain the updated initial invariant features. This can be achieved by first determining the attention features corresponding to the initial isovariant features, and then updating the initial invariant features based on these attention features, using the following formula:
[0116]
[0117]
[0118] In the above formula, S′ i and S" i This can be viewed as using different equivariant functions to perform two different updates on the initial invariant features; VN-MLP1 and VN-MLP2 represent the equivariant multilayer perceptron functions. Therefore, in the above formula... and These are two different isovariant functions.
[0119] For updating the initial isovariant features, an information update layer can be used to update the initial isovariant features based on the initial invariant features, resulting in the updated initial isovariant features. Specifically, this process can involve first determining the attention features corresponding to the initial isovariant features based on the initial isovariant features, and then updating the attention features corresponding to the initial isovariant features based on the initial invariant features. See the following formula for details:
[0120]
[0121] `diag` is used to represent an equivariant function, and `diag` is used to represent a function that transforms a matrix into a diagonal matrix form. That is, the attention features used to represent the initial isovariant features. For the initial invariant feature S i The updated initial isovariant features obtained after the update.
[0122] After obtaining the updated initial invariant features, the information update layer can determine the final invariant features using the updated initial invariant features and the determined attention features. See the following formula for details:
[0123]
[0124] Where ker1 is used to represent the Kelvin function, Used to represent the distance between atom i and atom j, while Then it is used to represent the updated initial invariant feature S′ i and attention features The invariant features that have been determined.
[0125] For isovariant features, the isovariant features can be determined based on the updated initial isovariant features, the updated initial invariant features, and the determined attention features. The specific formula is as follows:
[0126]
[0127] Where ker2 and ker3 are used to represent Kelvin functions, Used to represent the distance between atom i and atom j, while Then it is used to represent the updated initial invariant feature S". i Updated initial isovariant features And the isovariant features determined by y2 and y3 in the attention features.
[0128] In this specification, the prediction model may actually have multiple information update layers. These multiple information update layers communicate with each other through feature transfer to continuously update and optimize invariant and equivalent features, so that the updated and optimized invariant and equivalent features are finally input into the output layer of the prediction model.
[0129] Therefore, after determining the above invariant and equivariant features, they can be combined into the feature combination of the following formula.
[0130]
[0131] Then, the feature combination and the original input are fed into the next information update layer to obtain the output of the information update layer, and the output is then passed to the next information update layer, and so on. For details, please refer to the following formula:
[0132]
[0133] In this formula, m iX is the combination of invariant features and isovariant features of atom i. i f represents the atomic characteristic information of atom i. a Used to represent the attention mechanism applied to m i To process, This is used to represent the output of the information update layer. It should be noted that when determining the invariant characteristics of atom i, the atomic characteristic information of atom j connected to atom i is actually considered. However, atom i may be connected to more than one atom. Therefore, when determining the invariant characteristics of atom i, the atomic characteristic information of atom i and each connected atom j can actually be considered. By summarizing, the same applies to the isovariant characteristics of atom i, thus obtaining the characteristic combination m of atom i. i .
[0134] As can be seen from the above, the invariant features of the identified protein degradation targeting chimera molecules actually characterize the molecular features of the identified protein degradation targeting chimera molecules that do not change with the molecular structure (because the invariant features of each atom are actually determined by information such as atomic charge number and three-dimensional coordinate information, which are usually fixed under a given molecular structure). On the other hand, the isovariant features of the identified protein degradation targeting chimera can characterize the properties exhibited by some more noteworthy atoms in the molecular structure of the identified protein degradation targeting chimera molecules (because isovariant features are mainly determined through attention mechanisms).
[0135] Therefore, it can be understood that the subsequent output layer actually predicts the molecular properties of the specified protein degradation target chimera based on the isovariant and invariant characteristics of the specified protein degradation target chimera molecule, the characteristics of some atoms in the molecular structure of the specified protein degradation target chimera molecule that are more noteworthy, and some fixed characteristics of the specified protein degradation target chimera molecule in the molecular structure.
[0136] S104: The prediction model is trained based on the deviation between the predicted molecular property information and the actual molecular property information corresponding to the specified protein degradation target chimeric molecule.
[0137] After predicting the above molecular property information, the deviation between the molecular property information and the actual molecular property information corresponding to the specified protein degradation target chimeric molecule can be further determined. Then, the prediction model is trained with minimizing this deviation as the optimization objective.
[0138] The molecular property information mentioned in this specification may include at least one of the following: lipophilicity of the specified protein degradation targeting chimera molecule, pH of the specified protein degradation targeting chimera molecule, molecular weight of the specified protein degradation targeting chimera molecule, hydrogen bond donors and acceptors in the specified protein degradation targeting chimera molecule, solubility of the specified protein degradation targeting chimera molecule, and permeability of the specified protein degradation targeting chimera molecule.
[0139] As can be seen from the above methods, the model training method provided in this specification references the structural features of the molecular structure of the specified protein degradation target chimeric molecule from multiple perspectives when determining the three-dimensional molecular map data. Furthermore, when predicting the molecular properties of the specified protein degradation target chimeric molecule, since it is determined based on the invariant and isovariant characteristics of the molecule, this method can fully characterize the molecular structural properties of the molecule, thereby ensuring that the prediction model can subsequently make accurate and reasonable predictions of molecular properties.
[0140] After training the above prediction model, it can be used to predict the molecular properties of a specified molecule in practical applications. The specific process is shown in the figure below.
[0141] Figure 2 This is a schematic diagram illustrating the process of a method for predicting molecular property information provided in this specification.
[0142] S201: Obtain data on the degradation of the target protein targeting the chimeric molecule.
[0143] S202: Based on the data, construct three-dimensional molecular map data of the target protein degradation targeting chimeric molecule.
[0144] In this specification, the terminal device can receive a user-input command to predict the molecular properties of the target protein degradation-targeting chimeric molecule. Through this command, it acquires data on the target protein degradation-targeting chimeric molecule and constructs a three-dimensional molecular map of it. The determination of this three-dimensional molecular map data is essentially the same as the process used in the model training described above, and will not be elaborated upon here. The terminal device mentioned here can refer to a desktop computer, laptop computer, or similar device.
[0145] S202: Input the three-dimensional molecular map data of the target protein degradation targeting chimeric molecule into a pre-trained prediction model so that the prediction model can predict the molecular property information of the target protein degradation targeting chimeric molecule. The prediction model is trained using the model training method described above.
[0146] The terminal device can input the three-dimensional molecular map data of the target protein degradation targeting chimeric molecule into the prediction model deployed in the terminal device, and the prediction model will output the predicted molecular property information of the target protein degradation targeting chimeric molecule.
[0147] After obtaining molecular property information through the prediction model, this information can be displayed to the user, or information can be recommended based on this information, such as recommending drug compounds that can effectively integrate with the target protein degradation-targeting chimeric molecule. Alternatively, in practical applications, the three-dimensional molecular map data of the target protein degradation-targeting chimeric molecule can be saved in conjunction with the predicted molecular property information.
[0148] This specification also provides a system for predicting molecular properties, such as Figure 3 As shown.
[0149] Figure 3 This is a schematic diagram of the architecture of a molecular property prediction system provided in this specification.
[0150] from Figure 3 As can be seen from this, the system mainly consists of the following parts:
[0151] The storage subsystem is used to store the aforementioned dataset, as well as information on molecular properties and their pharmaceutical and chemical properties predicted by the prediction model in practical applications.
[0152] The control subsystem is used to predict the molecular properties of the target protein degradation-targeting chimeric molecule based on the three-dimensional molecular map data of the target protein degradation-targeting chimeric molecule input into the subsystem.
[0153] The control subsystem includes two units: a molecular feature extraction unit and a molecular property prediction unit. These two units are used to determine the atomic feature information of each atom from the three-dimensional molecular map data and to predict the molecular property information of the target protein degradation targeting chimeric molecule, respectively.
[0154] The above describes one or more implementations of the methods described in this specification. Based on the same approach, this specification also provides corresponding model training devices and molecular structure information recommendation devices, such as... Figure 4 , Figure 5 As shown.
[0155] Figure 4 A schematic diagram of a model training apparatus provided in this specification includes:
[0156] Acquisition module 401 is used to acquire data on chimeric molecules targeting the degradation of a specified protein;
[0157] Construction module 402 is used to construct three-dimensional molecular map data of the specified protein degradation target chimeric molecule based on the data;
[0158] The prediction module 403 is used to input the three-dimensional molecular map data of the specified protein degradation target chimeric molecule into the prediction model to be trained, so that the prediction model can predict the molecular property information of the specified protein degradation target chimeric molecule.
[0159] The training module 404 is used to train the prediction model based on the deviation between the predicted molecular property information and the actual molecular property information corresponding to the specified protein degradation target chimeric molecule.
[0160] Optionally, the molecular property information includes at least one of the following: the lipophilicity of the specified protein degradation targeting chimeric molecule, the pH of the specified protein degradation targeting chimeric molecule, the molecular weight of the specified protein degradation targeting chimeric molecule, the hydrogen bond donors and acceptors in the specified protein degradation targeting chimeric molecule, the solubility of the specified protein degradation targeting chimeric molecule, and the permeability of the specified protein degradation targeting chimeric molecule.
[0161] Optionally, the construction module 402 is specifically used to use each atom contained in the specified protein degradation targeting chimeric molecule as an atomic node, and to construct an adjacency matrix corresponding to each atom to represent the edges between each atomic node, so as to obtain the three-dimensional molecular graph data of the specified protein degradation targeting chimeric molecule, wherein the adjacency matrix is used to represent the bonding information between connected atoms.
[0162] Optionally, the node information corresponding to the atomic nodes of each atom contained in the three-dimensional molecular graph data includes at least one of atom type information, atom three-dimensional coordinate information, and atomic nuclear charge number.
[0163] Optionally, the prediction module 403 is specifically configured to: input the three-dimensional molecular map data of the specified protein degradation-targeting chimeric molecule into the embedding layer of the prediction model to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer; input the embedding vector into the information update layer of the prediction model to determine the invariant and isovariant features of the specified protein degradation-targeting chimeric molecule; and input the invariant and isovariant features into the output layer of the prediction model to output the predicted molecular property information of the specified protein degradation-targeting chimeric molecule through the output layer.
[0164] Optionally, the prediction module 403 is specifically used to determine the atomic feature information corresponding to each atom in the specified protein degradation target chimeric molecule based on the three-dimensional molecular map data. The atomic feature information includes at least one of the following: atomic type information of the atom, spacing information between the atom and other atoms, and neighboring atom information of the atom. The atomic feature information corresponding to each atom is input into the embedding layer in the prediction model to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer.
[0165] Optionally, the prediction module 403 is specifically configured to: input the embedding vector into the information update layer of the prediction model, so as to determine the initial invariant features of the specified protein degradation targeting chimeric molecule through the information update layer; determine the attention weights corresponding to the initial invariant features through the information update layer, and determine the attention features corresponding to the specified protein degradation targeting chimeric molecule based on the attention weights; and determine the invariant features and equivalent features of the specified protein degradation targeting chimeric molecule based on the attention features through the information update layer.
[0166] Optionally, the prediction module 403 is specifically configured to: determine initial isovariant features through the information update layer; update the initial invariant features according to the initial isovariant features through the information update layer to obtain updated initial invariant features, and update the initial isovariant features according to the initial invariant features to obtain updated initial isovariant features; determine the invariant features according to the updated initial invariant features and the attention features through the information update layer, and determine the isovariant features according to the updated initial isovariant features, the updated initial invariant features and the attention features.
[0167] Figure 5 A schematic diagram of a molecular property information prediction device provided in this specification includes:
[0168] Acquisition module 501 is used to acquire data on the degradation of the target protein targeting the chimeric molecule;
[0169] Construction module 502 is used to construct three-dimensional molecular map data of the target protein degradation targeting chimeric molecule based on the data;
[0170] The prediction module 503 is used to input the three-dimensional molecular map data of the target protein degradation targeting chimeric molecule into a pre-trained prediction model, so that the prediction model can predict the molecular property information of the target protein degradation targeting chimeric molecule. The prediction model is trained by the above-described model training method.
[0171] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 A method for model training or Figure 2 This provides a method for predicting molecular property information.
[0172] This instruction manual also provides Figure 6 One of the corresponding Figure 1 or Figure 2 A schematic diagram of the structure of an electronic device. (e.g.) Figure 6 As shown, at the hardware level, this electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above. Figure 1 The method for training the model or Figure 2 The method for predicting molecular property information.
[0173] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0174] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0175] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0176] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0177] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.
[0178] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0179] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0180] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0181] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0182] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0183] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0184] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0185] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0186] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0187] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0188] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0189] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for training a model, characterized in that, include: Obtain data on the degradation of a specified protein targeting a chimeric molecule; Based on the data, construct three-dimensional molecular map data of the specified protein degradation-targeting chimeric molecule; The three-dimensional molecular map data of the chimeric molecule targeted by the specified protein degradation is input into the embedding layer of the prediction model to be trained, so as to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer. The embedding vector is input into the information update layer of the prediction model, so that the initial invariant features and initial equivariant features of the specified protein degradation target chimeric molecule are determined through the information update layer. The process involves determining the attention weights corresponding to the initial invariant features, and based on these attention weights, determining the attention features corresponding to the specified protein degradation-targeting chimeric molecule; updating the initial invariant features based on the initial isovariant features to obtain updated initial invariant features, and updating the initial isovariant features based on the initial invariant features to obtain updated initial isovariant features; determining the invariant features based on the updated initial invariant features and the attention features, and determining the isovariant features based on the updated initial isovariant features, the updated initial invariant features, and the attention features. The invariant features and the isovariant features are input into the output layer of the prediction model, so that the molecular property information of the predicted protein degradation-targeting chimeric molecule is output through the output layer. The prediction model is trained based on the deviation between the predicted molecular property information and the actual molecular property information corresponding to the specified protein degradation target chimeric molecule.
2. The method as described in claim 1, characterized in that, The molecular property information includes at least one of the following: lipophilicity of the specified protein degradation targeting chimeric molecule, pH of the specified protein degradation targeting chimeric molecule, molecular weight of the specified protein degradation targeting chimeric molecule, hydrogen bond donors and acceptors in the specified protein degradation targeting chimeric molecule, solubility of the specified protein degradation targeting chimeric molecule, and permeability of the specified protein degradation targeting chimeric molecule.
3. The method as described in claim 1, characterized in that, Based on the data, a three-dimensional molecular map of the specified protein degradation-targeting chimeric molecule is constructed, specifically including: Using each atom in the specified protein degradation targeting chimeric molecule as an atom node, an adjacency matrix corresponding to each atom is constructed to represent the edges between each atom node, so as to obtain the three-dimensional molecular graph data of the specified protein degradation targeting chimeric molecule. The adjacency matrix is used to represent the bonding information between connected atoms.
4. The method as described in claim 3, characterized in that, The node information corresponding to the atomic nodes of each atom contained in the three-dimensional molecular graph data includes at least one of the following: atom type information, atom three-dimensional coordinate information, and atomic nuclear charge number.
5. The method as described in claim 1, characterized in that, The three-dimensional molecular map data of the chimeric molecule targeted by the specified protein degradation is input into the embedding layer of the prediction model to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer, specifically including: For each atom in the specified protein degradation targeting chimeric molecule, the atomic feature information corresponding to the atom is determined based on the three-dimensional molecular map data. The atomic feature information includes at least one of the following: atom type information of the atom, spacing information between the atom and other atoms, and neighboring atom information. The atomic feature information corresponding to each atom is input into the embedding layer in the prediction model, so as to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer.
6. A method for predicting molecular property information, characterized in that, include: Obtain data on the degradation of the target protein and the targeting chimeric molecule; Based on the data, construct three-dimensional molecular map data of the target protein degradation targeting chimeric molecule; The three-dimensional molecular map data of the target protein degradation targeting chimeric molecule is input into a pre-trained prediction model so that the prediction model can predict the molecular property information of the target protein degradation targeting chimeric molecule. The prediction model is trained by the method described in any one of claims 1 to 5.
7. A device for model training, characterized in that, Includes: The acquisition module is used to acquire data on chimeric molecules that target the degradation of a specified protein. A construction module is used to construct three-dimensional molecular map data of the specified protein degradation-targeting chimeric molecule based on the data; The prediction module inputs the three-dimensional molecular map data of the specified protein degradation target chimeric molecule into the embedding layer of the prediction model to be trained, so as to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer; and inputs the embedding vector into the information update layer of the prediction model, so as to determine the initial invariant features and initial isovariant features of the specified protein degradation target chimeric molecule through the information update layer. The process involves determining the attention weights corresponding to the initial invariant features, and based on these attention weights, determining the attention features corresponding to the specified protein degradation-targeting chimeric molecule; updating the initial invariant features based on the initial isovariant features to obtain updated initial invariant features, and updating the initial isovariant features based on the initial invariant features to obtain updated initial isovariant features; determining the invariant features based on the updated initial invariant features and the attention features, and determining the isovariant features based on the updated initial isovariant features, the updated initial invariant features, and the attention features; and inputting the invariant features and the isovariant features into the output layer of the prediction model to output the predicted molecular property information of the specified protein degradation-targeting chimeric molecule. The training module is used to train the prediction model based on the deviation between the predicted molecular property information and the actual molecular property information corresponding to the specified protein degradation target chimeric molecule.
8. The apparatus as claimed in claim 7, characterized in that, The molecular property information includes at least one of the following: lipophilicity of the specified protein degradation targeting chimeric molecule, pH of the specified protein degradation targeting chimeric molecule, molecular weight of the specified protein degradation targeting chimeric molecule, hydrogen bond donors and acceptors in the specified protein degradation targeting chimeric molecule, solubility of the specified protein degradation targeting chimeric molecule, and permeability of the specified protein degradation targeting chimeric molecule.
9. The apparatus as claimed in claim 7, characterized in that, The construction module is specifically used to take each atom contained in the specified protein degradation targeting chimeric molecule as an atomic node, and to construct an adjacency matrix corresponding to each atom to represent the edges between each atomic node, so as to obtain the three-dimensional molecular graph data of the specified protein degradation targeting chimeric molecule. The adjacency matrix is used to represent the bonding information between connected atoms.
10. The apparatus as claimed in claim 9, characterized in that, The node information corresponding to the atomic nodes of each atom contained in the three-dimensional molecular graph data includes at least one of the following: atom type information, atom three-dimensional coordinate information, and atomic nuclear charge number.
11. The apparatus as claimed in claim 10, characterized in that, The prediction module is specifically used to determine the atomic feature information corresponding to each atom in the specified protein degradation target chimeric molecule based on the three-dimensional molecular map data. The atomic feature information includes at least one of the following: the atom type information of the atom, the distance information between the atom and other atoms, and the neighboring atom information of the atom. The atomic feature information corresponding to each atom is input into the embedding layer in the prediction model to obtain the embedding vector corresponding to the three-dimensional molecular map data through the embedding layer.
12. A device for predicting molecular property information, characterized in that, include: The acquisition module is used to acquire data on the degradation of the target protein and the targeted chimeric molecule. A construction module is used to construct three-dimensional molecular map data of the target protein degradation targeting chimeric molecule based on the data; A prediction module is used to input the three-dimensional molecular map data of the target protein degradation targeting chimeric molecule into a pre-trained prediction model, so that the prediction model predicts the molecular property information of the target protein degradation targeting chimeric molecule, wherein the prediction model is trained by the method described in any one of claims 1 to 5.
13. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 6.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Three-dimensional protein-ligand activity prediction method based on attention mechanism
CN115512785A
Method and apparatus for determining drug molecule property, and storage medium
US20220415452A1