Model training method, molecular property information prediction method, and apparatus
The model training method for proteolytic chimeric molecules constructs three-dimensional graph data and trains predictive models to accurately predict molecular properties, addressing the inefficiencies in current methods and improving prediction accuracy.
Patent Information
- Application Number
- JP2024565024
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-06-15
- Filing Date
- 2024-02-04
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2044-02-04
AI Technical Summary
Current methods lack an effective way to accurately predict molecular properties of proteolytic chimeric molecules, crucial for drug discovery and protein design, which affects the efficiency and precision of technical design.
A model training method that constructs three-dimensional molecular graph data of proteolytic chimeric molecules, inputs this data into a predictive model, and trains the model based on the difference between predicted and actual molecular property information, using techniques like adjacency matrices, embedding layers, and information update layers to determine invariant and equivariant features.
This approach allows for quick and accurate prediction of molecular properties, enhancing the efficiency and accuracy of molecular property information determination.
Smart Images

Figure 0007911087000039 
Figure 0007911087000040 
Figure 0007911087000041
Abstract
Description
[Technical Field]
[0001] This invention relates to the fields of artificial intelligence and biotechnology, and more particularly to model training methods, molecular property information prediction methods, and apparatus. [Background technology]
[0002] Ubiquitinating enzyme systems are a classic method for degrading proteins in the body. This classic method can be used to design proteolysis-targeting chimeras (PROTACs) containing bifunctional fragments, thereby reducing or eliminating pathogenic proteins in the body. Here, proteolysis-targeting chimeras are a targeted proteolysis technique that modulates protein levels using small molecules. They have a different mode of action than conventional small molecule drugs and can deliver proteins to the proteasome using unique targets to directly mediate the protein ubiquitination and degradation processes.
[0003] Currently, proteolytic chimeric technology is a field attracting attention in related pharmaceutical research. Such molecules are being used in the development of therapeutic drugs for diseases including cancer, immune system disorders, and neurological disorders. In particular, in anti-cancer applications, proteolytic chimeric technologies can target proteins that induce the formation of cancer cells, thereby reducing the side effects of chemotherapy drugs.
[0004] Molecular property prediction is crucial for applications such as drug discovery and protein design. By analyzing a molecular structure model, estimating its associated physical and chemical properties, and calculating accurate molecular characteristics, the efficiency and precision of technical design can be significantly improved. However, currently, there is no effective method for accurately predicting molecular properties. [Overview of the Initiative]
[0005] This invention provides a model training method, a method for predicting molecular property information, and an apparatus to solve some of the problems that exist in the prior art. This invention employs the following technical proposals.
[0006] In a first aspect, the present invention provides a model training method comprising: acquiring data of a designated proteolytic chimeric molecule; constructing a three-dimensional molecular graph data of the designated proteolytic chimeric molecule based on the data; inputting the three-dimensional molecular graph data of the designated proteolytic chimeric molecule into a predictive model to be trained, causing the predictive model to predict molecular property information of the designated proteolytic chimeric molecule; and training the predictive model based on the difference between the predicted molecular property information and the actual molecular property information corresponding to the designated proteolytic chimeric molecule.
[0007] Optionally, the molecular property information includes at least one of the following: the lipophilicity of the designated proteolytic chimeric molecule, the pH of the designated proteolytic chimeric molecule, the molecular weight of the designated proteolytic chimeric molecule, the hydrogen bond donors and acceptors in the designated proteolytic chimeric molecule, the solubility of the designated proteolytic chimeric molecule, and the permeability of the designated proteolytic chimeric molecule.
[0008] Optionally, the step of constructing three-dimensional molecular graph data of the specified proteolytic chimeric molecule based on the aforementioned data includes representing the edges between each atomic node by constructing an adjacency matrix corresponding to each atom, with each atom in the specified proteolytic chimeric molecule as an atomic node, thereby obtaining the three-dimensional molecular graph data of the specified proteolytic chimeric molecule, wherein the adjacency matrix is used to represent the bonding information between each connected atom.
[0009] Optionally, each atom node in the three-dimensional molecular graph data corresponds to node information, and the node information includes at least one of the following: atom type information, atom three-dimensional coordinate information, and nuclear charge number.
[0010] Optionally, for each atom included in the three-dimensional molecular graph data, the atom type information is used to indicate which atom the atom specifically belongs to and is encoded using a one-hot encoding method; and the atom three-dimensional coordinate information is determined by projecting the atom onto a predetermined three-dimensional coordinate system, and the three-dimensional coordinate system includes a three-dimensional Cartesian coordinate system constructed with the molecular mass center of the designated proteolytic chimeric molecule as the coordinate origin.
[0011] As an option, the step of inputting the three-dimensional molecular graph data of the designated proteolytic chimeric molecule into a prediction model to be trained, and having the prediction model predict the molecular property information of the designated proteolytic chimeric molecule, includes the steps of inputting the three-dimensional molecular graph data of the designated proteolytic chimeric molecule into the embedding layer of the prediction model to obtain an embedding vector corresponding to the three-dimensional molecular graph data using the embedding layer; inputting the embedding vector into the information update layer of the prediction model to determine the target invariant features and target equivariant features of the designated proteolytic chimeric molecule; and inputting the final invariant features and final equivariant features into the output layer of the prediction model to output the molecular property information of the designated proteolytic chimeric molecule predicted by the output layer.
[0012] Optionally, the step of inputting the three-dimensional molecular graph data of the designated proteolytic chimeric molecule into the embedding layer of the prediction model to obtain an embedding vector corresponding to the three-dimensional molecular graph data by the embedding layer includes the step of determining atomic feature information corresponding to each atom in the designated proteolytic chimeric molecule based on the three-dimensional molecular graph data, wherein the atomic feature information includes at least one of the atom type information of the atom, spacing information between the atom and other atoms, and neighboring atom information of the atom; and the step of inputting the atomic feature information corresponding to each atom into the embedding layer of the prediction model to obtain an embedding vector corresponding to the three-dimensional molecular graph data by the embedding layer.
[0013] Optionally, the embedding vector includes initial invariant features, and the step of inputting the embedding vector into the information update layer in the prediction model to determine the target invariant and target equivariant features of the designated proteolytic chimeric molecule includes the step of the information update layer determining attention weights corresponding to the initial invariant features and determining attention features corresponding to the designated proteolytic chimeric molecule based on the attention weights, and the step of the information update layer determining the target invariant and target equivariant features of the designated proteolytic chimeric molecule based on the attention features.
[0014] Optionally, the step of determining the target invariant feature and the target covariant feature of the designated proteolytic induction chimeric molecule based on the attention feature by the information update layer includes: the step of obtaining, by the information update layer, an attention feature corresponding to an initial covariant feature and an attention feature corresponding to the initial invariant feature, where the initial covariant feature is a zero vector; the step of updating the attention feature corresponding to the initial invariant feature based on the attention feature corresponding to the initial covariant feature to obtain an updated initial invariant feature; the step of updating the attention feature corresponding to the initial covariant feature based on the attention feature corresponding to the initial invariant feature to obtain an updated initial covariant feature; the step of determining the target invariant feature by the information update layer based on the updated initial invariant feature and the attention feature corresponding to the initial invariant feature, and determining the target covariant feature based on the updated initial covariant feature, the updated initial invariant feature, and the attention feature corresponding to the initial invariant feature.
[0015] Optionally, in response to the number of the information update layers being 1, the target covariant feature and the target invariant feature are respectively used as the final covariant feature and the final invariant feature, and in response to the number of the information update layers being greater than 1, the target covariant feature and the target invariant feature determined by the last information update layer of the prediction model are respectively used as the final covariant feature and the final invariant feature.
[0016] In a second aspect, the present invention provides a method for predicting molecular property information, including the steps of obtaining data of a target proteolytic induction chimeric molecule, constructing three-dimensional molecular graph data of the target proteolytic induction chimeric molecule based on the data, and inputting the three-dimensional molecular graph data of the target proteolytic induction chimeric molecule into a pre-trained prediction model to cause the prediction model to predict the molecular property information of the target proteolytic induction chimeric molecule, where the prediction model is trained by the method described in the first aspect.
[0017] In a third aspect, the present invention provides a model training apparatus including: an acquisition module for acquiring data of a designated proteolysis-inducing chimeric molecule; a construction module for constructing three-dimensional molecular graph data of the designated proteolysis-inducing chimeric molecule based on the data; a prediction module for inputting the three-dimensional molecular graph data of the designated proteolysis-inducing chimeric molecule into a prediction model to be trained, and causing the prediction model to predict molecular characteristic information of the designated proteolysis-inducing chimeric molecule; and a training module for training the prediction model based on a difference between the predicted molecular characteristic information and actual molecular characteristic information corresponding to the designated proteolysis-inducing chimeric molecule.
[0018] In a fourth aspect, the present invention provides an apparatus for predicting molecular characteristic information, including: an acquisition module for acquiring data of a target proteolysis-inducing chimeric molecule; a construction module for constructing three-dimensional molecular graph data of the target proteolysis-inducing chimeric molecule based on the data; and a prediction module for inputting the three-dimensional molecular graph data of the target proteolysis-inducing chimeric molecule into a pre-trained prediction model, and causing the prediction model to predict molecular characteristic information of the target proteolysis-inducing chimeric molecule, wherein the prediction model is trained by the model training method of the first aspect.
[0019] In a fifth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the model training method of the first aspect is implemented.
[0020] In a sixth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the method for predicting molecular characteristic information of the second aspect is implemented.
[0021] In a seventh embodiment, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the model training method of the first embodiment is performed.
[0022] In an eighth embodiment, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a method for predicting molecular property information according to a second embodiment is performed.
[0023] The at least one of the above-described technical solutions adopted by the present invention can achieve the following beneficial effects.
[0024] As can be seen from the above method, the present invention can construct three-dimensional molecular graph data of a specified proteolytic chimera from the acquired data of the specified proteolytic chimera molecule. This three-dimensional molecular graph data can adequately represent various features of the molecular structure of the specified proteolytic chimera molecule. After inputting the three-dimensional molecular graph data into a prediction model, the prediction model predicts the molecular property information of the specified proteolytic chimera molecule based on the three-dimensional molecular graph data. By training the prediction model based on the difference between the predicted molecular property information and the actual molecular property information corresponding to the specified proteolytic chimera molecule, the prediction model can quickly and accurately obtain molecular property information in the subsequent molecular property information prediction process, thereby improving the efficiency and accuracy of determining molecular property information. [Brief explanation of the drawing]
[0025] The drawings described herein are for the purpose of providing a further understanding of the present invention and constitute part of the present invention. The schematic embodiments and descriptions thereof are for interpretation purposes only and do not constitute an unreasonable limitation of the present invention. [Figure 1]This is a schematic diagram showing the flow of the model training method according to the present invention. [Figure 2] This is a schematic diagram illustrating the process of the molecular property information prediction method according to the present invention. [Figure 3] This is a schematic diagram showing the architecture of the molecular property prediction system according to the present invention. [Figure 4] This is a schematic diagram showing a model training device according to the present invention. [Figure 5] This is a schematic diagram showing a molecular property information prediction device according to the present invention. [Figure 6] This is a schematic diagram showing the structure of an electronic device according to the present invention. [Modes for carrying out the invention]
[0026] To further clarify the object, technical proposal, and advantages of the present invention, the technical proposal of the present invention will be clearly and completely described below with reference to specific embodiments of the present invention and corresponding drawings. Clearly, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments that a person skilled in the art could obtain without creative effort based on the embodiments described in the present invention are all within the scope of protection of the present invention.
[0027] The technical concepts provided by each embodiment of the present invention will be described in detail below with reference to the drawings.
[0028] Figure 1 is a schematic diagram showing the flow of the model training method according to the present invention, and includes the following steps.
[0029] S101: Obtain data on a specified protein degradation-inducing chimeric molecule.
[0030] The entity executing the model training method provided by the present invention may be a terminal device such as a desktop computer or laptop computer, or it may be a server. For the sake of convenience, the present invention will describe the model training method provided using only a terminal device as the executing entity.
[0031] In this invention, the terminal device may acquire structural raw data of various proteolytic chimeric molecules to construct a dataset for training a model, and this data may be collected from an external network or acquired in the form of file input. Exemplaryly, the raw data may include SMILES (Simplified Molecular-Input Line-Entry System) datasets of class proteolytic chimeric molecules and actual proteolytic chimeric molecules disclosed on the network, where the SMILES dataset is a one-dimensional structural representation describing a compound.
[0032] After obtaining the above dataset, the terminal device needs to search for data suitable for training a predictive model. In other words, not all data on various proteolytic chimeric molecules included in the dataset are suitable as training samples; some data contains poorly labeled molecular fragments or belongs to the category of "dirty data."
[0033] Therefore, the terminal device needs to search for a suitable proteolytic chimeric molecule as a training sample from the dataset, i.e., determine a designated proteolytic chimeric molecule. A specific implementation method may involve creating an expandable three-dimensional molecular graph structure data generator to determine the three-dimensional molecular graph data of the designated proteolytic chimeric molecule in a subsequent process, and then determining the data of the designated proteolytic chimeric molecule as a training sample by performing washing, reconstruction, and optimization on the data of the class of proteolytic chimeric molecules. The designated proteolytic chimera obtained after washing, reconstruction, and optimization of the data of the class of proteolytic chimeric compounds has logP and PK values related to the following invariant and equivariant features. logP represents the oil-water partition coefficient of the compound, and the PK value represents the pharmacokinetic quality.
[0034] S102: Based on the above data, a three-dimensional molecular graph of the specified protein degradation-inducing chimera is constructed.
[0035] In this invention, the terminal device may determine data for a specified proteolytic chimeric molecule from the above dataset, and then determine three-dimensional molecular graph data for the specified proteolytic chimeric molecule. The three-dimensional molecular graph data referred to herein is used to represent structural features of the specified proteolytic chimeric molecule, such as the type of each atom, the position of each atom (e.g., Cartesian coordinates), the number of nuclear charges of each atom, and the connection relationships between each atom.
[0036] For example, a portion of the three-dimensional molecular graph data of a specified protein degradation-inducing chimeric molecule may be constructed by representing the edges between each atomic node by constructing an adjacency matrix corresponding to each atom, with each atom in the specified protein degradation-inducing chimeric molecule being an atomic node.
[0037] Here, the adjacency matrix described above is used to represent the bonding information between each connected atom. For example, for two connected atoms, the relationship between them may be represented as "0" if it is a single bond, "1" if it is a double bond, "2" if it is a triple bond, and "3" if it is a quadruple bond.
[0038] Furthermore, with respect to the three-dimensional molecular graph data described above, each atom's atomic node corresponds to corresponding node information, and for any one atom, the node information of that atom's atomic node may represent several unique characteristics of that atom. For example, the node information of that atom's atomic node may include at least one of the following: atomic type information, atomic three-dimensional coordinate information, and nuclear charge number.
[0039] Atomic type information is used to indicate which atom a given atom belongs to, such as carbon (C), oxygen (O), sulfur (S), nitrogen (N), fluorine (F), or chlorine (Cl). There are various methods for determining atomic type information; for example, it may be determined in the form of one-hot encoding, or it may be determined by feature vectors corresponding to each predetermined element symbol.
[0040] The three-dimensional coordinate information of each atom may be determined by projecting each atom onto a predetermined three-dimensional coordinate system. In some examples, the predetermined three-dimensional coordinate system may be a three-dimensional Cartesian coordinate system constructed with the molecular mass center of the designated proteolytic chimeric molecule as the coordinate origin. In other examples, the predetermined three-dimensional coordinate system may be a three-dimensional coordinate system constructed with the first molecule of the designated proteolytic chimeric molecule as the coordinate origin. The present invention is not limited thereto.
[0041] As can be seen from the above, the three-dimensional molecular graph data of the designated proteolytic chimeric molecule determined in this invention makes it possible to consider the structural characteristics of the designated proteolytic chimeric molecule from multiple perspectives, thereby guaranteeing the accuracy and rationality of the prediction results of subsequent prediction models.
[0042] S103: The three-dimensional molecular graph data of the designated protein degradation-inducing chimeric molecule is input into the prediction model to be trained, causing the prediction model to predict the molecular property information of the designated protein degradation-inducing chimeric molecule.
[0043] In this invention, an embedding layer is provided in the above-mentioned prediction model, and the input three-dimensional molecular graph data is passed through the embedding layer to obtain an embedding vector, thereby allowing molecular property information to be predicted in subsequent processes based on the embedding vector.
[0044] Furthermore, the prediction model in the present invention may be further provided with an information update layer. The main function of this information update layer is to obtain invariant and equivalent characteristics of a specified protein degradation-inducing chimeric molecule through multiple data update operations performed by the information update layer after the above-mentioned embedding vector has been input to the information update layer. Next, these invariant and equivalent characteristics are input to an output layer in the prediction model, and the output layer outputs molecular property information of the predicted specified protein degradation-inducing chimeric molecule.
[0045] Here, for each atom in the specified protein degradation-inducing chimeric molecule, the terminal device may first determine the atomic feature information corresponding to that atom based on three-dimensional molecular graph data. The atomic feature information referred to here represents the properties of the atom and may specifically include atomic type information of the atom (atomic type information is information that indicates which atom it belongs to, as referred to above), spacing information between the atom and other atoms, and information on neighboring atoms of the atom.
[0046] In this invention, the information about the distance between two atoms may be determined in various ways. For example, the three-dimensional coordinate information of each atom may be determined, and then the distance between the three-dimensional coordinates of the two atoms may be determined and the determined distance may be used directly as the information about the distance between the two atoms. Alternatively, the information about the distance between two atoms may be determined using a radial basis function (RBF). Specifically, determining the information about the distance using a radial basis function may refer to the following equation (1).
number
[0047] The neighboring atom information mentioned above may be determined by the following embedding formula (2).
number
number
[0048] After determining the atomic feature information corresponding to each atom, the atomic feature information corresponding to each atom may be input into an embedding layer in a prediction model, and an embedding vector corresponding to the three-dimensional molecular graph data may be obtained by the embedding layer. Exemplarily, the embedding vector may include the initial invariant features shown in the following formula (3-1).
[0049] As mentioned in the above content, in the present invention, the information update layer mainly inputs an embedding vector into the information update layer, and then, through a plurality of data update operations performed by the information update layer, is finally used to obtain the invariant features and equivariant features of the proteolysis-inducing chimeric molecule. In the present invention, the information update layer mainly uses an attention mechanism and a message passing mechanism to continuously update the invariant features and equivariant features, and after completing the feature update, predicts the molecular property information based on the updated features.
[0050] Specifically, after inputting the embedding vector into the information update layer in the prediction model, the information update layer determines the attention weights corresponding to the initial invariant features, and based on the attention weights, determines the attention features corresponding to the initial invariant features of the designated proteolysis-inducing chimeric molecule. Similarly, after inputting the embedding vector into the information update layer in the prediction model, the information update layer determines the attention weights corresponding to the initial equivariant features, and based on the attention weights, determines the attention features corresponding to the initial equivariant features of the designated proteolysis-inducing chimeric molecule. This entire process may be considered to be performed in the information update layer.
[0051] When determining the initial invariant features of the designated proteolysis-inducing chimeric molecule, it may be determined by the following formula (3-1).
Number
number
number
[0052] The initial equivalence features may be determined by numerical initialization, and specifically, refer to equation (3-2) below.
number
[0053] As can be seen from the above formula, the initial equivalence features are
number
[0054] After determining the initial equivariant and invariant features described above, these two features may be updated to obtain the target equivariant and invariant features.
[0055] Here, the prediction model may first obtain the embeddings of query (Q) and key (K) by the attention mechanism in the information update layer, and specifically, they may be determined by the following equations (4-1) and (4-2).
number
number
[0056] Subsequently, the attention features corresponding to the designated proteolytic chimeric molecule may be determined based on the attention weights determined by the attention mechanism in the information update layer, and specifically, the following equations (5-1) and (5-2) may be referred to.
number
number
[0057] Regarding the initial equivalence features, the present invention may update them using the mechanism of a VN-MLP equivalence multilayer perceptron, and specifically refer to the following equations (6-1) and (6-2).
number
number
number
[0058] Subsequently, attention features for the initial equivalence features may be determined using the attention weights corresponding to the determined equivalence features, and specifically, refer to equation (7) below.
number
[0059] After obtaining attention features for initial invariant features (which may also be abbreviated as invariant features) and attention features for initial equivariant features (which may also be abbreviated as equivariant features) using the above method, the invariant features and equivariant features may be further updated by an information update layer.
[0060] Specifically, the information update layer may update the invariant features based on the equivariant features and obtain the updated invariant features. Here, using the above method, first the attention features corresponding to the initial equivariant features are obtained.
number
number
number
number
number
[0061] Regarding the updating of equivariant features, the information update layer may update the equivariant features based on invariant features to obtain the updated equivariant features. Specifically, this process first determines the attention features corresponding to the initial equivariant features based on the initial equivariant features, and then the invariant features S i Based on this, attention features corresponding to the initial equivalence features
number
number
number
number
number
[0062] After obtaining the updated invariant features as described above, the information update layer retrieves the updated invariant features S'. i and attention features corresponding to the determined initial invariant features
number
number
number
number
[0063] For identical characteristics, see the updated identical characteristics.
number
number
number
number
[0064] Note that the target equivariant and invariant features obtained here are only those output by this information update layer. In some examples, in response to the number of information update layers being 1, these may be directly output to the model's output layer as the final equivariant and invariant features. In some examples, in response to the number of information update layers being greater than 1, these may be output to another information update layer for further updating and optimization of these features.
[0065] In this invention, the prediction model may be provided with multiple information update layers. Between the multiple information update layers, feature transfer enables continuous updating and optimization of invariant and equivalent features, and finally, the updated and optimized invariant and equivalent features are input to the output layer of the prediction model.
[0066] Therefore, after determining the above-mentioned target-invariant and target-equivalent features, they may be combined into a feature combination like the one shown in equation (10-1) below.
number
[0067] Next, the feature combination and the original input are input to the next information update layer, the output result of the information update layer is obtained, and the output result is transmitted to the next information update layer. Specifically, you may refer to the following equation (10-2).
number
number
number
[0068] In the final information update layer of the prediction model, the outputted target equivariant and target invariant features are output to the model's output layer as the final equivariant and invariant features. The output layer combines the final invariant and equivariant features to perform the final molecular drug formation characteristic prediction and obtain molecular characteristic information of the specified protein degradation-inducing chimeric molecule.
[0069] As can be seen from the above, the determined invariant features of the designated proteolytic chimeric molecule actually represent molecular features of the designated proteolytic chimeric molecule that do not change with the molecular structure (the invariant features of each atom are actually determined by information such as the charge number of the atom, and this information usually does not change when the molecular structure is given), and the equivariant features of the designated proteolytic chimeric molecule can represent the properties exhibited by several notable atoms in the molecular structure of the designated proteolytic chimeric molecule (the equivariant features of each atom are actually determined by information such as the three-dimensional coordinate information of each atom, and for the sake of simplifying calculations, the initial equivariant features are set to 0, and the equivariant features are determined by the attention mechanism).
[0070] Therefore, the subsequent output layer may be understood to actually predict the molecular property information of the specified proteolytic chimeric molecule based on the final equivariant and invariant characteristics of the specified proteolytic chimeric molecule, as well as the properties of several notable atoms in the molecular structure of the specified proteolytic chimeric molecule and several fixed properties in the molecular structure of the specified proteolytic chimeric molecule.
[0071] S104: The prediction model is trained based on the difference between the predicted molecular property information and the actual molecular property information corresponding to the designated protein degradation-inducing chimeric molecule.
[0072] After predicting the molecular characteristics described above, the difference between the molecular characteristics and the actual molecular characteristics corresponding to the specified protein degradation-inducing chimeric molecule may be determined, and the prediction model may be trained with the optimization goal of minimizing this difference.
[0073] The molecular property information referred to in this invention may include at least one of the following: lipophilicity of the specified proteolytic chimeric molecule, pH of the specified proteolytic chimeric molecule, molecular weight of the specified proteolytic chimeric molecule, hydrogen bond donors and hydrogen bond acceptors in the specified proteolytic chimeric molecule, solubility of the specified proteolytic chimeric molecule, and permeability of the specified proteolytic chimeric molecule.
[0074] As can be seen from the above method, the model training method provided by the present invention refers to the structural features of the molecular structure of the specified proteolytic chimeric molecule from multiple perspectives in the process of determining the three-dimensional molecular graph data of the specified proteolytic chimeric molecule. Furthermore, when predicting the molecular property information of the specified proteolytic chimeric molecule, it is determined based on the final invariant and equivariant features of the specified proteolytic chimeric molecule. Therefore, this method can adequately represent the characteristics of the molecular structure of the specified proteolytic chimeric molecule, and ensures that the prediction model subsequently makes accurate and reasonable predictions of molecular property information.
[0075] After training the above prediction model, the trained model can be used to predict the molecular properties of a specified molecule in practical applications. The specific process is shown in the following diagram.
[0076] Figure 2 is a schematic diagram illustrating the process of the molecular property information prediction method according to the present invention.
[0077] S201: Obtain data on target protein degradation-inducing chimeric molecules.
[0078] S202: Based on the above data, a three-dimensional molecular graph of the target protein degradation-inducing chimeric molecule is constructed.
[0079] In this invention, the terminal device may receive a molecular property prediction command for a target proteolytic chimeric molecule input by the user, acquire data for the target proteolytic chimeric molecule based on the molecular property prediction command, and construct three-dimensional molecular graph data for the target proteolytic chimeric molecule. Here, the molecular property prediction command for the target proteolytic chimeric molecule may include two fragment molecules provided by the user. The determination of the three-dimensional molecular graph data here is essentially the same as the process for determining the three-dimensional molecular graph data in the model training described above, so a detailed explanation is omitted here. The terminal device mentioned herein may be a device such as a desktop computer or a laptop computer.
[0080] S203: The three-dimensional molecular graph data of the target protein degradation-inducing chimeric molecule is input into a pre-trained prediction model, causing the prediction model to predict the molecular characteristics information of the target protein degradation-inducing chimeric molecule. The prediction model is trained by the model training method described above.
[0081] The terminal device may input three-dimensional molecular graph data of the target protein degradation-inducing chimeric molecule into a prediction model located on the terminal device, and the prediction model outputs molecular property information of the predicted target protein degradation-inducing chimeric molecule.
[0082] After obtaining molecular property information using a predictive model, the information may be displayed to the user, or recommendations may be made to the user based on the molecular property information, such as recommending pharmaceutical compounds that can bind well to the target protein degradation-inducing chimeric molecule. Of course, in actual applications, the three-dimensional molecular graph data of the target protein degradation-inducing chimeric molecule and the predicted molecular property information may be stored in association.
[0083] The present invention further provides a molecular property prediction system, as shown in Figure 3.
[0084] Figure 3 is a schematic diagram showing the architecture of the molecular property prediction system according to the present invention.
[0085] As can be seen from Figure 3, the system mainly consists of the following components. The memory subsystem stores the above dataset and is used to store molecular property information and its pharmaceutical and chemical properties predicted by the prediction model in actual applications. The control subsystem is used to predict the molecular property information of the target proteolytic chimeric molecule based on the three-dimensional molecular graph data of the target proteolytic chimeric molecule input into the subsystem. The control subsystem includes two units: a molecular feature extraction unit and a molecular property prediction unit. These two units are used to determine the atomic feature information of each atom from the three-dimensional molecular graph data and predict the molecular property information of the target proteolytic chimeric molecule.
[0086] The above describes one or more embodiments of the present invention, and based on the same idea, the present invention further provides corresponding model training devices and molecular property information prediction devices, as shown in Figures 4 and 5.
[0087] Figure 4 is a schematic diagram showing a model training apparatus according to the present invention, and includes an acquisition module 401 for acquiring data of a designated proteolytic chimeric molecule; a construction module 402 for constructing three-dimensional molecular graph data of the designated proteolytic chimeric molecule based on the data; a prediction module 403 for inputting the three-dimensional molecular graph data of the designated proteolytic chimeric molecule into a prediction model to be trained, causing the prediction model to predict molecular characteristic information of the designated proteolytic chimeric molecule; and a training module 404 for training the prediction model based on the difference between the predicted molecular characteristic information and the actual molecular characteristic information corresponding to the designated proteolytic chimeric molecule.
[0088] Optionally, the molecular property information includes at least one of the following: the lipophilicity of the designated proteolytic chimeric molecule, the pH of the designated proteolytic chimeric molecule, the molecular weight of the designated proteolytic chimeric molecule, the hydrogen bond donors and acceptors in the designated proteolytic chimeric molecule, the solubility of the designated proteolytic chimeric molecule, and the permeability of the designated proteolytic chimeric molecule.
[0089] Optionally, the construction module 402 is used to represent the edges between each atomic node by constructing an adjacency matrix corresponding to each atom, with each atom in the designated proteolytic chimeric molecule as an atomic node, in order to obtain the three-dimensional molecular graph data of the designated proteolytic chimeric molecule, and the adjacency matrix is used to represent the bonding information between each connected atom.
[0090] Optionally, each atom node in the three-dimensional molecular graph data corresponds to node information, and the node information includes at least one of the following: atom type information, atom three-dimensional coordinate information, and nuclear charge number.
[0091] Optionally, for each atom included in the three-dimensional molecular graph data, the atom type information is used to indicate which atom the atom specifically belongs to and is encoded using a one-hot encoding method; and the atom three-dimensional coordinate information is determined by projecting the atom onto a predetermined three-dimensional coordinate system, and the three-dimensional coordinate system includes a three-dimensional Cartesian coordinate system constructed with the molecular mass center of the designated proteolytic chimeric molecule as the coordinate origin.
[0092] Optionally, the prediction module 403 is used to input the three-dimensional molecular graph data of the designated proteolytic chimeric molecule into the embedding layer of the prediction model, obtain an embedding vector corresponding to the three-dimensional molecular graph data using the embedding layer, input the embedding vector into the information update layer of the prediction model to determine the target invariant and target equivariant features of the designated proteolytic chimeric molecule, input the final invariant and final equivariant features into the output layer of the prediction model, and output the molecular characteristic information of the designated proteolytic chimeric molecule predicted by the output layer.
[0093] Optionally, the prediction module 403 specifically determines atomic feature information corresponding to each atom in the designated proteolytic chimeric molecule based on the three-dimensional molecular graph data, wherein the atomic feature information includes at least one of the atom type information, spacing information between the atom and other atoms, and neighboring atom information, and is used to input the atomic feature information corresponding to each atom into the embedding layer of the prediction model to obtain the embedding vector corresponding to the three-dimensional molecular graph data using the embedding layer.
[0094] Optionally, the embedding vector includes initial invariant features, and the prediction module 403 is used to determine attention weights corresponding to the initial invariant features by the information update layer, to determine attention features corresponding to the designated proteolytic chimeric molecule based on the attention weights, and to determine the target invariant and target equivariant features of the designated proteolytic chimeric molecule by the information update layer based on the attention features.
[0095] Optionally, the prediction module 403 is used to determine the target invariant feature based on the updated initial equivalence feature, the updated initial invariant feature, and the initial equivalence feature, which is a zero vector. Specifically, the information update layer obtains an attention feature corresponding to the initial equivalence feature and an attention feature corresponding to the initial invariant feature, the initial equivalence feature is a zero vector, the attention feature corresponding to the initial invariant feature is updated based on the attention feature corresponding to the initial invariant feature, and the updated initial equivalence feature is obtained. The information update layer then determines the target invariant feature based on the updated initial invariant feature and the attention feature corresponding to the initial invariant feature, and the target equivalence feature is determined based on the updated initial equivalence feature, the updated initial invariant feature, and the attention feature corresponding to the initial invariant feature.
[0096] Optionally, the prediction module 403 is used to set the target equivalent features and target invariant features as the final equivalent features and final invariant features, respectively, in response to the number of information update layers being 1, and to set the target equivalent features and target invariant features determined by the last information update layer of the prediction model as the final equivalent features and final invariant features, respectively, in response to the number of information update layers being greater than 1.
[0097] Figure 5 is a schematic diagram showing a molecular property information prediction device according to the present invention, and includes an acquisition module 501 for acquiring data of a target proteolytic chimeric molecule, a construction module 502 for constructing three-dimensional molecular graph data of the target proteolytic chimeric molecule based on the data, and a prediction module 503 for inputting the three-dimensional molecular graph data of the target proteolytic chimeric molecule into a pre-trained prediction model and causing the prediction model to predict molecular property information of the target proteolytic chimeric molecule, the prediction model being trained by the model training method described above.
[0098] The present invention further provides a computer-readable storage medium in which a computer program is stored, the computer program may be used to perform a model training method provided by Figure 1 or a molecular property information prediction method provided by Figure 2. The computer-readable storage medium may be a non-volatile storage medium.
[0099] The present invention further provides a schematic diagram showing the structure of an electronic device corresponding to Figure 1 or Figure 2, as shown in Figure 6. As shown in Figure 6, the electronic device includes a processor, internal memory, and non-volatile memory, and may of course include other hardware necessary for operation, such as a network interface and an internal bus. The processor loads a corresponding computer program from the non-volatile memory into the internal memory and executes it, performing the model training method described in Figure 1 or the molecular property information prediction method described in Figure 2.
[0100] Of course, in addition to implementation by software, the present invention does not exclude other implementation methods, such as logical devices or combinations of hardware and software. In other words, the entity executing the following processing steps is not limited to each logical unit, but may also be hardware or a logical device.
[0101] In the 1990s, improvements in a technology could be clearly distinguished into hardware improvements (such as improvements to circuit structures like diodes, transistors, and switches) and software improvements (improvements to method flow). However, with technological advancements, many current method flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that method flow improvements cannot be realized by hardware physical modules. For example, programmable logic devices (PLDs) (e.g., field programmable gate arrays (FPGAs)) are such integrated circuits, and their logic functions are determined by programming by the device user. Instead of chip manufacturers designing and manufacturing dedicated integrated circuit chips, designers program and "integrate" digital systems onto a single PLD.Currently, instead of manually manufacturing integrated circuit chips, this programming is almost always done using software called a "logic compiler." This is similar to the software compiler used when writing programs, and in order to compile the original code, it needs to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many types such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. It will be obvious to those skilled in the art that by simply programming the method flow logically in one of the above hardware description languages and programming it into an integrated circuit, a hardware circuit that realizes that logical method flow can be easily obtained.
[0102] The controller may be implemented in any suitable way, for example, a controller may take the form of a microprocessor or processor, a computer-readable storage medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, and logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, microcontrollers such as the ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and a memory controller may further be implemented as part of the memory control logic. It will also be apparent to those skilled in the art that, in addition to implementing the controller with pure computer-readable program code, it is entirely possible to make the controller perform the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming method steps. Thus, such a controller may be considered a hardware component, and the devices included therein for implementing various functions may also be considered structures within the hardware component. Alternatively, the devices for realizing various functions may be considered to be software modules that implement the methods, or structures within hardware components.
[0103] The systems, devices, modules, or units described in the above embodiments may be specifically implemented by computer chips, entities, or products having some function. A typical implementing device is a computer. Specifically, a computer may be, for example, a personal computer, a laptop computer, a mobile phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet, a wearable device, or any combination of these devices.
[0104] For the sake of clarity, the above-mentioned device will be described by dividing it into various units according to their function. Of course, when implementing the present invention, it is also possible to realize the functions of each unit using the same or multiple software and / or hardware.
[0105] As those skilled in the art will see, embodiments of the present invention may be provided as methods, systems, or computer program products. Accordingly, the present invention may be provided in the form of embodiments consisting only of hardware, embodiments consisting only of software, or embodiments combining software and hardware. Furthermore, the present invention may be provided in the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0106] The present invention will be described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a dedicated computer, an embedded processor, or another programmable data processing device to generate a machine, thereby generating a device for implementing one or more flows in a flowchart and / or one or more blocks in a block diagram, through instructions executed by the processor of the computer or other programmable data processing device.
[0107] These computer program instructions may be stored in computer-readable memory that can instruct a computer or other programmable data processing device to work in a particular way, resulting in a product that includes an instruction unit that implements a function specified in one or more flows of a flowchart and / or one or more blocks of a block diagram using the instructions stored in the computer-readable memory.
[0108] These computer program instructions may be loaded onto a computer or other programmable data processing device, so that a series of operational steps are executed on the computer or other programmable device, generating processing to be performed by the computer, thereby providing steps for implementing a function specified in one or more flows of a flowchart and / or one or more blocks of a block diagram, by instructions executed on the computer or other programmable device.
[0109] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0110] Memory can include various forms of computer-readable storage media, such as volatile memory, random-access memory (RAM), and / or non-volatile memory, for example, read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable storage medium.
[0111] Computer-readable storage media include non-volatile and volatile media, movable and immovable media, and information storage can be realized by any method or technique. Information may be computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, Phase Change RAM (PRAM), Static Random-Access Memory (SRAM), Dynamic Random Access Memory (DRAM), other types of Random Access Memory (RAM), Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Flash Memory or other memory technologies, Compact Disc Read-Only Memory (CD-ROM), Digital Versatile Disc (DVD) or other optical storage, Magnetic Cassette Tape, Magnetic Tape Disk Storage or other magnetic storage devices, or any other non-transmission media that may be used to store information accessible from a computing device. According to the definition herein, computer-readable storage media does not include temporary storage computer-readable storage media (transitory media), such as modulated data signals and carriers.
[0112] Furthermore, the terms “include,” “contain,” or any other variation thereof are intended to include non-exclusive inclusion, thereby including not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. Unless further restrictions are imposed, an element limited by the phrase “include one…” is not precluded from the existence of other identical elements in a process, method, article, or device containing such element.
[0113] As those skilled in the art will see, embodiments of the present invention may be provided as methods, systems, or computer program products. Accordingly, the present invention may be provided in the form of embodiments consisting only of hardware, embodiments consisting only of software, or embodiments combining software and hardware. Furthermore, the present invention may be provided in the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0114] The present invention can be described in the general context of computer executable instructions executed by a computer, such as program modules. Generally, a program module includes routines, programs, objects, components, data structures, etc., that perform a specific task or realize a specific abstract data type. The present invention can also be implemented in a distributed computing environment in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules may reside on local and remote computer storage media, including storage devices.
[0115] Each embodiment of the present invention is described in a gradual manner, and any identical or similar parts between embodiments may be referenced to one another. The emphasis of each embodiment is on the differences from the other embodiments. In particular, the system embodiment is basically similar to the method embodiment, so it is described briefly, and relevant parts may be referenced to the description of part of the method embodiment.
[0116] The above are merely examples of the present invention and are not intended to limit the invention. To those skilled in the art, the present invention is subject to various modifications and changes. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. Steps to obtain data on a specified protein degradation-inducing chimeric molecule, Based on the aforementioned data, the step of constructing three-dimensional molecular graph data of the specified protein degradation-inducing chimeric molecule, The steps include inputting the three-dimensional molecular graph data of the designated protein degradation-inducing chimeric molecule into a prediction model to be trained, and causing the prediction model to predict the molecular properties of the designated protein degradation-inducing chimeric molecule, The process includes the step of training the prediction model based on the difference between the predicted molecular property information and the actual molecular property information corresponding to the designated proteolytic chimeric molecule, The step of inputting the three-dimensional molecular graph data of the designated protein degradation-inducing chimeric molecule into a prediction model to be trained, and having the prediction model predict the molecular property information of the designated protein degradation-inducing chimeric molecule, is: The steps include inputting the three-dimensional molecular graph data of the specified protein degradation-inducing chimeric molecule into the embedding layer of the prediction model, and obtaining an embedding vector corresponding to the three-dimensional molecular graph data using the embedding layer, The steps include inputting the aforementioned embedding vector into the information update layer of the prediction model to determine the target invariant and target equivariant features of the specified protein degradation-inducing chimeric molecule, The step includes inputting the final invariant features and final equivariant features into the output layer of the prediction model and outputting the molecular property information of the specified protein degradation-inducing chimeric molecule predicted by the output layer, The step of inputting the three-dimensional molecular graph data of the specified protein degradation-inducing chimeric molecule into the embedding layer in the prediction model, and obtaining an embedding vector corresponding to the three-dimensional molecular graph data using the embedding layer, A step of determining atomic feature information corresponding to each atom in the specified protein degradation-inducing chimeric molecule based on the three-dimensional molecular graph data, wherein the atomic feature information includes at least one of the following: atomic type information of the atom, spacing information between the atom and other atoms, and information on neighboring atoms of the atom. The step includes inputting atomic feature information corresponding to each atom into the embedding layer in the prediction model, and obtaining the embedding vector corresponding to the three-dimensional molecular graph data using the embedding layer, and / or, The embedding vector includes initial invariant features, and the step of inputting the embedding vector into the information update layer in the prediction model to determine the target invariant and target equivariant features of the specified proteolytic chimeric molecule is: The information update layer determines the attention weights corresponding to the initial invariant features, and based on the attention weights, determines the attention features corresponding to the designated protein degradation-inducing chimeric molecule. The information update layer includes the step of determining the target invariant and target equivariant characteristics of the designated protein degradation-inducing chimeric molecule based on the attention features, and / or, In response to the number of information update layers being 1, the target identical feature and the target invariant feature are set to the final identical feature and the final invariant feature, respectively. In response to the number of information update layers being greater than 1, the target equivalent features and target invariant features determined by the last information update layer of the prediction model are set as the final equivalent features and the final invariant features, respectively. A model training method characterized by the following features.
2. The molecular property information includes at least one of the following: the lipophilicity of the designated proteolytic chimeric molecule, the pH of the designated proteolytic chimeric molecule, the molecular weight of the designated proteolytic chimeric molecule, the hydrogen bond donors and hydrogen bond acceptors in the designated proteolytic chimeric molecule, the solubility of the designated proteolytic chimeric molecule, and the permeability of the designated proteolytic chimeric molecule. The method according to claim 1, characterized by the features described above.
3. Based on the aforementioned data, the step of constructing three-dimensional molecular graph data of the specified protein degradation-inducing chimeric molecule is: The process includes the step of representing the edges between each atomic node by constructing an adjacency matrix corresponding to each atom, with each atom in the specified protein degradation-inducing chimeric molecule as an atomic node, thereby obtaining the three-dimensional molecular graph data of the specified protein degradation-inducing chimeric molecule. The aforementioned adjacency matrix is used to represent the bonding information between each connected atom. Each atomic node of an atom included in the three-dimensional molecular graph data corresponds to node information, and the node information includes at least one of the following: atomic type information, atomic three-dimensional coordinate information, and nuclear charge number. For each atom included in the aforementioned three-dimensional molecular graph data, The aforementioned atomic type information is intended to indicate which specific atom the atom belongs to, and is encoded using a one-hot encoding method. The three-dimensional coordinate information of the atom is determined by projecting the atom onto a predetermined three-dimensional coordinate system, and the three-dimensional coordinate system includes a three-dimensional Cartesian coordinate system constructed with the molecular mass center of the designated protein degradation-inducing chimeric molecule as the coordinate origin. The method according to claim 1, characterized by the features described above.
4. The step of determining the target invariant and target equivariant characteristics of the designated protein degradation-inducing chimeric molecule based on the attention features using the information update layer is as follows: The information update layer provides an attention feature corresponding to the initial equivalence feature and an attention feature corresponding to the initial invariance feature, wherein the initial equivalence feature is a zero vector. The steps include updating the attention feature corresponding to the initial invariant feature based on the attention feature corresponding to the initial equivalence feature, and obtaining the updated initial invariant feature, The steps include updating the attention feature corresponding to the initial equivalence feature based on the attention feature corresponding to the initial invariant feature, and obtaining the updated initial equivalence feature, The information update layer includes the steps of determining the target invariant features based on the updated initial invariant features and the attention features corresponding to the initial invariant features, and determining the target equivalence features based on the updated initial equivalence features, the updated initial invariant features, and the attention features corresponding to the initial invariant features. The method according to claim 1, characterized by the features described above.
5. Steps to obtain data on target protein degradation-inducing chimeric molecules, Based on the aforementioned data, the step of constructing three-dimensional molecular graph data of the target protein degradation-inducing chimeric molecule, The step includes inputting three-dimensional molecular graph data of the target protein degradation-inducing chimeric molecule into a pre-trained predictive model, thereby causing the predictive model to predict molecular characteristic information of the target protein degradation-inducing chimeric molecule. The prediction model is trained by the method described in any one of claims 1 to 4. A method for predicting molecular property information, characterized by the following features.
6. A data acquisition module for obtaining data on specified protein degradation-inducing chimeric molecules, Based on the aforementioned data, a construction module for constructing three-dimensional molecular graph data of the specified protein degradation-inducing chimeric molecule, A prediction module that inputs the three-dimensional molecular graph data of the specified protein degradation-inducing chimeric molecule into a prediction model to be trained, and causes the prediction model to predict the molecular property information of the specified protein degradation-inducing chimeric molecule, The system includes a training module for training the prediction model based on the difference between the predicted molecular property information and the actual molecular property information corresponding to the designated proteolytic chimeric molecule, Specifically, the prediction module is used to input the three-dimensional molecular graph data of the designated protein degradation-inducing chimeric molecule into the embedding layer of the prediction model, obtain an embedding vector corresponding to the three-dimensional molecular graph data using the embedding layer, input the embedding vector into the information update layer of the prediction model to determine the target invariant and target equivariant features of the designated protein degradation-inducing chimeric molecule, input the final invariant and final equivariant features into the output layer of the prediction model, and output the molecular characteristic information of the designated protein degradation-inducing chimeric molecule predicted by the output layer. Specifically, the prediction module determines atomic feature information corresponding to each atom in the designated protein degradation-inducing chimeric molecule based on the three-dimensional molecular graph data, and the atomic feature information includes at least one of the following: atomic type information of the atom, spacing information between the atom and other atoms, and information on neighboring atoms of the atom, and the atomic feature information corresponding to each atom is input into the embedding layer of the prediction model, and the embedding layer is used to obtain the embedding vector corresponding to the three-dimensional molecular graph data. and / or, The embedding vector includes initial invariant features, and the prediction module is used, specifically, by the information update layer to determine attention weights corresponding to the initial invariant features, by determining attention features corresponding to the designated proteolytic chimeric molecule based on the attention weights, and by the information update layer to determine the target invariant and target equivariant features of the designated proteolytic chimeric molecule based on the attention features. and / or, Specifically, the prediction module is used to set the target equivalent features and target invariant features as the final equivalent features and final invariant features, respectively, in response to the number of information update layers being 1, and to set the target equivalent features and target invariant features determined by the last information update layer of the prediction model as the final equivalent features and final invariant features, respectively, in response to the number of information update layers being greater than 1. A model training device characterized by the following features.
7. An acquisition module for obtaining data on target protein degradation-inducing chimeric molecules, Based on the aforementioned data, a construction module for constructing three-dimensional molecular graph data of the target protein degradation-inducing chimeric molecule, The system includes a prediction module that inputs three-dimensional molecular graph data of the target protein degradation-inducing chimeric molecule into a pre-trained prediction model, causing the prediction model to predict molecular property information of the target protein degradation-inducing chimeric molecule. The prediction model is trained by the method described in any one of claims 1 to 4. A device for predicting molecular property information, characterized by the following features.
8. A computer-readable storage medium in which a computer program is stored, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is performed. A computer-readable storage medium characterized by the following features.
9. A computer-readable storage medium in which a computer program is stored, wherein when the computer program is executed by a processor, the method according to claim 5 is performed. A computer-readable storage medium characterized by the following features.
10. An electronic device comprising memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1 to 4 is performed. An electronic device characterized by the following features.
11. An electronic device comprising memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to claim 5 is carried out. An electronic device characterized by the following features.
Citation Information
Patent Citations
Small sample learning method, system and equipment for graph structure enhancement and storage medium
CN113314188A
Protein degradation targeted chimera connector generation method based on deep reinforcement learning
CN114171125A
PROTAC target molecule generation method, computer system and storage medium
CN115050429A
Predicting Adverse Drug Reactions
JP2020530158A
Method and apparatus for determining drug molecule property, and storage medium
US20220415452A1