A method for predicting molecular properties by combining crystal morphology characteristics
Through deep learning methods combining crystal morphological characteristics and molecular structure information, the time-consuming problem in the existing technology is solved, efficient molecular properties prediction is achieved, and the speed and accuracy of material and drug screening design is improved.
Patent Information
- Application Number
- CN202111282971.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-01
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-11-01
AI Technical Summary
The prior art ignores crystal morphological characteristics information in molecular properties prediction, which makes the calculation time-consuming and inefficient, making it difficult to complete the screening and design of a large number of candidate compounds in a short time.
Combining crystal morphological feature information and molecular structure information, a molecular property prediction model is designed through deep learning methods, and the structural information extraction network is used to extract features, fused into fusion molecular features, and predicted through the property discrimination network.
Improves the accuracy and efficiency of molecular properties prediction, reduces time and calculation costs, and shortens the time for material and drug screening design.
Smart Images

Figure CN116092591B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically to a method for predicting molecular properties by combining crystal morphology features. Background Art
[0002] Molecular property prediction is one of the key tasks in the processes of computer-aided materials and drug discovery. By extracting the characteristic information of the molecules to be detected and making predictions on physical and chemical properties, people can find compounds that meet the expected properties among a large number of candidate compounds, accelerating the screening and design speed of materials and drugs, which has important research significance.
[0003] Traditional quantum chemical methods based on Density Functional Theory (DFT) can accurately predict various molecular properties, but they require huge time costs and computing power consumption, and it takes several hours to complete the calculation of a single molecule. In addition, the number of candidate compounds is often quite large, and it is difficult for people to complete molecular property prediction in a short time. Recently, many researchers have adopted the method of treating molecules as graph data and combined deep learning methods to complete the molecular property prediction task. Since the forward propagation of the neural network takes extremely short time, its computing time is far less than that of traditional quantum chemical calculation methods such as DFT. Currently, the molecular property prediction models designed based on deep learning methods only consider the structural information of molecules. Research shows that the crystal morphology information of molecules is also an important factor determining molecular properties, but this characteristic information is ignored by existing methods. Summary of the Invention
[0004] To solve the deficiencies of the existing technology, the present invention provides a method for predicting molecular properties by combining crystal morphology features while considering the molecular structure information and introducing the crystal morphology feature information ignored by current methods.
[0005] The technical solution adopted by the present invention to achieve the above object is: a method for predicting molecular properties by combining crystal morphology features, the method comprising the following steps:
[0006] A method for predicting molecular properties by combining crystal morphology features, the method comprising the following steps:
[0007] (1) Obtain training samples, where the training samples include the structural data, morphology data, and property annotation data of molecules;
[0008] (2) Use a structure information extraction network to extract features from the structural data to obtain a first molecular feature;
[0009] (3) Use a morphology information extraction network to extract features from the morphology data to obtain a second molecular feature;
[0010] (4) Combine the described first molecular feature and the described second molecular feature to obtain the fused molecular feature of the training sample;
[0011] (5) Use the property discrimination network to process the fused molecular feature to obtain the predicted property of the training sample;
[0012] (6) Adjust the network parameters of the molecular property prediction model according to the predicted property of the training sample and the property annotation data;
[0013] (7) Use the optimized model for testing to output the molecular property prediction result of the molecule to be predicted.
[0014] The structural data acquisition in the step (1) includes the following steps:
[0015] (1-1) Take the atoms in the molecule as nodes and the chemical bonds between atoms as edges to transform the molecule into graph structure data;
[0016] (1-2) Extract the nodes, edges, node attributes and edge attributes as structural data, where the node attributes include at least one of the following: atom type, atomicity, number of free electrons, and the edge attributes include at least one of the following: atom bond type, bond length, bond angle, torsion angle, whether it is a cyclic structure.
[0017] The morphological data in the step (1) is the SEM image obtained by using an electron microscope with a fixed resolution.
[0018] The property annotation data in the step (1) is determined according to the property to be predicted. When the property is quantitatively represented by a numerical value, the property annotation data is a numerical value; when the property is used to describe the presence or absence, the property annotation data is the 0, 1 encoding.
[0019] The training samples in the step (1) are divided into a training set, a validation set and a test set according to a preset ratio.
[0020] The structural information extraction network in the step (2) is a neural network for processing graph data, which maps the structural data into a feature vector in a low-dimensional space to capture the topological structure of the graph, the relationship between nodes, and the relevant information about the graph, subgraphs and vertices, and takes the obtained feature vector as the first molecular feature.
[0021] The morphological information extraction network in the step (3) is a neural network for processing image data, which maps the morphological data into a feature vector in a low-dimensional space to obtain the crystal size, aspect ratio, and coverage rate in the morphological data as the second molecular feature.
[0022] In step (4), the manner of combining the first molecular feature and the second molecular feature includes splicing, addition, and attention weighting. A feature vector with both molecular structure information and crystal morphology information is obtained through feature fusion and used as the fused molecular feature.
[0023] In step (5), the property discrimination network is determined according to the predicted property and includes a non-linear transformation layer and a linear transformation layer.
[0024] Step (6) includes the following steps:
[0025] (6-1) Calculate the loss function using the predicted property and the property annotation data, where the loss function includes cross-entropy loss and mean squared error loss;
[0026] (6-2) Update the model parameters using the gradient backpropagation algorithm according to the loss function value calculated in (6-1), where the model includes a structure information extraction network, a morphology information extraction network, and a property discrimination network;
[0027] (6-3) Repeat steps (6-1) to (6-2) until the validation set loss is lower than the preset value, and stop updating the model parameters to obtain the final model.
[0028] Compared with the prior art, the beneficial effects of the present invention are reflected in: fusing the structural feature information and morphological feature information of molecules, and combining neural network extraction and analysis, applying deep learning to molecular property prediction, thereby accelerating the screening and design speed of materials and drugs. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic framework diagram of a method for predicting molecular properties by combining crystal morphology features of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] To make the above objects, features, and advantages of the present invention more obvious and understandable, the specific implementation methods of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the invention. Therefore, the present invention is not limited by the specific implementations disclosed below.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention.
[0032] In chemical research, predicting molecular properties before screening materials and drugs can reduce the number of compounds actually screened and improve the efficiency of discovering lead compounds. Therefore, the present invention combines the structural feature information and morphological feature information of molecules and uses deep learning methods to predict molecular properties, improving the accuracy of molecular property prediction, significantly reducing time costs and computational losses, and thus accelerating the screening and design speed of materials and drugs.
[0033] As Figure 1 shown in the structural schematic diagram of a molecular property prediction method combining crystal morphological features, which specifically includes the following steps:
[0034] (1) Obtain training samples, where the training samples include the structural data, morphological data, and property annotation data of molecules. Among them, the acquisition method of structural data includes the following steps:
[0035] (1-1) Taking the atoms in the molecule as nodes and the chemical bonds between atoms as edges, convert the molecule into graph structure data;
[0036] (1-2) Extract the nodes, edges, node attributes, and edge attributes as structural data. Among them, the node attributes include at least one of the following: atom type, atomicity, number of free electrons, and the edge attributes include at least one of the following: atom bond type, bond length, bond angle, torsion angle, and whether it is a cyclic structure.
[0037] The acquisition method of morphological data is to obtain the SEM image of the molecule using an electron microscope with a fixed resolution; the property annotation data is the value of the property to be predicted, which is annotated by artificial experience. When the property can be quantified by a certain value, the property annotation data is the value. For example, for the adsorption property of a molecule, the annotation data is the adsorption amount of the molecule for the adsorbate under specific temperature and pressure conditions; when the property is only used to describe whether it exists, the property annotation data is 0, 1 encoding. For example, the water solubility of a molecule can be divided into poorly soluble, slightly soluble, soluble, and easily soluble, and the annotation data is one-hot encoded according to the solubility of the molecule.
[0038] At the same time, divide the training samples into a training set, a validation set, and a test set according to a preset ratio to prepare for the subsequent training and testing of the model. Among them, the preset ratio is not mandatory, such as 8:1:1.
[0039] (2) Use a structure information extraction network to extract features from the structure data. The structure information extraction network is a neural network for processing graph data, and the specific network structure is determined according to the actual task. A graph convolutional neural network or a graph attention neural network can be selected. For the structure data, that is, the graph structure type data converted from the molecular structure, further feature extraction is performed and aggregated into a feature vector in a low-dimensional space to capture the topological structure of the graph, the relationships between nodes, and other relevant information about the graph, subgraphs, and vertices. The obtained feature vector is used as the first molecular feature.
[0040] (3) Use a morphology information extraction network to extract features from the morphology data. The morphology information extraction network is a neural network for processing image data, and the specific network structure is determined according to the actual task. The target detection network RCNN or the image segmentation network U-Net can be selected to extract features from the morphology data, that is, the SEM images obtained by electron microscopy, and obtain features such as the size, aspect ratio, and coverage rate of the crystals in the morphology data and other relevant information about the crystal morphology, and encode them into a feature vector in a low-dimensional space as the second molecular feature.
[0041] (4) Fuse the first molecular feature and the second molecular feature. The fusion method can be to splice the two feature vectors in dimension; or to align the two feature vectors in dimension and perform a summation operation on the corresponding elements; or to learn an attention coefficient through a network and apply it to the first molecular feature and the second molecular feature, and apply different weight coefficients to different elements in the two feature vectors to complete attention-weighted fusion. The fused feature is used as the fused molecular feature of the training sample;
[0042] (5) Use a property discrimination network to process the fused molecular feature to obtain the predicted property of the training sample. The predicted property is consistent with the labeled data, such as adsorption characteristics, water solubility, etc. Among them, the property discrimination network is determined according to the predicted property. When the predicted property can be quantified by a certain value, such as the adsorption characteristics of a molecule, the task property is a regression task, that is, to predict the adsorption amount of the molecule for the adsorbate under specific temperature and pressure conditions; when the property is only used to describe whether it exists, such as the water solubility of a molecule, the task property is a classification task, that is, to determine whether the water solubility of the molecule to be predicted is poorly soluble, slightly soluble, soluble, or easily soluble. The property discrimination network can be a perceptron model composed of a fully connected layer with a non-linear transformation and a fully connected layer with a linear transformation.
[0043] (6) According to the predicted property of the training sample and the property labeled data, adjust the network parameters of the molecular property prediction model, which specifically includes the following steps:
[0044] (6-1) Calculate the loss function by using the predicted properties and the property-labeled data, where the loss function is determined according to the nature of the property prediction task. When the task nature is a classification task, cross-entropy loss can be used; when the task nature is a regression task, mean squared error loss can be used.
[0045] (6-2) Update the model parameters according to the loss function value calculated in (6-1) by using the gradient backpropagation algorithm. The model includes a structure information extraction network, a morphology information extraction network, and a property discrimination network. Additionally, if the first molecular feature and the second molecular feature are fused in an attention-weighted manner, the updated model parameters also include the hyperparameters in the attention module.
[0046] (6-3) Repeat steps (6-1) to (6-2) for model training and parameter update. During the training process, record the training set loss, validation set loss, and model weight file in the form of a log for each round until the training stops when the preset conditions are met. The preset conditions are set according to the actual situation, and only two examples are given: (1) When the validation set loss is lower than a specific value, stop updating the model parameters and use the last round of training results as the final model; (2) When the number of training rounds reaches a specific value, stop updating the model parameters and use the round with the lowest validation set loss during the training process as the final model.
[0047] (7) Use the final model obtained during the training process to test on the test set of the samples to evaluate the ability of the present invention to predict molecular properties. The evaluation indicators include but are not limited to root mean square error, mean absolute error, and prediction accuracy.
[0048] The above specific embodiments have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
[0049] The above specific embodiments have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for predicting molecular properties by combining crystal morphology features, characterized in that, The method includes the following steps: (1) Obtain training samples, where the training samples include structural data, morphological data, and property annotation data of molecules; in step (1), the morphological data is an SEM image obtained using an electron microscope with a fixed resolution; (2) Use a structure information extraction network to extract features from the structural data to obtain first molecular features; (3) Use a morphological information extraction network to extract features from the morphological data to obtain second molecular features; in step (3), the morphological information extraction network is a neural network for processing image data, which maps the morphological data into a feature vector in a low-dimensional space to obtain the crystal size, aspect ratio, and coverage rate in the morphological data as the second molecular features; (4) Combine the first molecular features and the second molecular features to obtain fused molecular features of the training samples; (5) Use a property discrimination network to process the fused molecular features to obtain the predicted properties of the training samples; (6) Adjust the network parameters of the molecular property prediction model according to the predicted properties of the training samples and the property annotation data; (7) Use the optimized model for testing to output the molecular property prediction results of the molecules to be predicted.
2. The molecular property prediction method combining crystal morphology features according to claim 1, wherein: In step (1), the acquisition of the structural data includes the following steps: (1-1) Take atoms in the molecule as nodes and chemical bonds between atoms as edges to transform the molecule into graph structure data; (1-2) Extract nodes, edges, node attributes, and edge attributes as structural data, where the node attributes include at least one of the following: atom type, atomicity, number of free electrons, and the edge attributes include at least one of the following: atom bond type, bond length, bond angle, torsion angle, and whether it is a ring structure.
3. The molecular property prediction method combining crystal morphology characteristics according to claim 1, characterized in that: In step (1), the property annotation data is determined according to the property to be predicted. When the property is quantitatively represented by a numerical value, the property annotation data is a numerical value; when the property is used to describe the presence or absence, the property annotation data is a 0, 1 encoding.
4. A method for predicting molecular properties by combining crystal morphology characteristics according to claim 1, characterized in that: In step (1), the training samples are divided into a training set, a validation set, and a test set according to a preset ratio.
5. A method for predicting molecular properties by combining crystal morphology features according to claim 1, characterized in that: In step (2), the structure information extraction network is a neural network for processing graph data, which maps the structural data into a feature vector in a low-dimensional space to capture the topological structure of the graph, the relationship between nodes, and the relevant information about the graph, subgraphs, and vertices, and takes the obtained feature vector as the first molecular feature.
6. A method for predicting molecular properties in combination with crystal morphology features according to claim 1, characterized in that: In step (4), the method of combining the first molecular features and the second molecular features includes splicing, addition, and attention weighting, and a feature vector with both molecular structure information and crystal morphology information is obtained through feature fusion as the fused molecular feature.
7. A method for predicting molecular properties in combination with crystal morphology characteristics according to claim 1, characterized in that: In step (5), the property discrimination network is determined according to the property to be predicted and includes a non-linear transformation layer and a linear transformation layer.
8. A method for predicting molecular properties in combination with crystal morphology characteristics according to claim 1, characterized in that: Step (6) includes the following steps: (6-1) Calculate a loss function using the predicted property and the property annotation data, where the loss function includes cross-entropy loss and mean squared error loss; (6-2) Update the model parameters using the gradient backpropagation algorithm based on the loss function value calculated in (6-1), where the model includes a structure information extraction network, a morphology information extraction network, and a property discrimination network; (6-3) Repeat steps (6-1) to (6-2) until the validation set loss is lower than the preset value, stop updating the model parameters, and use it as the final model.