A molecular property prediction method based on local graph features and global relationships
By combining local topological feature networks and global attention networks, the global relationship between molecules is captured, and the problem of insufficient local and global information integration in traditional methods is solved, which significantly improves the prediction and generalization ability of molecular properties.
Patent Information
- Application Number
- CN202411873944.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Traditional molecular properties prediction methods are difficult to effectively integrate local structural characteristics with global relationships, resulting in poor prediction accuracy and robustness.
A hybrid neural network method based on local graph features and global relationships is proposed. By combining local topological feature networks and global attention networks, a global relationship encoder module is introduced to capture the global relationship between molecules and improve the prediction and generalization capabilities of the model.
It significantly improves the predictive ability and generalization ability of complex molecular properties, and overcomes the problem of insufficient local and global information integration in traditional methods.
Smart Images

Figure CN119339827B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics and relates to a molecular property prediction method based on local graph features and global relationships, which includes technologies such as deep learning, contrastive learning, and attention mechanism. Background Art
[0002] Molecular property prediction is a core task in modern chemistry, biology and materials science, aiming to infer the physical, chemical or biological properties of molecules based on their structure or related properties, such as solubility, toxicity, biological activity and reaction rate. The formation of molecular properties depends on the complex atomic bond interactions within the molecules. Therefore, accurately capturing and modeling the complex interactions within molecules is crucial to revealing the mechanisms behind molecular properties. This task is of great significance to accelerate the development of new drugs, promote the development of green chemical technology and design high-performance materials. Although traditional methods such as quantum chemical calculations and quantitative structure-activity relationship analysis are widely used, their application scope is limited due to the high experimental cost, long time consumption and difficulty in handling complex molecular structures. In addition, the high-dimensional sparsity of molecular data further increases the difficulty of prediction.
[0003] The development of deep learning technology, especially the introduction of graph neural networks, has brought new opportunities for the prediction of molecular properties. Graph neural networks represent molecules as graph structures and capture the interaction information between atoms (nodes) and chemical bonds (edges), significantly improving the ability to model the complex properties of molecules. However, most existing methods use supervised learning and rely heavily on labeled data, which is expensive to label and unevenly distributed. In this regard, contrastive learning provides an effective solution. By constructing positive and negative sample pairs, data features can be learned without explicit labeling, enabling the model to more comprehensively characterize the physical and chemical properties of molecules. Summary of the invention
[0004] Traditional molecular property prediction methods often find it difficult to effectively integrate local structural features with global relationships when processing molecules, resulting in poor prediction accuracy and robustness. In order to solve this problem, the present invention proposes a hybrid neural network method based on local graph features and global relationships, which extracts feature structures by combining two graph convolution branches, namely the local topological structure feature network and the global attention network, and introduces a global relationship encoder module to capture the global relationship between molecules, thereby improving the model's prediction and generalization capabilities for complex molecular properties, overcoming the problem of insufficient integration of local and global information in traditional methods.
[0005] A molecular property prediction method based on local graph features and global relationships includes four processes: data processing, feature extraction through neural networks, contrastive learning optimization, and feature fusion and prediction. The specific steps are as follows:
[0006] Step 1: Generate a molecular graph using the tool from the molecular SMILES (Simplified Molecular Input Line Entry System) sequence data;
[0007] Step 2: Based on the molecular graph generated in step 1, feature extraction is performed through a local topological structure feature network, and the global relationship between distant nodes in the molecular graph is further captured through a local relationship encoder;
[0008] Step 3: Extract features from the molecular graph generated in step 1 through a global attention network, and further capture the global relationships between distant nodes in the molecular graph through a global relationship encoder;
[0009] Step 4. Finally, the graph features output from steps 2 and 3 are dynamically weighted and fused with the molecular fingerprint features, and the attention mechanism is introduced to optimize the multimodal feature combination for molecular property prediction.
[0010] A molecular property prediction method based on local graph features and global relationships. The implementation process of step 1 is as follows:
[0011] First, the input molecular data is parsed through RDKit (an open source software library for chemical informatics) to generate a molecular graph, which consists of node features (representing atomic information), edge features (representing chemical bond properties), and the connection relationship of the graph (the topological structure between nodes); secondly, the edge features are embedded and mapped using a fully connected layer to convert them into high-dimensional feature representations, and the feature dimensions are adjusted to adapt to the input requirements of the graph neural network; finally, the combination of node and edge features provides high-quality input data for feature extraction of subsequent models.
[0012] A molecular property prediction method based on local graph features and global relationships. The implementation process of step 2 is as follows:
[0013] First, through the local topological structure feature network, combined with the node features output in step 1, the embedded edge features and the connection relationship of the graph, the local topological features of the molecular graph are extracted through a multi-layer graph convolutional network, and the nonlinear expression ability is enhanced through the activation function to fully capture the local structural information of the molecular graph; secondly, the relationship between global nodes is captured through the global encoder; finally, the node features are aggregated into a graph-level feature vector through global average pooling for subsequent molecular property prediction tasks.
[0014] A molecular property prediction method based on local graph features and global relationships, the implementation process of step 3 is as follows:
[0015] First, the data output from step 1 is passed through a local graph feature extraction network, which includes three layers of graph isomorphic convolution layers. It combines node features, edge features, and graph structure information to extract local topological features of the molecular graph. In each layer of graph isomorphic convolution, the interaction between edge features and node features helps capture the details of the local structure in the graph, and the ELU activation function is used to enhance the nonlinear ability of the model, so that the network can learn more complex graph structure features. Secondly, the model uses a global encoder module and a multi-head attention mechanism to further fuse node features, edge features, and graph structure, thereby effectively capturing the global dependencies between remote nodes in the molecular graph.
[0016] A molecular property prediction method based on local graph features and global relationships, the implementation process of step 4 is as follows:
[0017] First, the graph features output from step 2 and step 3 are dynamically weighted and fused with the molecular fingerprint features to fully utilize the complementary information of different feature sources. Secondly, the attention mechanism optimization is introduced to enable the weights of different features to be adaptively adjusted according to their importance during the fusion process; finally, the node-level features are aggregated into graph-level representations using a global pooling method to obtain a high-dimensional representation of the entire molecule, and the aggregated graph features are used to predict molecular properties. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flow chart of a molecular property prediction method based on local graph features and global relationships.
[0019] Figure 2 It is a network structure diagram with local topological characteristics.
[0020] Figure 3 It is the global attention network structure diagram. DETAILED DESCRIPTION
[0021] The present invention is described in detail below with reference to the accompanying drawings and examples.
[0022] The purpose of this invention is to propose a molecular property prediction method based on local graph features and global relationships. Figure 1 It is a flow chart of a hybrid neural network molecular property prediction method based on local graph features and global relationships, including four processes: data processing, feature extraction through neural networks, contrast learning optimization, and feature fusion and prediction. The specific steps of the process are as follows:
[0023] Step 1: Preprocess the molecular SMILE data, including the following: First, represent the chemical structure of the input molecule through SMILES, and use the Chem.MolFromSmiles() function to convert it into an RDKit molecule object. This object contains the atomic information, chemical bond information and other structural features of the molecule, which is suitable for subsequent graph structure processing. Traverse each atom in the molecule and use the atom_features() function to extract the features of each atom, including atom type, degree, number of hydrogen atoms, implicit valence, formal charge, chiral label and hybridization state, and convert the discrete features into numerical vectors through one-hot encoding. Extract each chemical bond feature in the molecule, use the get_edge_features() function to obtain the edge type, and process these edge features through one-hot encoding to obtain the numerical representation of each edge. Secondly, use the RDKit function to generate the Morgan fingerprint of the molecule, which is a 1024-dimensional bit vector representing the global structural features of the molecule. Finally, all extracted node features, edge features, and molecular fingerprints are standardized to ensure that the scales of different feature dimensions are consistent, and these features are integrated into the final input data for use by the graph neural network for subsequent model training.
[0024] Step 2: Extract features from the data in step 1 through a local topological structure feature network, and further capture the global relationship between distant nodes in the molecular graph through a local relationship encoder; Figure 2 The local topological structure feature network structure diagram includes the following contents: First, the input data includes the node features, edge features and topological structure of the molecular graph. After normalization preprocessing, the data can be adapted to the input requirements of the graph neural network; on this basis, the model maps the edge features to 93 dimensions, 512 dimensions and 1024 dimensions respectively through three fully connected layers, completes the alignment of edge features and node features, and provides consistent input for subsequent feature interactions. Secondly, the molecular graph data passes through three layers of graph isomorphism networks in turn to extract local graph features respectively, and the input feature dimension of each layer of the network matches the dimension of the mapped edge features; the first layer of the network maps the node features from 93 dimensions to 512 dimensions, the second layer of the network expands from 512 dimensions to 1024 dimensions, and the third layer of the network shrinks back to 512 dimensions, and uses edge features in each layer to enhance the interactive information between nodes, and uses the activation function ELU to improve the nonlinear expression ability; at the same time, the model introduces a multi-head attention mechanism through two layers of global relationship encoders to capture the global relationship between long-distance nodes in the molecular graph and contextually reorganize local features. Finally, the processed node features are aggregated into fixed-length 512-dimensional graph-level features through global average pooling, completing the feature fusion from local to global; these graph-level features are ultimately used for the task of predicting molecular properties, achieving accurate modeling and analysis of molecular properties.
[0025] Step 3: Combine the features extracted by the global attention network from the data in step 1 to extract features, and further capture the global relationship between distant nodes in the molecular graph through the global relationship encoder; Figure 3 This is the global attention network structure diagram. It includes the following: First, the input data includes the node features, edge features and topological structure of the molecular graph, and is normalized and preprocessed. On this basis, the model completes the alignment of edge features and node features through two fully connected layers to provide consistent input for subsequent feature interactions. Secondly, the molecular graph data is sequentially passed through two layers of graph attention networks (GAT) for local feature extraction. The first layer of GATConv maps node features from 93 dimensions to 512 dimensions, and uses 10 attention heads to capture multi-dimensional information, and then applies the ELU activation function and Dropout ( p = 0.2) to prevent overfitting. The second layer GATConv maps the node features from 512×10 back to 512 dimensions, and uses the ELU activation function and Dropout again. Secondly, the model introduces a multi-head attention mechanism through a global encoder to capture the global relationship between long-distance nodes in the graph and reorganize contextual features. The global encoder contains two layers, each with an output dimension of 512, and uses 8 attention heads to handle global dependencies. Finally, the globally processed node features are aggregated into a fixed-length 512-dimensional graph-level feature through global average pooling, completing feature fusion from local to global.
[0026] Step 4: Dynamically weighted fusion of graph features extracted by graph neural network and molecular fingerprint features, and introduce attention mechanism for molecular property prediction. First, through the local topological structure feature network and the global attention network, the model extracts local and global features of the graph and generates node-level or graph-level representations. Graph features usually contain node attributes and neighborhood information, which can effectively capture the structure of molecular graphs. The dimension of graph features is usually 512. Secondly, a weighted mechanism is used to combine the graph features extracted by the network with the molecular fingerprint features to ensure that the two features can complement each other during the fusion process and give play to their respective advantages. The weighting coefficient is optimized through the training process and can be adjusted dynamically. In the weighted fusion process, the attention mechanism is introduced to dynamically adjust the fusion ratio of graph features and molecular fingerprint features according to the contribution of each feature. By training and optimizing the attention weight, the model can focus on the most important features, thereby improving the accuracy of prediction. The dimension after feature fusion is 1536 dimensions. Finally, the fused feature vector is processed by three fully connected layers; the first fully connected layer maps the fused input features to 512 dimensions through a linear transformation and performs a nonlinear transformation through a ReLU activation function, followed by a batch normalization and dropout layer ( p= 0.2) to prevent overfitting; the second fully connected layer maps the output 512-dimensional features to 256 dimensions, and then uses the ReLU activation function for nonlinear transformation, batch normalization and Dropout operations; the third fully connected layer maps the output of the second fully connected layer to 128 dimensions, and uses the ReLU activation function for nonlinear mapping, and finally applies the Dropout layer for regularization, and finally outputs the molecular property prediction value through linear transformation. The optimizer used in training is Adam, with a weight parameter of 0.02, a learning rate of 0.0004, and a learning rate decay factor of 0.1.
[0027] The dataset used for pre-training of this method was downloaded from ZINC 15, including SMILES descriptors of 306,347 small molecules with biological activity. The datasets used for downstream training include BBBP, SIDER, ClinTox, Tox21, BACE and HIV. The average AUC value of the method proposed in this invention on the six datasets is 85.29%, which is 2.09% higher than that of the previous method.
[0028] The detailed description of the above examples is a further detailed description of the present invention, but it cannot be determined that the present invention is limited to the scope described in the above examples. Within the scope of the present invention, ordinary technicians in this field can also make several related simple deductions or substitutions for other examples, which are all considered to be within the protection scope of the present invention.
Claims
1. A molecular property prediction method based on local graph features and global relationships includes four steps, which are as follows: Step 1. First, use the Chem.MolFromSmiles() function to convert the SMILES string into a molecular object, call the atom_features() function to extract the atomic features, and digitize them through one-hot encoding; then, embed the edge features through the fully connected layer to obtain a high-dimensional feature representation, and adjust the dimension to adapt to the input requirements of the graph neural network; then, extract the features of each chemical bond, and obtain the edge type through the get_edge_features() function, perform one-hot encoding, and finally combine the node and edge features to provide high-quality graph structure input data; Step 2: Based on the molecular graph generated in step 1, a graph isomorphism network is used for feature extraction, and the relationship between long-distance nodes is captured through the attention mechanism; the graph isomorphism network is referred to as GIN; the data is extracted layer by layer through a three-layer GIN network, and the input feature dimension of each layer matches the edge feature dimension; the first layer maps the node features from 93 dimensions to 512 dimensions, the second layer expands to 1024 dimensions, and the third layer shrinks to 512 dimensions; each layer uses the ELU activation function and enhances node interactions through edge features; a multi-head attention mechanism is introduced to capture long-distance node relationships, and finally the node features are aggregated into 512-dimensional graph-level features through global average pooling; Step 3: Based on the molecular graph data generated in step 1, use the graph attention network to extract features and capture global relationships; the graph attention network is referred to as GAT; first, the data is normalized, and then the edge features and node features are aligned through two fully connected layers; the data passes through two layers of GAT network, the first layer maps the node features from 93 dimensions to 512 dimensions, using 10 attention heads, ELU activation function and Dropout to prevent overfitting; the second layer maps the node features from 512×10 back to 512 dimensions, and uses ELU and Dropout again; finally, the global node relationship is captured through the multi-head attention mechanism, and aggregated into graph-level features through the global pooling layer; Step 4: Use GIN and GAT to extract local and global features and generate node-level or graph-level representations with a dimension of 512. Then, the graph features extracted by the graph neural network are fused with the molecular fingerprint features through a weighted mechanism to ensure that the two features complement each other. The weighting coefficients are optimized during the training process and can be adjusted dynamically. The attention mechanism is introduced to dynamically adjust the fusion ratio of graph features and fingerprint features according to the contribution of each feature to the final prediction. Finally, the fused feature vector is input into three fully connected layers, and after linear transformation, ReLU activation and Dropout processing, the molecular property prediction value is finally output.
Citation Information
Patent Citations
Bioactive peptide function prediction method based on multi-view multi-modal representation learning
CN119108018A