A method for explaining molecular properties based on deep learning
By introducing graph Transformer and hierarchical edge selector, combining reinforcement learning and contrast learning, optimizing the molecular property interpretation model, the problem that existing models are difficult to identify key substructures is solved, and efficient and interpretable molecular property interpretation is achieved, which is suitable for chemical and drug development.
Patent Information
- Application Number
- CN202510467052.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-15
AI Technical Summary
Existing molecular properties prediction models such as GNN models are difficult to effectively identify functional groups, conjugated systems or active centers that affect molecular properties, resulting in limited interpretability and generalization capabilities of the model, and high computational cost, making it difficult to automate large-scale applications.
The self-attention mechanism and hierarchical edge selector based on graph Transformer are adopted, combined with reinforcement learning and contrast learning, reward function and loss function are designed, molecular property interpretation model is optimized, and the global interaction between nodes and edges is enhanced through self-attention mechanism, hierarchical edge selector optimizes the selection of key edges, and the design reward mechanism encourages the selection of high confidence edges, and a trained interpreter is obtained through iterative training.
It improves the accuracy and stability of molecular properties interpretation, improves the interpretability and automation level of the model, and can automatically identify key substructures, which are suitable for the fields of chemical and drug development.
Smart Images

Figure CN119993292B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of molecular property interpretation, and particularly relates to a method for interpreting molecular properties based on deep learning. Background Art
[0002] The physicochemical properties of molecules (such as solubility, reactivity, biological activity, etc.) are of great significance in the fields of drug discovery, materials science, bioinformatics, and fine chemical engineering. Accurately interpreting molecular properties not only helps to understand the structure-property relationship (SPR) of molecules, but also can accelerate the design of new materials, the screening of new drugs, and the prediction of chemical reactions. Traditionally, the research on molecular properties mainly relies on methods such as quantum chemical calculations (such as density functional theory DFT), molecular dynamics simulations, and quantitative structure-activity relationships (QSAR). Although these methods can provide high accuracy, they usually have high computational costs, especially when dealing with large-scale molecular libraries, consuming a large amount of computational resources. In addition, these methods have strong interpretability but rely on expert experience and are difficult to be applied automatically on a large scale. In recent years, machine learning (ML) and deep learning (DL) have made remarkable progress in the field of molecular property prediction. Researchers have tried to use methods such as graph neural networks (GNN) to directly learn features from molecular graphs, avoiding the complexity of manually constructing features. Although GNN models perform well in molecular property prediction tasks, GNN models are usually regarded as "black box" models and it is difficult to directly understand which molecular substructures play a key role in the prediction results; moreover, existing GNN models are difficult to effectively identify functional groups, conjugated systems, or active centers that affect molecular properties, resulting in limited interpretability and generalization ability of GNN models.
[0003] Currently, there are many methods for the interpretability of molecular property prediction models at home and abroad, and certain results have been achieved in practical applications. Among them, the Grad-CAM method proposed by Selvaraju et al. is a technique commonly used for molecular property interpretation, which identifies the key molecular substructures that determine molecular properties by calculating gradient weights. The GNNExplainer designed by Ying et al. uses an edge masking mechanism to determine the key substructures that affect molecular properties and their corresponding feature subsets. The PGExplainer proposed by Luo et al. introduces a parameterized explanation generator and uses an independent multi-layer perceptron (MLP) to learn the substructures in the molecular structure that contribute most significantly to the properties. The FlowX method proposed by Gui et al. introduces the concept of message flow and determines the substructures that contribute most to molecular property prediction by tracing the message passing process.
[0004] In the task of molecular property interpretation, reinforcement learning (RL) and contrastive learning (CL) are widely used to improve the interpretability and robustness of molecular property prediction, playing an important role in key substructure identification, feature learning, and optimization. Contrastive learning trains the representation of molecular substructures by constructing positive and negative sample pairs, making the substructure features more stable, more discriminative, and enhancing the generalization ability of the model. During the molecular property interpretation process, contrastive learning can help the model more effectively identify the key structures affecting molecular properties and reduce the influence of noise. Reinforcement learning, on the other hand, performs optimal substructure selection through an exploration and exploitation mechanism, using a reward mechanism to optimize the extraction strategy of key substructures, thereby further enhancing the interpretability, consistency, and stability of molecular property interpretation. Compared with traditional attention- or gradient-based methods, reinforcement learning can more effectively mine the core substructures with chemical or physical significance and improve the reliability of the model. Wang et al. proposed RCExplainer, which combines reinforcement learning and contrastive learning to construct an end-to-end interpretation framework. In this method, contrastive learning is used to learn more robust molecular substructure representations, while reinforcement learning is used to optimize the substructure selection process, thus significantly enhancing the stability, interpretability, and generalization ability of molecular property interpretation, providing an efficient and interpretable analysis tool for fields such as chemistry, drug discovery, and materials science. Summary of the Invention
[0005] In view of the above problems existing in the prior art, the present invention proposes a deep learning-based method for molecular property interpretation, which is reasonably designed, solves the deficiencies of the prior art, and has good effects.
[0006] A deep learning-based method for molecular property interpretation, comprising the following steps:
[0007] S1. Predict the molecular properties based on a pre-trained GNN model to obtain molecular node embeddings;
[0008] S2. Design an interpreter to update the molecular node embeddings based on a graph Transformer to obtain node features, use a hierarchical edge selector to determine the edges in the molecular substructure, and design a reward function to limit and guide the substructure generation process;
[0009] S3. Iteratively train the interpreter through the designed loss function until the loss function meets the set loss value, complete the training and learning, and obtain a trained interpreter;
[0010] S4. Embed the molecular node into the trained interpreter, and output the substructure that plays a key role in the molecular properties.
[0011] Further, in the above S1, the molecule is formally represented as an undirected graph , where is the set of nodes, corresponding to the atoms in the molecule, is the set of edges, corresponding to the chemical bonds in the molecule;
[0012] The initial feature representation of the node is , where represents the feature vector of the th node, is the number of nodes, is the dimension of each node feature; after being encoded by the GNN model, the initial feature generates the node embedding representation , is the dimension of the node embedding.
[0013] Further, in the above S2, in the self-attention mechanism of the graph Transformer, calculate the query matrix , key matrix and value matrix , and the expressions are:
[0014] ; (1)
[0015] where , , is the learnable weight matrix, is the dimension of the key matrix, query matrix and value matrix, , is the number of attention heads;
[0016] For each attention head, use the query and key to calculate the attention weight , and the expression is:
[0017] ; (2)
[0018] Use the attention weight to perform weighted summation on the value matrix to generate the output of the attention head, and the expression is:
[0019] ; (3)
[0020] In the multi-head attention mechanism, the outputs of all attention heads will be concatenated together, and the expression is:
[0021] ; (4)
[0022] Among them, is the output of the spliced multi-head attention, is the Concatenate operation: indicating that the outputs of all attention heads are spliced together, is the output of the th attention head,
[0023] Perform a linear transformation on and add a residual connection and an activation function to obtain the updated node feature , and the expression is:
[0024] ; (5)
[0025] Among them, is the non-linear activation function, is the trainable parameter matrix of the linear transformation.
[0026] Furthermore, in the above S2, the substructure used to explain the model's prediction of molecular properties is determined by maximizing the mutual information metric,
[0027] ; (6)
[0028] Among them, is the mutual information, used to measure the information sharing degree between the substructure and , is the set of subgraphs of the graph with edges, is the prediction result of the graph in the prediction property model GNN, = 0 indicates that the molecule has no activity, = 1 indicates that the molecule has activity;
[0029] According to the relevant knowledge of mutual information, the above formula becomes:
[0030] ; (7)
[0031] Among them, is the th edge in, represents the probability distribution of the edge appearing under the condition of the given predicted molecular property ;
[0032] Regarding the generation of substructures as a Markov chain process, the th edge of the substructure is obtained using the following formula:
[0033] ; (8)
[0034] where is the optimal choice for the th edge in the generated substructure, is any edge in the candidate edge set at the th step, is the conditional entropy, representing the uncertainty of choosing edge and given , represents the substructure after the th step, represents the set of candidate edges added to the current interpreted subgraph ;
[0035] Design a hierarchical edge selector to find the th edge of the substructure. At the th step, calculate the probability of each candidate edge and select the edge with the highest score. For edge , the first MLP layer edge representation generator performs the following processing:
[0036] ; (9)
[0037] where , respectively represent the features of the th node and the th node in , and , represents the edge representation of edge ;
[0038] The second MLP layer edge possible generator is used to generate the possibility of an edge, that is, the probability score of the edge. The expression is:
[0039] ; (10)
[0040] where is the concatenation operation, is the representation of the molecular graph obtained through the prediction model GNN, is the edge For the score, the edge with the highest score will be added to the sub-structure; the edges that have been selected into the sub-structure will not enter the candidate edge set again; repeat continuously until the termination condition is met until
[0041] Furthermore, in the above S2, the reward function is as follows:
[0042] ;(16)
[0043] where is a hyperparameter used to balance the two rewards;
[0044] is and the evaluation of the combined contribution of the sub-structure to the molecular graph by the GNN model, and the expression is:
[0045] ;(15)
[0046] where represents the model for predicting molecular properties;
[0047] is the evaluation of the contribution to the interpretation of molecular activity, and the expression is:
[0048] ;(14)
[0049] where is the th node, is the th node, is the th sub-structure in the set it belongs to, is the set of sub-structures with the same prediction label as , is the set of sub-structures with a label different from .
[0050] Furthermore, in the above S3, the designed loss function is as follows:
[0051] ;(17)
[0052] where is a hyperparameter used to ensure that the input of the logarithmic function is always greater than 0;
[0053] During the training process, first, the loss function is optimized using stochastic gradient descent or the Adam optimizer, enabling the interpreter to learn the optimal substructure selection strategy. The training data consists of molecular graphs and their corresponding property labels. In each iteration, the model selects the most representative substructures based on the current molecular structure and calculates the loss function to update the parameters. To improve the generalization ability and interpretability of the model, hyperparameters are fine-tuned, and grid search or Bayesian optimization is used to find the optimal parameters. To prevent overfitting of the model, Dropout, weight decay, or data augmentation techniques are introduced to enhance the stability and interpretability of the model.
[0054] The beneficial technical effects brought by the present invention:
[0055] The present invention introduces Graph Transformer into the interpretation of molecular properties, enhances the global interaction modeling of nodes and edges in the molecular graph through the self-attention mechanism, and improves the accuracy of molecular property interpretation. The hierarchical edge selector combines the multi-level structure of deep learning to optimize the selection of key edges, with technical uniqueness; it can enhance the interpretability of molecular graphs and is widely applied in fields such as chemistry and drug development. Through the substructure reward mechanism, it encourages the preferential selection of high-confidence edges, enhancing the stability and interpretability of the model; combining the self-attention mechanism and reinforcement learning optimizes the molecular graph interpretation process, reduces manual intervention, improves the automation and intelligence level, and has high implementation feasibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a flowchart of feature extraction based on Graph Transformer in the present invention.
[0057] Figure 2 It is a flowchart of the operation of the hierarchical edge selector in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0058] The following further describes the specific implementation manners of the present invention with specific embodiments:
[0059] A method for interpreting molecular properties based on deep learning, comprising the following steps:
[0060] S1. Predict the molecular properties based on a pre-trained GNN model to obtain molecular node embeddings;
[0061] In chemoinformatics and computer science, molecules are usually represented as graph structures, where atoms are nodes and chemical bonds are edges. This representation can intuitively depict the topological relationships and chemical characteristics of molecules. Compared with traditional vector representations, graph structures can not only flexibly adapt to molecules of different scales and complexities but also effectively capture key information such as ring structures, aromaticity, and branches in molecules, providing a basis for molecular property prediction and interpretation.
[0062] Formalize the molecule as an undirected graph , where is the set of nodes, corresponding to the atoms in the molecule, is the set of edges, corresponding to the chemical bonds in the molecule;
[0063] The initial feature of the node is represented as , where represents the feature vector of the th node, is the number of nodes, is the dimension of each node feature; after the initial feature is encoded by the GNN model, the node embedding representation is generated, is the dimension of the node embedding. Although the node embedding can effectively capture local structure and feature information in the neural network, the dependence on the local neighborhood during the update process may lead to the loss of global information. At the same time, as the number of layers of the model network increases, the node embedding may face the problem of over-smoothing of information, making the features of different nodes become too similar and reducing the discrimination ability of the model.
[0064] S2. Design an interpreter to update the node embedding of the molecule based on the graph Transformer to obtain node features, use a hierarchical edge selector to determine the edges in the molecular substructure, and design a reward function to limit and guide the generation process of the substructure;
[0065] To overcome the above limitations, the graph Transformer is introduced when processing molecular node features. As Figure 1 shown, the self-attention mechanism of the graph Transformer allows each node to pay attention to all nodes in the graph when updating features, and the structure of the Transformer also allows the use of deeper networks without worrying about the problem of over-smoothing of information. Given the embedding representation of a node, calculate the query matrix , the key matrix and the value matrix in the self-attention mechanism. The expressions are as follows:
[0066] ; (1)
[0067] where , , is a learnable weight matrix, is the dimension of the key matrix, query matrix and value matrix, , is the number of attention heads;
[0068] For each attention head, use the query and the key to calculate the attention weights , with the expression:
[0069] ;(2)
[0070] Utilize the attention weights to perform a weighted sum on the value matrix to generate the output of the attention head, with the expression:
[0071] ;(3)
[0072] In the multi-head attention mechanism, the outputs of all attention heads are concatenated together, with the expression:
[0073] ;(4)
[0074] Among them, is the concatenated multi-head attention output, is the Concatenate operation: which means concatenating the outputs of all attention heads together, is the output of the th attention head, is the number of attention heads;
[0075] Perform a linear transformation on and add a residual connection and an activation function to obtain the updated node feature , with the expression:
[0076] ;(5)
[0077] Among them, is the non-linear activation function, is the trainable parameter matrix of the linear transformation.
[0078] The node feature That is, the GNN network to be explained is utilized to capture the local structure in the graph, aggregate the information of neighboring nodes, enabling the interpreter to better understand the relationships and patterns among local nodes. Meanwhile, the graph Transformer is also used to focus on the global node information, effectively capturing long-range dependencies and important global context information through the self-attention mechanism. This combination not only enhances the expressive ability of node features but also facilitates the interpretation of the molecular property prediction model, making the identification of key substructures more accurate and efficient. At the same time, this method helps to reasonably determine the order of candidate edges in the substructures that affect molecular properties, thereby more precisely extracting the structural information that has an important impact on molecular properties. By comprehensively considering global and local features, it can provide a more abundant and comprehensive molecular representation, thus improving the accuracy, stability, and interpretability of molecular property interpretation.
[0079] Usually, the substructures for explaining the model's prediction of molecular properties are determined by maximizing the mutual information metric , and the expression is:
[0080] ; (6)
[0081] where is the mutual information, used to measure the degree of information sharing between the substructure and , is the set of subgraphs of the graph with edges, is the prediction result of the graph in the prediction property model GNN, = 0 indicates that the molecule has no activity, = 1 indicates that the molecule has activity;
[0082] According to the relevant knowledge of mutual information, the above formula becomes:
[0083] ; (7)
[0084] where is the -th edge in , represents the probability distribution of the appearance of the edge given the predicted molecular property ;
[0085] It can be foreseen that as the number of edges included in the explanatory subgraph increases, the possible candidate subgraphs The quantity will increase in a super-exponential relationship, meaning that it is difficult to directly optimize this formula.
[0086] To solve this problem, the generation of substructures is regarded as a Markov chain process, and the following formula is used to obtain the th edge of the substructure:
[0087] ; (8)
[0088] where is the optimal choice for the th edge in the generated substructure, is any edge in the set of candidate edges at the th step, and is the conditional entropy, representing the uncertainty of choosing edge given and ; represents the substructure after the th step,
[0089] Design a hierarchical edge selector to find the th edge of the substructure, calculate the probability of each candidate edge at the th step, and select the edge with the highest score. For edge , the first MLP layer edge represents the generator and processes it as follows:
[0090] ; (9)
[0091] where , respectively represent the features of the th node and the th node in , and , represents the edge representation of edge ;
[0092] The second MLP layer edge possible generator is used to generate the possibility of the edge, that is, the probability score of the edge, and the expression is:
[0093] ; (10)
[0094] where is the concatenation operation, is the representation of the molecular graph obtained through the prediction model GNN, is the edge Score;
[0095] For a molecular graph with 5 nodes and 6 edges, the hierarchical edge selector process is as Figure 2 shown. By fusing the representation of the edge, the graph representation of the current substructure, and the graph representation of the entire molecular graph, the probability score of each candidate edge is calculated. The edge with the highest score will be added to the substructure; edges that have already been selected into the substructure will not enter the candidate edge set again; repeat until the termination condition is met.
[0096] To restrict and guide the generation process of the substructure, the following reward mechanism is designed.
[0097] Edge validity: For the edge selected by the substructure reward mechanism, it should be contributing to the prediction result of the molecular property prediction model with respect to the molecular graph . Regarding the molecular property prediction as a classification task, the classification category of the molecular graph is , where is the number of classification categories.
[0098] The molecular graph can be partitioned into subgraphs using a clustering method:
[0099] ; (11)
[0100] ; (12)
[0101] ; (13)
[0102] where is the set of substructures with the same prediction label as , and on the other hand, is the set of substructures with a label different from . Obviously, for a given molecular graph , it is expected that the selected edge belongs to . Therefore, the following reward is designed:
[0103] For evaluating the contribution to the interpretation of molecular activity, the expression is:
[0104] ; (14)
[0105] Among them, is the th node, is the th node, is the th sub-structure in the belonging set;
[0106] Sub-graph and edge The effectiveness of their cooperation: and sub-graph For the combined contribution evaluation of the GNN model prediction of the molecular graph is as follows: As follows:
[0107] ;(15)
[0108] Among them, represents the model for predicting molecular properties;
[0109] Ensure that the edge has a unique contribution to the explanation, Ensure that and the previously selected explanatory sub-graph also have a positive impact on the explanation. The final reward function is:
[0110] (16).
[0111] S3. Through the designed loss function, the interpreter is iteratively trained until the loss function meets the set loss value, completing the training and learning, and obtaining the trained interpreter;
[0112] In the molecular property explanation task, in order to improve the interpretability and stability of the model, it is usually necessary to design a special loss function to optimize the interpreter so that it can effectively identify the sub-structures that play a key role in the molecular properties. The designed loss function is:
[0113] ;(17)
[0114] Among them, is a hyperparameter used to ensure that the input of the logarithmic function is always greater than 0;
[0115] During the training process, the loss function is first optimized using Stochastic Gradient Descent or the Adam optimizer, enabling the interpreter to learn the optimal substructure selection strategy; the training data consists of molecular graphs and their corresponding property labels. In each iteration, the model selects the most representative substructures based on the current molecular structure and calculates the loss function to update the parameters; to improve the generalization ability and interpretability of the model, hyperparameters are fine-tuned, and grid search or Bayesian optimization is used to find the optimal parameters; to prevent model overfitting, Dropout, weight decay, or data augmentation techniques are introduced to enhance the stability and interpretability of the model. Finally, the optimized interpretation model can automatically identify the key substructures affecting molecular properties and provide stable and reliable molecular property interpretations, providing strong support for fields such as molecular design, drug discovery, and materials science.
[0116] Meanwhile, we split the molecular dataset, using 80% of it for training and 20% for testing. The model is trained through the above methods to achieve the interpretation of molecular properties.
[0117] S4. Embed the molecular nodes into the trained interpreter and output the substructures that play a key role in molecular properties.
[0118] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for explaining molecular properties based on deep learning, characterized in that, It includes the following steps: S1. Predict the molecular properties based on a pre-trained GNN model to obtain molecular node embeddings; S2. Design an interpreter, update the molecular node embeddings based on the graph Transformer to obtain node features, determine the edges in the molecular substructure using a hierarchical edge selector, and design a reward function to limit and guide the generation process of the substructure; S3. Iteratively train the interpreter through the designed loss function until the loss function meets the set loss value, complete the training and learning, and obtain the trained interpreter; S4. Input the molecular node embeddings into the trained interpreter and output the substructure that plays a key role in the molecular properties; In S2, the substructure for explaining the model's prediction of molecular properties is determined by maximizing the mutual information metric , and the expression is as follows: ;(6) Among them, is the mutual information, which is used to measure the and degree of information sharing between them, is the set of subgraphs that have edges of the graph ; is the prediction result of the graph in the prediction property model GNN, = 0 indicates that the molecule has no activity, = 1 indicates that the molecule has activity; According to the relevant knowledge of mutual information, the above formula becomes: ;(7) Among them, is the th edge in , indicating the probability distribution of the occurrence of edge under the condition of a given predicted molecular property ; Regarding the generation of substructures as a Markov chain process, the th edge of the substructure is obtained using the following formula: ;(8) Among them, is the optimal selection of the th edge in the generated substructure, is any edge in the th step candidate edge set, is the conditional entropy, indicating the uncertainty of selecting edge and under the condition of given , represents the substructure after the th step, represents the set of candidate edges added to the current interpreted subgraph ; Design a hierarchical edge selector to find the th edge of the substructure, calculate the probability of each candidate edge at the th step, and select the edge with the highest score. For the edge , the first MLP layer edge representation generator processes as follows: ;(9) Among them, , respectively represent the -th node and the -th node in features, and , represents the edge edge representation; Second MLP layer edge probability generator Used to generate the probability of an edge, i.e., the probability score of the edge, with the expression: ;(10) Among them, is a connection operation, is the representation of the molecular graph obtained through the prediction model GNN, is an edge score, and the edge with the highest score will be added to the sub-structure; the edges that have been selected into the sub-structure will not enter the candidate edge set again; repeat continuously until the termination condition is met.
2. The method for explaining molecular properties based on deep learning according to claim 1, characterized in that In S1, the molecule is formally represented as an undirected graph , where is a set of nodes corresponding to the atoms in the molecule, is a set of edges corresponding to the chemical bonds in the molecule; The initial feature of the node is represented as , where represents the feature vector of the -th node, is the number of nodes, is the dimension of each node feature; after the initial feature is encoded by the GNN model, the node embedding representation is generated, is the dimension of the node embedding.
3. The method for interpreting molecular properties based on deep learning according to claim 1, wherein In S2, calculate the query matrix , key matrix and value matrix in the self-attention mechanism of the graph Transformer. The expression is as follows: ;(1) Among them, , , are learnable weight matrices, is the dimension of the key matrix, query matrix, and value matrix, , is the number of attention heads; For each attention head, use the query and the key to calculate the attention weights , with the expression: ;(2) Using attention weights Perform weighted summation on the value matrix to generate the output of the attention head, and the expression is: ;(3) In the multi-head attention mechanism, the outputs of all attention heads will be concatenated together, and the expression is: ;(4) Among them, is the output of the multi-head attention after splicing, is the Concatenate operation: which means concatenating the outputs of all attention heads together, is the output of the th attention head, is the number of attention heads; Pair Perform a linear transformation and add a residual connection and an activation function to obtain the updated node features , and the expression is: ;(5) Among them, is a non-linear activation function, is a trainable parameter matrix for linear transformation.
4. The method for explaining molecular properties based on deep learning according to claim 3, characterized in that In S2, the reward function is as follows: ;(16) wherein is a hyperparameter for balancing two rewards; For and substructures to evaluate the combined contribution predicted by the GNN model for the molecular graph , the expression is: ;(15) Among them, represents a model for predicting molecular properties; For evaluating the contribution to the interpretation of molecular activity, the expression is: ;(14) Among them, is the th node, is the th node, is the th sub-structure in the set it belongs to, is a set of sub-structures having the same prediction label as , is a set of sub-structures having a label different from .
5. The method for interpreting molecular properties based on deep learning according to claim 4, wherein In S3, the designed loss function is as follows: ;(17) Among them, is a hyperparameter used to ensure that the input of the logarithmic function is always greater than 0; During the training process, first use stochastic gradient descent or the Adam optimizer to optimize the loss function so that the interpreter can learn the optimal substructure selection strategy; the training data consists of molecular graphs and their corresponding property labels. The model will select the most representative substructure according to the current molecular structure at each iteration and calculate the loss function to update the parameters; to improve the generalization ability and interpretability of the model, fine-tune the hyperparameters and use grid search or Bayesian optimization to find the optimal parameters; to prevent the model from overfitting, introduce Dropout, weight decay, or data augmentation techniques to enhance the stability and interpretability of the model.
Citation Information
Patent Citations
Molecular property prediction method based on molecular line graph structure
CN119418821A
Drug discovery via reinforcement learning with three-dimensional modeling
US20240105277A1