A method to enhance the accuracy of small molecule hydration free energy prediction

By combining graph comparative learning with graph neural networks, using the QM9 database and RDKit to standardize and parse molecular samples, and performing graph enhancement operations and feature encoding, we solved the problems of insufficient adaptability and accuracy in the prediction of small molecule hydration free energy in existing technologies, and achieved efficient end-to-end prediction and good model adaptability.

CN120412824BActive Publication Date: 2025-09-09UESTC (SHENZHEN) ADVANCED RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510918124.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-09
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Existing graph contrast learning methods have limited ability to understand and express molecular structure information when predicting the hydration free energy of small molecules, and their adaptability and accuracy in prediction tasks are poor.

Method used

Combining graph comparative learning with graph neural networks, pre-training is performed using the QM9 database, molecular samples are processed using Gaussian calculations and RDKit standardized parsing, graph enhancement operations and feature encoding are performed, positive and negative sample pairs are constructed, and pre-training and fine-tuning are performed through graph neural networks to improve the adaptability and accuracy of the model.

Benefits of technology

It realizes end-to-end graph representation learning and attribute prediction, improves the model's adaptability and prediction accuracy on small target datasets, has good versatility and transferability, and is suitable for other molecular property prediction tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412824B_ABST
    Figure CN120412824B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for enhancing the accuracy of prediction of small molecule hydration free energy, relates to the technical field of computational chemistry, and solves the technical problem of poor adaptability and accuracy of graph contrast learning prediction. The method comprises: based on the QM9 database, using Gaussian calculation to obtain a pre-training data set; performing graph enhancement operations on the pre-training data set, constructing positive and negative sample pairs, and obtaining a pre-trained graph neural network model; obtaining a target task data set based on the FreeSolv data set, and performing molecular graph construction and feature encoding on all data; using the pre-trained graph neural network model as an initialization model, using the training set for training, and fine-tuning the model through supervised retraining; using the test set to evaluate the model performance to obtain a prediction model. The present invention utilizes graph contrast learning to improve the robustness and generalization ability of graph structure representation, and can be extended to other molecular property prediction tasks. It has good versatility and transferability, and improves the adaptability, stability and accuracy of prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computational chemistry, and in particular to a method for enhancing the accuracy of prediction of small molecule hydration free energy. Background Art

[0002] The hydration free energy of small molecules refers to the change in Gibbs free energy when a single solute molecule transfers from a vacuum (or gas phase) into an aqueous solution. This physical quantity is an important thermodynamic parameter for evaluating the behavior of molecules in solvents and has wide applications in drug design, materials science, chemical reaction kinetics, and other fields. Traditionally, this parameter has been calculated primarily using quantum chemistry and molecular dynamics methods. Although these methods offer high accuracy, they are computationally expensive and difficult to apply to large-scale molecular screening tasks.

[0003] To this end, machine learning methods have been widely used in molecular property prediction in recent years. Current mainstream machine learning methods rely on manually constructed molecular fingerprints or molecular descriptors during the feature engineering phase to predict the hydration free energy of small molecules, requiring manual selection and extraction of molecular properties based on specialized knowledge. However, this manual label selection approach not only limits the model's ability to fully learn molecular structural information, but also easily introduces subjective biases and prior assumptions, thereby affecting the model's training and generalization performance. Because manual features often struggle to capture the complex structural relationships and nonlinear patterns between molecules, the model exhibits instability when faced with unknown data and may even overlook key structure-property relationships, leading to systematic errors in the prediction results.

[0004] To address this issue, recent efforts have seen the introduction of graph contrastive learning (GraphCL) into self-supervised learning mechanisms. This approach can fully leverage large-scale, unlabeled molecular data to learn universal, discriminative structural representations, effectively improving the model's generalization performance in downstream tasks. Compared to traditional methods that manually construct features, graph contrastive learning is more automated and adaptable, retaining more structural semantic information and demonstrating stronger representational and transfer capabilities.

[0005] However, existing graph contrastive learning methods have not yet been combined with graph neural networks, resulting in limited understanding and representation of molecular structural information, and poor adaptability and accuracy in prediction tasks. Therefore, there is an urgent need to build a prediction framework that integrates graph contrastive learning and graph neural networks to provide a more versatile, scalable, and accurate solution for modeling small molecule hydration free energy prediction. Summary of the Invention

[0006] The present invention aims to provide a method for enhancing the accuracy of small molecule hydration free energy predictions, addressing the technical issues in the prior art of graph contrast learning, which suffer from limited understanding and representation of molecular structure information, and poor adaptability and accuracy in prediction tasks. The various technical effects achieved by the preferred technical solutions provided by the present invention are detailed below.

[0007] To achieve the above objectives, the present invention provides the following technical solutions:

[0008] The present invention provides a method for enhancing the accuracy of prediction of small molecule hydration free energy, comprising the following steps: S100: based on the QM9 database, using Gaussian to calculate the hydration free energy of small molecules in the database as a pre-training data source, performing standardized parsing processing on the original molecular representations in the pre-training data source to obtain molecular samples, and converting each molecular sample into a molecular graph representation to obtain a pre-training data set; S200: performing different types of graph enhancement operations on the molecular graphs in the pre-training data set, constructing positive and negative sample pairs, and performing feature encoding, comparative learning and pre-training through a graph neural network to obtain a pre-trained graph neural network model ; S300: Obtain the target task dataset of the target task based on the FreeSolv dataset, randomly divide the target task dataset into training set, validation set and test set, and perform molecular graph construction and feature encoding on all data; S400: Use the pre-trained graph neural network model as the initialization model, use the training set in the target task dataset for training, and fine-tune the model through the validation set; S500: Adjust the model's hyperparameters and early stopping strategy on the validation set, use the test set in the target task dataset to evaluate the model performance, and obtain the small molecule hydration free energy prediction model by measuring the model's prediction effect on the small molecule hydration free energy.

[0009] Preferably, in step S100, the original molecular representation is subjected to standardized analysis processing by RDKit, and the standardized analysis processing includes at least one of the following operations: completing explicit hydrogen atoms, eliminating non-standard functional group representations, unifying aromaticity labels, and correcting atomic valences.

[0010] Preferably, in step S100, at least one of molecular filtering and screening, graph normalization, and feature normalization is performed on the pre-trained dataset; wherein, molecular filtering and screening are used to exclude molecules containing metal atoms, unidentified elements or abnormal graph structures, graph normalization is used to perform consistency processing on the topological structures of all molecular graphs, and feature normalization is used to normalize or standardize numerical atomic or bond attributes.

[0011] Preferably, in step S200, graph enhancement operations are performed through node masking, edge perturbation, and subgraph sampling; wherein, node masking randomly selects some atomic nodes and sets their feature vectors to zero or replaces them with predefined average features; edge perturbation randomly adds or deletes some chemical bond connections in the molecular graph to create slight structural changes; subgraph sampling randomly selects a substructure of the molecule, extracts local information of the corresponding molecular graph, and forms molecular graph representations from different semantic perspectives.

[0012] Preferably, in step S200, each pair of two enhancement results generated from the original image is used as a positive sample pair, and the enhancement results of other different images are used as negative sample pairs.

[0013] Preferably, in step S200, feature encoding is performed through a graph convolutional network GCN, a graph isomorphism network GIN, or a message passing neural network MPNN.

[0014] Preferably, in step S300, there are no overlapping cross samples between the training set, validation set, and test set, which account for 80%, 10%, and 10% of the data volume of the target task data set, respectively.

[0015] Preferably, in step S400, during the retraining process, all layer parameters of the model participate in the gradient update, and the root mean square error (RMSE) or the mean absolute error (MAE) is selected as the loss function.

[0016] Preferably, in step S400, the model is fine-tuned by learning rate adjustment, regularization and early stopping operations; wherein, the learning rate adjustment adopts a preheating and cosine annealing scheduling strategy; regularization uses random dropout regularization Dropout or L2 regularization in the fully connected layer; the early stopping operation monitors the loss change on the validation set, and terminates the training early when there is no significant improvement within several rounds.

[0017] Preferably, in step S500, the prediction effect of the model on the hydration free energy of small molecules is measured by any one of the mean absolute error MAE, the root mean square error RMSE, and the goodness of fit R².

[0018] Implementing one of the above technical solutions of the present invention has the following advantages or beneficial effects:

[0019] The present invention utilizes a graph contrast learning mechanism to enhance the robustness and generalization capability of graph structure representation, combines pre-training and fine-tuning processes to enhance the model's adaptability on small target datasets, uses existing databases to avoid manual feature construction, and implements end-to-end graph representation learning and attribute prediction. This approach can be extended to other molecular property prediction tasks, possesses good versatility and transferability, and improves the adaptability, stability, and accuracy of predictions. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work. In the drawings:

[0021] Figure 1 The present invention is a flowchart of a method for enhancing the accuracy of small molecule hydration free energy prediction. DETAILED DESCRIPTION

[0022] In order to make the objects, technical solutions and advantages of the present invention clearer, the various exemplary embodiments to be described below will refer to the corresponding drawings, which constitute a part of the exemplary embodiments, in which various exemplary embodiments that may be used to implement the present invention are described. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation methods described in the following exemplary embodiments do not represent all implementation methods consistent with the present disclosure. It should be understood that they are only examples of processes, methods and devices that are consistent with some aspects of the present disclosure as detailed in the appended claims, and other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present invention.

[0023] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse" and the like indicate the orientation or positional relationship based on the figures, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the elements referred to must have a specific orientation, be constructed and operated in a specific orientation. The terms "first", "second" and the like are only used for descriptive purposes and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. The term "multiple" means two or more. The terms "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, an integral connection, a mechanical connection, an electrical connection, a communication connection, a direct connection, an indirect connection through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.

[0024] In order to illustrate the technical solution of the present invention, a specific embodiment is provided below, in which only the parts related to the embodiment of the present invention are shown.

[0025] Example:

[0026] like Figure 1As shown, the present invention provides a method for enhancing the prediction accuracy of small molecule hydration free energy, comprising the following steps: S100: based on the QM9 database (the database contains a large number of organic molecule samples with different structures and functional groups, with sufficient data volume and diverse molecular structures, suitable as an unsupervised pre-training basis for graph neural networks, thereby facilitating the training of a highly accurate prediction model), using Gaussian calculations (quantum chemical calculations performed using Gaussian software, which can be achieved using existing technical processes) The hydration free energy of small molecules in the database is used as a pre-training data source, and the original molecular representation in the pre-training data source (such as a SMILES string, Simplified Molecular Input Line Entry System, SMILES is a simplified molecular linear input specification, which is a specification that clearly describes the molecular structure using an ASCII string, a simplified representation method for representing the molecular structure using a string, realizes the digital expression of molecular information, can be imported by most molecular editing software and converted into a two-dimensional graph or a three-dimensional model of the molecule, can concisely represent complex molecular structures, enable the computer to obtain more accurate and useful information, and is easy to process and store in the computer, thereby facilitating machine learning operations) is subjected to standardized parsing processing to obtain molecular samples, and each molecular sample is converted into a molecular graph representation to obtain a pre-training data set. This stage of pre-training not only provides a good foundation for structural perception capabilities for subsequent tasks, but also significantly improves the performance stability of the model in low-resource tasks. S200: Perform different types of graph enhancement operations on the molecular graphs in the pre-training dataset, construct positive and negative sample pairs, and perform feature encoding, contrastive learning, and pre-training through a graph neural network to obtain a pre-trained graph neural network model. Pre-training can fully utilize large-scale unlabeled data to learn universal molecular structure representations. In the contrastive learning stage, the enhanced samples are input into the graph neural network with shared parameters for embedded representation learning. Based on the pre-training mechanism of Graph Contrastive Learning (GraphCL), the discriminability and generalization capabilities of molecular graph representations are improved. This stage of pre-training not only provides a good foundation for structural perception capabilities for subsequent tasks, but also significantly improves the performance stability of the model in low-resource tasks. S300: Based on the FreeSolv dataset (which contains experimentally measured values ​​of several small molecules and their free energies of solubility in water and is widely used to evaluate the performance of molecular representation models in predicting physicochemical properties), the target task dataset is randomly partitioned into training, validation, and test sets. A molecular graph is constructed and feature encoded for all the data. For each molecular sample, the same graph structure construction and feature encoding process as in the pre-training phase is used. Specifically, each molecule is converted into a graph representation, with atoms as nodes and chemical bonds as edges.Node features include basic properties such as atom type, electronegativity, hybrid orbital type, whether it is in a ring, formal charge, etc.; edge features include bond type (single bond, double bond, triple bond, aromatic bond), whether the bond is a conjugated structure, whether it is in a ring, etc. The above-mentioned atomic and bond-level properties are encoded as multidimensional vectors as the initial features of the graph neural network input. S400: The pre-trained graph neural network model is used as the initialization model and trained using the training set in the target task dataset. The training task is a typical molecular regression operation, and the model is fine-tuned using the validation set. In this embodiment, the graph neural network in the fine-tuning stage encodes the molecular graph into a low-dimensional vector embedding, and then maps the embedding into a scalar output through one or more fully connected layers (which may include nonlinear activations such as ReLU). The operation in the fine-tuning stage can adapt the model to the specific small molecule hydration free energy prediction task. S500: Adjust the model's hyperparameters and early stopping strategy on the validation set, evaluate the model performance using the test set in the target task dataset, obtain a small molecule hydration free energy prediction model by measuring the model's prediction effect on the small molecule hydration free energy, and perform small molecule hydration free energy based on the prediction model. In this embodiment, a graph contrast learning mechanism is used to improve the robustness and generalization ability of the graph structure representation, and the pre-training and fine-tuning process is combined to improve the model's adaptability on small target datasets. The existing database avoids manual feature construction, realizes end-to-end graph representation learning and attribute prediction, and can be extended to other molecular property prediction tasks. It has good versatility and transferability, and improves the adaptability, stability and accuracy of the prediction.

[0027] As an optional embodiment, in step S100, when the molecular sample is converted to a molecular graph representation, ① Node (Atom): Each node in the graph corresponds to an atom. The node feature vector includes, but is not limited to, atom type (element symbol), atomic mass, electronegativity, hybridization state, aromaticity, charge, ring location, polarity, number of covalent bonds, etc. Some features are represented using one-hot encoding, while others are numerical. ② Edge (Chemical Bond): Edges in the graph represent chemical bonds between atoms. Edge features include bond type (single, double, triple, aromatic), conjugated bond, ring structure, and bond length (optional). ③ For bidirectional undirected graphs, bidirectional edge encoding can also be used to ensure symmetry in message passing. The graph structure is modeled as an undirected graph and input into the graph neural network using an adjacency matrix or sparse graph representation. To standardize input dimensionality and format, node and edge features in all graph samples are encoded as fixed-length vectors.

[0028] As an optional implementation, in step S100, the original molecular representation is subjected to standardized parsing processing through RDKit, which is an open source chemical informatics toolkit that provides rich chemical information processing functions and is widely used in drug discovery, materials science and chemical data analysis. It is used to process and analyze information such as molecular structure, chemical reactions, chemical properties, etc., and can convert SMILES string sequences into structurally standardized and unique molecular representations. Standardized parsing processing includes at least one of the following operations: completing explicit hydrogen atoms, eliminating non-standard functional group representations, unifying aromaticity labels, and correcting atomic valences, which are used to ensure the consistency and accuracy of the subsequent graph construction process. In step S100, the pre-training dataset is also subjected to at least one of the following operations: molecular filtering and screening, graph standardization, and feature normalization, thereby further enhancing the training efficiency and pre-training effect. Among them, molecular filtering and screening are used to exclude molecules containing metal atoms, unidentified elements, or abnormal graph structures to ensure the validity of the graph structure. Graph normalization is used to ensure consistency in the topological structure of all molecular graphs, such as renumbering atomic indices, normalizing bond order, and fixing graph connectivity requirements to avoid isolated nodes or broken links. Feature normalization is used to normalize or standardize numerical atomic or bond properties (such as mass and electronegativity) to adapt them to the initial weight range of the model and improve training convergence. After the above processing, the final pre-training dataset consists of molecular graph samples with consistent structure and clear semantics, which can be directly input into the graph neural network for unsupervised graph comparative learning training.

[0029] As an optional implementation, in step S200, graph augmentation operations are performed through node masking, edge perturbation, and subgraph sampling. Node masking simulates scenarios of information loss or perturbation by randomly selecting some atomic nodes and setting their feature vectors to zero or replacing them with predefined average features. Edge perturbation creates slight structural changes by randomly adding or removing some chemical bonds in the molecular graph, enhancing the model's adaptability to small variations in the molecular graph. Subgraph sampling randomly selects a substructure of the molecule (for example, a motif pattern, random walk, or BFS subgraph for identifying conserved fragment patterns with specific biological functions in DNA / RNA / protein sequences), extracts local information from the corresponding molecular graph, and forms molecular graph representations from different semantic perspectives. Using graph augmentation operations to construct the positive and negative sample pairs required for contrastive learning can improve the model's robustness and discriminative ability to molecular structural changes, while preserving key chemical semantic information.

[0030] As an optional implementation, in step S200, each pair of enhanced results generated from the original image is considered a positive sample pair, while the enhanced results of other different images are considered negative sample pairs. The model optimizes the structure of the embedding space by maximizing the representation similarity between positive sample pairs while minimizing the similarity between negative sample pairs. This training method effectively improves the model's ability to capture key features of molecular structure under unsupervised conditions.

[0031] As an optional implementation, in step S200, feature encoding is performed through a graph convolutional network GCN (which has strong expressive power, strong scalability, supports irregular grid data, and can simultaneously process node features and structural information), a graph isomorphism network GIN (which has strong representational capabilities, flexibility, and versatility), or a message passing neural network MPNN (which has the advantages of being applicable to complex graph structured data, efficiently processing sequence data, and having a wide range of application scenarios), thereby improving the adaptability of the model. Of course, other methods of feature encoding can also be selected as needed. At the same time, the pre-trained graph encoder can be directly used for downstream tasks such as molecular property prediction, drug screening, and molecule generation, and has extremely strong transferability and generalization capabilities. Compared with traditional supervised training methods, it does not rely on manually labeled data, significantly reduces data costs, and can still demonstrate excellent performance in small sample learning scenarios.

[0032] As an optional implementation, in step S300, there is no overlapping sample between the training set, validation set, and test set. That is, a molecule can only appear in one of the sets, accounting for 80%, 10%, and 10% of the target task dataset, respectively. This ratio is appropriate for facilitating model training and testing. Of course, other ratios can also be adjusted based on the actual data volume and training results. At the same time, to enhance the robustness of the partitioning results, the partitioning experiment can be repeated multiple times by setting a fixed random seed to evaluate the consistency of the model under different sample combinations.

[0033] As an optional implementation, in step S400, during the retraining process, all layer parameters of the model participate in the gradient update, allowing the model network to gradually adjust from a general representation to a specific representation that is optimal for the specific prediction task, and selecting either the root mean square error (RMSE) or the mean absolute error (MAE) as the loss function. The root mean square error (RMSE) penalizes larger deviations between the predicted value and the true value and is suitable for tasks sensitive to extreme errors; the mean absolute error (MAE) is more robust to outliers and can stably optimize the overall error distribution. In actual training, the two can be selected or a weighted combination can be made based on model stability and prediction accuracy.

[0034] As an optional implementation, in step S400, the model is also fine-tuned through learning rate adjustment, regularization and early stopping operations; among them, the learning rate adjustment adopts a warm-up and cosine annealing scheduling strategy, first slowly increasing the learning rate to help the model transition stably, and then gradually attenuating it to prevent overfitting; regularization uses random dropout regularization or L2 regularization in the fully connected layer to alleviate the risk of overfitting during fine-tuning; early stopping operation (Early Stopping) monitors the loss changes on the validation set, and terminates training early when there is no significant improvement within several rounds, so as to avoid resource waste and enhance generalization ability.

[0035] As an optional implementation, in step S500, the applicability of the model for predicting the hydration free energy of small molecules can be improved by using any of the following metrics: mean absolute error (MAE), root mean square error (RMSE), or goodness of fit (R²). These evaluation metrics provide a generalizable prediction framework applicable to other property prediction tasks, with good scalability and cross-task transferability, facilitating rapid deployment in diverse fields.

[0036] The embodiment is only a special example and does not represent only one way of implementing the present invention.

[0037] The foregoing is merely a preferred embodiment of the present invention. Those skilled in the art will appreciate that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. Furthermore, under the guidance of the present invention, these features and embodiments may be modified to suit specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are intended to be within the scope of the present invention.

Claims

1. A method for enhancing the accuracy of prediction of small molecule hydration free energy, characterized in that: The following steps are involved: S100: Based on the QM9 database, the hydration free energy of small molecules in the database is calculated using Gaussian as the pre-training data source. The original molecular representations in the pre-training data source are standardized and parsed to obtain molecular samples. Each molecular sample is then converted into a molecular graph representation to obtain a pre-training dataset. S200: Perform different types of graph augmentation operations on the molecular graphs in the pre-training dataset, construct positive and negative sample pairs, and perform feature encoding, contrastive learning, and pre-training through a graph neural network to obtain a pre-trained graph neural network model; S300: Obtain the target task dataset based on the FreeSolv dataset, randomly divide the target task dataset into training set, validation set and test set, and perform molecular graph construction and feature encoding on all data; S400: Use the pre-trained graph neural network model as the initialization model, train it using the training set in the target task dataset, and fine-tune the model using the validation set; S500: Adjust the model's hyperparameters and early stopping strategy on the validation set, evaluate the model performance using the test set in the target task dataset, and obtain a small molecule hydration free energy prediction model by measuring the model's prediction effect on small molecule hydration free energy; In step S100, the original molecular representation is subjected to standardized parsing processing by RDKit, wherein the standardized parsing processing includes at least one of the following operations: filling in explicit hydrogen atoms, eliminating non-standard functional group representations, unifying aromaticity labels, and correcting atomic valences; In step S200, graph enhancement operations are performed through node masking, edge perturbation, and subgraph sampling. Node masking randomly selects some atomic nodes and sets their feature vectors to zero or replaces them with predefined average features. Edge perturbation randomly adds or deletes some chemical bonds in the molecular graph to create slight structural changes. Subgraph sampling randomly selects a substructure of the molecule and extracts local information of the corresponding molecular graph to form molecular graph representations from different semantic perspectives. In step S400, during the retraining process, all layer parameters of the model participate in the gradient update, and the root mean square error (RMSE) or the mean absolute error (MAE) is selected as the loss function; In step S400, the model is fine-tuned through learning rate adjustment, regularization, and early stopping. The learning rate adjustment uses a warm-up and cosine annealing scheduling strategy. Regularization uses Dropout or L2 regularization in the fully connected layer. The early stopping operation monitors the loss changes on the validation set and terminates the training early when there is no significant improvement within several rounds.

2. A method for enhancing the accuracy of prediction of small molecule hydration free energy according to claim 1, characterized in that: In step S100, at least one of molecular filtering and screening, graph normalization, and feature normalization is performed on the pre-trained dataset; wherein, molecular filtering and screening are used to exclude molecules containing metal atoms, unidentified elements, or abnormal graph structures, graph normalization is used to perform consistency processing on the topological structures of all molecular graphs, and feature normalization is used to normalize or standardize numerical atomic or bond attributes.

3. The method for enhancing the accuracy of prediction of small molecule hydration free energy according to claim 1, characterized in that: In step S200, each pair of two enhanced results generated from the original image is used as a positive sample pair, and the enhanced results of other different images are used as negative sample pairs.

4. The method for enhancing the accuracy of prediction of small molecule hydration free energy according to claim 1, characterized in that: In step S200, feature encoding is performed through a graph convolutional network GCN, a graph isomorphism network GIN, or a message passing neural network MPNN.

5. The method for enhancing the accuracy of prediction of small molecule hydration free energy according to claim 1, characterized in that: In step S300, there are no overlapping cross samples between the training set, validation set, and test set, which account for 80%, 10%, and 10% of the data volume of the target task data set, respectively.

6. The method for enhancing the accuracy of prediction of small molecule hydration free energy according to claim 1, characterized in that: In step S500, the prediction effect of the model on the hydration free energy of small molecules is measured by any one of the mean absolute error (MAE), root mean square error (RMSE), and goodness of fit (R²).

Citation Information

Patent Citations

  • Absolute binding free energy calculation method without precise compound structure

    CN118039001A

  • A deep learning framework to predict on-target and off-target activity of crispr guide rnas

    WO2024259103A1