Molecular solvation free energy prediction method based on machine learning

Through a machine learning-based method, the SMILES formula is converted into molecular descriptors and combined with experimental data to train the model, the problem of high computational resources and time consumption of molecular dynamics methods is solved, and the rapid and accurate prediction of molecular solvation free energy is achieved.

CN120260737AActive Publication Date: 2025-07-04UESTC (SHENZHEN) ADVANCED RES INST +1

Patent Information

Application Number
CN202510739407.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In the prior art, the calculation of molecular solvation free energy depends on molecular dynamics methods, resulting in large consumption of computing resources and time, and the calculation results are not accurate enough and the cost is high.

Method used

Using a machine learning-based method, the SMILES formula is used to convert one-dimensional, two-dimensional and three-dimensional descriptors into molecules, and the machine learning model is trained in combination with open source experimental data sets to predict the free energy of solvation of molecules.

Benefits of technology

Significantly shortens the calculation time, improves the accuracy and stability of prediction results, reduces costs, is suitable for high-throughput screening of large-scale molecular libraries, and simplifies calculation steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260737A_ABST
    Figure CN120260737A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting molecular solvation free energy based on machine learning, relates to the technical field of computational chemistry, and solves the technical problem that a molecular dynamics method consumes a large amount of computing resources and time. The method comprises the following steps: using SMILES formulas of solute molecules and solvent molecules as input data; converting the SMILES formula into a one-dimensional descriptor, a two-dimensional descriptor and a three-dimensional descriptor of a molecule based on an RDKit tool; based on an open source experiment data set, training a machine learning model through an association relationship between a one-dimensional descriptor, a two-dimensional descriptor and a three-dimensional descriptor of a molecule and experiment solvation free energy to obtain a prediction model; and inputting molecular descriptors of solute molecules and solvent molecules into the prediction model to obtain a result of predicting solvation free energy of the solute molecules. Existing molecular dynamics simulation is replaced by the machine learning model, and meanwhile, model training is performed by adopting a machine learning algorithm and experimental data, so that the accuracy and the stability of a prediction result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computational chemistry, and particularly to a method for predicting molecular solvation free energy based on machine learning. Background Art

[0002] Molecular solvation free energy is an important field in chemistry, materials science, and biology. Computational methods for molecular solvation free energy have been widely applied in fields such as drug design, materials development, and biomolecular simulation. In the prior art, the calculation of molecular solvation free energy mainly relies on the Molecular Dynamics method. This method is based on classical force field theory and calculates the solvation free energy by simulating the interactions of molecules in a solvent through molecular dynamics. However, due to the computational complexity of the molecular dynamics method and the limitations of the model itself, the calculated results of solvation free energy often have errors compared with experimental values. These errors may stem from factors such as the complexity of the molecular topology, the inaccuracy of force field parameters, and the insufficient treatment of solvent effects.

[0003] To reduce these errors, it is usually necessary to optimize the relevant parameters. However, this optimization process usually requires a large amount of computational resources and time, and even after optimization, the improvement in computational accuracy is still limited, especially in complex molecular systems. This makes the molecular dynamics method still have deficiencies in practical applications.

[0004] In recent years, machine learning technology has developed rapidly. Machine learning methods can learn from a large amount of data and capture the complex relationships between molecular features and their physical and chemical properties, obtaining more accurate prediction results. There is an urgent need for a machine learning method combined with molecular descriptors to calculate and predict solvation free energy.

[0005] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art: Calculating the molecular solvation free energy by the molecular dynamics method requires a large amount of computational resources and time, the calculation results are not accurate enough, and the cost is also high. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for predicting molecular solvation free energy based on machine learning, so as to solve the technical problems in the prior art that calculating the molecular solvation free energy by the molecular dynamics method requires a large amount of computational resources and time, the calculation results are not accurate enough, and the cost is also high. The many technical effects that can be produced by the preferred technical solutions provided by the present invention are described in detail below.

[0007] To achieve the above purpose, the present invention provides the following technical solutions: A method for predicting the solvation free energy of molecules based on machine learning provided by the present invention includes the following steps: S100: Using the SMILES formulas of solute molecules and solvent molecules as input data; S200: Based on the RDKit tool, converting the SMILES formulas into one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors of molecules; S300: Based on an open-source experimental dataset, training a machine learning model through the correlation between the one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors of molecules and the experimental solvation free energy to obtain a prediction model; S400: Inputting the molecular descriptors of solute molecules and solvent molecules into the prediction model to obtain the result of predicting the solvation free energy of the solute molecules.

[0008] Preferably, in the step S200, the one-dimensional descriptors include molecular weight and number of bonds, the two-dimensional descriptors include molecular graph features, topological polar surface area TPSA, and molecular connectivity index, and the three-dimensional descriptors include molecular surface area and molecular volume.

[0009] Preferably, in the step S200, the missing values in the conversion of the SMILES formulas are estimated and filled by the K-nearest neighbor classification algorithm to keep the datasets of the one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors complete.

[0010] Preferably, in the step S200, the datasets of the one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors are also normalized to create a consistent scale for the entire dataset.

[0011] Preferably, in the step S300, the open-source experimental data is from the FreeSolv database. The molecules in the FreeSolv database include a training set and a test set. The training set is used to train the machine learning model, and the test set is used to verify the prediction model.

[0012] Preferably, in the step S300, the machine learning model is trained by two strategies: the first training strategy and the second training strategy. The first training strategy is trained based on the offset of the directly predicted experimental values from the molecular features, and the second training strategy is trained based on the difference between the predicted experimental values and the simulated values.

[0013] Preferably, in the step S300, the machine learning model includes support vector machine SVM, random forest RF, deep neural network DNN, multiple linear regression MLR, and extreme gradient boosting XGB. The optimal hyperparameter configuration of each model is identified by the Bayesian optimization method.

[0014] Preferably, in the step S300, during the process of training the machine learning model, an ensemble method and a dimensionality reduction method are also used for model training; in the ensemble method, the prediction result of the prediction model is obtained by combining the predicted outputs of each machine learning model on average; the dimensionality reduction method performs feature dimensionality reduction through the principal component analysis algorithm PCA to improve the model training efficiency.

[0015] Preferably, the solvent molecule is a polar solvent, a non-polar solvent or an ionic solvent, and the solute molecule is an organic small molecule.

[0016] Preferably, the solvent is water, and the solute is an alcohol, an ester, a ketone, an aromatic compound, or a compound containing nitrogen, sulfur, or oxygen functional groups.

[0017] Implementing one of the above technical solutions of the present invention has the following advantages or beneficial effects: By replacing the existing molecular dynamics simulation with a machine learning model, the present invention significantly shortens the calculation time, is applicable to high-throughput screening of large-scale molecular libraries. At the same time, advanced machine learning algorithms and large-scale experimental data are used for model training, significantly improving the accuracy and stability of the prediction results, reducing costs, and providing a multi-faceted and efficient solution for the rapid prediction of the solvation free energy of molecules. At the same time, based on the SMILES formula and molecular descriptors, the cumbersome molecular dynamics simulation process in the traditional method is avoided, and the calculation steps of the solvation free energy are simplified. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings: Figure 1 is a flowchart of a method for predicting the solvation free energy of molecules based on machine learning according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, various exemplary embodiments to be described below will refer to the corresponding drawings, which form a part of the exemplary embodiments and describe various exemplary embodiments that may be adopted to implement the present invention. Unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. It should be understood that they are merely examples of processes, methods, devices, etc. consistent with some aspects of the present invention disclosed in detail in the appended claims. Other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present invention.

[0020] In order to illustrate the technical solutions described in the present invention, specific examples will be used for illustration below, only showing the parts related to the embodiments of the present invention.

[0021] Embodiment 1: As Figure 1As shown, the present invention provides a method for predicting the solvation free energy of molecules based on machine learning, comprising the following steps: S100: Using the SMILES (Simplified Molecular Input Line Entry System) formulas of solute molecules and solvent molecules as inputs. The SMILES formula, namely the simplified molecular linear input specification, is a specification that clearly describes the molecular structure with ASCII strings, a simplified representation method for expressing molecular structures with strings, which realizes the digital expression of molecular information. It can be imported by most molecular editing software and converted into two-dimensional graphics or three-dimensional models of molecules, can concisely represent complex molecular structures, enables the computer to obtain more accurate and useful information, and is easy to process and store in the computer, thus facilitating machine learning operations, avoiding the complex simulation process and high computational cost in traditional calculations, and greatly improving the prediction efficiency. S200: Based on the RDKit tool, convert the SMILES formula into one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors of molecules. The one-dimensional descriptors (such as molecular weight, number of hydrogen bond donors, number of hydrogen bond acceptors, etc.), two-dimensional descriptors (such as topological polar surface area TPSA, molecular connectivity index, etc.), and three-dimensional descriptors (such as three-dimensional geometric properties such as molecular surface area and volume) are comprehensive and can cover the topological structure, geometric features, and electronic properties of molecules. Due to the inherent complexity of molecular structures, directly using three-dimensional molecular structures for prediction often involves complex spatial information and multi-dimensional data, and it is difficult for machine learning algorithms to effectively capture these features. In order to systematically study the relationship between input features and solvent free energy, as many features as possible must be generated and different types of features must be studied to identify those features highly correlated with the solvation free energy of molecules. Therefore, converting these molecular structures into numerical features of one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors suitable for machine learning facilitates more accurate prediction calculations. The RDKit tool provides rich chemical information processing functions for processing and analyzing information such as molecular structures, chemical reactions, and chemical properties, facilitating the obtaining of one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors of molecules. S300: Based on an open-source experimental data set, train a machine learning model through the correlation between the one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors of molecules and the experimental solvation free energy to obtain a prediction model. The trained prediction model can quickly and accurately predict the solvation free energy of target molecules. S400: Input the molecular descriptors of solute molecules and solvent molecules into the prediction model to obtain the result of the predicted solvation free energy of the solute molecules. The present invention replaces the existing molecular dynamics simulation with a machine learning model, greatly shortening the calculation time. It is estimated that it takes 500 hours to calculate the solvation free energy of the above 1000 compounds by the statistical molecular dynamics method, while the method of the present invention only takes 10 minutes, significantly saving the calculation time.Therefore, the present invention is applicable to high-throughput screening of large-scale molecular libraries. At the same time, advanced machine learning algorithms and large-scale experimental data are used for model training, significantly improving the accuracy and stability of prediction results, reducing costs. The method provided in this embodiment provides a multi-faceted and efficient solution for the rapid prediction of the solvation free energy of molecules. At the same time, based on SMILES formulas and molecular descriptors, the cumbersome molecular dynamics simulation process in traditional methods is avoided, and the calculation steps of solvation free energy are simplified.

[0022] As an alternative implementation, in step S200, the one-dimensional descriptors include molecular weight, number of bonds (such as number of hydrogen bond donors, number of hydrogen bond acceptors), two-dimensional descriptors include molecular graph features, topological polar surface area TPSA, molecular connectivity index, and three-dimensional descriptors include molecular surface area, molecular volume. Through one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors, the topological features, geometric features, and electronic features of molecules can be obtained, thus facilitating the molecular solvation free energy. Further, the descriptors specifically used in this embodiment include APFP, ECFP6, TOPOL, MolProps, and their combined combinations MolPropsAPFP and MolPropsECFP6. Among them, a single descriptor such as APFP is only a way to describe molecular features. Using only this one molecular descriptor may ignore some physicochemical features of the molecule, so there may be a problem of incomplete molecular features. However, a combined descriptor such as MolPropsAPFP combines two types of molecular descriptors, thus providing more comprehensive molecular features.

[0023] As an alternative implementation, in step S200, the K-nearest neighbor classification algorithm (KNN, k-NearestNeighbor, each sample can be represented by its K closest neighboring values, suitable for automatic classification of relatively large sample capacity domains) is used to estimate and fill in the missing values in the conversion of SMILES formulas, that is, the KNN algorithm is used to process feature interpolation to keep the datasets of one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors complete, thus facilitating the training of machine learning models. At the same time, since machine learning models cannot process non-numerical features, during the conversion process of SMILES formulas, some features are non-numerical features. Deleting these non-numerical features to ensure that only relevant numerical data is retained helps to perform effective data analysis. In step S200, the datasets of one-dimensional descriptors, two-dimensional descriptors, and three-dimensional descriptors are also normalized to create a consistent scale for the entire dataset, so that more accurate training results can be obtained. These two comprehensive preprocessing methods, the K-nearest neighbor classification algorithm and normalization processing, produce a highly robust dataset, which is very suitable for analysis and subsequent model training.

[0024] As an alternative implementation, in step S300, the open-source experimental data is sourced from the FreeSolv database, preferably version 0.52 of the FreeSolv database. The FreeSolv database is a widely used benchmark dataset for predicting solvation free energies, providing experimental measurements and theoretical calculations of solvation free energies for 642 small neutral organic molecules. It is an ideal resource for building and evaluating computational models aimed at predicting solvation properties. The molecules in the FreeSolv database include a training set and a test set. The training set is used to train machine learning models, and the test set is used to validate the prediction models. For each individual model, the optimal set of hyperparameters that minimizes the MUE during cross-validation is first determined. This step ensures that each model is fine-tuned to achieve its best performance on the validation data, balancing the trade-off between model complexity and accuracy. After determining the optimal hyperparameters, the model is retrained using the complete training set, which enables the model to fully utilize all available data and is crucial for ensuring robust prediction performance. After training, the model is applied to the test set, where it generates predictions based on the newly trained configuration, and these predictions represent the final output of each model. Out of the 642 molecules in the FreeSolv database, 47 were specifically selected for the SAMPL4 blind test, an influential benchmarking exercise aimed at evaluating and comparing different prediction models in the context of solvation and molecular interactions. These 47 molecules were designated as the test set because their free energy values were excluded from the training process to allow unbiased model validation. This blind test ensures that the performance of the model can be evaluated based on unseen data, simulating real-world prediction scenarios. The remaining 595 molecules out of the 642 molecules serve as the training set for developing and optimizing the computational model. These molecules provide a rich dataset that enables the model to learn the relevant patterns and relationships between molecular structure and its solvation energy. To avoid data leakage and ensure generalizability, no molecules from the test set were used during the model construction phase. By separating the molecules in the FreeSolv database into a test set and a training set, the present invention implements best practices in machine learning and statistical analysis, ensuring that the reported performance metrics reflect the model's generalization ability beyond the training data.

[0025] As an alternative implementation, in step S300, the machine learning model is trained by two strategies: the first training strategy and the second training strategy. The first training strategy is trained based on directly predicting the offset of the experimental value from the molecular features, and the second training strategy is trained based on the difference between the predicted experimental value and the simulated value. The results of the two training strategies can be used to verify the advantages of this method compared with existing methods. The first training strategy involves directly predicting the experimental value from the molecular features, thereby generating the molecular solvation free energy. The second training strategy incorporates the calculated value as an additional molecular feature, and the model predicts the offset between the calculated value and the experimental value, allowing for the correction of the calculated free energy. Based on the first training strategy, the solvation free energy of the molecule can be directly predicted. Since there is a relatively direct physicochemical relationship between the features and the results, the results obtained by this method maintain a high degree of interpretability. In addition, the trained model exhibits strong generalization ability because it focuses on the direct mapping between the molecular features and the solvation free energy. This characteristic of the model's generalization ability makes it very suitable for predicting the solvation free energy of new molecules, even if they are different from the training data, as long as the molecular feature space remains similar. In this solution, the specific expression of the first training strategy is preferably: , where A represents any sample in the training set, is the calculated free energy obtained by performing molecular dynamics (MD) simulation using the General Amber ForceField (GAFF), is the experimental value obtained by the machine learning model in this method, represents the predicted offset value corresponding to the first training strategy. The second training strategy involves correcting the simulated value by predicting the difference between the experimental value and the simulated value. This method retains the correct trend set by the simulation calculation and avoids wasting effort in areas where the calculated value is already highly accurate. This strategy depends to a large extent on the consistency between the data sources of the training set and the test set, especially the calculation method used to generate the molecular simulation data. The expression of the second training strategy is: , for each training set defined by its descriptors in the first training strategy, the machine learning model is applied using 5-fold cross-validation, which results in a total of N = 5 training models. Each individual model in the set predicts its own offset value, is the arithmetic mean of these predicted offset values, and the standard deviation of the mean can be used as a measure of precision.

[0026] As an alternative implementation, in step S300, the machine learning models include Support Vector Machine (SVM), Random Forest (RF), Deep Neural Network (DNN), Multiple Linear Regression (MLR), and Extreme Gradient Boosting (XGB). Table 1 defines the hyperparameter space for each model. For small datasets, traditional machine learning algorithms have an advantage over complex deep learning methods. They can be trained relatively quickly on standard hardware without requiring extensive and intensive computations, thus significantly reducing the computational cost. The above-mentioned centralized model algorithms have different advantages in dealing with different aspects of prediction tasks. By combining these models, it is convenient to integrate the advantages of various models to obtain a more accurate prediction model. For example, DNN and XGB have better consistency for continuous variables, while RF and SVM have better consistency for discrete variables. Considering both the physical meaning of features and the inherent capabilities of the models can achieve the best prediction performance. These models all identify the optimal hyperparameter configuration for each model through the Bayesian optimization method. The Bayesian optimization method uses the SciKit-Optimize (SKOPT) library, version 0.5.2. The Bayesian optimization method utilizes the Expected Improvement acquisition function. Compared with traditional methods such as random or grid search, it allows for a more strategic and effective exploration of the hyperparameter space. By focusing the search on promising regions in the space, Bayesian optimization helps to identify high-performance configurations with fewer iterations. In this method, the optimization step is set to a maximum of 40 steps because previous experience has shown that model convergence usually occurs before this limit. For each iteration during the optimization process, the Mean Unsigned Error (MUE) (average across cross-validation folds) on the validation set is calculated and returned to the SKOPT routine. This metric serves as a cost function to guide the selection of the next set of hyperparameters. With each successive call, SKOPT further improves its search, aiming to minimize the cost function and achieve the best performance. This iterative process continues until convergence is observed or the step limit is reached.

[0027] Table 1 Definition of the hyperparameter space for each model in the machine learning models As an alternative implementation, in step S300, during the process of training the machine learning model, an ensemble method and a dimensionality reduction method are also adopted for model training; in the ensemble method, the prediction results of the prediction model are obtained by combining the outputs predicted by each machine learning model on average; the dimensionality reduction method performs feature dimensionality reduction through the principal component analysis algorithm PCA to improve the model training efficiency. In the ensemble method strategy, the predictions of all models are combined by averaging their outputs. This method takes advantage of the unique strengths of each model because different models may be good at capturing different patterns or relationships in the data. By integrating their predictions, the ensemble method mitigates the potential weaknesses or biases inherent in any single model. The result is usually an improvement in overall performance because the ensemble benefits from the diversity of model predictions, leading to better generalization to new data. Finally, the ensemble prediction is compared with the actual values of the test set to calculate the average value MUE of the final cross-validation. This metric provides a comprehensive measure of the ensemble accuracy, indicating the degree of agreement between the combined prediction and the true result. The ensemble learning framework not only enhances the prediction ability of the system but also helps reduce the risk of overfitting and ensures more powerful performance compared to any single model. Specifically, after determining the best predictions of each individual model, it is further enhanced through the ensemble method. By averaging the predictions of all models, a unified ensemble prediction is generated, which is then compared with the test set to evaluate the final MUE. The overall prediction is usually closer to the experimental results, effectively reducing the number of outliers. Through this ensemble method, a higher level of reliability and accuracy can be achieved in the final result. In model training, a large feature set may introduce noise during the training process, potentially reducing model accuracy and leading to overfitting. In addition, many features may be highly correlated with each other. Highly correlated features usually express redundant information, either affecting the target label in the same way or having similar physical and chemical properties. This redundancy of features increases the computational complexity and may lead to the sparsity of the feature matrix, ultimately reducing the training efficiency and wasting computational resources. Therefore, the reduction of feature dimensionality is crucial for the training result. To address the problems of redundant and highly correlated features, the principal component analysis algorithm PCA is applied in this application for dimensionality reduction. The basic principle of the principal component analysis algorithm PCA is to analyze the correlation between features by calculating the covariance matrix of the data, and then extract the direction of the maximum variance (i.e., the principal component direction) through eigenvalue decomposition. These principal components are linear combinations of the original features, sorted according to their variances. The components with higher variances capture most of the information in the data, and the components with lower variances represent less significant changes and can usually be ignored. In the implementation process, first, the original feature data is standardized to ensure that each feature has the same scale, making their contributions to the model equivalent. After standardization, the PCA function in scikit-learn is applied to reduce the dimensionality of the data.During the reduction process, a threshold is set for the proportion of variance to be retained. 95% and 75% are selected as the variance thresholds to be retained, which can both reduce the number of features while retaining as much important information as possible in the original data. Then, the reduced features are used as the input for training the machine learning model. In this application, the changes in the feature dimensions after dimensionality reduction and the corresponding data are shown in Table 2.

[0028] Table 2 Changes in the feature dimensions after dimensionality reduction and the corresponding data for each model in the machine learning model As an optional implementation, the solvent molecule is a polar solvent, a non-polar solvent, or an ionic solvent. Thus, the method of the present invention has stronger adaptability and can achieve good prediction effects for different types of solvents, and can thus be widely applied in fields such as drug research and development, material design, and green chemistry.

[0029] As an optional implementation, in this method, the preferred solvent is water and the solute is acetone. Using the method of the present invention, the predicted value of the solvation free energy of acetone in water is -6.0 kcal / mol. Compared with the experimental value of -6.1 kcal / mol, the error is only 0.1 kcal / mol, indicating that the model of the present invention has high prediction accuracy.

[0030] The embodiment is only a special case and does not indicate that the present invention has only such an implementation.

[0031] The above are only the preferred embodiments of the present invention. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent substitutions can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the protection scope of the present invention.

Claims

1. A method for predicting the solvation free energy of molecules based on machine learning, characterized in that, It includes the following steps: S100: Using the SMILES formulas of solute molecules and solvent molecules as input data; S200: Based on the RDKit tool, converting the SMILES formulas into one-dimensional descriptors, two-dimensional descriptors and three-dimensional descriptors of molecules; S300: Based on an open-source experimental dataset, training a machine learning model through the correlation between the one-dimensional descriptors, two-dimensional descriptors and three-dimensional descriptors of molecules and the experimental solvation free energy to obtain a prediction model; S400: Inputting the molecular descriptors of solute molecules and solvent molecules into the prediction model to obtain the result of predicting the solvation free energy of the solute molecules.

2. The method for predicting the solvation free energy of a molecule based on machine learning according to claim 1, wherein In the step S200, the one-dimensional descriptors include molecular weight and number of bonds, the two-dimensional descriptors include molecular graph features, topological polar surface area TPSA, and molecular connectivity index, and the three-dimensional descriptors include molecular surface area and molecular volume.

3. A method for predicting the solvation free energy of a molecule based on machine learning according to claim 1, characterized in that, In the step S200, the K-nearest neighbor classification algorithm is used to estimate and fill in the missing values in the conversion of the SMILES formulas to keep the datasets of the one-dimensional descriptors, two-dimensional descriptors and three-dimensional descriptors complete.

4. The method for predicting the solvation free energy of a molecule based on machine learning according to claim 1, wherein In the step S200, the datasets of the one-dimensional descriptors, two-dimensional descriptors and three-dimensional descriptors are also normalized to create a consistent scale for the entire dataset.

5. A method for predicting the solvation free energy of a molecule based on machine learning according to claim 1, characterized in that, In the step S300, the open-source experimental data comes from the FreeSolv database. The molecules in the FreeSolv database include a training set and a test set. The training set is used to train the machine learning model, and the test set is used to verify the prediction model.

6. The method for predicting the solvation free energy of a molecule based on machine learning according to claim 1, wherein In the step S300, the machine learning model is trained through two strategies: the first training strategy is based on directly predicting the offset of the experimental value based on molecular features for training, and the second training strategy is based on the difference between the predicted experimental value and the simulated value for training.

7. A method for predicting the solvation free energy of a molecule based on machine learning according to claim 1, characterized in that, In the step S300, the machine learning model includes support vector machine SVM, random forest RF, deep neural network DNN, multiple linear regression MLR, and extreme gradient boosting XGB. The optimal hyperparameter configuration of each model is identified through the Bayesian optimization method.

8. A method for predicting the solvation free energy of molecules based on machine learning according to claim 1, characterized in that, In the step S300, during the process of training the machine learning model, an ensemble method and a dimensionality reduction method are also used for model training; in the ensemble method, the prediction result of the prediction model is obtained by combining the outputs predicted by each machine learning model on average; the dimensionality reduction method performs feature dimensionality reduction through the principal component analysis algorithm PCA to improve the model training efficiency.

9. A method for predicting the solvation free energy of a molecule based on machine learning according to claim 1, characterized in that, The solvent molecule is a polar solvent, a non-polar solvent or an ionic solvent, and the solute molecule is an organic small molecule.

10. A method for predicting the solvation free energy of a molecule based on machine learning according to claim 1, characterized in that, The solvent is water, and the solute is an alcohol, an ester, a ketone, an aromatic compound, or a compound containing nitrogen, sulfur, or oxygen functional groups.

Citation Information

Patent Citations

  • Molecular screening method based on deep learning technology qualitative and quantitative model

    CN113674807A

  • Data processing method and device, model training method and free energy prediction method

    CN114218869A

  • System and method for estimating solubility

    CN115547426A

  • Method for predicting solvation energy of small molecule compound based on graph convolutional neural network

    CN115938501A

  • Method for predicting solubility of carbon dioxide in alcohol amine absorbent based on machine learning model

    CN118645174A

Cited By

  • Ionic liquid-organic solvent system density prediction method and system based on Bayesian optimization and SHAP interpretation mechanism

    CN121171404A