Inplanatable molecular property prediction method using neural feature analysis in combination with global fragment features
Through the method of neural feature analysis combined with global fragment characteristics, problems such as insufficient interpretability and high computing resource consumption of molecular properties prediction models in the prior art are solved, and higher prediction accuracy, stronger interpretability and lower computing resource consumption are achieved.
Patent Information
- Application Number
- CN202510254326.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
Existing machine learning models have problems such as insufficient interpretability, high computing resources consumption, long training time and fluctuations in molecular properties prediction.
The molecular properties are predicted by neural feature analysis combined with global fragment characteristics. By combining global physical and chemical properties with local fragment vectors, neural feature analysis is used for prediction, and the prediction results are obtained through nuclear methods.
It improves prediction accuracy, enhances the interpretability of the model, reduces computing resource consumption and training time, and improves the stability and robustness of the model.
Smart Images

Figure CN120183547A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of chemical and biological information technology, and particularly relates to an interpretable molecular property prediction method using neural feature analysis combined with global fragment features. Background Art
[0002] During the drug development process, certain key properties of molecules play a decisive role. Precise prediction of these properties can significantly improve the efficiency of new drug design, shorten the development cycle, optimize the formulation process, and improve the bioavailability of drugs. Although reliable data can be obtained through traditional experimental methods, their time-consuming, laborious, and costly characteristics often restrict the research progress. In recent years, with the breakthrough development of computing power, machine learning technology has made important progress in molecular property prediction, especially showing unique advantages in the pharmaceutical field. The computational prediction model based on machine learning has developed into the core method for virtual screening and candidate compound optimization. Compared with traditional experimental methods, such models provide chemists with efficient digital tools by analyzing the complex relationship between molecular structure and its physical and chemical properties, helping to optimize experimental schemes and focus research directions.
[0003] It is worth noting that although some machine learning algorithms show excellent prediction ability, many models (especially property prediction models based on deep learning) still face challenges in interpretability. The continuously enhanced computing performance has not alleviated the "black box" characteristics of neural networks. Instead, due to their complex parameter system and non-linear feature transformation in high-dimensional space, it has further deepened the difficulty of understanding the model mechanism. To achieve the development goal of "scientific intelligence", enhancing the transparency and interpretability of models has become a key issue. In this context, the research community has introduced game theory-based interpretation techniques (such as LIME and SHAP) to clarify the molecular structure-activity relationship by revealing the decision-making mechanism of neural networks, which has important scientific value for deeply understanding the formation mechanism of compound properties.
[0004] The invention with the application publication number CN119170138A constructs hybrid features by integrating rational descriptors and molecular graph embedding vectors, and uses a multi-layer perceptron for prediction. However, it completely relies on neural network training, often accompanied by high computational resource consumption and long training time, and lacks sufficient globality when explaining model decisions.
[0005] The invention with the application publication number CN119418821A adopts a molecular wire graph structure based on Transformer and makes predictions by using molecular fingerprints and descriptors as node features. Although it has made breakthroughs in feature capture, it also faces problems of insufficient interpretability and low computational efficiency.
[0006] The invention with the publication number CN119358591A uses a graph neural network combined with frequent subgraph mining and an edge masking mechanism to extract key subgraphs. Although certain results have been achieved in local interpretation, overall, it is still limited by the result fluctuations caused by random initialization and the defect of global interpretability.
[0007] In contrast, the present invention innovatively introduces a molecular representation method that combines global physicochemical properties and local fragment vectors, and uses a method of neural feature analysis for prediction, effectively making up for the above deficiencies. Specifically:
[0008] (1) Higher prediction accuracy: By leveraging the complementary advantages of global and local information, it can capture key information in the molecular structure more comprehensively and precisely, thus achieving excellent prediction results on multiple datasets.
[0009] (2) Stronger interpretability: The proposed interpretability method can not only perform local analysis on the contributions of the features of a single molecule, but also statistically analyze key features globally, helping researchers understand the model decision-making process more intuitively.
[0010] (3) Improved computational efficiency: By reasonably designing the feature extraction and interpretation processes, the present invention avoids a large number of non-linear operations in traditional deep neural networks while significantly reducing the consumption of computing resources and training time.
[0011] (4) Enhanced stability and robustness: The combination of neural feature analysis and kernel methods effectively overcomes the problem of result fluctuations caused by different random number seeds, demonstrating better stability and generalization ability in experiments.
[0012] In summary, the present invention has achieved breakthroughs in aspects such as accuracy, interpretability, computational efficiency, and model stability, directly solving the deficiencies of the existing technology in molecular property prediction, and providing efficient and transparent technical support for new drug research and development and molecular design. Summary of the Invention
[0013] In order to overcome the problem of poor interpretability of neural networks, make a reasonable explanation for the model prediction results, and further improve the prediction accuracy of the model for molecular properties. The present invention proposes an interpretability method for neural feature analysis; proposes a molecular representation method that combines global physicochemical properties and local fragment vectors, greatly improving the prediction accuracy compared with traditional fingerprints MACCS and Morgan; finally, generalizes the kernel function in kernel ridge regression.
[0014] Combining the neural feature mechanism with kernel ridge regression enables traditional kernel learning to learn the data features learned in the neural network. Without the complex network structure and a large number of non-linear operations of the neural network, the interpretability is greatly enhanced. The present invention discovers that the importance analysis of sample features can be obtained by double mapping the data samples with a well-trained average gradient outer product matrix.
[0015] The described interpretability method can locally analyze the features that contribute positively and negatively to the predicted property for a single molecule. The formula is as follows, which reflects the contribution of each feature in a single sample to the target variable;
[0016]
[0017]
[0018] The described interpretability method can globally analyze the features most relevant to the predicted property in the molecular set, which is obtained by summing up the feature analyses of all samples through statistical methods. The formula is as follows, revealing the key features in the molecular set;
[0019]
[0020] The described global fragment features are jointly composed of the global physicochemical properties of the molecule and the frequency fragment vector. The global descriptors are calculated by the commonly used chemical software rdkit package, and after calculation, they are screened by variance threshold, correlation threshold, and zero ratio threshold. The formula is as follows:
[0021]
[0022]
[0023]
[0024] The described fragment fingerprints are obtained by means of molecular fragmentation (SMILES Pair Encoding, SPE) (Li, Xinhao, and Denis Fourches. "SMILES pair encoding: a data-driven substructure tokenization algorithm for deep learning." Journal of chemical information and modeling 61.4 (2021)), and the functional group information of the molecule is supplemented to reduce information loss. The functional group supplementation is completed by steps such as defining the functional group pattern, identifying and merging rings, labeling functional groups, merging functional groups, removing single-atom functional groups in rings, and returning the SMILES representation. Finally, a set of SMILES representations of functional groups is output: fg smiles ={FG SMILES i |for all FG i ∈FGs}. After obtaining all the fragments, screening is carried out through the correlation threshold and the zero ratio threshold.
[0025] The generalization of the described kernel function generalizes the Laplace kernel in the recursive feature machine to the Matern kernel, the Rational Quadratic kernel, and the Gaussian kernel. The formulas of the three kernels are as follows:
[0026]
[0027] where ||x - x'|| is the Euclidean distance between the input vectors x and x'; v is the smoothness parameter, controlling the smoothness of the function; l is the scale parameter, controlling the scalability of the kernel function, and K v is the modified Bessel function; Γ(v) is the gamma function.
[0028]
[0029] where l is the scale parameter, controlling the scalability of the kernel function; α is the smoothness parameter, controlling the smoothness of the kernel function.
[0030]
[0031] where l is the scale parameter, determining the width of the kernel function. Brief Description of the Drawings
[0032] Figure 1 is a block diagram of the global fragment feature generation of the present invention and the prediction of molecular properties by combining neural feature analysis;
[0033] Figure 2It is the single - molecule interpretation diagram of the present invention;
[0034] Figure 3 It is the global - molecule interpretation diagram of the present invention;
[0035] Figure 4 It is the density scatter plot for predicting molecular properties. Detailed implementation manners
[0036] The content of the present invention will be further elaborated below with reference to the accompanying drawings.
[0037] As Figure 1 shown is a block diagram for predicting molecular properties through global fragment feature generation and combined with neural feature analysis. The molecular set first undergoes SPE fragmentation and functional - group supplementation to generate fragment - set information. Another channel calculates physicochemical properties such as molecular weight and the number of halogen atoms through rdkit software to generate global information. Both pass through threshold screening to remove redundant information. The molecule generates a frequency vector and a descriptor vector respectively through a data generator. After the two - type feature information is fused, feature learning is carried out through neural feature analysis, and finally the prediction result is obtained through a kernel method.
[0038] As Figure 1 shown, the upper - right corner is a schematic diagram of the interpretability method. By performing two - matrix mappings on the feature vector representing the sample and the outer - product matrix of the trained average gradient, the importance evaluation information of each feature in the sample is obtained. Different from the interpretation method in SHAP that assumes each feature is independent, the two - matrix mappings consider the correlation between different features and make a more reasonable interpretation.
[0039] As Figure 2 It is the single - molecule interpretation diagram of the present invention. Taking the solubility degree of the molecule "CCCOP(=S)(OCCC)SCC(=O)N1CCCCC1C" in water as an example, the top 10 most influential features in the global fragment features are selected. Table 4 provides the specific information of the 10 features. Through interpretability analysis, it can effectively help biologists and chemists discuss molecular properties, thus accelerating the research and development of new drugs.
[0040] As Figure 3 The global - molecule interpretation diagram of the present invention obtains the features most relevant to the predicted property in the overall data by analyzing the feature importance of all molecules in the data set, providing guiding significance for the design of molecules.
[0041] As Figure 4 It is the regression density scatter plot of the molecular properties predicted by the present invention. Taking the lipophilicity data set as an example, the octanol - water partition coefficient of the predicted molecule is predicted. Consistent with general regression tasks, the greater the density near the diagonal, the more accurate the prediction effect.
[0042] The results show that the global fragmentation features have great superiority and can better characterize the structure of molecules, exceeding the traditional fingerprints MACCS and Morgan. Table 1 below shows the results of different machine learning methods under the global fragmentation features for 10 datasets. The results show that the combination of neural feature parsing and global fragmentation features has the best prediction performance. Among the 10 datasets, the target values of ESOL, Arash, AqSolD B, wide, and narrow are water solubility; the target value of FreeSolv is hydration free energy; the target value of lipophilicity is octanol-water partition coefficient; the target values of acetone, benzene, and ethanol are the solubility values of the compounds in acetone, benzene, and ethanol, respectively. It can be seen that the present invention exceeds other methods, including traditional machine learning methods and deep learning based on neural networks. For the data in Table 1, the dataset is randomly divided into training set: validation set: test set = 8:1:1, randomly divided and run 10 times, and the mean (variance) is recorded. The "-" in Table 1 indicates that the obtained RMSE is too large and has no practical prediction significance.
[0043] Table 1 Comparison of different modeling methods of the present invention (evaluation index RMSE)
[0044]
[0045]
[0046] Table 2 is a comparison between the features proposed by the present invention and the classical fingerprints (MACCS, Morgan) (taking the Arash dataset as an example). It can be found that the method of the present invention comprehensively exceeds the classical feature methods.
[0047] Table 2 Comparison of different features (evaluation index RMSE)
[0048] Methods / representation MACCS Morgan Hybrid RFM 0.7809(0.0009) 0.7909(0.0010) 0.6013(0.0007) XGBoost 0.7853(0.0006) 0.8992(0.0009) 0.6332(0.0003) Gradient Boosting Tree 0.7882(0.0007) 0.9252(0.0011) 0.6215(0.0003) Random Forest 0.8129(0.0013) 0.8765(0.0015) 0.6572(0.0005) Linear - - 0.8489(0.0012) FT-transformer 0.8971(0.0008) 1.0598(0.0005) 0.7007(0.0005) ResNet 0.8742(0.0011) 0.9185(0.0010) 0.7297(0.0017)
[0049] In addition, the method of the present invention can be comparable to the state-of-the-art graph neural network (GNN). GNN uses the topological graph of molecules as features and can perform end-to-end feature learning. Most of the current advanced molecular prediction methods are derived from GNN. Table 3 shows the comparison between the effect of the present invention and the advanced GNN on the public dataset. The results show that the present invention can be comparable to the current optimal GNN, requires much less running time than the deep neural network, and has good interpretability.
[0050] Table 3 Comparison between the present invention and advanced GNN (evaluation index RMSE)
[0051]
[0052]
[0053] Table 4 corresponds to Figure 2 the specific detailed information of 10 features, including the location where the feature is located, the name of the feature, the value of the feature, and the category to which the feature belongs.
[0054] The corresponding features of the single molecule interpretation diagram in Table 4
[0055] Feature Index Feature Name Feature Value Feature class Feature_69 C 14 fragments Feature_99 C(C) 10 fragments Feature_266 O 3 fragments Feature_100 C(C)(C) 6 fragments Feature_5 ATS0p 2.257132 descriptors Feature_102 C(C)(O) 2 fragments Feature_38 ATSC1p 3.920552 descriptors Feature_26 ATSC1v 2.536397 descriptors Feature_4 ATSOZ 0.972922 descriptors Feature_110 C(N) 3 fragments
[0056] In summary, compared with other networks, the present invention performs excellently in terms of accuracy, robustness and generalization ability. This not only improves the reliability of data analysis, but also can extract key information more effectively, which is of great significance for subsequent research.
Claims
1. The molecular property prediction method using neural feature analysis and global fragment features is unique in that: the method first uses advanced neural feature analysis technology to perform deep nonlinear mapping on molecular data, thereby extracting the multi-dimensional feature information implicit in the molecule, and combining the multiple representations of the molecule in terms of structure, physicochemistry and reactivity; then, by introducing global fragment features, the overall physicochemical properties and local topological structure of the molecule are quantitatively described, thereby making up for the information missing problem that may exist in a single representation method. Specifically, the neural feature analysis part uses a multi-layer deep network structure to realize feature extraction and dimensionality reduction of the original molecular data. This process not only improves the model's ability to understand complex data, but also provides rich and detailed input information for subsequent predictions. At the same time, in order to improve the robustness of the model, this method introduces a feature importance analysis mechanism for the prediction process, which can quantitatively evaluate the importance of each feature from both the local and global levels, which not only ensures the accuracy of the prediction results, but also provides a theoretical basis for the interpretability of the model. In order to adapt to the characteristics of different molecular systems, the traditional Laplace kernel in the neural feature parsing kernel ridge regression is further expanded to the Mattern kernel, rational quadratic kernel and Gaussian kernel, which greatly expands the scope of application and generalization ability of the model.
2. According to claim 1, it is characterized in that: The neural feature analysis is combined with molecular fingerprint characterization, wherein the molecular fingerprint includes the traditional MACCS fingerprint, Morgan fingerprint and the custom molecular feature designed by the present invention. In the specific implementation, the above-mentioned multiple fingerprint information is fused, which can not only capture the global structural information of the molecule, but also accurately reflect the local structural characteristics of the molecule. Custom molecular features are special descriptors designed based on specific chemical reactivity or target properties, which play a complementary role in different molecular systems, thereby improving the overall prediction accuracy. This fusion strategy effectively solves the information limitation problem of a single characterization method, and provides a more comprehensive and accurate basis for the property analysis of complex molecular systems.
3. According to claim 1, it is characterized in that A feature importance analysis method for neural feature analysis is proposed. Local importance analysis analyzes the impact of different features in a molecule on the properties of a single molecule. Global importance analysis selects the features in the molecule library that are most relevant to a certain property. The local and global importance analysis formulas are as follows (1) and (2), respectively, where x i is a sample, and M is the average gradient outer product matrix of the learning.
4. According to claim 1, it is characterized in that: A global-fragment fingerprint (GFF) construction scheme is proposed to further improve the characterization ability and prediction accuracy of molecules. The global fragment feature consists of two parts: on the one hand, the overall structural information is obtained by quantifying the global physicochemical properties of the molecule (such as molecular weight, polarity, solubility, etc.); on the other hand, detailed local topological information is obtained by locally fragmenting the molecule and counting the frequency of occurrence of each fragment. The characteristic vector formed by the combination of the two not only reflects the overall physicochemical properties of the molecule, but also captures the subtle differences in the local structure, so that the prediction model can have higher adaptability and accuracy when dealing with different types of molecules.
5. According to claim 1, it is characterized in that: In the above-mentioned global fragmentation features, the construction of fragment fingerprints not only relies on traditional molecular fragmentation technology, but also comprehensively utilizes molecular functional group information. Specifically, by utilizing advanced molecular mapping analysis technology, the molecules are systematically segmented, representative structural fragments are extracted, and these fragments are combined with the chemical properties of molecular functional groups to form a newly designed fingerprint description method. This method can more comprehensively reflect the local activity and global trends of molecules in chemical reactions, and provide more sophisticated theoretical support for the prediction of molecular properties. Experimental results show that this method exhibits high prediction accuracy and stability in various complex molecular systems, showing broad application prospects and significant practical value.
Citation Information
Patent Citations
Compound property prediction method based on molecular structure and physical information fusion
CN119170138A
Graph neural network-based interpretable molecular property prediction method
CN119358591A
Molecular property prediction method based on molecular line graph structure
CN119418821A
Cited By
Polymer property prediction method and device based on substructure and knowledge enhancement
CN121331270A
Multi-feature fusion oral bioavailability prediction method, device, equipment and medium
CN121885236A
Molecular design constraint condition generation method, system and equipment based on interpretable artificial intelligence and medium
CN121922245A