Chemical active structure identification method based on Morga fingerprint and optimized Shapley value

By introducing optimized Shapley values ​​and activity index into the SHAP method, the method of identifying the active structure of chemicals is solved by the problem of existing methods ignoring a small number of high-active structures, and improving the reliability of the QSAR model and the ability to discover the active structure of new chemicals.

CN119993324AActive Publication Date: 2025-05-13TIANJIN UNIV

Patent Information

Application Number
CN202510064630.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The existing SHAP method easily ignores a small number of highly active structures when identifying chemical active structures, resulting in incomplete reliability analysis of QSAR model and discovery of new chemical active structures.

Method used

Using a chemical active structure recognition method based on Morgan fingerprint and optimized Shapley values, the activity index is defined to discover highly active molecular fragments by generating a feature matrix, calculating the Shapley value matrix and removing interfering molecular structures.

Benefits of technology

The analysis and interpretation methods of the QSAR model are improved, and a small number of highly active structures that are present can be more accurately identified, improving the reliability of the model and the ability to discover the active structure of new chemicals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993324A_ABST
    Figure CN119993324A_ABST
Patent Text Reader

Abstract

The invention discloses a chemical active structure identification method based on a Morga fingerprint and an optimized Shapley value. The method comprises the following steps of: generating a feature matrix for a training data set for a model established by using Morga molecular fingerprints as features, loading a model weight by taking the feature matrix as an explanation background, explaining the model by using an SHAP method, generating a Shapley matrix, searching and deleting an interference column, finding a substructure which appears at low frequency and is highly related to an activity end point of a chemical through definition of an activity index, and carrying out chemical activity detection on the substructure. Highly active molecular fragments are found.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of QSAR model application oriented to the environment field, and specifically is a method for identifying the active structure of chemicals based on Morgan fingerprints and optimized Shapley values. Background Art

[0002] The quantitative structure-activity relationship (QSAR) model is a method that associates the activity of a molecule with its chemical structure. When the chemical structure is known, the properties of the molecule can be predicted. Compared with traditional experimental measurement methods, the QSAR method has the advantages of low cost, fast screening, and strong extrapolation. Morgan fingerprint is an efficient and widely used molecular structure description tool in cheminformatics. It is generally generated with the help of simplified molecular linear input specification (SMILES). It generates a fixed-length binary or count vector to record molecular features by iteratively expanding the atomic neighborhood in the molecular graph and hashing technology. Each bit of the vector corresponds to a unique molecular fragment, and its existence or quantity is equal to the characteristic value of the corresponding position in the Morgan fingerprint. In recent years, with the improvement of computing power, algorithms and data sets, more and more QSAR models based on Morgan fingerprints have been established and widely used in various fields such as drug discovery, reaction rate prediction, and hazardous materials screening.

[0003] It is necessary to interpret QSAR models and identify the active structures of chemicals because it can be used to verify whether the QSAR model has really learned the chemical laws and is reliable. In addition, the interpretation of the discovered active structures of new chemicals can also be used to discover new reaction patterns. The SHAP (SHapley Additive exPlanations) method is a method based on the Shapley value theory in game theory to interpret machine learning models. It can capture the nonlinear relationship between features by evaluating the impact of removing specific features on model performance and is widely used to interpret QSAR models. However, this method has certain limitations. The principle of the SHAP method is to prioritize the discovery of features that lead to greater differences in the predicted values ​​of the entire data set, which requires that the feature in the data set has greater differences. Taking the binary Morgan fingerprint composed only of 0 and 1 as an example, when the ratio of 0 and 1 is half each, that is, when the existence of the substructure is closer to 50%, it is easier to be discovered by the SHAP method and evaluated as an important substructure. However, the active structure of chemicals does not necessarily exist in large quantities. Taking the screening of highly toxic substructures of chemicals as an example, the screening target is the substructures whose overall toxicity of the compound will also change significantly when the existence of the substructure changes slightly. The content of such substructures is usually small, but their discovery is necessary. At this time, the SHAP method will assign higher toxicity rankings to low-toxic structures that appear frequently, while the truly highly toxic substructures will be ignored by the traditional SHAP method. Therefore, the active structures discovered using the SHAP method are not comprehensive, and there is a limitation that it is difficult to discover highly active structures that exist in small quantities.

[0004] In summary, there is an urgent need for a chemical active structure identification method based on Morgan fingerprints and optimized Shapley values, which is of great significance for the reliability analysis of QSAR models and the discovery of new chemical active structures. Summary of the invention

[0005] In order to solve the technical problems raised in the background technology, the present invention proposes a method for identifying active structures of chemicals based on Morgan fingerprints and optimized Shapley values, and discovers highly active structures that exist in small quantities, which provides a basis for reliability analysis of QSAR models and discovery of new active structures of chemicals.

[0006] The technical solution of the present invention is: a method for identifying the active structure of chemicals based on Morgan fingerprints and optimized Shapley values, comprising the following steps:

[0007] 1) Data collection: Collect public quantitative structure-activity relationship models based on Morgan fingerprints or use public data sets to build quantitative structure-activity relationship models based on Morgan fingerprints, and obtain model-related files, including loadable model weight files, the radius and length of the Morgan fingerprint used in the model, and the SMILES code of the training data set used to build the model.

[0008] 2) Generate feature matrix: With the help of RDKit library, generate feature matrix according to the SMILES encoding of molecules in the training data set and the Morgan fingerprint radius and length used by the model.

[0009] 3) Load the model and calculate the Shapley value matrix: According to the model type, use the shap library (https: / / github.com / slundberg / shap) to explain the model, use the feature matrix generated by the training data set as the explanation background, and calculate the Shapley matrix of the training data set. Specifically, when the explanation object is a neural network model, use the DeepExplainer module to explain the model. When the explanation object is a tree-based model such as a decision tree, random forest, or gradient boosting tree, use the TreeExplainer module to explain the model. When the explanation object is a model established by other methods such as k-nearest neighbor and support vector machine, use the KernelExplainer module to explain the model.

[0010] 4) Remove molecular structures that may interfere: To avoid the model's overfitting of very few existing features that affects the analysis results, and to avoid the situation where the denominator is 0 when calculating the activity index in columns where all features are 0, the interference structure is deleted. The method for determining the interference structure is that if the ratio of the number of non-zero elements in the feature column corresponding to the structure in the feature matrix to the total number of samples in the training data set is less than a preset ratio, such as 1-3%, then the column is judged to be the column where the interference structure is located, and the column index is added to the position index set of the interference column. According to the index set, the interference columns of the Shapley matrix and the feature matrix are deleted: when the column position index of the Shapley matrix and the feature matrix exists in the position index set of the interference column, it is judged to be an interference column, and the column of the two matrices is deleted at the same time.

[0011] 5) By defining the activity index, highly active molecular fragments are discovered. The Shapley value of any j-th feature of any sample i can be approximated as follows: randomly extract sample k with replacement from the training set; for sample i, keep the j-th feature unchanged, and randomly select the remaining features in sample k and sample i to generate a virtual sample i'; calculate the difference between the activity endpoints of sample i and virtual sample i' predicted by the model. Repeat this process 1000 times and calculate the average of these differences, which is the Shapley value of the j-th feature of sample i;

[0012] If we ignore the random replacement of other features and the possible mutual influence between feature j and other features, the Shapley value is completely caused by the difference of feature j between sample i and sample k. After repeated sampling, the mean of feature j in the set generated by sample k will approach the mean of feature j in the entire training set, which means that the difference between the value of feature j in sample i and the mean of feature j in the training set leads to its Shapley value. Assuming that this change is linear and using division to measure the steepness of this change, the definition of activity index is proposed. For feature j, its activity index AI j The formula is expressed as:

[0013]

[0014] Where n is equal to the number of samples in the training set, S ij represents the value of the i-th sample and the j-th feature in the Shapley value matrix of the training data set, F ij represents the value of the i-th sample and the j-th feature in the feature matrix of the training data set, and f ave j Represents the average value of the jth feature in the training set. Activity Index AI j The value of is approximately the average change in activity value after adding a j feature to each sample in the training set.

[0015] Beneficial Effects

[0016] The present invention further identifies the active structure using the optimized Shapley value, and improves the analysis and interpretation method of the QSAR model. The small amount of highly active structures discovered by this method can provide a basis for the reliability analysis of the QSAR model and the discovery of new active mechanisms.

[0017] 1. The analysis and interpretation methods of the model were improved. By introducing the average value of the number of substructures into the activity index formula, the disadvantage of the original SHAP method that the low-frequency active structures of chemicals are easily ignored was overcome.

[0018] 2. Using the training dataset as the explanation background reduces the possibility of underfitting the original model in other datasets and improves the accuracy of the analysis method.

[0019] 3. By deleting the interference columns, the possibility of interference caused by overfitting of the original model is reduced, the risk of discovering irrelevant structures when using the activity index method is reduced, and the reliability of the analysis method is improved.

[0020] 4. The highly active substructures discovered through the activity index are conducive to the reliability analysis of the model and the discovery of new reaction mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a bar chart of the average absolute values ​​of the SHAP values ​​of each feature in the fish toxicity training set, that is, the feature importance evaluated by the original SHAP method.

[0022] Figure 2 This is the information related to the top 12 substructures of the fish toxicity activity index.

[0023] Figure 3 It is a bar chart of the average absolute values ​​of the SHAP values ​​of each feature of the secondary rate constants dataset for the reaction of free chlorine and organic matter, that is, the feature importance evaluated by the original SHAP method.

[0024] Figure 4 It is the information related to the top 12 substructures of the secondary rate constant activity index of the reaction between free chlorine and organic matter. DETAILED DESCRIPTION

[0025] The present invention is further described below by specific examples and drawings. The examples of the present invention are intended to enable those skilled in the art to better understand the present invention, and do not limit the present invention in any way.

[0026] Example 1

[0027] The feasibility of using the model to determine the toxicity of the two substructures represented by feature 1380 (carbon on the benzene ring) and feature 226 (nitrogen-nitrogen double bond) to fish was determined. Based on the activity coefficient, the steps are as follows:

[0028] Collect data. Includes the neural network model weight file for predicting the median lethal concentration (LC50, mol / L) of fish exposed to 908 different compounds (Cassotti, M.; Ballabio, D.; Todeschini, R.; Consonni, V. A similarity-based QSAR model for predicting acute toxicity towards the fathead minnow (Pimephales promelas). SAR QSAR Environ. Res. 2015, 26, 217-243.), the radius and length of the Morgan fingerprint used are 1 and 2048 respectively, and the SMILES encoding of the training set consisting of 726 (908*0.8=726) data.

[0029] Calculate the feature matrix: Generate Morgan fingerprints with a radius and length of 1 and 2048 respectively for the 726 chemicals in the training set to form a feature matrix.

[0030] Calculate the Shapley matrix: Because the weight file is a neural network model, use the shap.DeepExplainer module to interpret the feature matrix and calculate the Shapley matrix of the training data set.

[0031] Determine whether the structure is an interfering substructure that should be removed: When the training set size is 726, the number of molecules with substructures should be greater than or equal to 8 (726*0.01=7.26), that is, the number of non-zero elements in the corresponding column of the feature matrix generated by the training set should be greater than or equal to 8. The 1380th and 226th columns of the feature matrix are queried, and the number of non-zero elements in the 1380th column of the feature matrix is ​​counted. The number of non-zero elements in the 1380th column of the feature matrix is ​​greater than 8, so the activity coefficient of the carbon on the benzene ring can be calculated, while the number of non-zero elements in the 226th column of the feature matrix is ​​less than 8, which is judged to be an interfering structure, so the activity coefficient of the nitrogen-nitrogen double bond cannot be calculated.

[0032] Calculate the activity index of feature 1380: intercept the 1380th column of the feature matrix and the Shapley matrix respectively, and obtain two lists, which respectively mean the number of structures represented by feature 1380 in the 726 samples of the training set and the Shapley value of this feature in different samples; average the intercepted feature list to obtain 1.71, that is, each molecule in the training set contains 1.71 structures on average; after querying, in the first sample, the Shapley value is 0.6 and the number is 4. At this time, according to the definition of the activity index, the calculated activity value is 0.6 / (4-1.71)=0.26; after calculating in this way in all 726 samples, the average of the 726 activity values ​​is calculated, and the final activity index is 0.225. This means that after adding one feature to all molecules in the training set, the average value of the increase in toxicity is about 0.225 (in terms of -lgLC50, in mol / L).

[0033] Example 2

[0034] Find molecular structures that are highly toxic to fish. The method based on activity index is as follows:

[0035] According to the definition of activity index, the activity index of all molecular structures with a number greater than 7 in the training set was calculated, and the first 12 of them were selected. Figure 1 As shown in the figure, the information related to the top 12 substructures in terms of activity index includes the position of the feature in the molecular fingerprint, the feature meaning, the importance based on the basic SHAP method, and the value of its activity index. For comparison, a bar chart based on the importance ranking based on the original SHAP method is drawn, as shown in Figure 2 As shown. After the definition of the activity index, toxic structures that were originally given lower importance by the SHAP method due to their rare presence were found. These structures include substructures containing sulfur and phosphorus elements (features 97, 116, 192 and 1729), bromine atoms (feature 728), lipid parts (feature 1386), polycyclic substituents (feature 352), unsaturated carbon bonds (features 915, 1366 and 1645), nitrogen-containing ring substituents (feature 1145) and nitro groups (feature 715). They need to be focused on, continue to require subsequent experimental verification, need to be handled with caution in industrial applications and daily use, and actively seek safer alternatives. Among these structures, sulfur and phosphorus elements are widely present in pesticides and insecticides, halogen atoms, unsaturated carbon, etc. are prone to nucleophilic substitution reactions due to high electronegativity, and polycyclic aromatic hydrocarbons have also been shown to be associated with high carcinogenicity. They were ranked very low in the initial SHAP method, but they were discovered by the activity index method of the present invention, indicating the effectiveness of the analytical method of the present invention.

[0036] Example 3

[0037] Find molecular structures that significantly increase the secondary rate constant for the reaction of free chlorine with organic matter. The activity index-based method is as follows:

[0038] Obtain data. The data used to establish the model include 177 data collected from the literature (Deborde, M.; von Gunten, U. Reactions of chlorine with inorganic and organic compounds during water treatment-Kinetics and mechanisms: a critical review. Water Res 2008, 42, (1-2), 13-51.) and 177 data generated by experimental methods, a total of 354 data. The reaction rate is specified to be greater than 100M -1 s -1 The reaction rate is fast, less than 1M -1 s -1 The reaction rate is slow, and the reaction rate between the two is medium. The training data set and the test set are divided into 8:2 ratios, and the optimal classification model is established. Among them, the final model established is a random forest model, using the radius of Morgan fingerprint as 2 and the length as 2048. The training set consists of 283 (354*0.8=283) data.

[0039] Calculate the feature matrix. Generate Morgan fingerprints with a radius of 2 and a length of 2048 for the molecules in the training set as the feature matrix.

[0040] Calculate the Shapley matrix, load the random forest model, use the feature matrix of the training set as the explanation background, use the shap.TreeExplainer function to explain the model, and calculate the Shapley matrix of the model.

[0041] According to the position index set of the interference column, the columns of the Shapley matrix and the characteristic matrix are deleted. According to the judgment method of the interference column, the number of non-zero features of each column in the characteristic matrix should be greater than or equal to 3 (283*0.01=2.83). A total of 1681 interference columns with column position index [0,2,3,4...,2038] are screened out. Finally, 367 (2048-1681=367) columns are retained in both the Shapley matrix and the characteristic matrix.

[0042] Active structure discovery was performed according to the definition of activity index. The activity index values ​​of 367 unique substructures were calculated, and these activity index values ​​were sorted from large to small, and the top 12 were selected. The results are as follows Figure 3As shown in the figure, the information related to the top 12 substructures in terms of activity index includes the position of the feature in the molecular fingerprint, the feature meaning, the importance based on the basic SHAP method, and the value of its activity index. For comparison, a bar chart based on the importance ranking based on the original SHAP method is drawn, as shown in Figure 4 As shown. After the definition of the activity index, active structures that were originally given less importance by the SHAP method due to their rare existence were found, including amine structures (features 444, 735, 150, 193), thioether structures (feature 116) and isopropylbenzene substructures (feature 598). Studies have shown that chlorine reacts rapidly with electron-donating groups (Deborde, M.; von Gunten, U. Reactions of chlorine with inorganic and organic compounds during watertreatment-Kinetics and mechanisms: a critical review. Water Res 2008, 42, (1-2), 13-51.), and amine structures and thioether structures meet this feature; under light conditions, the α carbon of the side chain of isopropylbenzene can also react with chlorine radicals to generate C6H5-CCl(CH3)2. They were ranked very low in the initial SHAP method, but they were discovered by the activity index method of the present invention, indicating the effectiveness of the analytical method of the present invention.

[0043] It should be understood that the embodiments and examples discussed here are for illustrative purposes only and may be modified or altered by those skilled in the art, and all such modifications and alterations should fall within the scope of protection of the appended claims of the present invention.

Claims

1. A method for identifying active structures of chemicals based on Morgan fingerprints and optimized Shapley values, characterized in that: The following steps are involved: 1) Collect data; 2) Generate feature matrix: Generate feature matrix according to SMILES encoding in training data set and Morgan fingerprint radius and length used by the model; 3) Load the model and calculate the Shapley value matrix: According to the model type, use the shap library to interpret the model, use the feature matrix generated by the training data set as the interpretation background, and calculate the Shapley matrix of the training data set; 4) Remove molecular structures that may interfere: search for the columns where the interfering structures exist in the feature matrix according to the preset ratio, record the positions of these columns, form a position index set of the interfering columns, and delete the interfering columns of the Shapley matrix and the feature matrix; 5) Discover highly active molecular fragments by defining the activity index: For feature j, the activity index AI of the structure is represented j , its formula is expressed as: Where n is equal to the number of samples in the training data set, S ij represents the value of the i-th sample and the j-th feature in the Shapley value matrix of the training data set, F ij represents the value of the i-th sample and the j-th feature in the feature matrix of the training data set, f avej Represents the average value of the jth feature in the feature matrix of the training dataset.

2. The method according to claim 1, characterized in that In the step 1), publicly available quantitative structure-activity relationship models based on Morgan fingerprints are collected or publicly available data sets are used to establish quantitative structure-activity relationship models based on Morgan fingerprints, and model-related files are obtained, including loadable model weight files, the radius and length of the Morgan fingerprint used in the model, and the SMILES code of the training data set used to establish the model.

3. The method according to claim 1, characterized in that In the step 2), with the help of the RDKit library, a feature matrix is ​​generated according to the SMILES encoding of the molecules in the training data set and the Morgan fingerprint radius and length used by the model.

4. The method according to claim 1, characterized in that: In the step 3), according to the model type, the shap library is used to explain the model, and the feature matrix generated by the training data set is used as the explanation background to calculate the Shapley matrix of the training data set.

5. The method according to claim 1, characterized in that: In the step 4), the molecular structures that may interfere are removed. The method for determining the interfering structure is that, in the feature column corresponding to the structure in the feature matrix, if the ratio of the number of non-zero elements to the total number of samples in the training data set is less than a preset ratio, the column is determined to be the column where the interfering structure is located, and the column index is added to the position index set of the interfering column. According to the index set, the interfering columns of the Shapley matrix and the feature matrix are deleted: when the column position index of the Shapley matrix and the feature matrix exists in the position index set of the interfering column, the columns of the two matrices are deleted at the same time.

6. The method according to claim 1, characterized in that In step 5), the Shapley value of any j-th feature of any sample i can be approximated as follows: randomly extract sample k with replacement in the training set; for sample i, keep the j-th feature unchanged, and randomly select values ​​of the remaining features in sample k and sample i to generate a virtual sample i'; calculate the difference between the active endpoints of sample i and sample i' predicted by the model, repeat this process 1000 times, and calculate the average of these differences, which is the Shapley value of the j-th feature of sample i; If the random replacement of other features and the possible mutual influence between feature j and other features are ignored, the Shapley value is completely caused by the difference in feature j between sample i and extracted sample k; after repeated sampling, the mean of feature j in the set generated by sample k will approach the mean of feature j in the entire training set, which means that the difference between the value of feature j in sample i and the mean of feature j in the training set leads to its Shapley value; assuming that this change is linear and using division to measure the steepness of this change, the activity index is proposed to be the average value of the change in sample activity in the training set after adding a feature j to each sample in the training set.

Citation Information

Patent Citations

  • Method for identifying types of rare wood

    CN117871650A

  • Millimeter wave radar point cloud class incremental learning method based on diversified sample playback

    CN118171717A

  • Sample-difference-based method and system for interpreting deep-learning model for code classification

    US20240192929A1

Cited By

  • Method for screening phytogenic saponin with lipase inhibitory activity and derivative thereof based on R group contribution value and application of phytogenic saponin and derivative thereof

    CN120617274A