A method for constructing a carcinogenicity prediction model for predicting carcinogenicity
By combining the three-dimensional molecular graph structure and mass spectrometry data characteristics, a comprehensive prediction model of carcinogenicity was constructed, which solved the problem of limited predictive ability of carcinogenicity in the existing technology, and achieved efficient and accurate analysis of the carcinogenicity of compounds.
Patent Information
- Application Number
- CN202211396892.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-11-09
AI Technical Summary
The prior art has limited ability to distinguish between carcinogens and non-carcinogens, and deep learning models also have limitations in predicting carcinogenicity, making it difficult to achieve efficient and accurate assessment of the carcinogenicity of compounds.
By combining three-dimensional molecular graph structure characterization and mass spectrometry data characteristics, a carcinogenic-graph convolutional neural network model and autoencoder model are constructed, and feature matrix fusion and comprehensive training are carried out to build a comprehensive prediction model for carcinogenicity.
Accurate analysis of the carcinogenicity of compounds is achieved, and the performance and accuracy of the carcinogenicity prediction model is improved.
Smart Images

Figure CN115565624B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for analyzing compounds, and specifically to a method for constructing a carcinogenicity prediction model for predicting carcinogenicity. Background Art
[0002] Due to the development of technology, the synthesis speed of new compounds has accelerated, and tens of thousands of compounds are born every year. It is impossible for traditional evaluation methods to efficiently evaluate all compounds; and in recent years, the number of cancer patients has increased sharply, and it is still unclear to which carcinogenic compounds most cancers are caused by exposure. Traditional evaluation of compound carcinogenicity is mainly carried out through experimental tests, with a long test cycle, high cost, and too many uncertain factors. Therefore, there is an urgent need to develop alternative methods and tools to evaluate the carcinogenicity of compounds.
[0003] At present, many methods for predicting compound carcinogenicity based on chemical structure characteristics have been proposed in the prior art. These methods can be roughly divided into: expert rule models, structure-activity relationship (SAR) models, and quantitative structure-activity relationship (QSAR) models. Although the above methods have achieved reasonable prediction capabilities, their abilities to distinguish carcinogens and non-carcinogens are still limited, and there is still much room for improvement. Due to the complexity of carcinogens, the prediction of carcinogenicity by existing deep learning models still has limitations. Therefore, constructing a carcinogenicity prediction model by integrating molecular structures and improving the performance of the carcinogenicity prediction model are urgent problems to be solved at present. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for constructing a carcinogenicity prediction model for predicting carcinogenicity. The construction method combines three-dimensional molecular graph structure characterization and mass spectrometry data features to construct a carcinogenicity prediction model and achieve precise analysis of the carcinogenicity of compounds.
[0005] The technical solution of the present invention to solve the above technical problems is:
[0006] A method for constructing a carcinogenicity prediction model for predicting carcinogenicity, comprising the following steps:
[0007] S1. Collect data of various compounds, screen the data of various compounds, and separately screen out carcinogenicity data and non-carcinogenicity data to construct a carcinogenicity prediction data set;
[0008] S2. Construct a carcinogenicity-graph convolutional neural network model through the molecular graph structure data in the carcinogenicity prediction data set; construct an autoencoder model through the mass spectrometry data in the carcinogenicity prediction data set;
[0009] S3. Feature fusion is performed on the feature matrices finally output by the carcinogenicity - graph convolutional neural network model and the auto - encoder model respectively. The fused feature matrix is comprehensively trained, and finally the prediction result is output. By analyzing the prediction result, when the accuracy requirement is met, a comprehensive carcinogenicity prediction model is obtained.
[0010] Preferably, in step S1, after screening the carcinogenic data and non - carcinogenic data, 463 kinds of compound data are obtained, among which there are 248 carcinogenic data and 215 non - carcinogenic data. The carcinogenic data is used as the positive sample, and the non - carcinogenic data is used as the negative sample. The names of the 463 compounds obtained are saved in a txt file, and the cid number of each compound is separated by a line.
[0011] Preferably, in step S2, the construction steps of the carcinogenicity - graph convolutional neural network are as follows:
[0012] S211. Extract the cid number of each compound in the carcinogenicity dataset and convert it into an sfd file to construct a dataset.
[0013] S212. Divide the dataset into a training set, a validation set, and a test set. Among them, the training set accounts for 80%, the validation set accounts for 10%, and the test set accounts for 10%.
[0014] S213. Input the training set, the validation set, and the test set into the graph convolutional neural network model respectively to construct the carcinogenicity - graph convolutional neural network model.
[0015] Preferably, in step S213, the input of the graph convolutional neural network model is the feature matrix, the adjacency matrix, and the relative position matrix in the compound. Among them, the feature matrix extracts 51 - dimensional features of the atoms in the compound.
[0016] Preferably, the carcinogenicity - graph convolutional neural network includes an embedding module, a feature construction module, an aggregation module, and a fully - connected module, where
[0017] In the embedding module, the adjacency matrix, the relative position matrix, and the feature matrix of the compound molecule are extracted from the sdf file through the rdkit toolkit; self - connection normalization processing is performed on the adjacency matrix; the feature matrix is used as the scalar feature V and embedded into the feature construction module.
[0018] In the feature construction module, the scalar feature V is embedded into the first convolutional layer, and the initial vector feature S is set to zero; in the first convolutional layer, two adjacent scalar features V on each node are combined, that is, the scalar vectors V are pairwise connected to generate an intermediate feature, where the node is an atom in the compound molecule; in the second convolutional layer, all the generated intermediate features are collected and summarized along the neighborhood to produce higher-level features; through two convolutional layers, the scalar feature V and the vector feature S are updated using neighborhood information to achieve the fusion of neighbor information; subsequently, all scalar features V are output using the Relu function, and all vector features S are output using the tanh function.
[0019] In the aggregation module, the MAX aggregation mechanism is used to select the scalar feature V with the maximum value as the molecular feature; when the maximum value of the vector feature S is found, the norm is used for comparison; the MEAN aggregation mechanism is used to calculate the mean of the scalar features V distributed on the nodes to generate new scalar features V and vector features S; the SUM aggregation mechanism is used to sum the scalar features V distributed on the nodes; the aggregation features generated by the MAX aggregation mechanism, the MEAN aggregation mechanism, and the SUM aggregation mechanism are respectively input into the full aggregation mechanism; the global aggregation layer in the full aggregation mechanism performs weighted summation on the three aggregation features to obtain a global embedding, and the generated global embedding enters the fully connected module.
[0020] In the fully connected module, the generated global embedding passes through the first fully connected layer to obtain the final molecular feature. After this molecular feature passes through the second fully connected layer, the generated molecular feature is sent to a fully connected neural network with a sigmoid activation function for prediction; the scalar feature V is sent into a two-layer fully connected neural network, while the vector feature S is sent into a hierarchical fully connected neural network to suppress the separation between axes during the linear combination process.
[0021] Preferably, in step S2, the autoencoder model adopts a stacked autoencoder structure, and this autoencoder model consists of an input layer, a hidden layer, and an output layer; and a deep autoencoder with 3 hidden layers is adopted in the autoencoder model.
[0022] Preferably, in step S2, the construction method of the autoencoding model is as follows:
[0023] S221. Perform wavelet denoising on the mass spectrometry data in the carcinogenicity dataset, then perform data baseline correction on it, and extract the corresponding mass spectrometry dataset.
[0024] S222. Preprocess the mass spectrometry dataset.
[0025] S223. Select a stacked autoencoder and use the sigmoid function as the activation function for activation.
[0026] S224. Feature extraction is performed through three hidden layers, and the node parameters of the hidden layers are adjusted simultaneously to achieve the expected performance, and the feature matrix Z is output.
[0027] Preferably, in step S222, the processing of the mass spectrometry data set includes cleaning of dirty data and filling of missing values.
[0028] Preferably, in step S224, the feature extraction through three hidden layers includes the following steps:
[0029] (1) Based on the initially given input X, the first autoencoder performs unsupervised training on the first hidden layer V, and through the minimization operation of the reconstruction error, the set value of the input X and the reconstructed output X' is kept unchanged;
[0030] (2) The hidden layer V extracted from the first autoencoder is used as the input of the next autoencoder, and the next hidden layer is trained in the same steps;
[0031] (3) Step (2) is cycled until all autoencoders complete the initialization task;
[0032] (4) The output of the finally trained hidden layer is extracted as the input of the fully connected layer.
[0033] Preferably, in step S4, the feature matrix X output by the carcinogenicity - graph convolutional neural network and the feature matrix Z output by the autoencoder model are input into the first fully connected layer for feature fusion. The feature matrices X and Z are fused by using a concate layer, and then the fused feature matrix is input into the second fully connected layer for comprehensive training. The prediction result of the comprehensive molecular structure prediction model is output, and the performance of the carcinogenicity comprehensive prediction model is analyzed through three indicators: precision, recall, and Auc - Roc.
[0034] The present invention has the following beneficial effects compared with the prior art:
[0035] (1). The construction method of the carcinogenicity prediction model for predicting carcinogenicity of the present invention is used to construct a carcinogenicity comprehensive prediction model. The carcinogenicity comprehensive prediction model combines the three - dimensional molecular graph structure representation and the mass spectrometry data features to construct a carcinogenicity prediction model, realizing accurate analysis of the carcinogenicity of compounds.
[0036] (2). The carcinogenicity comprehensive prediction model constructed by the construction method of the carcinogenicity prediction model for predicting carcinogenicity of the present invention can realize accurate analysis of the carcinogenicity of compounds and has higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic flowchart of a method for constructing a carcinogenicity prediction model for predicting carcinogenicity according to the present invention.
[0038] Figure 2 It is a schematic flowchart of the construction steps of a carcinogenicity - graph convolutional neural network model in the present invention.
[0039] Figure 3 It is a schematic flowchart of the construction steps of an auto - encoder model in the present invention. Detailed implementation manners
[0040] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.
[0041] See Figures 1 - 3 , the method for constructing a carcinogenicity prediction model for predicting carcinogenicity according to the present invention includes the following steps:
[0042] S1. Collect data of multiple compounds, screen the data of multiple compounds, and separately screen out carcinogenicity data and non - carcinogenicity data to construct a carcinogenicity prediction data set.
[0043] S2. Construct a carcinogenicity - graph convolutional neural network model through the molecular graph structure data in the carcinogenicity prediction data set; construct an auto - encoder model through the mass spectrometry data in the carcinogenicity prediction data set.
[0044] S3. Perform feature fusion on the feature matrices finally output by the carcinogenicity - graph convolutional neural network model and the auto - encoder model respectively, perform comprehensive training on the fused feature matrix, and finally output a prediction result; by analyzing the prediction result, when the accuracy requirement is met, a comprehensive carcinogenicity prediction model is obtained.
[0045] Among them, in step S1, 463 kinds of carcinogenicity and non - carcinogenicity data are screened out from authoritative institutions; after screening the carcinogenicity data and non - carcinogenicity data, 463 kinds of compound data are obtained, among which there are 248 carcinogenicity data and 215 non - carcinogenicity data; the carcinogenicity data is used as positive samples and the non - carcinogenicity data is used as negative samples; the names of the 463 compounds obtained are saved in a txt file, and the cid numbers of each compound are separated by rows.
[0046] See Figures 1 - 3 , in step S2, the construction steps of the carcinogenicity - graph convolutional neural network are as follows:
[0047] S211. Extract the cid number of each compound in the carcinogenicity data set and convert it into an sfd file to construct a data set.
[0048] S212. Divide the data set into a training set, a validation set, and a test set, where the training set accounts for 80%, the validation set accounts for 10%, and the test set accounts for 10%.
[0049] S213. Input the training set, the validation set, and the test set into the graph convolutional neural network model respectively to construct a carcinogenicity-graph convolutional neural network model.
[0050] See Figures 1 - 3 , in step S213, the inputs of the graph convolutional neural network model are the feature matrix, the adjacency matrix, and the relative position matrix in the compound, where the feature matrix extracts 51-dimensional features of the atoms in the compound.
[0051] See Figures 1 - 3 , the carcinogenicity-graph convolutional neural network includes an embedding module, a feature construction module, an aggregation module, and a fully connected module, where
[0052] In the embedding module, extract the adjacency matrix, the relative position matrix, and the feature matrix of the compound molecule from the sdf file through the rdkit toolkit; perform self-connection normalization on the adjacency matrix. If normalization is not performed, the original feature distribution will be changed; use the feature matrix as the scalar feature V and embed it into the feature construction module;
[0053] In the feature construction module, embed the scalar feature V into the first convolutional layer and set the initial vector feature S to zero; in the first convolutional layer, combine the two adjacent scalar features V on each node, that is, connect the scalar vectors V pairwise to generate an intermediate feature, where the node is an atom in the compound molecule; in the second convolutional layer, collect all the generated intermediate features and summarize them along the neighborhood to generate higher-level features; through two convolutional layers, and use neighborhood information to update the scalar feature V and the vector feature S to achieve the fusion of neighborhood information; then use the Relu function to output all the scalar features V and use the tanh function to output all the vector features S;
[0054] In the aggregation module, adopt the following aggregation mechanism:
[0055] (1). Adopt Max Pooling (i.e., the maximum pooling layer) to select the scalar feature V with the maximum value as the molecular feature; when finding the maximum value of the vector feature S, use the norm for comparison, that is, in the experiment, feature selection is performed according to the Pearson correlation coefficient; the Pearson correlation coefficient is equivalent to the standardization of the covariance. The range of the Pearson correlation coefficient is between -1 and 1. When the value is closer to -1 and 1, it indicates that there is an obvious linear relationship between the two variables, and only one of them needs to be retained. When the value is closer to 0, it indicates that the correlation between the two variables is weaker;
[0056] (2) Use Mean Pooling (i.e., average pooling layer) to calculate the mean of the scalar features V distributed on the nodes, generating new scalar features V and vector features S;
[0057] (3) Use Sum Pooling (summation pooling layer) to sum the scalar features V distributed on the nodes; to make up for the information loss generated after each aggregation mechanism, a new hybrid aggregation mechanism is formed by splicing the above three aggregation mechanisms. First, the aggregation features generated by Max Pooling, Mean Pooling, and Sum Pooling are input into the full aggregation mechanism; the global aggregation layer in the full aggregation mechanism performs weighted summation on the three aggregation features to obtain the global embedding, and the generated global embedding enters the fully connected module;
[0058] In the fully connected module, the generated global embedding passes through the first fully connected layer and outputs the final molecular features, and these molecular features pass through the second fully connected layer; the generated molecular features are sent to a fully connected neural network with a sigmoid activation function for prediction; the scalar features V are sent into a two-layer fully connected neural network, while the vector features S are sent into a hierarchical fully connected neural network to suppress the separation between axes during the linear combination process.
[0059] See Figures 1 - 3 , in step S2, the autoencoder model adopts a stacked autoencoding structure, and this autoencoder model consists of an input layer, a hidden layer, and an output layer; and a deep autoencoder with 3 hidden layers is adopted in the autoencoder model.
[0060] See Figures 1 - 3 , in step S2, the construction method of the autoencoding model is as follows:
[0061] S221. Perform wavelet denoising on the mass spectrometry data in the carcinogenicity dataset, then perform data baseline correction on it, and extract the corresponding mass spectrometry dataset;
[0062] S222. Preprocess the mass spectrometry dataset;
[0063] S223. According to the constructed dataset and the requirements for the training results, select a stacked autoencoder and use the sigmoid function as the activation function for activation; using a stacked autoencoder for mass spectrometry data can train each layer separately while ensuring controllable dimensionality reduction, simplifying complex problems, which is conducive to accelerating the completion of tasks and has excellent feature extraction effects;
[0064] S224. Perform feature extraction through 3 hidden layers, and at the same time adjust the node parameters of the hidden layer to achieve the expected performance, and output the feature matrix Z.
[0065] See Figures 1 - 3 , in step S222, the processing of the mass spectrometry data set includes cleaning dirty data and filling missing values.
[0066] See Figures 1 - 3 , in step S224, the feature extraction through three hidden layers includes the following steps:
[0067] (1) Based on the initially given input X, the first autoencoder performs unsupervised training on the first hidden layer V, and through the minimization operation of the reconstruction error, the set value of the input X and the reconstructed output X' is kept unchanged;
[0068] (2) Take the hidden layer V extracted from the first autoencoder as the input of the next autoencoder, and then train the next hidden layer in the same steps;
[0069] (3) Loop through step (2) until all autoencoders complete the initialization task;
[0070] (4) Extract the output of the finally trained hidden layer (i.e., the feature matrix Z) as the input of the fully connected layer.
[0071] See Figures 1 - 3 , in step S4, input the feature matrix X output by the carcinogenicity - graph convolutional neural network and the feature matrix Z output by the autoencoder model into the first - layer fully connected layer for feature fusion. By using a concate layer to fuse the feature matrix X and the feature matrix Z, and then input the fused feature matrix into the second - layer fully connected layer for comprehensive training, output the prediction result of the comprehensive molecular structure prediction model, and analyze the performance of the comprehensive molecular structure prediction model through three indicators: Precision, Recall, and Auc - Roc (the area under the ROC curve, Area Under Curve).
[0072] The above is a preferred embodiment of the present invention. However, the embodiments of the present invention are not limited by the above content. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for constructing a carcinogenicity prediction model for predicting carcinogenicity, characterized in that, It includes the following steps: S1. Collect data of multiple compounds, screen the data of multiple compounds, and separately screen out carcinogenic data and non-carcinogenic data, so as to construct a carcinogenicity prediction data set; S2. Construct a carcinogenicity-graph convolutional neural network model through the molecular graph structure data in the carcinogenicity prediction data set; construct an autoencoder model through the mass spectrometry data in the carcinogenicity prediction data set; The autoencoder model adopts a stacked autoencoding structure, and this autoencoder model consists of an input layer, a hidden layer, and an output layer; and a deep autoencoder with 3 hidden layers is adopted in the autoencoder model, and the construction method of this autoencoder model is: S221. Perform wavelet denoising on the mass spectrometry data in the carcinogenicity prediction data set, then perform data baseline correction on it, and extract the corresponding mass spectrometry data set; S222. Preprocess the mass spectrometry data set; S223. Select a stacked autoencoder and use the sigmoid function as the activation function for activation; S224. Extract features through 3 hidden layers, and at the same time adjust the node parameters of the hidden layer to achieve the expected performance, and output the feature matrix Z; S3. Perform feature fusion on the feature matrices finally output by the carcinogenicity-graph convolutional neural network model and the autoencoder model respectively, perform comprehensive training on the fused feature matrix, and finally output the prediction result; by analyzing the prediction result, when the accuracy requirement is met, a comprehensive carcinogenicity prediction model is obtained.
2. The method for constructing a carcinogenicity prediction model for predicting carcinogenicity according to claim 1, characterized in that In step S1, after screening the carcinogenic data and non-carcinogenic data, 463 kinds of compound data are obtained, among which there are 248 carcinogenic data and 215 non-carcinogenic data; the carcinogenic data is used as the positive sample, and the non-carcinogenic data is used as the negative sample; the names of the 463 compounds obtained are saved to a txt file, and the cid numbers of each compound are separated in rows.
3. The method for constructing a carcinogenicity prediction model for predicting carcinogenicity according to claim 1, characterized in that, In step S2, the construction steps of the carcinogenicity-graph convolutional neural network are: S211. Extract the cid number of each compound in the carcinogenicity data set and convert it into an sfd file, so as to construct a data set; S212. Divide the data set into a training set, a validation set, and a test set, where the training set accounts for 80%, the validation set accounts for 10%, and the test set accounts for 10%; S213. Input the training set, the validation set, and the test set into the graph convolutional neural network model respectively, so as to construct a carcinogenicity-graph convolutional neural network model.
4. The method for constructing a carcinogenicity prediction model for predicting carcinogenicity according to claim 3, characterized in that, In step S213, the input of the graph convolutional neural network model is the feature matrix, adjacency matrix, and relative position matrix in the compound, where the feature matrix extracts 51-dimensional features of the atoms in the compound.
5. The method for constructing a carcinogenicity prediction model for predicting carcinogenicity according to claim 4, characterized in that, The carcinogenicity-graph convolutional neural network includes an embedding module, a feature construction module, an aggregation module, and a fully connected module, where In the embedding module, extract the adjacency matrix, relative position matrix, and feature matrix of the compound molecule from the sdf file through the rdkit toolkit; perform self-connection normalization processing on the adjacency matrix; use the feature matrix as the scalar feature V and embed it into the feature construction module; In the feature construction module, the scalar feature V is embedded into the first convolutional layer, and the initial vector feature S is set to zero; in the first convolutional layer, two adjacent scalar features V on each node are combined, that is, the scalar vectors V are pairwise connected to generate an intermediate feature, where the node is an atom in the compound molecule; in the second convolutional layer, all the generated intermediate features are collected and summarized along the neighborhood to produce higher-level features; through two convolutional layers, the scalar feature V and the vector feature S are updated using neighborhood information to achieve the fusion of neighbor information; subsequently, all scalar features V are output using the Relu function, and all vector features S are output using the tanh function; In the aggregation module, the MAX aggregation mechanism is used to select the scalar feature V with the maximum value as the molecular feature; when the maximum value of the vector feature S is found, the norm is used for comparison; the MEAN aggregation mechanism is used to calculate the mean of the scalar features V distributed on the nodes to generate new scalar features V and vector features S; the SUM aggregation mechanism is used to sum the scalar features V distributed on the nodes; the aggregation features generated by the MAX aggregation mechanism, the MEAN aggregation mechanism, and the SUM aggregation mechanism are respectively input into the full aggregation mechanism; the global aggregation layer in the full aggregation mechanism performs weighted summation on the three aggregation features to obtain a global embedding, and the generated global embedding enters the fully connected module; In the fully connected module, after the generated global embedding passes through the first fully connected layer, the final molecular feature is obtained. After this molecular feature passes through the second fully connected layer, the generated molecular feature is sent to a fully connected neural network with a sigmoid activation function for prediction; the scalar feature V is sent into a two-layer fully connected neural network, while the vector feature S is sent into a hierarchical fully connected neural network to suppress the separation between axes during the linear combination process.
6. The method for constructing a carcinogenicity prediction model for predicting carcinogenicity according to claim 1, characterized in that, In step S222, the processing of the mass spectrometry dataset includes dirty data cleaning and missing value filling.
7. The method for constructing a carcinogenicity prediction model for predicting carcinogenicity according to claim 1, characterized in that In step S224, the feature extraction through 3 hidden layers includes the following steps: (1) Based on the initially given input X, the first autoencoder performs unsupervised training on the first hidden layer V, and through the minimization operation of the reconstruction error, the set value of the input X and the reconstructed output X' is kept unchanged; (2) The hidden layer V extracted from the first autoencoder is used as the input of the next autoencoder, and the next hidden layer is trained in the same steps; (3) Step (2) is looped until all autoencoders complete the initialization task; (4) The output of the finally trained hidden layer is extracted as the input of the fully connected layer.
8. The method for constructing a carcinogenicity prediction model for predicting carcinogenicity according to claim 1, wherein In step S4, the feature matrix X output by the carcinogenicity-graph convolutional neural network and the feature matrix Z output by the autoencoder model are input into the first fully connected layer for feature fusion. The feature matrices X and Z are fused by using a concatenation layer, and then the fused feature matrix is input into the second fully connected layer for comprehensive training to output the prediction result of the comprehensive molecular structure prediction model. The performance of the carcinogenicity comprehensive prediction model is analyzed through three indicators: precision, recall, and Auc-Roc.
Citation Information
Patent Citations
Residual map convolutional neural network-based RNA-protein binding site discrimination method
CN113241117A
Anticancer drug screening method based on multichannel neural network
CN114496303A