A deep learning method, device and storage medium for evaluating carcinogenic risk of a compound based on multi-modal data fusion

By constructing a cross-modal knowledge graph and using deep learning methods, and integrating multimodal data feature representations, the accuracy and efficiency issues of compound carcinogenic risk assessment were solved. This enabled efficient assessment and mechanism explanation of compound carcinogenic risk, and improved the model's generalization ability and accuracy.

CN120108566BActive Publication Date: 2025-11-28ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510050561.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-11-28
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Existing methods for assessing the carcinogenic risk of chemicals are time-consuming, inefficient, and raise ethical concerns. Traditional QSAR models and machine learning models have low accuracy and high false negative rates, making it difficult to efficiently assess the carcinogenic risk of a large number of compounds and explain their mechanisms of action.

Method used

By constructing a cross-modal knowledge graph, combining convolutional neural networks and deep neural networks, and integrating multimodal data feature representations, a carcinogenicity prediction model is built. Through cross-modal feature transformation and feature fusion, the carcinogenic risk of compounds can be accurately assessed and the mechanism explained.

Benefits of technology

It improves the accuracy and generalization ability of compound carcinogenic risk assessment, performs well on imbalanced datasets, and can explain the carcinogenic mechanism of positive compounds. The cross-modal transformation model achieves an accuracy of 0.8579 and a PRC of 0.8908 on the external validation set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108566B_ABST
    Figure CN120108566B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multimodal data fusion evaluation compound carcinogenic risk deep learning method, device and storage medium, belong to chemical health risk assessment technical field, comprising: (1) establish carcinogenicity prediction dataset;(2) construct cross-modal knowledge graph, adopt convolutional neural network to learn cross-modal knowledge graph and obtain cross-modal feature representation;(3) extract molecular fingerprint feature representation and molecular graph feature representation, obtain fusion feature representation by fusing molecular fingerprint feature representation with cross-modal feature representation;(4) construct fusion model, first deep neural network, second deep neural network in fusion model are used to process different feature representations, and classifier is used for prediction;(5) the parameters of fusion model are optimized by supervision training;(6) the carcinogenic risk of to-be-tested compound is evaluated using the optimized fusion model.The method can comprehensively and accurately evaluate the carcinogenic risk of compound and explain the carcinogenic mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of chemical health risk assessment, and particularly relates to a deep learning method for evaluating carcinogenic risk of compounds based on multi-modal data fusion, a device and a storage medium. BACKGROUND

[0002] High exposure of carcinogenic risk chemicals is one of the important factors leading to the increase of cancer incidence. Screening and controlling potential carcinogenic chemicals is of great importance to reduce the incidence of cancer and protect human life safety and health. Traditional methods for evaluating the carcinogenic risk of chemicals mainly rely on animal experiments. However, this method is time-consuming, low in efficiency and may have ethical problems, and it is difficult to evaluate the carcinogenic risk of a large number of chemicals one by one, so it is necessary to develop high-throughput screening technology.

[0003] Alternative methods for carcinogenicity testing, such as in vitro experiments and computer simulation methods, all face a series of challenges. In vitro methods such as cell transformation and stem cell proliferation experiments face bottlenecks in evaluating non-genotoxic carcinogens and interpreting their action modes. Quantitative structure-activity relationship (QSAR) methods based on structure-activity principles can quickly evaluate the potential carcinogenicity of various compounds and help researchers more efficiently screen chemical carcinogens. However, there are limitations in model generalization ability and limited evaluation of complex mechanisms.

[0004] Chinese patent document with publication number CN114743614A discloses an integrated learning method for screening carcinogenic chemicals, which includes: (1) constructing a carcinogenicity data set of chemicals; (2) calculating molecular PubChem fingerprints and performing preprocessing and feature selection; (3) integrated model training; (4) selecting accuracy and other indicators to evaluate the performance of the model; (5) referring to the OECD guide to represent the application domain of the model.

[0005] Chinese patent document with publication number CN115565624A discloses a method for constructing a carcinogenicity prediction model for predicting carcinogenicity. The invention first collects data of various compounds, screens out carcinogenicity data and non-carcinogenicity data, and constructs a carcinogenicity prediction data set. Then, a carcinogenicity-graph convolutional neural network model is constructed based on the molecular graph structure data in the carcinogenicity prediction data set. An autoencoder model is constructed based on the mass spectrometry data in the carcinogenicity prediction data set. Subsequently, the feature matrices output by the carcinogenicity-graph convolutional neural network model and the autoencoder model are fused and trained, and finally the prediction results are output. By analyzing the prediction results, when the accuracy requirement is met, a comprehensive carcinogenicity prediction model is obtained. The obtained carcinogenicity comprehensive prediction model can be used to predict the carcinogenicity of compounds.

[0006] However, the performance of the above machine learning / deep learning model does not linearly improve with the complexity, number of input descriptors, complexity of algorithm or complexity of architecture, and is limited by low accuracy and high false negatives. Therefore, it is necessary to establish a carcinogenicity prediction model based on multi-modal data fusion to improve the accuracy of data-driven carcinogenic risk assessment and clarify the toxicity mechanism based on harmful outcome pathways. SUMMARY

[0007] The present application provides a deep learning method for evaluating the carcinogenic risk of compounds based on multi-modal data fusion, which obtains cross-modal feature representation by constructing a cross-modal knowledge graph combined with rich molecular representation, constructs a carcinogenicity prediction model combined with molecular fingerprint feature representation, cross-modal feature representation and molecular graph feature representation, can more comprehensively and accurately evaluate the carcinogenic risk of compounds, and explain the mechanism of action at the molecular level, and perfect the carcinogenic risk assessment of compounds.

[0008] The specific technical solutions adopted are as follows:

[0009] A deep learning method for evaluating the carcinogenic risk of compounds based on multi-modal data fusion, comprising:

[0010] (1) establishing a carcinogenicity prediction dataset, wherein the carcinogenicity prediction dataset includes a plurality of compounds and their real molecular labels indicating whether they are carcinogenic or not;

[0011] (2) constructing a cross-modal knowledge graph using compound-protein association data, protein-protein association data, protein-biological process data, protein-cell component data and protein-molecular function data obtained from a database; and obtaining cross-modal feature representation by learning the cross-modal knowledge graph using a convolutional neural network;

[0012] (3) extracting molecular fingerprint feature representation of the compounds in the carcinogenicity prediction dataset, fusing the molecular fingerprint feature representation with the cross-modal feature representation to obtain fusion feature representation, and extracting molecular graph feature representation of the compounds in the carcinogenicity prediction dataset;

[0013] (4) constructing a fusion model combining a first deep neural network, a second deep neural network and a classifier, wherein the first deep neural network is used to extract a fusion feature vector based on the fusion feature representation, the second deep neural network is used to extract a graph feature vector based on the molecular graph feature representation of the compounds, and the classifier is used to perform prediction and classification based on the fusion feature vector and the graph feature vector to obtain a molecular label prediction result;

[0014] (5) performing supervised training of the fusion model based on the molecular label prediction result and the real molecular label using the carcinogenicity prediction dataset to optimize the parameters of the fusion model;

[0015] (6) Using the optimized fusion model to evaluate the test compound, predict the carcinogenic risk of the compound, and if the prediction result is positive, explain the potential carcinogenic mechanism of the compound.

[0016] Toxicology information of chemicals is usually scattered in various databases, including active proteins, affected pathways, animal models, and potential carcinogenic effects. Chemical decisions based on existing computational methods are not transparent, and require labor-intensive experimental screening by experts in the field of toxicology, consuming a lot of time to determine the toxicological correlation between chemicals and carcinogenic toxicity endpoints, and determining the mechanism of action is crucial for risk assessment. In order to solve the problems of scattered data and non-uniform identifiers, cross-modal knowledge graphs are used to interconnect multi-modal data, and further combined with multiple feature representations of compounds to build a carcinogenicity prediction fusion model to predict compounds that may have carcinogenic effects. By integrating compounds in multiple databases including CPDB (Consensus Pathway Database), carcinogenicity data and related toxicology data of multiple compounds are obtained.

[0017] Specifically, in step (1), the compounds collected from the database are used to establish a carcinogenicity prediction dataset after cleaning, and the cleaning steps include: removing inorganic substances, organometallic compounds, salt compounds, mixtures and repeated substances, and the obtained carcinogenicity prediction dataset includes a total of 3762 sample compounds, including 2306 negative compounds and 1456 positive compounds.

[0018] In step (2), the cross-modal knowledge graph includes compound nodes, protein (Protein) nodes, biological process (Biological process, BP) nodes, cellular component (Cellular component, CC) nodes and molecular function (Molecular function, MF) nodes;

[0019] The compound node includes a compound SMILES string;

[0020] The protein node includes a protein sequence;

[0021] The biological process node, the cellular component node and the molecular function node include a text description;

[0022] The different types of nodes are connected by edges, and the edges between the compound nodes and the protein nodes represent the action activity relationship of the compounds and the proteins, the edges between the protein nodes and the protein nodes represent the interaction relationship between the proteins and the proteins, the edges between the protein nodes and the biological process nodes represent the participation or regulation relationship of the proteins in the biological processes, the edges between the protein nodes and the cell component nodes represent the localization or action relationship of the proteins in the specific cell components, and the edges between the protein nodes and the molecular function nodes represent the relationship that the proteins perform specific molecular functions.

[0023] The different types of nodes are connected by edges through the graph structure to form a complete multi-modal knowledge graph, and subsequent deep learning methods are used to mine hidden relationships and predict possible interactions between chemicals and proteins, and reveal potential associations between certain proteins and specific biological processes, cell components or molecular functions.

[0024] Specifically, in step (2), when learning the cross-modal knowledge graph using a convolutional neural network, the cross-modal knowledge graph can be learned by constructing a cross-modal conversion model based on a convolutional neural network. In the learning process, the embedding vectors of entities and relationships in the cross-modal knowledge graph are initialized, a projection layer is designed for each type of node and relationship, embedding vectors of different feature dimensions are mapped to a unified hidden space through the projection layer, deep features of the head entity in the hidden space are extracted using the convolution layer, batch normalization and activation function in the convolutional neural network, the feature mapping of the head entity in the tail entity space is represented, and cross-modal feature representation is obtained; the paired and unpaired node relationships are optimized by the InfoNCE loss function.

[0025] Preferably, in step (3), the molecular fingerprint feature representation of the compound includes MACCS molecular fingerprint, Morgan molecular fingerprint and mol2vec descriptor.

[0026] Specifically, according to the SMILES expression of the sample compound, the MACCS molecular fingerprint and the Morgan fingerprint are calculated using the RDKit software package, the mol2vec descriptor is calculated using the mol2vec pre-training model, and then the features are fused, the missing values and the features with a variance of 0 and a Pearson correlation coefficient greater than 0.95 are removed to obtain the molecular fingerprint feature representation of the compound.

[0027] In step (3), the SMILES expression of the sample compound is converted into a molecular graph by using a Deep Graph Library (DGL) software package to represent atoms, bonds, and local environments, specifically, the molecular graph feature representation of the compound includes node features for describing atomic information of the compound and edge features for describing the bonding connection relationship between atoms; the node features include atomic type, hybridization type, formal charge, number of free radical electrons, atomicity, atomic implicit valence, or chirality; the bond features include bond type, whether it is a conjugated bond, whether it is a ring, or whether it is an aromatic bond.

[0028] Preferably, in step (4), the first deep neural network is constructed based on the KAN framework, and the second deep neural network is constructed based on the GCN framework.

[0029] Preferably, in step (5), the accuracy, precision, recall, F1 score, false negative rate, area under the receiver operating characteristic curve (AUROC), or balanced accuracy is used as an index to optimize the parameters of the fusion model.

[0030] The application also provides a device, which comprises:

[0031] The data set construction module is configured to construct a carcinogenicity prediction data set based on the collected compounds and their real molecular labels indicating whether they are carcinogenic or not;

[0032] The knowledge graph construction module is configured to construct a cross-modal knowledge graph containing compound-protein correlation data, protein-protein correlation data, protein-biological process data, protein-cell component data, and protein-molecular function data for the compounds in the carcinogenicity prediction data set;

[0033] The knowledge representation learning module is configured to learn entities and relationships in the cross-modal knowledge graph by using a convolutional neural network to obtain cross-modal feature representations;

[0034] The feature extraction and fusion module is configured to extract molecular graph feature representations and molecular fingerprint feature representations for the compounds in the carcinogenicity prediction data set, and fuse the molecular fingerprint feature representations and the cross-modal feature representations to obtain fused feature representations;

[0035] The fusion model construction module is configured to construct a fusion model combining a first deep neural network, a second deep neural network, and a classifier according to the fused feature representations and the molecular graph feature representations;

[0036] The model training module is configured to train and optimize the fusion model;

[0037] The evaluation identification module is used for evaluating the to-be-tested compound by using the optimized fusion model, predicting the carcinogenic risk of the compound, and explaining the potential carcinogenic mechanism of the compound if the prediction result is positive.

[0038] The application further provides a computer readable storage medium storing computer instructions, which are executed by a processor to implement the deep learning method for evaluating the carcinogenic risk of a compound based on multi-modal data fusion.

[0039] Compared with the prior art, the application has the following beneficial effects:

[0040] (1) The cross-modal conversion model based on the convolutional neural network is used to learn the cross-modal knowledge graph, so that the key carcinogenic feature information can be captured, the carcinogenic mechanism of 1548 compounds (with complete related information) is learned by the cross-modal conversion model, and the cross-modal feature representation of 3672 compounds can be obtained, the accuracy of the cross-modal conversion model based on the convolutional neural network on the external validation set reaches 0.8579, the PRC reaches 0.8908, and the overall optimal, which shows that the cross-modal conversion model has strong generalization ability and can accurately capture the carcinogenic feature information of the compound, and the performance on the unbalanced data set is very good.

[0041] (2) The method fuses multi-modal data, constructs a cross-modal knowledge graph, and evaluates the carcinogenic risk of a compound by using a deep learning method, so that not only the potential carcinogenicity of the compound can be evaluated, but also the action mechanism of the compound with a positive prediction of the carcinogenic risk can be explained. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 The flowchart of the deep learning method for evaluating the carcinogenic risk of a compound based on multi-modal data fusion is shown.

[0043] Figure 2 The effect diagram of learning the cross-modal knowledge graph by using the convolutional neural network is shown.

[0044] Figure 3 The effect diagram of predicting the carcinogenicity of a compound by using the optimized fusion model is shown. DETAILED DESCRIPTION

[0045] The application will be further illustrated below in conjunction with examples. It should be understood that the examples are only used to illustrate the application, and are not used to limit the scope of the application.

[0046] Example 1

[0047] Specifically, the flowchart of the deep learning method for evaluating the carcinogenic risk of a compound based on multi-modal data fusion is shown in Figure 1 ;

[0048] (1) Establish a carcinogenicity prediction dataset;

[0049] Based on the list of International Agency for Research on Cancer (IARC) and the like, the true molecular labels of a plurality of compounds and their carcinogenicity or not are obtained; the collected compounds are cleaned to remove inorganic compounds, organic metal compounds, salt compounds, mixtures and repeated substances, and after cleaning the data and unifying the classification standard, a total of 3762 sample compounds are obtained, including 2306 negative compounds and 1456 positive compounds, to establish a carcinogenicity prediction dataset.

[0050] (2) Construct a cross-modal knowledge graph, and learn the cross-modal knowledge graph by using a convolutional neural network to obtain a cross-modal feature representation;

[0051] The BioAssay Results information of the compounds on Pubchem is obtained by crawling, and the data with taxid of 9606 and activity of Active is screened to remove repeated records, eliminate abnormal data, correct error data and the like, to obtain compound-protein association data (compound and action protein pair data set);

[0052] The node types of protein, BP, CC and MF are selected from PrimeKG, and the features of different nodes are recorded to obtain protein sequences and text descriptions of biological processes, cell components and molecular functions, to obtain protein-protein association data (protein and protein pair data set), protein-biological process data (protein and biological process pair data set), protein-cell component data (protein and cell component pair data set), and protein-molecular function data (protein-molecular function pair data set) for constructing a cross-modal knowledge graph.

[0053] Specifically, the cross-modal knowledge graph includes compound nodes, protein nodes, biological process nodes, cell component nodes and molecular function nodes; the compound nodes include compound SMILES strings; the protein nodes include protein sequences; the biological process nodes, the cell component nodes and the molecular function nodes include text descriptions; different types of nodes are connected by edges, the edges between the compound nodes and the protein nodes represent the action activity relationship between the compounds and the proteins, the edges between the protein nodes represent the interaction relationship between the proteins, the edges between the protein nodes and the biological process nodes represent the participation or regulation relationship of the proteins in the biological processes, the edges between the protein nodes and the cell component nodes represent the positioning or action relationship of the proteins in the specific cell components, and the edges between the protein nodes and the molecular function nodes represent the relationship that the proteins perform specific molecular functions; and the specific embodiments are shown in Table 1.

[0054] Table 1 Cross-modal knowledge graph entity, relationship and quantity statistics

[0055]

[0056] The cross-modal knowledge graph is learned by constructing a cross-modal conversion model based on a convolutional neural network. In the learning process, the embedding vectors of entities and relationships in the cross-modal knowledge graph are initialized, a projection layer is designed for each node and relationship, the embedding vectors of different feature dimensions are mapped to a unified hidden space through the projection layer, the deep features of the head entity in the hidden space are extracted using the convolution layer, batch normalization and activation function in the convolutional neural network, the feature mapping of the head entity in the tail entity space is represented, and the cross-modal feature representation is obtained; the paired and unpaired node relationships are optimized by the InfoNCE loss function.

[0057] Specifically, the projected head node embedding, head node type, tail node type and relationship type embedding are stacked, and the input embedding is dimensionally rearranged to adapt to the 1D convolution operation. Feature extraction is performed through three convolution layers and the batch normalization layer and ReLU activation function after the convolution layers. The convolution layers extract local features of different ranges with different kernel sizes (kernel_size=7, kernel_size=5 and kernel_size=3) and padding parameters, then the dimension of the result is rearranged and the first position feature is taken, which is added to the head node embedding input as the output to retain the original information and ensure convergence and stability, and also represent the features of the head node converted to the target modality space. The loss is calculated using the InfoNCE loss function according to the tail node and negative tail node embedding. The convolutional neural network extracts and converts node embedding features, reorganizes and maps different modality information in the feature space, adjusts the parameters combined with InfoNCE, provides more discriminative and representative feature representations for cross-modal tasks, and improves the performance of the model.

[0058] After learning the cross-modal knowledge graph by constructing a cross-modal conversion model based on a convolutional neural network, cross-modal retrieval can be performed by matching the converted embedding with candidate samples in the target modality embedding space. Cross-modal matching of chemicals and proteins, proteins and BP, MF and CC is performed on the validation set, and the recall rate is used to evaluate the matching effect (i.e., the effect of learning the cross-modal knowledge graph using the convolutional neural network) as shown in Table 2. Figure 2 As shown in Table 2, the cross-modal conversion model established shows good effects on target prediction of compounds and description of biological functions of proteins, with reall@5 reaching 0.503, 0.46, 0.5 and 0.78, respectively. It is proved that through pairwise cross-modal search conversion, the mechanism of compound action can be explained.

[0059] Cross-modal knowledge graph connects chemicals and proteins, enriching the feature representation of chemicals. Protein and biological attributes enhance protein representation learning. Learning cross-modal knowledge graph with convolutional neural network converts protein embedding into biomolecular function embedding, better aligning protein sequence with functional semantics.

[0060] (3) extract molecular fingerprint feature representation and molecular graph feature representation, fuse the molecular fingerprint feature representation and the cross-modal feature representation to obtain the fused feature representation;

[0061] The SMILES expression of the sample compound is derived, and the 166 substructure MACCS molecular fingerprints of the chemical based on SMART and the 2048-bit Morgan molecular fingerprints with a diameter of 2 are calculated using the RDKit software package. A 300-dimensional vector representation is constructed for the sample compound using the mol2vec pre-training model to obtain the mol2vec descriptor. Then the mol2vec descriptor and the RDKit fingerprint are fused, and the missing values and features with a variance of 0 are further removed. The Pearson correlation coefficient between each feature is calculated, and features greater than 0.95 are removed to prevent overfitting. Finally, 3189 molecular descriptors are obtained, i.e. a feature matrix of 3762 rows and 3189 columns, obtaining the molecular fingerprint feature representation of the compound. The molecular fingerprint feature representation and the cross-modal feature representation are fused to obtain the fused feature representation;

[0062] The DGL package with PyTorch port is used to encode 9 kinds of atomic features and 5 kinds of bond features, and the molecule is converted into molecular graph feature representation to represent atoms, bonds and local environment. The specific encoding rules of atoms and bonds are shown in Table 2.

[0063] Table 2 Feature encoding of atoms and bonds

[0064]

[0065] (4) Construct a fusion model, the first deep neural network and the second deep neural network in the fusion model are used to process different feature representations, and the classifier in the fusion model is used for prediction.

[0066] The first deep neural network is based on the KAN framework, and the second deep neural network is based on the GCN framework.

[0067] KAN is implemented by KANLinear class, which combines B-spline function and linear transformation and learnable activation function SiLU on weights to map input features to a grid space and combine them nonlinearly through learnable spline coefficients, aiming to enhance the model's expressive power for non-graph structured data (fingerprint data and bioactivity features) through complex nonlinear mapping. KAN class forms a deep network by stacking multiple KANLinear layers, each layer further nonlinearly transforms the input, and dynamically updates the grid during training. The model's activation and entropy are controlled through a regularization loss function to enhance the stability and generalization ability of the model.

[0068] GCN processes node features in the graph layer by layer through the GraphConv layer in the DGL library and introduces nonlinearity through the ReLU activation function to gradually extract high-order features in the graph structure. To prevent overfitting, a dropout layer with a dropout rate of 0.41 is added after each convolutional layer to randomly discard part of the node features during training to enhance the model's generalization ability. After the convolution operation, the mean_nodes function of DGL is used for global pooling operation to average all node features in the graph to generate a fixed-length feature vector representing the features of the entire graph. The pooled features are mapped to the output space through a linear layer for predicting the class of the graph.

[0069] The outputs from GNN and KAN are added to generate a comprehensive output for the binary classification task.

[0070] (5) Supervised training of the fusion model to optimize the fusion model parameters;

[0071] Data is processed in batches during training, and DataLoader is used for batch loading to improve training efficiency and effectively utilize GPU resources. The training process uses a cross-entropy loss function to minimize the difference between the model's predicted value and the actual label. The adam optimization algorithm is used to continuously iterate and update the weights for model learning and training to obtain the optimal weights. ReduceLROnPlateau learning rate scheduler is used, which automatically reduces the learning rate when the model's performance on the validation set does not improve, to better perform gradient descent. When the performance on the validation set reaches the best, the model is saved. Early stopping mechanism is used to monitor the performance of the validation set, and if the performance of the validation set does not improve within 10 epochs, the training is terminated early to prevent overfitting and make the model more stable when converging. The test set is used to evaluate the model's generalization ability. The performance of the model is evaluated through multiple indicators, including accuracy, precision, recall, f1_score, or ROC-AUC.

[0072] The hyperparameters of the model adjustment are as follows: in the GCN, the number of hidden layers of the graph convolution layer is 3, the number of hidden units is 128, 64 and 32 respectively, and the dropout of the graph convolution layer is 0.41; in the KAN, the number of hidden layers is 2, and the number of hidden units is 64 and 32 respectively. The optimizer is Adam, the loss function is cross-entropy loss CrossEntropyLoss, and the class weight is torch.tensor([1.0, 6.9]) with class weight. The loss function weight class_weights of the unbalanced class is higher. The learning rate lr is 0.001, the weight decay coefficient weight_decay is 0.0001, the batch size is 64, the total number of training iterations epochs is 50. The patience parameter patience of the early stopping (Early Stopping) method is 10, and the training is stopped after 10 rounds without improvement.

[0073] The data set is randomly divided into a training set, a validation set and a test set in a ratio of 7:2:1. The method of the application is evaluated according to the test set, the test set has 753 substances, and after deriving the SMILES expression of the test set compound, the fusion feature representation and the molecular graph feature representation are extracted and input into the fusion model to obtain the carcinogenicity test result. In the test result, 635 molecular label prediction results are consistent with the true molecular label, the method accuracy is 0.85, and the balanced accuracy is 0.83. The balanced accuracy of the fusion model compared with other models is shown in Table 1. Figure 3

[0074] (6) using the optimized fusion model to evaluate the carcinogenic risk of the test compound;

[0075] Given a chemical Butyl(3-hydroxypropyl)nitrosamine (BBN) (CAS No.: 51938-13-7), the test compound is evaluated using the above method and the optimized fusion model, the carcinogenic risk of the compound is predicted, and the result is that the activity is 1 (the prediction result is positive), which is consistent with the conclusion of the paper. The mechanism of action of the compound is explained: CYP family (CYP2C9, CYP2C19, CYP2D6, CYP1A2) is directly involved in the phase I oxidative metabolism of BBN, and NR1I2 may indirectly affect the metabolism of BBN by regulating the expression of CYP family enzymes to generate active intermediates.

[0076] The above examples have described the technical solutions of the application in detail, and it should be understood that the above examples are only specific embodiments of the application and are not used to limit the application. Any modification, supplement or similar replacement within the principle range of the application should be included in the protection scope of the application.​

Claims

1. A deep learning method for assessing the carcinogenic risk of compounds based on multimodal data fusion, characterized in that, include: (1) Establish a carcinogenicity prediction dataset, which includes a variety of compounds and their real molecular labels of whether they are carcinogenic or not; (2) Using the compound-protein association data, protein-protein association data, protein-biological process data, protein-cell component data, and protein-molecular function data obtained from the database, a cross-modal knowledge graph is constructed; a convolutional neural network is used to learn the cross-modal knowledge graph to obtain cross-modal feature representations; (3) Extract the molecular fingerprint feature representation of compounds in the carcinogenicity prediction dataset, fuse the molecular fingerprint feature representation with the cross-modal feature representation to obtain the fused feature representation; extract the molecular graph feature representation of compounds in the carcinogenicity prediction dataset; (4) Construct a fusion model combining a first deep neural network, a second deep neural network, and a classifier. The first deep neural network is used to extract fusion feature vectors based on fusion feature representations, the second deep neural network is used to extract graph feature vectors based on molecular graph feature representations of compounds, and the classifier is used to perform prediction and classification based on fusion feature vectors and graph feature vectors to obtain molecular label prediction results. (5) Use the carcinogenicity prediction dataset to conduct supervised training of the fusion model based on the prediction results of molecular labels and the real molecular labels, and optimize the parameters of the fusion model; (6) Use the optimized fusion model to evaluate the test compound and predict the carcinogenic risk of the compound. If the prediction result is positive, explain the potential carcinogenic mechanism of the compound. In step (2), the cross-modal knowledge graph includes compound nodes, protein nodes, biological process nodes, cellular component nodes, and molecular function nodes; The compound node includes the compound SMILES string; Protein nodes include protein sequences; Biological process nodes, cellular component nodes, and molecular function nodes include text descriptions; Different types of nodes are connected by edges. The edge between compound nodes and protein nodes represents the interaction activity relationship between compounds and proteins. The edge between protein nodes represents the interaction relationship between proteins. The edge between protein nodes and biological process nodes represents the participation or regulation relationship of proteins in biological processes. The edge between protein nodes and cell component nodes represents the location or role of proteins in specific cell components. The edge between protein nodes and molecular function nodes represents the relationship of proteins performing specific molecular functions. In step (4), the first deep neural network is built based on the KAN framework, and the second deep neural network is built based on the GCN framework.

2. The deep learning method for assessing the carcinogenic risk of compounds based on multimodal data fusion according to claim 1, characterized in that, In step (2), the embedding vectors of entities and relations in the cross-modal knowledge graph are initialized, and a projection layer is designed for each node and relation. The embedding vectors of different feature dimensions are mapped to a unified hidden space through the projection layer. The deep features of the head entity in the hidden space are extracted using the convolutional layer, batch normalization and activation function in the convolutional neural network. The feature mapping of the head entity in the tail entity space is represented to obtain the cross-modal feature representation. The paired and unpaired node relations are optimized through the InfoNCE loss function.

3. The deep learning method for assessing the carcinogenic risk of compounds based on multimodal data fusion according to claim 1, characterized in that, In step (3), the molecular fingerprint features of the compound include the MACCS molecular fingerprint, the Morgan molecular fingerprint, and the mol2vec descriptor.

4. The deep learning method for assessing the carcinogenic risk of compounds based on multimodal data fusion according to claim 1, characterized in that, In step (3), the molecular graph feature representation of the compound includes node features used to describe the atomic information of the compound and edge features used to describe the bonding relationships between atoms; Node characteristics include atom type, hybridization type, formal charge, number of free radical electrons, atomicity, and implicit valence state or chirality of atoms; Bond characteristics include the type of bond, whether it is a conjugated bond, whether it is a cyclic bond, or whether it is an aromatic bond.

5. The deep learning method for assessing the carcinogenic risk of compounds based on multimodal data fusion according to claim 1, characterized in that, In step (5), the parameters of the fusion model are optimized using accuracy, precision, recall, F1 score, false negative rate, area under the receiver operating characteristic curve (AUROC), or balanced accuracy as indicators.

6. An apparatus, characterized in that, The device includes: Dataset construction module: Constructs a carcinogenicity prediction dataset based on the collected compounds and their real molecular labels indicating whether they are carcinogenic or not; Knowledge Graph Construction Module: For compounds in the carcinogenicity prediction dataset, a cross-modal knowledge graph is constructed, including compound-protein association data, protein-protein association data, protein-biological process data, protein-cellular component data, and protein-molecular function data. The cross-modal knowledge graph includes compound nodes, protein nodes, biological process nodes, cellular component nodes, and molecular function nodes. Compound nodes include the compound SMILES string; protein nodes include protein sequences; biological process nodes, cellular component nodes, and molecular function nodes include text descriptions. Different types of nodes are connected by edges. The edges between compound nodes and protein nodes represent the interaction activity relationship between compounds and proteins; the edges between protein nodes represent the interaction relationship between proteins; the edges between protein nodes and biological process nodes represent the participation or regulation relationship of proteins in biological processes; the edges between protein nodes and cellular component nodes represent the location or role of proteins in specific cellular components; and the edges between protein nodes and molecular function nodes represent the relationship of proteins performing specific molecular functions. Knowledge Representation Learning Module: Used to learn entities and relationships in cross-modal knowledge graphs using convolutional neural networks to obtain cross-modal feature representations; Feature extraction and feature fusion module: Extracts molecular graph feature representation and molecular fingerprint feature representation for compounds in the carcinogenicity prediction dataset, and fuses the molecular fingerprint feature representation with the cross-modal feature representation to obtain the fused feature representation; Fusion model construction module: Constructs a fusion model combining a first deep neural network, a second deep neural network, and a classifier based on fusion feature representation and molecular graph feature representation; the first deep neural network is built based on the KAN framework, and the second deep neural network is built based on the GCN framework; The model training module is used to train and optimize the fusion model. The evaluation and identification module is used to evaluate the test compound using an optimized fusion model, predict the carcinogenic risk of the compound, and if the prediction result is positive, explain the potential carcinogenic mechanism of the compound.

7. A computer-readable storage medium storing computer instructions, characterized in that, When executed by the processor, this instruction implements the deep learning method for assessing the carcinogenic risk of compounds based on multimodal data fusion as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Integrated learning method for screening carcinogenic chemicals

    CN114743614A

  • Method for constructing carcinogenicity prediction model for predicting carcinogenicity

    CN115565624A

  • Rapid prediction and correlation mechanism analysis method for molecular targets

    CN115631808A

  • Drug target interaction prediction model training method and device, equipment and medium

    CN116486936A