Neural Network Model Prediction Method for Gene Expression Profiles under Drug Exposure
Through the graph convolutional neural network model combined with the drug-drug similarity network and the two-dimensional structure of the drug, the gene expression profile under drug exposure was predicted, and the problem of low utilization efficiency of gene expression profile data in drug development in the prior art was solved, and the reduction of drug development time and cost and the improvement of drug reuse efficiency was achieved.
Patent Information
- Application Number
- CN202310157307.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-02-20
AI Technical Summary
The prior art is difficult to efficiently utilize gene expression profile data under drug exposure in drug development, resulting in excessive drug development time and cost.
The graph convolutional neural network model (GraphDrug) is used, combining the drug-drug similarity network and the drug two-dimensional structure as input to predict the genome expression profile of organisms under drug exposure, and the correlation coefficient with the disease-related gene expression profile is calculated by analyzing gene expression data to perform drug reuse prediction.
Accurate prediction of gene expression profiles under drug exposure is achieved, reducing the time and cost of drug development and improving the efficiency of drug reuse.
Smart Images

Figure CN116092626B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and particularly relates to a neural network model prediction method for gene expression profiles under drug exposure. Background Art
[0002] The development of drugs plays a crucial role in the treatment of diseases. However, traditional drug development requires a large amount of time and money.
[0003] The gene expression profiles of organisms have been widely used to identify cell characteristics and organism phenotypes. Systematically analyzing the gene expression of the genomes of human cell lines under the exposure of bioactive compounds (potential drug compounds) can significantly improve the success rate of drug development. In particular, the gene expression profiles under drug exposure play an important guiding role in drug reutilization, drug mechanism research, identification of lead compounds, and prediction of side effects of preclinical drugs. The main purpose of drug reutilization is to use approved drugs to treat new diseases. Since the safety, therapeutic effects, and appropriate management plans of approved drugs have been studied clearly, the development costs of drugs in terms of money and time can be significantly reduced, and developers can successfully apply them to the treatment of another disease at a lower cost.
[0004] Predicting potential drug targets by analyzing the gene expression profiles of the human body under drug exposure is an important research direction in current drug reutilization. At present, the gene expression data under drug exposure mainly still remains at the stage of using biochemical methods for experiments and transcriptome sequencing, which requires a large amount of time and money. How to accurately predict the gene expression profile data under drug exposure by combining the existing gene expression profile data with the drug molecular structure has become an important research topic in the current computer era.
[0005] In addition, databases of some drugs have been developed and established, such as the Pubchem database, the DrugBank database, etc. These databases contain the chemical formulas, three-dimensional structures, aliases, protein targets of drugs, drug-drug side effects, drug approval status, etc. of various chemical small molecules. In addition, the Library of Integrated Network-Based Cellular Signatures (LINCS) project conducts drug (bioactive small molecules, ligands, etc.) exposure gene expression experiments on the main cells of the human body (Subramanian, A. et. al., (2017). A next-generation connectivity map: L1000 platform and the first 1,000,000 profiles. Cell, 171(6), 1437-1452.). A large amount of gene expression data of drug exposure generated by this project can be found in the LINCS L1000 Phase I (GSE92742) and Phase II (GSE70138) databases. Generally, it is believed that drugs and their indications usually share related genes. The more perturbed genes shared by drugs and diseases, the more relevant they are, and the more likely this disease is to be the target of related drugs. Therefore, analyzing the gene expression data under the influence of these drugs can better design or screen drugs. The data accumulated in these databases can provide rich data support for the training and evaluation of neural network models. Summary of the Invention
[0006] In view of the above deficiencies in the prior art, the present application proposes a graph convolutional neural network model (GraphDrug) using a drug-drug similarity network and the two-dimensional structure of drugs as inputs to predict the genomic expression profile of organisms under drug exposure. Further, by analyzing the gene expression data of these drug exposures, the correlation coefficient with the gene expression profile related to the disease can be calculated to predict drug reutilization.
[0007] A neural network model prediction method for gene expression profiles under drug exposure, comprising the following steps:
[0008] (1) Select a drug to be analyzed, and obtain the self-information and interaction information of the drug to be analyzed.
[0009] (2) Construct a drug-drug similarity network according to the drug interaction information obtained in step (1).
[0010] (3) Convert the SMILES structure information in the self-information of the drug molecule to be analyzed into a "graph representation" containing an atomic feature matrix, a chemical bond matrix, and an adjacency matrix, and an "image representation" with different channels using chemical bond types as images; calculate the "graph representation" and "image representation" data of the drug using a graph convolutional neural network and a swim-transformer model (both are modules in neural networks) respectively, and concatenate the results obtained from these two models (modules) to form the feature vector of the drug molecule to be analyzed. (4) Use the feature vector of the drug obtained in step (3) as a node on the drug-drug similarity network, and use the drug-drug similarity network obtained in step (2) to aggregate information to obtain the feature vector of the drug molecule to be analyzed after aggregating similar drug information.
[0011] (5) Convert the feature vector of the drug molecule to be analyzed after aggregating similar drug information obtained in step (4) into the potential feature vector of the downstream gene.
[0012] (6) Calculate the influence of drug dosage on the potential feature representation of the downstream gene, and integrate the drug dosage influence into the potential feature vector of the downstream gene obtained in step (5) to obtain the gene feature vector integrated with the drug dosage influence.
[0013] (7) Calculate the influence of gene-gene interaction on the potential feature representation of the gene, and integrate it into the gene feature vector obtained in step (6) to obtain the gene feature vector integrated with the influence of gene-gene interaction.
[0014] (8) Convert the gene feature vector obtained in step (7) into the final gene expression matrix, which corresponds to the gene expression profile under drug exposure of the drug molecule to be analyzed.
[0015] Preferably, in step (1), the self-information of the drug molecule to be analyzed includes the CID serial number and the SMILES sequence, and the interaction information includes the drug-gene interaction information in the DGIdb database, the drug-gene interaction information in the CTD database, the drug-disease association information, and the drug comprehensive similarity score data in the STITCH database.
[0016] More preferably, in step (2), different drug-drug similarity networks are constructed according to the drug interaction information in different databases; in step (4), for each drug-drug similarity network, information aggregation is performed respectively to obtain multiple feature vectors after information aggregation based on different drug-drug similarity networks. After inputting them into the transformer model for calculation, different importance coefficients are attached to each input feature vector, and they are weighted and averaged to aggregate the information unique to each drug-drug similarity network.
[0017] Further preferably, in step (2), the method for constructing the drug-drug similarity network of different databases is as follows:
[0018] i. Method for constructing the drug-drug similarity network based on the drug-gene interaction network of the DGIdb database (Drug Gene Interaction database, https: / / www.dgidb.org) or the CTD database (Comparative Toxicogenomics Database, https: / / ctdbase.org): Use the common genes as a bridge to connect two drugs. If there are 3 or more identical genes shared between the drug to be analyzed and a certain drug, then the drug to be analyzed and the certain drug are similar drugs to each other, and the value in the adjacency matrix is 1; otherwise, it is 0.
[0019] ii. Method for constructing the drug-drug similarity network based on the drug-disease interaction network of the CTD database: Use the disease as a bridge to connect two drugs. First, map the disease information to gene information using the "DisGeNET database", and then use the common genes as a bridge to connect two drugs. If there are 3 or more identical genes shared between the drug to be analyzed and a certain drug, then the drug to be analyzed and the certain drug are similar drugs to each other, and the value in the adjacency matrix is 1; otherwise, it is 0.
[0020] iii. Method for constructing the drug-drug similarity network based on the comprehensive drug similarity score of the STITCH database (http: / / stitch.embl.de). Connect two drugs with the drug-drug similarity scores in the STITCH database that have been summarized. If the similarity score between the drug to be analyzed and a certain drug is higher than 150, then there is a connection edge between these two drugs in the "drug comprehensive similarity network based on the STITCH database".
[0021] Constructing the drug-drug similarity network using different types of drug interaction information can more comprehensively cover the drugs similar to the drug to be analyzed, so as to more comprehensively aggregate and supplement the characteristic information of the drug to be analyzed. Preferably, in step (3), the SMILES structure is converted into an atomic adjacency matrix, an atomic feature matrix, and a chemical bond feature matrix, and a one-dimensional drug feature vector is obtained by learning and combining these three matrices through a graph convolutional neural network as the processed "graph representation" data.
[0022] Preferably, in step (3), the swim-transformer model is used to learn the "image representation" data and obtain another one-dimensional drug feature vector as the processed "image representation" data.
[0023] In step (3), when concatenating to form the feature vector of the drug to be analyzed, the processed "graph representation" data and the processed "image representation" data are concatenated.
[0024] Further preferably, a graph convolutional neural network and a swim-transformer model are used to optimize the "graph representation" and "image representation" data of the drug respectively, to obtain the potential feature vector of each atom in the drug to be analyzed; then the potential feature vectors of each atom are combined to obtain the processed "graph representation" and "image representation" data of the drug to be analyzed.
[0025] Preferably, in step (5), the feature vector of the drug to be analyzed obtained in step (4) after aggregating similar drug information is transformed into the potential feature vector of the downstream gene through two non-linear transformations (i.e., fully connected layer: multiplying the input matrix by a matrix of learnable parameters, and then multiplying by a non-linear activation function such as the ReLU function, etc.).
[0026] In step (8), the gene feature vector obtained in step (7) is transformed into the final gene expression matrix using two non-linear transformations.
[0027] Preferably, in step (6), when calculating the influence of drug dose on the potential feature representation of the downstream gene, a logistic activation function is used to simulate the influence of drug dosage on gene expression, and the logistic activation function is:
[0028]
[0029] G v =δ·G d
[0030] In the formula, v is the dosage of drug exposure, e is the natural constant, k, a, and b are all trainable parameters and the initial value is 1, δ is the influence coefficient of drug dosage on gene expression, which is an intermediate variable, G d is the potential feature vector of the downstream gene obtained in step (5), and G v is the gene feature vector integrating the influence of drug dosage.
[0031] Calculating the influence of drug dosage on the potential feature representation of genes using continuous variables in this step can better learn the influence of different drug dosages on the potential features of genes and improve the learning ability of the model in this study.
[0032] Preferably, in step (7), when calculating the influence of gene-gene interaction on expression, a graph convolutional network is used for calculation.
[0033] Here, the change in gene expression caused by external perturbations is attributed to two independent factors. One is the change caused by drugs at different doses, and the other is the change in gene expression caused by gene-gene interactions. Such a breakdown can clarify the positions and contributions of different inputs in the neural network model, enabling better optimization of the model structure of this study and obtaining more generalized model prediction results.
[0034] The present invention performs convolutional learning on multiple graph networks (the "graph representation" of drug molecules, the multi-drug-drug similarity network, and the gene-gene interaction network), fully learning the characteristics of drug molecules and the influence of downstream gene-gene interactions. Through training on a large amount of drug data, it predicts the changes in the expression levels of a large number of biomarker genes in the human body under drug exposure, providing a new research method for subsequent drug repositioning prediction. Brief Description of the Drawings
[0035] Figure 1 It is a flowchart of the technical route of the present invention.
[0036] Figure 2 It is the input data of the "graph representation" of the drug of the present invention.
[0037] Figure 3 It is to construct a connection graph between all atoms of aspirin with the chemical bond type as different channels of the image. Detailed Embodiments
[0038] To describe the present invention more specifically, the following takes the prediction of the gene expression profile under aspirin exposure as an example and combines with the drawings ( Figure 1 ) to elaborate on the technical solution of the present invention in detail:
[0039] (1) Obtain various drug information from the PubChem public data using the constructed automated process
[0040] Use the workflow written in the Python language, with the BRD serial number of the aspirin molecule in L1000 as the retrieval serial number, to obtain various information (CSV format data) of aspirin in the drug public database PubChem. These information include the CID serial number (2244), the SMILES sequence (CC(=O)OC1=CC=CC=C1C(=O)O), the drug-gene interaction information in the DGIdb database, the drug-gene interaction information in the CTD database, and the drug-disease association information. These information are the drug similarity data, and the composed data set is the drug similarity data set.
[0041] In addition, the drug comprehensive similarity score data in the STITCH database is also integrated into the drug-drug similarity network.
[0042] (2) Construct a drug-drug similarity network from the obtained drug data
[0043] Use the drug interaction information obtained in step (1) (drug-gene interaction information in the DGIdb database, drug-gene interaction information in the CTD database, drug-disease association information, and drug comprehensive similarity score data in the STITCH database) to construct different drug-drug similarity networks. The following provides three different methods for constructing drug-drug similarity networks:
[0044] i. Construct a drug-drug similarity network using the drug-gene interaction networks of the "DGIdb database" and the "CTD database": Use common genes as bridges to connect two drugs. If aspirin and another drug share 3 or more identical genes, then aspirin and that drug are similar drugs to each other, and the value in the adjacency matrix is 1 (the value between non-similar drugs is 0).
[0045] ii. Method for constructing a drug-drug similarity network based on the connection between drugs and diseases: Use diseases as bridges to connect two drugs. First, map disease information to gene information (the gene set specifically expressed in this disease) using the "DisGeNET database", and then use the method in i. to construct a "drug-drug similarity network" based on drug-disease interaction information.
[0046] iii. Method for constructing a drug-drug similarity network "based on the drug comprehensive similarity score data in the STITCH database": Connect two drugs with the existing drug-drug interaction relationships in the summarized database. If the similarity score between aspirin and another drug is higher than 150, then there is a connection edge between these two drugs in the "drug comprehensive similarity network based on the STITCH database".
[0047] Through the above methods, a total of 4 "drug-drug similarity networks" are generated in this application: two networks are generated by method i (one drug-drug similarity network is generated by each of the two databases), and one is generated by each of methods ii and iii.
[0048] (3) Convert the SMILES structure of drug molecules into "graph representation" and "image representation"
[0049] i. "Graph representation" method of drugs: According to the atomic and chemical bond information contained in the aspirin molecule, convert the SMILES structure into an atomic adjacency matrix, an atomic feature matrix, and a chemical bond feature matrix ( Figure 2) Then, through the graph convolutional network in the subsequent steps, these three matrices representing the properties of the graph are aggregated into a drug feature vector. In this application, "graph representation" is "graph representation", where the "graph" (not "image") refers to "graph", which is a non-Euclidean abstract entity composed of nodes and edges between nodes. Here, the "graph representation of the drug" means taking three variables (atomic adjacency matrix, atomic feature matrix, and chemical bond feature matrix) used to describe the graph properties of the drug as inputs, and obtaining a one-dimensional vector result after a series of calculations. The "graph" in the "graph representation of the drug" refers to the source of this "representation" from the "graph properties of the drug", which is the source, and is distinguished from the "image representation of the drug" below.
[0050] ii. Method of "image representation" of drugs: Construct a connection graph between all atoms of aspirin with the chemical bond type as different channels of the image (SDTA is similar to the RGB three-color channels of ordinary pictures) ( Figure 3 ) and use it as the "image" representing all atoms of the drug (in the figure, S represents a single bond, D represents a double bond, T represents a triple bond, and A represents an aromatic bond); among them, the value between two connected atoms is the concatenation of the one-hot code representations of these two atoms (that is, connecting another drug vector at the end of a drug vector, and the length of the synthesized new vector is the sum of the lengths of the original two vectors). The "image representation" here is to obtain the matrix of the connection graph. The connection relationship is that if there is a chemical bond of the corresponding type between atoms, its value is 1, otherwise it is 0. The channels correspond to different chemical bond types. For example, if there is a single bond between the first C atom and the second C atom, then in the single bond channel, the values of the first C atom and the second C atom are the concatenation of the one-hot code representation of the "upper" atom in the connection graph and the one-hot code representation of the "left" atom. If two atoms are not connected, the value is zero.
[0051] (4) Use the graph convolutional network to learn the "graph representation" of drug molecules
[0052] To make the model calculation faster, this application uses the calculation method of the "relational graph neural network" to perform feature aggregation on all atoms in aspirin. The specific calculation method is as follows:
[0053]
[0054] In the formula represents the feature matrix of aspirin in the l-th layer regarding the j-th chemical bond relationship network (N is the number of atoms in the aspirin molecule, and D is the feature dimension of each atom), is the atomic adjacency matrix in the j-th chemical bond relationship network (such as a single-bond connection network). Since self-connections are added to the adjacency matrix of the drug, after aggregating its own information in addition to the node neighbor information in the above formula, the eigenvectors of each atom in aspirin are fused in the "bitwise addition" manner according to the corresponding chemical bond relationship network and merged in the way of expanding dimensions in the drug feature dimension:
[0055]
[0056] In the formula, E is an intermediate vector of the "graph representation" of the drug, obtained from formula 2, and it is transformed into the final drug vector representation through a non-linear transformation in the subsequent formula 3; J represents the number of different chemical bond channels, N is the number of atoms in the aspirin molecule, and h n represents the n-th atomic vector in aspirin.
[0057] H = σ(EW g ) (3)
[0058] Formula 3 is a non-linear transformation of the result E of formula 2. The non-linear transformation enables the neural network to fit any function. W g is a random digital matrix, and its value can be optimized through the backpropagation process of the neural network. The output vector H represents the final graph representation vector of the aspirin drug molecule, and σ(·) represents the ReLU activation function.
[0059] (5) Use swim-transformer to learn the "image representation" of drug molecules
[0060] To enable the neural network model to learn more comprehensive drug structure information, this application also uses the swim-transformer model (Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z.,... & Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE / CVF International Conference on Computer Vision (pp. 10012-10022).) to calculate the "image representation" of the drug, obtaining the feature vectors of each atom in aspirin. Since we need to obtain one-dimensional drug feature vectors, and the result obtained using the transformer model is one-dimensional vectors for each atom in the aspirin drug, and aspirin has 14 atoms, that is, 14 one-dimensional vectors, so we need to combine these 14 vectors into one vector, so we need a merging operation, and after merging, we obtain the final "image" representation of aspirin.
[0061] After obtaining the feature vectors of each atom in aspirin, use the method of "bitwise addition" to combine the feature vectors of all atoms in aspirin to obtain the final "image representation" vector H' of the aspirin drug molecule:
[0062]
[0063] In the formula, h' n represents the "image representation" vector of the nth atom in aspirin, and N is the number of atoms in aspirin.
[0064] (6) Use multiple drug-drug similarity networks to learn the interaction features between drugs
[0065] After simultaneously obtaining the "graph representation" vector H and the "image representation" vector H' of aspirin, concatenate them together to form the final drug feature vector L. This feature vector will then become a node on the drug-drug similarity network and aggregate information with other drugs through a graph convolutional network:
[0066]
[0067] In the formula, σ represents the activation function, represents the set of neighbor nodes (drugs similar to aspirin) of aspirin, generated by the method in step (2), L kis the feature vector of drugs similar to aspirin. In the graph convolutional network, the adjacency matrix adds self-connections to prevent gradient disappearance and add the information of its own node. Therefore, the set also contains the information of aspirin itself.
[0068] In this step, each drug-drug similarity network generates an aspirin feature representation that aggregates similar drug information, Z u is the aspirin feature vector for information aggregation using the u-th (1, 2,..., U) "drug-drug similarity network", and W u is the learnable parameter matrix corresponding to the drug-drug similarity network u. In this embodiment, U = 4, as specified by the method in step 2 (method i generates two networks, and methods ii and iii each generate one).
[0069] (7) Use the self-attention mechanism to add importance coefficients to each aspirin feature vector aggregated by "different drug-drug similarity networks"
[0070] After obtaining the aspirin feature vectors Z u (Z 1 , Z 2 ,..., Z U ) regarding multiple drug similarity matrices, use the transformer (Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N.,... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.) transformer (self-attention) model to add importance coefficients (formula 6) to different feature vectors of aspirin, corresponding to Z tu (Z t1 , Z t2 ,..., Z tU ), and then obtain the aspirin feature vector Zh that aggregates the information of different drug-drug similarity networks by calculating their average value (taking the average after adding the vectors bit by bit) (formula 7).
[0071] Z t1 , Z t2 , …, Z tU = transformer(Z 1 , Z 2 , …, Z u ) (6)
[0072]
[0073] (8) Project the aspirin feature vector into the potential feature vector of downstream genes
[0074] In this step, the aspirin feature vector Z h is transformed into the feature vector G of downstream genes through two non-linear transformations d , and this gene feature vector is the projection of the gene expression level under aspirin exposure in the gene feature space. The subsequent effects of aspirin dosage and time will also be calculated within this space.
[0075] (9) Calculate the effect of drug dosage on the potential feature representation of downstream genes using a custom logical activation function
[0076] This application designs a logical activation function to simulate the effect of drug dosage on gene expression levels (Equations 8 and 9). Here, the change in gene expression levels caused by gene-gene interactions is temporarily ignored, and its effect will be calculated in step (10). The logical activation function designed in this application is:
[0077]
[0078] G v = δ · G d (9)
[0079] In the formula, v is the dosage of drug exposure, e is the natural constant, and k, a, and b are all trainable parameters with an initial value of 1. Thus, when the dosage is 0, the value of the influence coefficient δ of drug dosage on gene expression is near 1, and the feature vector of the drug remains basically unchanged. G v is the gene feature vector integrating the influence of drug dosage.
[0080] (10) Use a graph convolutional network to calculate the influence of gene-gene interactions on the potential feature representation of genes
[0081] After obtaining the gene feature vector G v integrating the influence of drug dosage, this application uses convolution on the gene-gene interaction network to calculate the influence of gene-gene interactions on the potential feature representation of genes. The calculation method is graph convolutional calculation:
[0082]
[0083] In the formula is the standardized gene-gene interaction network, is 's degree matrix, σ is the ReLU activation function, and G u is the gene feature vector integrating the influence of gene-gene interactions. W uis a learnable parameter matrix.
[0084] (11) Calculate the gene expression matrix at different time periods of drug action
[0085] This application stipulates that the gene expression matrix calculated in step (10) for the first time is the potential feature vector of genes affected by drug action within a single time period (this application takes 3 hours as a fixed time period as an example. At this time, G u is the gene expression matrix G after 3 hours of drug administration u(3h) ). If you want to calculate the gene expression matrix after 6 hours of drug administration (in the next same time period), then use G after 3 hours of drug administration u(3h) to replace G in step (9) d and perform the integration of drug dose information calculation again to obtain the gene feature vector G after 6 hours of drug administration v(6h) , and then the method in step (10) will be used again for gene-gene interaction information integration to obtain the gene feature vector G after 6 hours of drug administration u(6h) . And so on, repeating the method in step (11) n times, the gene feature vector G after 3(n + 1) hours of drug administration can be obtained u(3(n+1)h) .
[0086] (12) Use the unified mapping module to convert the gene feature matrix into a gene expression matrix
[0087] After obtaining the gene feature vector Gu at the required drug administration time, use two non-linear transformations to convert it into the final model output, that is, the gene expression matrix P. The gene expression matrix P is a mathematical vector with 978 columns, and the values of the vector are the relative expression levels of 978 "representative genes" (https: / / clue.io / command?q= / gene-space%20lm) specified in the Connectivity Map project after 3(n + 1) hours of aspirin action (negative numbers represent a decrease in gene expression level, and positive numbers represent an increase in gene expression level). The relative expression level data of these 978 "representative genes" can largely replace the complete transcriptome and reflect the overall gene expression status of the measured samples.
[0088] The present invention integrates drug molecular structures, different types of drug-drug similarity networks, and gene-gene interaction networks, and uses graph convolutional algorithms to construct a prediction model for gene expression levels under drug exposure. By adding a drug-drug similarity network for complementary feature learning while retaining a complete analysis of drug molecular characteristics, it can more accurately predict gene expression levels at different drug doses and times. In addition, based on computer programs and algorithms, the present invention constructs a process for automatically generating integrated networks, enabling users to automatically obtain resources such as drug-drug similarity networks. Users only need to provide the drug CID number to perform the analysis and prediction of the complete process. It provides important tools and data support for research on human diseases, drug reutilization, etc.
Claims
1. Neural network model prediction method for gene expression profiles under drug exposure, characterized in that, it includes the following steps: (1) Select a drug to be analyzed, and obtain the self-information and interaction information of the drug to be analyzed, (2) Construct a drug-drug similarity network according to the drug interaction information obtained in step (1), (3) Convert the SMILES structure information in the self-information of the drug molecule to be analyzed into a graph representation, use the chemical bond type as the image representation of different channels of the image, and concatenate the data of the graph representation and the image representation together to form the feature vector of the drug to be analyzed, (4) Use the feature vector of the drug to be analyzed obtained in step (3) as a node on the drug-drug similarity network, and use the drug-drug similarity network obtained in step (2) to perform information aggregation to obtain the feature vector after aggregating the similar drug information of the drug to be analyzed, (5) Convert the feature vector after aggregating the similar drug information of the drug to be analyzed obtained in step (4) into the potential feature vector of the downstream gene, (6) Calculate the influence of the drug dose on the potential feature representation of the downstream gene, and integrate the drug dose influence into the potential feature vector of the downstream gene obtained in step (5) to obtain the gene feature vector integrated with the drug dose influence, When calculating the influence of the drug dose on the potential feature representation of the downstream gene in step (6), a logistic activation function is used to simulate the influence of the drug dosage on the gene expression level. The logistic activation function is: , , In the formula, v is the dose of drug exposure, e is the natural constant, k , , are all trainable parameters with an initial value of 1, is the influence coefficient of drug dose on gene expression, G d is the potential feature vector of the downstream gene obtained in step (5), G v is the gene feature vector integrating the influence of drug dose; (7) Calculate the influence of the gene-gene interaction pair on the expression level, and integrate it into the gene feature vector obtained in step (6) to obtain the gene feature vector after integrating the gene-gene interaction influence, (8) Convert the gene feature vector obtained in step (7) into the final gene expression matrix, which corresponds to the gene expression profile under drug exposure of the drug to be analyzed.
2. The neural network model prediction method for gene expression profiles under drug exposure according to claim 1, characterized in that, in step (1), the self-information of the drug to be analyzed includes the CID serial number and the SMILES sequence, and the interaction information includes the drug-gene interaction information in the DGIdb database, the drug-gene interaction information in the CTD database, the drug-disease association information, and the drug comprehensive similarity score data in the STITCH database.
3. The neural network model prediction method for gene expression profiles under drug exposure according to claim 2, characterized in that, in step (2), different drug-drug similarity networks are constructed according to the drug interaction information in different databases; in step (4), information aggregation is performed for each drug-drug similarity network respectively to obtain multiple feature vectors after information aggregation based on different drug-drug similarity networks. After inputting them into the transformer model for calculation, different importance coefficients are attached to each input feature vector, and they are weighted and averaged to aggregate the information unique to each drug-drug similarity network.
4. The neural network model prediction method for gene expression profiles under drug exposure according to claim 2, characterized in that, in step (2), the construction method of the drug-drug similarity networks in different databases is as follows: i. Method for constructing a drug-drug similarity network of a drug-gene interaction network based on the DGIdb database or the CTD database: Use common genes as a bridge to connect two drugs. If there are more than 3 identical genes shared between the drug to be analyzed and a certain drug, then the drug to be analyzed and that certain drug are similar drugs to each other, and the value in the adjacency matrix is 1; otherwise, it is 0. ii. Method for constructing a drug-drug similarity network based on a drug-disease interaction network of the CTD database: Use diseases as a bridge to connect two drugs. First, map disease information to gene information using the DisGeNET database, and then use common genes as a bridge to connect two drugs. If there are more than 3 identical genes shared between the drug to be analyzed and a certain drug, then the drug to be analyzed and that certain drug are similar drugs to each other, and the value in the adjacency matrix is 1; otherwise, it is 0. iii. Method for constructing a drug-drug similarity network based on the comprehensive similarity score of drugs in the STITCH database: Connect two drugs with the drug-drug similarity scores in the summarized STITCH database. If the similarity score between the drug to be analyzed and a certain drug is higher than 150, then there is a connection edge between these two drugs in the drug comprehensive similarity network based on the STITCH database.
5. The neural network model prediction method for gene expression profiles under drug exposure according to claim 1, characterized in that, in step (3), convert the SMILES structure into an atomic adjacency matrix, an atomic feature matrix, and a chemical bond feature matrix, and learn and merge these three matrices through a graph convolutional neural network to obtain a one-dimensional drug feature vector as the processed graph representation data.
6. The neural network model prediction method for gene expression profiles under drug exposure according to claim 5, characterized in that, in step (3), use the swim-transformer model to learn the data of the drug image representation and obtain another one-dimensional drug feature vector as the processed image representation data.
7. The neural network model prediction method for gene expression profiles under drug exposure according to claim 6, characterized in that, Use a graph convolutional neural network and a swim-transformer model to optimize the data of the drug graph representation and the image representation respectively, and obtain the potential feature vector of each atom in the drug to be analyzed; then merge the potential feature vectors of each atom to obtain the processed graph representation and image representation data of the drug to be analyzed.
8. The neural network model prediction method for gene expression profiles under drug exposure according to claim 1, characterized in that, in step (5), the feature vector aggregating the similar drug information of the drug to be analyzed obtained in step (4) is transformed into the potential feature vector of the downstream gene through two non-linear transformations; in step (8), the gene feature vector obtained in step (7) is transformed into the final gene expression matrix through two non-linear transformations.
9. The neural network model prediction method for gene expression profiles under drug exposure according to claim 1, characterized in that, When calculating the influence of gene-gene interaction on expression level in step (7), a graph convolutional network is used for the calculation.
Citation Information
Patent Citations
Systematic pharmacology method of personalized medication
CN106815486A
Drug prediction method, drug predication device and computer equipment
CN110310703A