HBV inhibitor screening method based on molecule-gene interaction constraint graph convolutional network
By applying a graph convolutional network model based on molecular-gene interaction constraints in HBV inhibitor screening, the problem of insufficient prediction accuracy and AUC values of compound molecular-gene interaction data in the prior art is solved, and efficient HBV inhibitor screening and prediction effects are achieved.
Patent Information
- Application Number
- CN202510115288.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The prior art has insufficient prediction accuracy and AUC values when processing complex compound molecule-gene interaction data, making it difficult to effectively screen out anti-HBV drug molecules.
Using a graph convolutional network (MGIC-GCN) model based on molecular-gene interaction constraints, a compound molecule-gene interaction map is constructed, and a graph convolution layer, global pooling layer and fully connected layer are used for model training to optimize hyperparameters to improve the prediction accuracy of the model.
The accuracy and AUC value of compound activity prediction have been significantly improved to reach 0.97, and new HBV inhibitors have been successfully screened out, providing new paths and ideas for the research and development of anti-HBV drugs.
Smart Images

Figure CN120108562A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of chemistry and biology, and relates to a method for predicting compound activity using a graph convolutional network (GCN) and its application, and in particular to a HBV inhibitor screening method based on a molecule-gene interaction-constrained graph convolutional network (MGIC-GCN), and its application in biopharmaceuticals, drug screening and precision medicine. Background Art
[0002] Hepatitis B virus infection is a serious public health problem worldwide. Although existing vaccines and antiviral drugs can effectively inhibit viral replication, it is difficult to completely eliminate the virus, and many patients require long-term treatment. Therefore, it is urgent to discover new anti-HBV drug molecules and study their mechanisms of action. Traditional drug discovery methods such as high-throughput screening and drug design based on known molecular structures face many challenges, such as high experimental costs, long cycles, and results affected by multiple factors. In recent years, graph convolutional networks have become an important tool for processing biological network data due to their powerful graph data processing capabilities. However, the existing GCN model still has limitations in processing complex compound molecule-gene interaction data. To this end, the present invention proposes a compound activity prediction method based on MGIC-GCN, which aims to improve the model prediction accuracy and AUC, and provide a new approach for the research and development of anti-HBV drugs. Summary of the invention
[0003] The purpose of the present invention is to address the deficiencies in the prior art and to propose a method for screening HBV inhibitors based on a compound molecule-gene interaction constrained graph convolutional network.
[0004] To achieve the above object, the present invention provides the following technical solutions: As a first aspect, a method for screening HBV inhibitors based on a molecule-gene interaction constrained graph convolutional network comprises the following steps:
[0005] S1, based on the existing database, collect compounds that can act on multiple HBV-related targets. Specifically, screen out compound molecules in the CHEMBL and Connectivity Map (CMAP) libraries that are related to HBV (hepatitis B virus) infection-related genes and their targets, and their related gene target indexes from the compound library. Ensure that the selected compounds are relevant to HBV infection. The multiple HBV-related targets are specifically: more than three.
[0006] S2, conducting in vitro anti-HBV activity tests and cytotoxicity tests on the above compounds, using the anti-HBV activity and cytotoxicity as indicators to construct activity labels for the compounds and establish a data set.
[0007] S3, data processing: test and extract the association data between compounds and genes and their targets, and construct a compound molecule-gene interaction graph; based on the interaction information between genes and their targets, construct a target interaction matrix, in which each element represents the interaction strength between two targets, and set a threshold to extract the edges of the graph. Extract the features of related genes and their targets from the gene and target feature matrix as node features of the graph convolutional network model. Construct a target feature matrix, in which each row represents the four-dimensional feature vector of the target.
[0008] S4, graph data generation: For each compound in the data set, construct a network graph of genes and their targets, including node features, edge indexes and edge weights; the nodes of the target network graph are targets related to the compound, the node features represent the four-dimensional features of the target, the edges represent the interaction between the two targets, and for targets whose interaction strength between the two targets is greater than the threshold, an edge index is established. The interaction strength between the two targets is used as the edge weight, and the edge weight is converted into an edge type. The generated node features, edge indexes, edge types and active labels are integrated into a graph data object. Specifically, the interaction strength between the two targets is measured by the interaction score.
[0009] S5, graph convolutional network model training: training based on the graph convolutional network model to obtain an anti-HBV activity prediction model. Specifically, the graph convolution layer, global pooling layer and fully connected layer are used for training to classify the activity of the compound. During the model training process, the cross entropy loss function is used to optimize the model parameters. The data set is randomly divided into a training set, a validation set and a test set. The training set is used to train the prediction model. The validation set verifies the performance of the model during the training process. The test set is used as an independent validation set to finally judge the performance of the model.
[0010] S6, Hyperparameter Optimization: Optuna is used to optimize hyperparameters, including the number of hidden layer neurons and learning rate. The accuracy of the model is improved by searching for the optimal combination of the number of hidden layer neurons and learning rate.
[0011] S7, model evaluation: The prediction models trained by various graph convolutional network models were verified and compared. The model performance was evaluated by indicators such as accuracy, F1 score, recall rate and AUC, and the optimal prediction model was selected. Finally, the advantages of the MGIC-GCN model in the task of compound activity prediction were verified.
[0012] S8, HBV inhibitor screening: Based on the anti-HBV activity prediction model, the test compound is input for prediction, and the test compound with anti-HBV activity and no cytotoxicity is screened to obtain HBV inhibitors.
[0013] Furthermore, in step S1, genetic data related to HBV infection are collected from databases such as The Comparative Toxicogenomics Database (CTD) and The National Center for Biotechnology Information (NCBI), and screened from the FDA_HY-L022 Library compound library to select compounds with more than three HBV infection-related targets in the CHEMBL and CMAP libraries and include them in the data set.
[0014] Furthermore, in step S2, the selected compounds were systematically evaluated for activity using HepG2.2.15 cells as an experimental model. The in vitro anti-HBV activity test specifically includes: determining the inhibitory effect (inhibition rate) of the compound on HBsAg (hepatitis B surface antigen) secretion and HBeAg (hepatitis B e antigen) secretion; the cytotoxicity test specifically includes: determining the cell survival rate of the compound using the in vitro MTT method.
[0015] (1) Dataset and preprocessing Compound-target association data: The compound-target association data file records the relationship between each compound and multiple targets. Let the compound set be C = {c1, c2, ..., cn}, and each compound ci corresponds to a target set G(ci) = {g1, g2, ..., gm}, where gj represents the target associated with compound ci.
[0016] (2) Activity signature: The activity signature of a compound is derived based on the activity data analysis and reflects its biological activity category.
[0017] The activity label of the compound is constructed by taking whether it has anti-HBV activity and whether it has cytotoxicity as indicators, specifically: HBsAg or HBeAg secretion inhibition rate ≥50% means it has anti-HBV activity; cell survival rate ≥50% means it has low cytotoxicity; the activity label of the compound is {0, 1, 2, 3}, wherein label 0 means HBsAg or HBeAg secretion inhibition <50% (no anti-HBV activity) and cell survival rate <50% (high cytotoxicity); 1 means HBsAg or HBeAg secretion inhibition ≥50% (anti-HBV activity) and cell survival rate <50% (high cytotoxicity); 2 means HBsAg or HBeAg secretion inhibition rate ≥50% (anti-HBV activity) and cell survival rate ≥50% (low cytotoxicity), 3 means HBsAg or HBeAg secretion inhibition rate <50% (no anti-HBV activity) and cell survival rate ≥50% (low cytotoxicity). Therefore, this problem is defined as a four-classification task.
[0018] (3) Gene-target interaction matrix
[0019] The target interaction matrix is obtained through compound gene and target interaction data (InteractionScore provided by StringDB database) and community analysis. The interaction matrix M is a p*p matrix, where p represents the number of all genes, and the element Mij in the matrix represents the interaction strength between genes gi and gj. In order to generate the graph data of each compound, we extract the submatrix MG(ci) of the gene set G(ci) related to the compound, which defines the target network of compound ci:
[0020] M G(ci) ={M jk |g j ,g k ∈G(ci)}
[0021] In the interaction matrix, the threshold T = 400 is set, that is, if Mjk>400, the target g j and g k There is an edge between them, otherwise there is no edge.
[0022] (4) Target feature data
[0023] After clustering the scattered targets, we set four-dimensional features for each target, including cytotoxicity and anti-HBV activity indicators. The target feature matrix F is a p*4 matrix, in which each row represents the four-dimensional features of a target. For each compound ci, we extract the feature submatrix F of its related targets from the target feature matrix G(ci) , each row of the matrix corresponds to a four-dimensional feature vector fj of a target.
[0024] Furthermore, in the step S4 of generating graph data, target network graph data is generated for each compound.
[0025] (1) Nodes and node features
[0026] In the graph data of each compound, the node represents the target related to the compound, and the node feature is the four-dimensional feature of the target. i The feature matrix of the relevant target is F G(ci) , then the matrix Where m is compound c i The number of associated targets, 4 represents the four-dimensional features of each target.
[0027] (2) Edge index and edge weight
[0028] Based on the target interaction matrix M G(ci) , we extract the edge index and edge weight of each compound graph data, set the threshold T, then the edge index Ei is defined as:
[0029] Ei={(g j ,g k )|M jk >T}
[0030] The edge weight Ew is M jk For the R-GCN model, we further convert the edge weight into the edge type Et. The edge weight is converted into the edge type using the threshold partitioning method. Specifically, let the edge type function be f ω (M jk ):
[0031]
[0032] The edge type assigned by this function is suitable for R-GCN functions, which handles the problem of multiple edge types in R-GCN.
[0033] (3) Graph Data Representation
[0034] The graph data can be represented as G = (X, Ei, Et, y), where is the target node feature matrix; Ei is the edge index, indicating that the interaction strength is greater than the threshold, or the target pair obtained after transformation; Et is the type of edge in the R-GCN model, and y∈(0,1,2,3) is the activity label of the compound.
[0035] Furthermore, in step S5, the graph convolutional network model training includes the steps of:
[0036] (1) Graph data generation: For each compound ci, extract the set of related targets G(ci); then extract the submatrix MG(ci) of G(ci) from the target interaction matrix, and then generate edge indices and edge types based on the interaction matrix; then extract the feature submatrix FG(ci) of the related targets from the target feature matrix; then integrate the generated node features, edge indices, edge types and activity labels into a graph data object. Through the above method, we constructed graph data for GCN, GIN, GAT and R-GCN models, and used them to train models to predict the activity categories of compounds.
[0037] (2) Model definition: Four graph convolutional network models are used, namely GCN (Graph Convolutional Network), GAT (Graph Attention Network), GIN (Graph Isomorphism Network) and RGCN (Relational Graph Convolutional Network). The four models share similar structures in the overall architecture, including graph convolution layer, global pooling operation and fully connected classification layer. Its basic operation can be summarized by the message passing mechanism, that is, the features of the node will interact with the features of the neighboring nodes, and the node representation will be updated through the convolution operation.
[0038] The node features in the graph convolutional network are represented as a matrix Where N is the number of nodes and F is the node feature dimension. The structure of the graph is represented by the adjacency matrix Description, where A ij Indicates whether there is a connection between node i and node j. In general, the basic convolution operation of the graph convolutional network can be abstracted as the following formula:
[0039] H (l+1) =f(H (l) ,A)
[0040] Among them, H (l) Represents the node feature matrix of the lth layer, and the initial feature matrix is H (0) =X, f is the convolution operation function, which aggregates the features of the node and its neighbors and updates the feature representation through nonlinear transformation. After several layers of graph convolution operations, we perform global pooling operations (such as global average pooling, global maximum pooling, etc.) on all nodes to generate the overall representation of the graph. Global pooling aggregates all node features into a single vector h g :
[0041] h g =GMP(H (L) )
[0042] Among them, H (L)is the node representation after the last layer of convolution. Finally, the pooled graph representation is classified through one or more fully connected layers, and the output of the model is optimized by the cross entropy loss function:
[0043]
[0044] Where C is the number of categories, y i is the true category label, p i The model predicts probability. Based on the above architecture, the main difference between the four models lies in the specific implementation of their convolution operations.
[0045] Furthermore, in the step S6 of hyperparameter optimization, the selection of hyperparameters has a crucial impact on the performance of the model. In this study, Optuna is used as a hyperparameter optimization tool to automatically search for the optimal model configuration through multiple experiments.
[0046] (1) Search space definition: Hidden Channels: The number of neurons in the hidden layer of the model. This value affects the feature extraction capability of each convolution layer and its value range is 8≤hidden_channels≤32.
[0047] Learning Rate: The learning rate controls the step size of the model weight update. Too large a learning rate may cause the model to be unstable, while too small a learning rate will slow down the model convergence. The search range of the learning rate is 10 -6 ≤learning_rate≤10 -3 , whose logarithm is defined as: learning_rate=10log_learning_rate,log_learning_rate∈[-6,-3].
[0048] (2) Objective function: In each experiment, Optuna extracts a set of hyperparameters from the search space, applies them to the model, and evaluates the model performance. This performance indicator (usually the accuracy of the test set) is used as the output of the objective function for adjusting the hyperparameters in subsequent iterations. The objective function is defined as: accuracy = f(hidden_channels, learning_rate) where f represents the process of model training and evaluation. Hyperparameter search process: Through the Optuna framework, multiple experiments are run to optimize the hyperparameters. In each experiment, Optuna selects a new set of hyperparameters from the search space and calculates the accuracy of the test set through the training process.
[0049] Furthermore, in the step S7 model evaluation, the performance of the MGIC series graph convolutional model and other mainstream machine learning models were comprehensively compared, and all models used the same sample data and adopted different data processing methods. The evaluation indicators included accuracy, F1 score, recall rate and AUC.
[0050] As a second aspect, the present invention provides HBV inhibitors obtained by screening based on a molecule-gene interaction constrained graph convolutional network, including Prostaglandin E2, Febuxostat, and Baricitinib, with the following structural formula:
[0051]
[0052] The beneficial effects of the present invention include: In view of the limitations of traditional drug discovery methods in processing complex biological data, the present invention utilizes the graph data processing capabilities of the graph convolutional network (GCN) to construct an MGIC-GCN model. The model is based on the constructed anti-HBV in vitro activity verification compound library, and the compound and HBV infection-related target index library, combined with the target protein interaction matrix, target protein feature matrix and compound activity label, to effectively predict the biological activity category of the compound. The specific steps include data processing, graph data generation, graph convolution network model training, hyperparameter optimization and model evaluation. The model parameter AUC value is 0.97, and the model effect is good. The present invention provides a new path and idea for the virtual screening of anti-HBV drugs, and has potential application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is the structure diagram of the MGIC-GCN model. DETAILED DESCRIPTION
[0054] The present invention establishes a data set of compound anti-HBV compound activity, collects association data between compounds and targets, interaction data between targets and target feature data, and constructs a target interaction network diagram; then selects three graph convolutional network models, GCN, GIN and RGCN, as the MGIC series graph convolutional network models for training to mine compound activity patterns; experimental verification shows that compared with traditional machine learning models, MGIC-GCN shows significant advantages in both accuracy and AUC indicators, and can more accurately process complex compound molecule-gene interaction graph structures, effectively improving the accuracy of compound activity prediction.
[0055] The following combination Figure 1 , the technical solutions and beneficial effects of the present invention are described in detail, wherein any unmentioned contents are prior art or contents that can be set or adjusted by those skilled in the art as needed.
[0056] The present invention proposes a method for predicting the activity of MGIC-GCN compounds, which specifically comprises the following steps:
[0057] S1, 12,200 genes related to HBV (hepatitis B virus) infection were collected from databases such as The Comparative Toxicogenomics Database (CTD) and The National Center for Biotechnology Information (NCBI). The FDA_HY-L022 Library compound library was screened to select compounds with more than three genes and targets related to HBV infection in the CHEMBL and CMAP libraries. Ensure that the selected compounds are related to HBV infection. 105 compounds that meet the criteria were obtained.
[0058] S2, in vitro anti-HBV activity test for these 105 compounds: The sample solution was prepared by dissolving the aforementioned 105 compounds in DMSO, and the solution concentration was 30 μM. The ELISA method was used to determine the effects of the compounds on the secretion of HBeAg and HBsAg antigens. The specific experimental operation steps are as follows: First, the HepG2.2.15 cells with appropriate cell status were washed with PBS, trypsin was added to disperse the cells, and culture medium (MEM+10% FBS+380mg / mL G418) was added and pipetted into a single cell suspension; then, the suspension was transferred to an EP tube (15mL), centrifuged to remove the supernatant, and then 2mL of complete culture medium (MEM+10% FBS+380μg / mL G418) was added and pipetted into a single cell suspension, and re-inoculated in a 48-well plate, with about 3000 cells per well, and placed in an incubator at 37°C, 5% CO 2 Finally, the culture medium was discarded, and the cells were treated with the prepared sample solution (final concentration 30 μM), lamivudine was used as a positive control, and cells without sample treatment were used as blank controls. The cells were placed in an incubator (37°C, 5% CO 2 ) and the culture fluid was collected after 72 hours of culture. The samples were tested using the HBeAg detection kit (Shanghai Kehua Bioengineering Co., Ltd., National Medical Device No. 20163400144) and the HBsAg detection kit (Shanghai Kehua Bioengineering Co., Ltd., National Medicine No. S10910113). The absorbance value (OD value) was measured using an ELISA reader at a detection wavelength of 450nm and a reference wavelength of 630nm. The inhibition rate of the compound on HBsAg and HBeAg was calculated according to the formula as follows:
[0059]
[0060] S3, in vitro MTT assay to determine compound cell viability
[0061] (1) HepG2.2.15 cells in appropriate cell states were washed with PBS, the cell culture medium was removed, 0.25% trypsin was added to digest the cells, culture medium (MEM+10% FBS+380 mg / mL G418) was added and the cells were pipetted into a single cell suspension; then, the suspension was transferred to an EP tube (15 mL), centrifuged to remove the supernatant, and 2 mL of complete culture medium (MEM+10% FBS+380 mg / mL G418) was added and the cells were pipetted into a single cell suspension, which was then re-seeded in a 96-well plate with approximately 3,000 cells per well;
[0062] (2) After 12 hours of adhesion, the culture medium was removed and 100 μL of culture medium (MEM + 10% EBS + 380 mg / mL G418) containing different concentrations of the test sample (diluted from below 200 μM to 1.56 μM) was added. Blank wells (containing only culture medium) and control wells (solvent control without adding the test sample) were also set up;
[0063] (3) Cells and drugs were maintained at 37°C, 5% CO 2 After 24 hours of co-incubation, 10 μL of MTT solution (10 mg / mL) was added to each well. After 4 hours of co-incubation, 100 μL of MTT solution was added. After incubation at 37°C overnight, the absorbance (A value) of each well was measured at a wavelength of 550 nm.
[0064] (4) Calculate the cell survival rate. Survival rate = (experimental well - blank well) / (control well - blank well) × 100%. The experiment was repeated three times.
[0065] The results of anti-HBV activity analysis (Table 1) showed that 38 small molecules had an inhibition rate of >50% on cell secretion of HBsAg, and 2 small molecules had an inhibition rate of >50% on HBeAg. Among them, there were 11 compounds with a cell survival rate greater than 50% detected by MTT, indicating that their anti-HBV activity did not affect cell survival rate.
[0066] Table 1. HBV inhibitor activity results
[0067]
[0068]
[0069] Table 2. Anti-HBV active compound molecules
[0070]
[0071]
[0072]
[0073]
[0074] S4, Dataset and preprocessing: In order to build a graph convolutional network (GCN) model for predicting compound activity, we first extracted the association information between compounds and targets, the interaction data between targets, and the target feature data from multi-source data. With these data, a target interaction network graph corresponding to each compound is generated. The nodes of each network graph represent targets, the edges represent the interactions between targets, and the node features are the four-dimensional features of the targets. The ultimate goal is to predict the activity category of the compound through graph data.
[0075] (1) Based on CHEMBL and CMAP libraries, compound (Table 3) and target association data: The compound and target association data file records the relationship between each compound and multiple targets. Let the compound set be C = {c1, c2, ..., cn}, and each compound ci corresponds to a target set G(ci) = {g1, g2, ..., gm}, where gj represents the target associated with compound ci.
[0076] (2) Activity label: The activity label of the compound is obtained based on the analysis of activity data and reflects its biological activity category. HBsAg secretion and HBeAg secretion inhibition rate ≥50% indicates an inhibitory effect; cell survival rate ≥50% indicates a low cytotoxic effect; among them, label 0 indicates HBsAg secretion and HBeAg secretion inhibition <50% and cell survival rate <50%; 1 indicates HBsAg secretion and HBeAg secretion inhibition ≥50% and cell survival rate <50%; 2 indicates HBsAg secretion and HBeAg secretion inhibition rate ≥50% and cell survival rate ≥50%; 3 indicates HBsAg secretion and HBeAg secretion inhibition rate <50% and cell survival rate ≥50%. Therefore, this problem is defined as a four-classification task.
[0077] (3) Target interaction matrix
[0078] We obtained the target interaction matrix through compound target community analysis. The interaction matrix M is a p*p matrix, where p represents the number of all targets, and the element Mij in the matrix represents the interaction strength between targets gi and gj. The present invention uses the interaction score provided by the stringDB database as a measure of the interaction strength between targets. In order to generate the graph data of each compound, we extract the submatrix MG(ci) of the target set G(ci) related to the compound, which defines the target network of compound ci:
[0079] M G(ci)={M jk |g j ,g k ∈G (ci)}
[0080] In the interaction matrix, the threshold T = 400 is set, that is, if Mjk>400, the target g j and g k There is an edge between them, otherwise there is no edge.
[0081] (4) Target feature data
[0082] After clustering the scattered targets, we set a four-dimensional feature for each target, including toxicity and activity indicators. The present invention uses four labels of anti-HBV activity data 0, 1, 2, and 3. The target feature matrix F is a p*4 matrix, in which each row represents the four-dimensional feature of a target. For each compound ci, we extract the feature submatrix F of its related target from the target feature matrix G(ci) , each row of the matrix corresponds to a four-dimensional feature vector fj of a target.
[0083] S5, in the generation of graph data, based on the above steps, target network graph data is generated for each compound.
[0084] (1) Nodes and node features
[0085] In the graph data of each compound, the node represents the target related to the compound, and the node feature is the four-dimensional feature of the target. i The feature matrix of the relevant target is F G(ci) , then the matrix Where m is compound c i The number of associated targets, 4 represents the four-dimensional features of each target.
[0086] (2) Edge index and edge weight
[0087] Based on the target interaction matrix M G(ci) , we extract the edge index and edge weight of each compound graph data, set the threshold T, then the edge index Ei is defined as:
[0088] Ei={(g j ,g k )|M jk >T}
[0089] The edge weight Ew is M jk For the R-GCN model, we further convert the edge weight into the edge type Et, and set the edge type function to be f ω (M jk ):
[0090]
[0091] The edge type assigned by this function is suitable for R-GCN functions, which handles the problem of multiple edge types in R-GCN.
[0092] (3) Graph Data Representation
[0093] The graph data can be represented as G = (X, Ei, Et, y), where is the target node feature matrix; Ei is the edge index, indicating that the interaction strength is greater than the threshold, or the target pair obtained after transformation; Et is the type of edge in the R-GCN model, and y∈(0,1,2,3) is the activity label of the compound.
[0094] S6, graph convolutional network model training includes the following steps:
[0095] (1) Graph data generation: First, for each compound ci, extract the relevant target set G(ci); then extract the submatrix MG(ci) of G(ci) from the target interaction matrix, and then generate edge indices and edge types based on the interaction matrix; then extract the feature submatrix FG(ci) of the relevant targets from the target feature matrix; then integrate the generated node features, edge indices, edge types and activity labels into a graph data object. Through the above method, we constructed graph data for GCN, GIN, GAT and R-GCN models, and used them to train models to predict the activity category of compounds.
[0096] (2) MGIC series graph convolutional network model definition: Four graph convolutional network models are used, namely GCN (Graph Convolutional Network), GAT (Graph Attention Network), GIN (Graph Isomorphism Network) and RGCN (Relational Graph Convolutional Network). Although these models have their own innovative features, they share similar structures in the overall architecture, including graph convolution layers, global pooling operations and fully connected classification layers. Its basic operations can be summarized by the message passing mechanism, that is, the features of a node will interact with the features of neighboring nodes, and the node representation will be updated through convolution operations.
[0097] The node features in the graph convolutional network are represented as a matrix Where N is the number of nodes and F is the node feature dimension. The structure of the graph is represented by the adjacency matrix Description, where A ij Indicates whether there is a connection between node i and node j. In general, the basic convolution operation of the graph convolutional network can be abstracted as the following formula:
[0098] H (l+1) =f(H (l) ,A)
[0099] Among them, H (l) Represents the node feature matrix of the lth layer, and the initial feature matrix is H (0) =X, f is the convolution operation function, which aggregates the features of the node and its neighbors and updates the feature representation through nonlinear transformation.
[0100] After several layers of graph convolution operations, we perform global pooling operations (such as global average pooling, global maximum pooling, etc.) on all nodes to generate an overall representation of the graph. Global pooling aggregates all node features into a single vector h g :
[0101] h g =GMP(H (L) )
[0102] Among them, H (L) is the node representation after the last layer of convolution. Finally, the pooled graph representation is classified through one or more fully connected layers, and the output of the model is optimized by the cross entropy loss function:
[0103]
[0104] Where C is the number of categories, y i is the true category label, p i The model predicts the probability. Based on the above architecture, the main difference between the four models lies in the specific implementation of their convolution operations. The following introduces the unique features of each model.
[0105] Graph Convolutional Network (GCN): uses convolution operations based on spectral graph theory. The convolution operation is implemented by normalizing the adjacency matrix. The normalized adjacency matrix A is defined as:
[0106]
[0107] in It is to add self-loops (i.e. the connection between the node itself and itself) to the original adjacency matrix, and D is the corresponding degree matrix. The convolution operation of GCN is expressed as:
[0108] H (l+1) =σ(AH (l) W (l) )
[0109] Where W (l)is the learnable weight matrix of the lth layer, and σ is the activation function (usually ReLU). GCN uses a normalized adjacency matrix to consider the features of the node itself and its neighbors when aggregating features, while preventing the gradient vanishing problem caused by excessive degree values.
[0110] Graph Attention Network (GAT) introduces a self-attention mechanism that allows the model to assign different weights to different neighbor nodes when aggregating neighbor node features. Unlike GCN, which uses a fixed adjacency matrix, GAT's attention mechanism enables each node to dynamically decide which neighbors it should pay attention to. Specifically, GAT uses a learnable attention coefficient α ij Represents the correlation between node i and node j. The coefficient is calculated by the following formula:
[0111]
[0112] in:
[0113]
[0114] z ij =a T [Wh i ||Wh j ]
[0115] Where a is the weight vector used to calculate the attention score, W is the feature transformation matrix, || represents the concatenation operation, and h i and h j is the feature vector of node i and node j, LeakyReLU is an activation function, and N(i) is the set of neighbor nodes of node i. The final node feature aggregation is completed through the multi-head attention mechanism:
[0116]
[0117] Where K is the number of attention heads and ∪ represents the concatenation operation.
[0118] Graph Isomorphic Network (GIN): Designed to have a stronger ability to distinguish graph structures, the goal is to make the model have unique representation capabilities when processing different graph structures. The convolution operation of GIN is similar to weighted summation, which updates the node representation by aggregating the features of the node itself and its neighboring nodes. The specific formula of GIN convolution is:
[0119]
[0120] Where ∈ is a learnable or fixed scalar parameter, and MLP (Multi-layer Perceptron) is used to perform nonlinear transformation on the aggregated features. The innovation of GIN is to use MLP to enhance the feature expression ability, and to adjust ∈ to enable the model to flexibly handle the combination of node features and neighbor features.
[0121] The Relational Graph Convolutional Network (RGCN) is used to handle graph structures with heterogeneous relations, that is, the edges between nodes can represent different types of relations. To this end, RGCN designs independent convolution kernel weights Wr for each relation type r. The convolution operation of RGCN is expressed as:
[0122]
[0123] S7,In hyperparameter optimization, the selection of hyperparameters has a crucial impact on the performance of the model. In this study, Optuna is used as a hyperparameter optimization tool to automatically search for the optimal model configuration through multiple experiments.
[0124] (1) Search space definition:
[0125] Hidden Channels: The number of neurons in the hidden layer of the model. This value affects the feature extraction capability of each convolution layer, and the value range is 8≤hidden_channels≤32.
[0126] Learning Rate: The learning rate controls the step size of the model weight update. Too large a learning rate may cause the model to be unstable, while too small a learning rate will slow down the model convergence. The search range of the learning rate is 10 -6 ≤learning_rate≤10 -3 , whose logarithm is defined as: learning_rate=10log_learning_rate,log_learning_rate∈[-6,-3].
[0127] (2) Objective function:
[0128] In each trial, Optuna extracts a set of hyperparameters from the search space, applies them to the model, and evaluates the model performance. This performance metric (usually the accuracy of the test set) is used as the output of the objective function for adjusting the hyperparameters in subsequent iterations. The objective function is defined as: accuracy = f(hidden_channels, learning_rate) where f represents the process of model training and evaluation.
[0129] Hyperparameter search process: Using the Optuna framework, we ran multiple experiments to optimize the hyperparameters. In each experiment, Optuna selected a new set of hyperparameters from the search space and calculated the accuracy of the test set through the training process.
[0130] Finally, we get the best hyperparameter combination. For example, the number of hidden channels after optimization is 19 and the learning rate is 7.244×10 -5 .
[0131] S8, Model Training Evaluation: The performance of the four MGIC graph convolutional network models was comprehensively compared with other mainstream machine learning models. All models used the same sample data and adopted different data processing methods. The evaluation indicators include accuracy, F1 score, recall rate and AUC. According to the experimental results, the GCIC series showed obvious advantages, especially in terms of accuracy and AUC. Specifically, the GCIC series model achieved an accuracy of 0.9137 and an AUC value of 0.9787, indicating that GCIC has a very strong predictive ability in this task, can effectively classify samples, and better distinguish between positive and negative samples. Compared with traditional machine learning models, GCIC significantly outperforms models such as XGBoost, support vector machine (SVM) and random forest (RF) in terms of accuracy and AUC indicators. For example, XGBoost has an accuracy of 0.7692 and an AUC of 0.5054; SVM has an accuracy of 0.7692 and an AUC of 0.5830; while random forest has the weakest performance, with an accuracy of 0.7308 and an AUC of 0.6182. This shows that the GCIC series can better handle complex data structures, especially when processing graph structured data, showing its strong advantages in this task. In contrast, the performance gap between GCIC and DeepBindGCN is small. GCIC has an accuracy of 0.9021 and an AUC of 0.9648, which is close to DeepBindGCN. This shows that the GeneNet series is competitive with DeepBindGCN in terms of accuracy and stability of task performance.
[0132] Table 3. Model parameter table
[0133]
[0134] S9 model prediction and activity verification: The MGIC-GCN model was used to predict the activity of data that did not appear in the molecular-gene network dataset to predict whether the compound was a label "2" condition, "2" indicating that the HBsAg secretion and HBeAg secretion inhibition rate ≥ 50% cell survival rate ≥ 50%. 13 compounds were tested, and the verification results are shown in Table 4. Five compounds had anti-HBV activity, of which three compounds had an inhibitory effect on HBeAg secretion, and three of them had a cell survival rate ≥ 50%. The relevant structures are shown in Table 5.
[0135] Table 4. MGIC-GCN model prediction and verification results
[0136]
[0137] Table 5. Compound structures predicted and verified by MGIC-GCN model
[0138]
[0139] In summary, the present invention proposes a HBV inhibitor screening method based on a molecule-gene interaction constrained graph convolutional network, with a model parameter AUC value of 0.97, good model effect, and successful screening of new HBV inhibitors. The present invention provides a new path and idea for the virtual screening of anti-HBV drugs and has potential application value.
Claims
1. A method for screening HBV inhibitors based on a molecule-gene interaction constrained graph convolutional network, characterized in that: The following steps are involved: Based on the existing database, compounds that can act on multiple HBV-related targets are collected for in vitro anti-HBV activity testing and cytotoxicity testing. The activity signature of the compound is constructed based on whether it has anti-HBV activity and cytotoxicity as indicators, and a data set is established; Construct a target interaction matrix, in which each element represents the interaction strength between two targets; construct a target feature matrix, in which each row represents the four-dimensional feature vector of the target; For each compound in the data set, the node is set as the target related to the compound, the node feature is the four-dimensional feature of the target, and an edge index is established for the target whose two-target interaction strength is greater than a threshold. The two-target interaction strength is used as the edge weight, and the edge weight is converted into an edge type; the generated node features, edge indexes, edge types and activity labels are integrated into a graph data object; training is performed based on the graph convolutional network model to obtain an anti-HBV activity prediction model, and the test compound is input for prediction, and the test compound with anti-HBV activity and no cytotoxicity is screened to obtain HBV inhibitors.
2. The HBV inhibitor screening method based on molecule-gene interaction constrained graph convolutional network according to claim 1, characterized in that: The multiple HBV-related targets are specifically: more than three; The in vitro anti-HBV activity test specifically includes: determining the inhibition rate of the compound on HBsAg and HBeAg; The cytotoxicity test specifically includes: using an in vitro MTT method to determine the cell viability of the compound; The activity label of the compound is constructed by taking whether it has anti-HBV activity and whether it has cytotoxicity as indicators, specifically: HBsAg or HBeAg inhibition rate greater than or equal to 50% is set as anti-HBV activity, and cell survival rate greater than or equal to 50% is set as low cytotoxicity; the activity label of the compound is {0, 1, 2, 3}, wherein 0 indicates no anti-HBV activity and high cytotoxicity, 1 indicates anti-HBV activity and high cytotoxicity, 2 indicates anti-HBV activity and low cytotoxicity, and 3 indicates no anti-HBV activity and low cytotoxicity.
3. The HBV inhibitor screening method based on molecule-gene interaction constrained graph convolutional network according to claim 1, characterized in that: The edge weight is converted into an edge type by using a threshold partitioning method.
4. The HBV inhibitor screening method based on molecule-gene interaction constrained graph convolutional network according to claim 1, characterized in that: The anti-HBV activity prediction model is obtained by training based on the graph convolutional network model, specifically: the graph convolutional network model is trained using a graph convolution layer, a global pooling layer and a fully connected layer, and the model parameters are optimized using a cross entropy loss function; the data set is randomly divided into a training set, a validation set and a test set, the training set is used to train the prediction model, the validation set is used to verify the performance of the model during the training process, and the test set is used as an independent validation set to ultimately judge the performance of the model.
5. The HBV inhibitor screening method based on molecule-gene interaction constrained graph convolutional network according to claim 1, characterized in that: The method of training the graph convolutional network model to obtain the anti-HBV activity prediction model also includes: using Optuna to optimize hyperparameters, and the optimized hyperparameters include the number of hidden layer neurons and the learning rate.
6. The HBV inhibitor screening method based on molecule-gene interaction constrained graph convolutional network according to claim 1, characterized in that: The method of training based on a graph convolutional network model to obtain an anti-HBV activity prediction model also includes: verifying and comparing the prediction models trained using multiple graph convolutional network models, and selecting the optimal prediction model based on indicators such as Accuracy, F1Score, Recall, and AUC.
7. The HBV inhibitor screening method based on molecule-gene interaction constrained graph convolutional network according to claim 1, characterized in that: The interaction strength between the two targets is measured using the interaction score.
8. A HBV inhibitor obtained by screening based on a molecule-gene interaction constrained graph convolutional network, characterized in that: Including Prostaglandin E2, Febuxostat, Baricitinib, the structural formula is as follows:
Citation Information
Patent Citations
New coronavirus target prediction and drug discovery method based on graph representation learning
CN111916145A
Drug ATCCode prediction method based on graph transformation network
CN114420310A
Drug-target affinity prediction method and system based on multilevel feature fusion
CN119007863A
Construction method for analyzing pulmonary fibrosis target caused by inflammation caused by action of medicinal components of phellinus igniarius
CN119296634A
System and method for identifying therapeutics for a given illness using machine learning
US20230377681A1