A method for predicting the inhibitory activity of chemicals on DNA gyrase

Through a multi-task graph neural network model, using chemical SMILES code input to extract molecular and molecular fragment features, the problem of low chemical prediction accuracy in existing technologies is solved, and high-precision inhibitory activity prediction of DNA gyrase is achieved, supporting the rapid screening of new environmental pollutants and bacterial resistance assessment.

CN119851808BActive Publication Date: 2025-10-03ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510047496.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-10-03
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Existing methods for predicting the inhibitory activity of chemicals on DNA gyrase have problems of low accuracy and narrow scope of applicability, making it difficult to quickly screen high-risk pollutants with the ability to induce bacterial quinolone resistance.

Method used

A multi-task graph neural network model was constructed with the chemical SMILES code as input. The molecular graph and molecular fragment graph embedding layer were used to extract features, and the neighborhood attention module and information aggregation module were combined to update the features. The graph convolutional attention layer was used to capture the intrinsic connections, and finally the task-specific predictor was used to output the scaling value of the inhibitory activity data of chemicals on GyrA and GyrB.

Benefits of technology

It has achieved high-precision prediction of the inhibitory activity of chemicals on DNA gyrase, improved the learning ability and predictive performance of the model, and can quickly discover new environmental pollutants that induce bacterial resistance, which has broad application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851808B_ABST
    Figure CN119851808B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting the inhibitory activity of chemicals against DNA gyrase. The method comprises: obtaining chemical structures and data on their inhibitory activity against GyrA and GyrB from public databases, performing structural cleaning and scaling, and then dividing the obtained known data into training and test sets; constructing a multi-task graph neural network model using the chemical's SMILES code as input; determining the optimal hyperparameters of the multi-task graph neural network model through cross-validation using the training set to obtain an optimal inhibitory activity prediction model; and using the inhibitory activity prediction model to obtain predicted values ​​for the scaled values ​​of the inhibitory activity data for the chemical to be evaluated against GyrA and GyrB. This method can simultaneously predict the inhibitory effects of environmental chemicals on both GyrA and GyrB using only the molecular structure of the compound as input, enabling rapid screening of molecular initiation events that may induce bacterial quinolone resistance, and has broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chemical-induced bacterial resistance assessment, and in particular to a method for predicting the inhibitory activity of chemicals on DNA gyrase. Background Art

[0002] Because DNA gyrase is a key target for inducing quinolone resistance, the interaction between pollutants and DNA gyrase also determines their ability to induce quinolone resistance. Pollutants with specific structures can inhibit DNA gyrase, thereby inhibiting DNA replication and inducing DNA damage. Subsequently, through a series of SOS and DNA replication and repair system responses, target mutations ultimately lead to quinolone resistance. Therefore, rapid screening of environmental pollutants with DNA gyrase inhibitory activity is key to identifying high-risk pollutants with the ability to induce bacterial quinolone resistance.

[0003] A large number of chemicals are added to industrial products and everyday items. Therefore, computational methods are needed to accelerate the screening of high-risk contaminants with the potential to induce bacterial quinolone resistance. Traditional methods for predicting the inhibitory activity of chemicals against proteins fall into two main categories: ligand-based and receptor-based approaches. Ligand-based approaches assess the potential effect of a chemical by comparing its structure with that of known active ligands. Ligand-based approaches are simple and effective when ligands have high similarity, but their applicability and accuracy are limited by the space of known ligand structures. Receptor-based approaches primarily utilize molecular docking, calculating parameters such as the three-dimensional structure and interaction energy of the ligand-receptor complex to determine the interaction mode between the ligand and the target protein. While receptor-based approaches fully consider the three-dimensional structural characteristics of the target protein, they still face several unresolved technical challenges. For example, the conformational flexibility of the protein, the accuracy of the scoring function, and the solvent effect of water molecules can all affect the accuracy of docking results.

[0004] Traditional methods for predicting the protein inhibitory activity of chemicals suffer from low accuracy and limited applicability. Therefore, there is an urgent need to develop new methods for assessing the DNA gyrase inhibitory activity of chemicals to rapidly screen for high-risk contaminants with the potential to induce bacterial quinolone resistance. Indeed, with the rise of deep learning methods, feature-based methods for predicting chemical properties have gradually developed. This is because feature-based methods can apply deep learning algorithms to learn the characteristics of chemicals with known DNA gyrase inhibitory activity and infer the DNA gyrase inhibitory activity of unknown chemicals. Furthermore, relevant studies have shown that deep learning models outperform molecular docking and other machine learning models.

[0005] In summary, a deep learning model that can fully learn the characteristics of chemicals is needed to achieve direct mapping of chemical structure to properties and to realize rapid screening of chemicals with DNA gyrase inhibitory activity. Limited by the amount of data from existing GyrB enzyme activity inhibition experiments, ordinary deep learning methods are difficult to further improve the prediction accuracy of the model and cannot accurately predict the inhibitory activity of chemicals on GyrB. Therefore, in order to further improve the prediction performance of the model on small samples, a multi-task model that shares part of the network structure was established, which can realize the inhibitory activity data (IC) of chemicals on GyrA and GyrB. 50 ) simultaneously, it can improve the model's learning ability and prevent overfitting by sharing parameters between tasks. Especially when the tasks are highly correlated, the GyrB prediction task can fully utilize the feature expression of the GyrA prediction task, achieving better results than single-task learning. Summary of the Invention

[0006] The purpose of the present invention is to address the deficiencies of the prior art and provide a method for predicting the inhibitory activity of chemicals on DNA gyrase.

[0007] The object of the present invention is achieved through the following technical solution: a method for predicting the inhibitory activity of a chemical on DNA gyrase, wherein the DNA gyrase includes GyrA and GyrB, the method comprising the following steps:

[0008] S1. Obtain chemical structures and inhibitory activity data of the chemicals on GyrA and GyrB, perform a structure cleaning operation on the chemical structures, perform a scaling operation on the inhibitory activity data of GyrA and GyrB to construct a data set, and randomly divide the data set into a training set and a test set according to the proportion;

[0009] S2. Build a multi-task graph neural network model with chemical SMILES codes as input, whose output is the predicted value of the inhibitory activity data scaled value of the chemical on GyrA and GyrB;

[0010] S3. Use cross-validation to determine the optimal hyperparameters of the multi-task graph neural network model, and train it using the training set to obtain the trained optimal multi-task graph neural network model, which is used as the DNA gyrase inhibitory activity prediction model to predict the inhibitory activity data scaling values ​​of chemicals on GyrA and GyrB;

[0011] S4. Input the SMILES code of the chemical to be tested in the test set or the SMILES code of the chemical to be evaluated into the DNA gyrase inhibitory activity prediction model to obtain the predicted value of the inhibitory activity data scaling value of the chemical on GyrA and GyrB.

[0012] Furthermore, step S1 includes the following sub-steps:

[0013] S1.1. Obtain chemical structures and inhibitory activity data on GyrA and GyrB from public databases. The chemical structures are represented by SMILES codes, and the inhibitory activity data are used to indicate the inhibitory activity of the chemical on GyrA and GyrB.

[0014] S1.2. Perform structure cleaning operations on chemical structures, specifically including: standardizing, removing solvents, charge correction, and deionizing chemical molecular structure cleaning operations on chemical SMILES codes to standardize and unify the form of chemical SMILES codes; wherein, these structure cleaning operations can perform molecular structure inspection to ensure that the structure of the molecule is chemically reasonable; correct drawing errors and standardize functional groups; hide hydrogen atoms; recalculate stereochemistry; remove covalently bonded impurities; ensure that the strongest acid group is ionized first; generate neutral molecules; standardize tautomers; and remove salts and covalently bound metals contained in the compound;

[0015] S1.3. performing a scaling operation on the numerical range of the inhibitory activity data of the chemical on GyrA and GyrB by taking the negative logarithm of the original numerical value of the inhibitory activity data of the chemical on GyrA and GyrB;

[0016] S1.4, constructing GyrA and GyrB datasets based on the chemical SMILES codes after the structure cleaning operation in step S1.2 and the inhibitory activity data of GyrA and GyrB after the scaling operation in step S1.3;

[0017] S1.5. Randomly divide the GyrA and GyrB datasets into training and test sets in proportion.

[0018] Furthermore, the multi-task graph neural network model includes a task-specific input layer, a molecular graph embedding layer, a molecular fragment graph embedding layer, multiple molecular graph neural network encoders, multiple molecular fragment graph neural network encoders, a dual-path combiner and a task-specific predictor, wherein the task-specific input layer includes a parallel GyrA task input layer and a GyrB task input layer, the molecular graph neural network encoder and the molecular fragment graph neural network encoder are both composed of a neighborhood attention module, an information aggregation module, a gating mechanism update node feature module and a feature attention module, the feature attention module includes a maximum pooling layer, a sum pooling layer, a multi-layer perceptron global pooling layer and an activation function; the dual-path combiner selects a graph convolutional attention layer, the graph convolutional attention layer includes multiple attention heads and a ReLU activation function; the task-specific predictor includes two prediction heads, each of which includes n l Layer fully connected layer.

[0019] Furthermore, the task-specific input layer extracts GyrA and GyrB data respectively through the parallel GyrA task input layer and the GyrB task input layer and inputs the data into the shared network of the multi-task graph neural network model for subsequent reasoning, wherein the shared network of the multi-task graph neural network model includes a molecular graph embedding layer, a molecular fragment graph embedding layer, multiple molecular graph neural network encoders, multiple molecular fragment graph neural network encoders and a dual-pathway combiner;

[0020] The molecular graph embedding layer calculates the molecular graph of the chemical according to the chemical SMILES code using RDKit to extract molecular graph features from the molecular graph; wherein the molecular graph features include a×d at The atomic feature matrix, the adjacency matrix of length b and the b×d bt The chemical bond characteristic matrix is ​​, where a is the number of atoms in the chemical, b is twice the number of chemical bonds in the chemical, and d is at is the number of atomic features, d bt is the number of chemical bond features, and the adjacency matrix is ​​used to represent the atomic numbers of the bonding atom pairs in the molecular graph;

[0021] The molecular fragmentation graph embedding layer calculates the molecular graph of the chemical according to the chemical SMILES code using RDKit, and extracts the molecular fragmentation features from the molecular graph based on the BRICS principle; wherein the molecular fragmentation features include a×d at The atomic feature matrix of length b f The molecular fragment adjacency matrix, b f ×d bt The chemical bond feature matrix of the molecular fragments and the molecular fragment index list of length a, where b f It is twice the number of chemical bonds remaining after the molecule is broken into molecular fragments. The molecular fragment index list is used to indicate the index of the molecular fragment where each atom in the chemical is located;

[0022] The various molecular graph features and molecular fragment features extracted from the molecular graph embedding layer and the molecular fragment graph embedding layer are first transformed into features of length d h Then, each type of feature vector is input into the molecular graph neural network encoder and the molecular fragment graph neural network encoder respectively to obtain the molecular graph feature sequence and the molecular fragment feature sequence respectively, where d hThe length of is the same as the size of the hidden layer; wherein, in the molecular graph neural network encoder and the molecular fragment graph neural network encoder, the molecular graph neural network encoder and the molecular fragment graph neural network encoder update the node features of the molecular graph and the molecular fragment graph through their respective neighborhood attention modules and information aggregation modules, and then update the node feature module through a gating mechanism to balance between retaining the original information and using new information. In the feature attention module, the global features of the molecular graph are extracted through the maximum pooling layer and the sum pooling layer, and the attention weights are generated by the multi-layer perceptron to perform weighted adjustment on the node features. Finally, the features of all nodes are aggregated into a global feature vector through the global pooling layer and the activation function, thereby obtaining the molecular graph feature sequence and the molecular fragment feature sequence.

[0023] The molecular graph feature sequence and the molecular fragment feature sequence are input into the dual-pathway combiner. The dual-pathway combiner uses the graph convolution attention layer to learn the relationship between the molecular graph feature sequence and the molecular fragment feature sequence. Specifically, it adopts n h Each attention head independently calculates the attention weights of different molecular fragments, and performs nonlinear transformation through the ReLU activation function to obtain the molecule-molecular fragment interaction feature sequence;

[0024] Then, the molecule-molecule fragment interaction feature sequence and the molecular graph feature sequence output by the molecular graph neural network encoder are spliced ​​and input into the task-specific predictor, which passes through the n prediction heads of the two prediction heads. l After the fully connected layer, the predicted values ​​of the inhibitory activity data scaling values ​​of the chemicals on GyrA and GyrB are output respectively.

[0025] Furthermore, step S3 includes the following sub-steps:

[0026] S3.1. Preset multiple sets of different model hyperparameters as {lr, weight-decay, batchsize_a, batchsize_b, dropout, depth, feature-dimension, n-slices, r}, where lr is the learning rate, weight-decay is the weight decay parameter, batchsize_a and batchsize_b are the number of GyrA and GyrB samples selected for one training, dropout is the random dropout regularization parameter, depth is the number of encoder layers, feature-dimension is the size of the hidden layer, n-slices is the number of slices, r is the decay rate, lr is used to control the progress of convergence to the local minimum, weight-decay and dropout are used to prevent model overfitting, depth, feature-dimension, n-slices and r are used to adjust model complexity, and the size of batchsize_a and batchsize_b affects the degree and speed of model optimization;

[0027] S3.2. For each set of model hyperparameters, use the GyrA and GyrB data in the training set to train a multi-task graph neural network model. f The multi-task graph neural network model corresponding to the set of model hyperparameters was evaluated by fold cross-validation based on statistical parameters to determine the optimal model hyperparameters. The trained multi-task graph neural network model corresponding to the optimal model hyperparameters was obtained, which was used as a DNA gyrase inhibitory activity prediction model to predict the inhibitory activity data scaling values ​​of chemicals on GyrA and GyrB.

[0028] Furthermore, the statistical parameter is one of the coefficient of determination, root mean square error, mean square error, and mean absolute error.

[0029] Furthermore, for each set of model hyperparameters, the multi-task graph neural network model is trained using the data in the training set, and the n f The fold-cross validation evaluates the multi-task graph neural network model corresponding to the set of model hyperparameters based on statistical parameters to determine the optimal model hyperparameters, specifically including:

[0030] S3.2.1. Train the multi-task graph neural network model based on a set of preset model hyperparameters, using n f The data in the training set are randomly divided into n f subsets, each time one of the subsets is used as the validation set, and the remaining n f-1 subsets of data were used to train the multi-task graph neural network model. During the training process, the statistical parameters were calculated based on the predicted values ​​of the inhibitory activity data scaling values ​​of the chemicals on GyrA and GyrB output by the multi-task graph neural network model and the corresponding true values ​​in the training set. With the optimal statistical parameters as the optimization goal, the parameters of the multi-task graph neural network model were adjusted, and then the multi-task graph neural network model was verified using the data in the validation set to calculate the corresponding statistical parameters. The above process was repeated for training n f times, and finally get n f Statistical parameters and calculate n f The average value of the statistical parameters;

[0031] S3.2.2. Repeat step S3.2.1 to obtain the average statistical parameters corresponding to all groups of model hyperparameters, and select a group of model hyperparameters corresponding to the optimal statistical parameter average as the optimal model hyperparameters.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] (1) The present invention can fully utilize existing high-throughput experimental big data and use chemical SMILES codes as input for the inhibitory activity prediction model. It does not require manually defined quantifiable structural parameters as molecular descriptors, thus saving time and computing resources for molecular descriptor calculation and descriptor selection, and its application requires less computational chemistry foundation.

[0034] (2) Compared with existing methods, the method of the present invention has high-precision prediction performance, which is helpful for the rapid discovery of new environmental pollutants that induce bacterial resistance and has broad application prospects in the fields of chemical risk assessment and environmental safety assessment;

[0035] (3) The structure of the inhibitory activity prediction model in the present invention is flexible, which greatly improves the predictive ability of the inhibitory activity prediction model and realizes high-precision prediction of the inhibitory activity of chemicals on GyrA and GyrB, which is conducive to the rapid discovery of new environmental pollutants that induce bacterial resistance and can also be extended to the study of different key molecular initiation events. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a flow chart of a method for predicting the inhibitory activity of a chemical substance of the present invention on DNA gyrase;

[0037] Figure 2 This is a structural diagram of the multi-task graph neural network model network in Example 1 of the present invention. DETAILED DESCRIPTION

[0038] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0039] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0040] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0041] The present invention will be described in detail below with reference to the accompanying drawings. Unless there is any conflict, the features of the following embodiments and implementations may be combined with each other.

[0042] The method for predicting the inhibitory activity of chemicals on DNA gyrase is based on the chemical SMILES (Simplified Molecular Input Line Entry System) code, avoiding the limitations of traditional methods such as low accuracy and small application range. Compared with other similar methods, the method of the present invention has the most superior predictive ability in current research.

[0043] It should be understood that SMILES is a specification that uses ASCII strings to clearly describe molecular structures, and SMILES codes can be used to describe the three-dimensional chemical structure of a compound.

[0044] The basic principle of the present invention is that the input of the model is the SMILES code of the chemical in the GyrA and GyrB data sets; first, the GyrA and GyrB data are obtained by the task-specific input layer, and then the data and model network parameters are shared through the shared network of the model, which is conducive to knowledge transfer and significantly improves the learning efficiency of new tasks. Among them, in the shared network of the model, the atomic feature matrix, adjacency matrix and chemical bond feature matrix are obtained through the molecular graph embedding layer, and the molecular fragment features are obtained through the molecular fragment graph embedding layer, specifically including the atomic feature matrix, the molecular fragment adjacency matrix, the chemical bond feature matrix of the molecular fragments and the molecular fragment index list; then the various feature vectors extracted from the molecular graph embedding layer and the molecular fragment graph embedding layer are respectively input into the molecular graph neural network encoder and the molecular fragment graph neural network encoder, and the molecular graph feature sequence and the molecular fragment feature sequence are obtained accordingly, and the intrinsic connection between the molecules and the molecular fragments is captured by the dual-path combiner. Finally, the task-specific predictor outputs the inhibitory activity data (IC) of the chemical on GyrA and GyrB. 50 ) scaled values.

[0045] See also Figure 1 The method for predicting the inhibitory activity of a chemical on DNA gyrase of the present invention, wherein the DNA gyrase includes GyrA and GyrB, specifically comprises the following steps:

[0046] S1. Obtain chemical structures and inhibitory activity data (IC 50 ), and performed structural cleaning operations on the chemical structures, and performed scaling operations on the inhibitory activity data of GyrA and GyrB to construct a dataset, which was randomly divided into training and test sets in proportion.

[0047] S1.1. Obtain chemical structures and inhibitory activity data of chemicals against GyrA and GyrB from public databases; the chemical structures are represented by chemical SMILES codes, and the inhibitory activity data are used to indicate the inhibitory activity of the chemicals against GyrA and GyrB.

[0048] S1.2. Perform structure cleaning operations on chemical structures, including standardization, solvent removal, charge correction, deionization, and other chemical molecular structure cleaning operations to standardize and unify the form of chemical SMILES codes, effectively avoiding interference in subsequent calculations. These structure cleaning operations can perform molecular structure inspection to ensure that the structure of the molecule is chemically reasonable; correct common drawing errors and standardize functional groups; hide hydrogen atoms; recalculate stereochemistry; remove covalently bonded impurities; ensure that the strongest acid group is ionized first; generate neutral molecules whenever possible; standardize tautomers; and remove salts and covalently bound metals contained in the compound.

[0049] S1.3. By taking the negative logarithm of the original numerical values ​​of the inhibitory activity data of the chemical on GyrA and GyrB, the numerical range of the inhibitory activity data of the chemical on GyrA and GyrB is scaled, thereby reducing the numerical range of the inhibitory activity data.

[0050] S1.4. Construct GyrA and GyrB datasets based on the chemical SMILES codes after the structure cleaning operation in step S1.2 and the inhibitory activity data of GyrA and GyrB after the scaling operation in step S1.3.

[0051] S1.5. Randomly split the GyrA and GyrB datasets into training and test sets according to a certain ratio. T% of the data is used as the training set for training and optimizing the multi-task graph neural network model, and w% of the data is used as the test set for evaluating the performance of the final model.

[0052] S2. Construct a multi-task graph neural network model with chemical SMILES codes as input, whose output is the predicted value of the inhibitory activity data scaling value of the chemical on GyrA and GyrB.

[0053] In this embodiment, the multi-task graph neural network model includes a task-specific input layer, a molecular graph embedding layer, a molecular fragment graph embedding layer, multiple molecular graph neural network encoders, multiple molecular fragment graph neural network encoders, a dual-pathway combiner and a task-specific predictor, wherein the task-specific input layer includes a parallel GyrA task input layer and a GyrB task input layer, and the molecular graph neural network encoder and the molecular fragment graph neural network encoder are both composed of a neighborhood attention module, an information aggregation module, a gating mechanism to update node feature module and a feature attention module, but their parameters are different, such as Figure 2 As shown in the figure, the feature attention module includes a maximum pooling layer, a sum pooling layer, a multi-layer perceptron global pooling layer and an activation function; the dual-path combiner uses a graph convolution attention layer, which includes multiple attention heads and a ReLU activation function; the task-specific predictor includes two prediction heads, each of which includes n l Layer fully connected layer.

[0054] Specifically, the task-specific input layer extracts the data of GyrA and GyrB respectively through the parallel GyrA task input layer and the GyrB task input layer and inputs them into the shared network of the multi-task graph neural network model for subsequent reasoning, wherein the shared network of the multi-task graph neural network model includes a molecular graph embedding layer, a molecular fragment graph embedding layer, multiple molecular graph neural network encoders, multiple molecular fragment graph neural network encoders and a dual-pathway combiner.

[0055] It should be understood that since the prediction of the inhibitory activity of chemicals on GyrA and GyrB has different emphases, it is necessary to first extract the data of GyrA and GyrB through the task-specific input layer, and then input it into the shared network of the multi-task graph neural network model to perform the subsequent reasoning process, and then use the task-specific predictor to output the corresponding inhibitory activity prediction results using the prediction heads corresponding to GyrA and GyrB.

[0056] The molecular graph embedding layer calculates the molecular graph of the chemical based on the chemical SMILES code using RDKit to extract molecular graph features from the molecular graph; the molecular graph features include a×d at The atomic feature matrix, the adjacency matrix of length b and the b×d bt The chemical bond characteristic matrix is ​​, where a is the number of atoms in the chemical, b is twice the number of chemical bonds in the chemical, and d is at is the number of atomic features, d bt is the number of chemical bond features, and the adjacency matrix is ​​used to represent the atomic numbers of the bonding atom pairs in the molecular graph.

[0057] The molecular fragmentation graph embedding layer calculates the molecular graph of the chemical based on the chemical SMILES code using RDKit, and extracts the molecular fragment features from the molecular graph based on the BRICS principle (Bonds Representing Interaction Clusters); wherein the molecular fragment features include a×d at The atomic feature matrix of length b f The molecular fragment adjacency matrix, b f ×d bt The chemical bond feature matrix of the molecular fragments and the molecular fragment index list of length a, where b f It is twice the number of chemical bonds remaining after the molecule is broken into molecular fragments. The molecular fragment index list is used to indicate the index of the molecular fragment where each atom in the chemical is located.

[0058] The various molecular graph features and molecular fragment features extracted from the molecular graph embedding layer and the molecular fragment graph embedding layer are first transformed into features of length d h Then, each type of feature vector is input into the molecular graph neural network encoder and the molecular fragment graph neural network encoder respectively to obtain the molecular graph feature sequence and the molecular fragment feature sequence respectively, where d hThe length of is the same as the size of the hidden layer. In the molecular graph neural network encoder and molecular fragment graph neural network encoder, the molecular graph neural network encoder and molecular fragment graph neural network encoder update the node features of the molecular graph and molecular fragment graph through their respective neighborhood attention modules and information aggregation modules. Then, the node feature module is updated through a gating mechanism to balance between retaining the original information and using new information. In the feature attention module, the global features of the molecular graph are extracted through the maximum pooling layer and the sum pooling layer. The multi-layer perceptron is used to generate attention weights and weighted adjustment of the node features. Finally, the features of all nodes are aggregated into a global feature vector through the global pooling layer and activation function, thereby obtaining the molecular graph feature sequence and the molecular fragment feature sequence.

[0059] It should be understood that feature transformation refers to the compression or amplification of feature dimensions, as well as the calculation of activation functions, to convert various molecular graph features and molecular fragment features into corresponding feature vectors. Furthermore, in deep learning models, all network layers other than the input and output layers are referred to as hidden layers.

[0060] The molecular graph feature sequence and the molecular fragment feature sequence are input into the dual-pathway combiner. The dual-pathway combiner first adopts the interactive attention strategy to capture the intrinsic connection between molecules and molecular fragments, and uses the graph convolutional attention layer to learn the relationship between the molecular graph feature sequence and the molecular fragment feature sequence. Specifically, it adopts n h Each attention head independently calculates the attention weights of different molecular fragments, which can better capture the characteristics of molecular fragments and obtain the molecular-molecular fragment interaction feature sequence through nonlinear transformation through ReLU activation function. Then, the molecule-molecular fragment interaction feature sequence and the molecular graph feature sequence output by the molecular graph neural network encoder are spliced ​​and input into the task-specific predictor, which is respectively passed through n prediction heads. l After the fully connected layer, the inhibitory activity data of chemicals on GyrA and GyrB (IC 50 ) scaled values.

[0061] S3. Use cross-validation to determine the optimal hyperparameters of the multi-task graph neural network model, and train it using the training set to obtain the trained optimal multi-task graph neural network model, which is used as a DNA gyrase inhibitory activity prediction model to predict the inhibitory activity data scaling values ​​of chemicals on GyrA and GyrB.

[0062] S3.1. Based on experience, we preset multiple sets of different model hyperparameters as {lr, weight-decay, batchsize_a, batchsize_b, dropout, depth, feature-dimension, n-slices, r}, where lr is the learning rate, weight-decay is the weight decay parameter, batchsize_a and batchsize_b are the number of GyrA and GyrB samples selected for one training session, dropout is the random dropout regularization parameter, depth is the number of encoder layers, feature-dimension is the size of the hidden layer, n-slices is the number of slices, r is the decay rate, lr is used to control the progress of convergence to the local minimum, weight-decay and dropout are used to prevent model overfitting, depth, feature-dimension, n-slices and r are used to adjust the model complexity, and the size of batchsize_a and batchsize_b affects the degree and speed of model optimization.

[0063] S3.2. For each set of model hyperparameters, use the GyrA and GyrB data in the training set to train a multi-task graph neural network model. f The multi-task graph neural network model corresponding to the set of model hyperparameters was evaluated by fold cross-validation based on statistical parameters to determine the optimal model hyperparameters. The trained multi-task graph neural network model corresponding to the optimal model hyperparameters was obtained, which was used as a DNA gyrase inhibitory activity prediction model to predict the inhibitory activity data scaling values ​​of chemicals on GyrA and GyrB.

[0064] Furthermore, the statistical parameter is the coefficient of determination (R 2 ), root mean square error (RMSE), mean square error (MSE), mean absolute error (MAE), and their calculation formulas are:

[0065]

[0066] Among them, n represents the number of samples, y i represents the true value of the scaled value of the inhibitory activity data of the chemical corresponding to the i-th sample on GyrA and GyrB, represents the predicted value of the inhibitory activity scaling value of the chemical on GyrA and GyrB corresponding to the i-th sample, Represents the average of all true values.

[0067] It should be noted that when using statistical parameters to evaluate the multi-task graph neural network model, the statistical parameters of the GyrA task and the statistical parameters of the GyrB task can be added together as the final method for evaluating the predictive ability of the multi-task graph neural network model.

[0068] Furthermore, for each set of model hyperparameters, the multi-task graph neural network model is trained using the data in the training set. f The fold-cross validation evaluates the multi-task graph neural network model corresponding to the set of model hyperparameters based on statistical parameters to determine the optimal model hyperparameters, specifically including:

[0069] S3.2.1. Train the multi-task graph neural network model based on a set of preset model hyperparameters, using n f The data in the training set are randomly divided into n f subsets, each time one of the subsets is used as the validation set, and the remaining n f -1 subsets of data were used to train the multi-task graph neural network model. During the training process, the statistical parameters were calculated based on the predicted values ​​of the inhibitory activity data scaling values ​​of the chemicals on GyrA and GyrB output by the multi-task graph neural network model and the corresponding true values ​​in the training set. With the optimal statistical parameters as the optimization goal, the parameters of the multi-task graph neural network model were adjusted, and then the multi-task graph neural network model was verified using the data in the validation set to calculate the corresponding statistical parameters. The above process was repeated for training n f times, and finally get n f Statistical parameters, take n f The average value of the statistical parameters is used to evaluate the predictive ability of the multi-task graph neural network model when different model hyperparameters are selected.

[0070] S3.2.2. Repeat step S3.2.1 to obtain the average statistical parameters corresponding to all groups of model hyperparameters, and select a group of model hyperparameters corresponding to the optimal statistical parameter average as the optimal model hyperparameters.

[0071] It should be understood that when different statistical parameters are selected, the requirements for the optimal statistical parameters are also different. For example, when the statistical parameter is RMSE, the smaller the RMSE, the better the statistical parameter; when the statistical parameter is R 2 When R 2 The larger the value, the better the statistical parameter.

[0072] S4. Input the SMILES code of the chemical to be tested in the test set or the SMILES code of the chemical to be evaluated into the DNA gyrase inhibitory activity prediction model to obtain the predicted value s of the inhibitory activity data scaling value of the chemical on GyrA and GyrB. out_a and s out_bThe larger the predicted value is, the stronger the inhibitory activity of the chemical on GyrA and GyrB is considered to be, and vice versa.

[0073] In summary, the present invention utilizes chemical SMILES codes to avoid the application limitations of traditional prediction models, greatly improves the predictive ability of the model, and realizes high-precision prediction of the inhibitory activity of chemicals on DNA gyrase; in addition, the present invention uses a shared network of a multi-task graph neural network model to enable the multi-task graph neural network model to learn the characteristics of chemicals more fully, so as to realize accurate and rapid screening of key molecular initiation events of pollutant-induced bacterial resistance, clarify the structural basis of pollutant-induced bacterial resistance, and has broad application prospects in the field of chemical health risk assessment.

[0074] The technical solution of the present invention is further described below by means of specific embodiments in combination with the accompanying drawings. It should be noted that the following specific embodiments are only for illustration and the protection scope of the present invention is not limited thereto.

[0075] Example 1

[0076] See also Figure 1-Figure 2 This embodiment implements a method for predicting the inhibitory activity of chemicals on DNA gyrase (GyrA and GyrB) based on a multi-task graph neural network model, specifically comprising the following steps:

[0077] (1) Acquisition and preprocessing of chemical structures and inhibitory activity data on GyrA and GyrB.

[0078] First, the UniProtIDs of bacterial GyrA and GyrB were obtained from the UniProt database. Based on these UniProtIDs, the inhibitory activity data of GyrA and GyrB were found in the UniProt database and ChEMBL database, which contained chemical SMILES codes and inhibitory activity data (IC 50 ) numerical value. The inhibitory activity data indicates the strength of the compound's inhibitory activity on the enzyme.

[0079] Subsequently, the chemical SMILES codes in the DNA gyrase inhibitory activity data of the chemicals were standardized; desolvation; charge correction; deionization and other chemical molecular structure cleaning operations were performed to avoid interference in subsequent calculations. These SMILES code cleaning operations can perform molecular structure inspection to ensure that the structure of the molecule is chemically reasonable; correct common drawing errors and standardize functional groups; hide hydrogen atoms; recalculate stereochemistry; remove covalently bonded impurities; ensure that the strongest acid group is ionized first; generate neutral molecules as much as possible; standardize tautomers; and remove salts and covalently bound metals contained in the compound. By using the half-maximal inhibitory concentration (IC50) of each chemical 50) to scale the inhibitory activity range.

[0080] Finally, the preprocessed chemical SMILES codes and the inhibitory activity data of the chemicals on DNA gyrase A subunit were randomly divided into training set and test set in an 8:2 ratio.

[0081] (2) Construction, training and hyperparameter search of multi-task graph neural network models.

[0082] The constructed multi-task graph neural network model includes a task-specific input layer, a molecular graph embedding layer, a molecular fragment graph embedding layer, three molecular graph neural network encoders, three molecular fragment graph neural network encoders, a dual-pathway combiner and a task-specific predictor, such as Figure 2 shown.

[0083] The chemical SMILES codes in the GyrA and GyrB training sets are used as input to the multi-task graph neural network model.

[0084] The task-specific input layer can extract the data of GyrA and GyrB respectively and input them into the shared network of the model (molecular graph embedding layer, molecular fragment embedding layer, molecular graph neural network encoder, molecular fragment graph neural network encoder and dual-pathway combiner) for training.

[0085] RDKit is used in the molecular graph embedding layer to extract an a×46 atomic feature matrix, a b-length adjacency matrix (used to represent the atomic numbers of the bonding atom pairs in the molecule), and a b×10 chemical bond feature matrix from the molecular graph, where a is the number of atoms in the chemical, b is twice the number of chemical bonds in the chemical, 46 is the number of atomic features, and 10 is the number of chemical bond features.

[0086] RDKit is used in the molecular fragment embedding layer to extract molecular fragment features from the molecular graph. The extracted molecular fragment features include: a×46 atomic feature matrix with a length of b f The molecular fragment adjacency matrix, b f ×10 chemical bond feature matrix of molecular fragments and a molecular fragment index list of length a (indices of the molecular fragments where each atom in the chemical is located), where b f It is twice the number of chemical bonds remaining after the molecule is broken into molecular fragments.

[0087] The feature vectors of each class extracted from the molecular graph embedding layer and the molecular fragment embedding layer are first transformed into a length of d h The feature vectors are then input into the molecular graph neural network encoder and the molecular fragment graph neural network encoder respectively to obtain the molecular graph feature sequence and the molecular fragment feature sequence respectively, where d hThe length of is the same as the size of the hidden layer. The molecular graph neural network encoder and molecular fragment graph neural network encoder are composed of a neighborhood attention module, an information aggregation module, a gating mechanism node feature update module, and a feature attention module, respectively. In the molecular graph neural network encoder and molecular fragment graph neural network encoder, the node features of the molecular graph and molecular fragment graph are updated through the neighborhood attention module and the information aggregation module. The node feature update module then uses the gating mechanism to balance the retention of original information and the use of new information. In the feature attention module, the global features of the graph are extracted through the maximum pooling layer and the sum pooling layer. The multi-layer perceptron is used to generate attention weights and weightedly adjust the node features. Finally, the features of all nodes are aggregated into a global feature vector through the global pooling layer and activation function, thereby obtaining the molecular graph feature sequence and the molecular fragment feature sequence.

[0088] The dual-pathway combiner first employs an interactive attention strategy to capture the intrinsic connections between molecules and molecular fragments. Using a graph attention convolutional layer, the network learns the relationship between the molecular graph feature sequence and the molecular fragment feature sequence, calculates the attention weights for different molecular fragments, and then performs a nonlinear transformation using a ReLU activation function to obtain a molecule-fragment interaction feature sequence. The graph attention convolutional layer uses four independent attention heads for computation, enabling better capture of molecular fragment features. The molecular graph feature sequence and the molecule-fragment interaction feature sequence are then concatenated and fed into the task-specific predictor.

[0089] The task-specific predictor consists of two prediction heads, each of which includes a fully connected layer and outputs the IC of the chemical pair GyrA and GyrB respectively. 50 The predicted value of the scaled value s out_a and s out_b .

[0090] During training, the adaptive momentum optimizer Adam is used to update the neural network parameters based on the gradient, and the learning rate lr is 0.000525. In addition, in order to control the progress of convergence to the local minimum, batchsize_a and batchsize_b are added to the model, and the parameters set the batch size to 32 and 32 respectively. In order to improve the generalization ability of the model and prevent the model from overfitting, the weight decay parameter weight-decay and the random activation regularization parameter dropout are set. The weight decay parameter weight-decay is set to 0.00001, and the random activation regularization parameter dropout is set to 0.0. In addition, the number of encoder layers depth, the size of the hidden layer feature-dimension, the number of slices n-slices, and the decay rate r are used to adjust the model complexity. The depth, feature-dimension, n-slices, and r are 4, 64, 2, and 1 respectively. 10-fold cross-validation is used, and the training is repeated 10 times. Each time, the validation set and training set are randomly sampled according to a certain ratio for training. Each training is performed for a maximum of 500 iterations. When the validation set R of the current iteration is 2 Less than the optimal validation set R 2 And the current number of iterations is greater than the optimal validation set R 2 The training is terminated when the number of iterations is greater than 50. The prediction ability of the model with different hyperparameters is evaluated based on the average value of the statistical parameters of the model on the validation set during the 10 training cycles. Finally, the model is retrained on the entire training set based on the optimal hyperparameter combination, and the test set R is calculated. 2 , RMSE, MSE, and MAE statistical parameters are used to characterize the model's predictive ability.

[0091] In order to further avoid overfitting of the model and improve the generalization ability of the model, the hyperparameters used are searched and optimized within a certain range and at a certain step size:

[0092] Select the learning rate parameter lr and get R corresponding to different learning rates 2 , lr is selected as 0.000525;

[0093] Select the weight decay parameter weight-decay and get the R corresponding to different weight decay parameters 2 , select λ as 0.00001;

[0094] Select batch size batchsize_a and batchsize_b to get the R values ​​corresponding to different batch sizes. 2 , batchsize_a and batchsize_b are selected as 32 and 32 respectively;

[0095] Select the random loss regularization parameter dropout and get the R corresponding to different random loss regularization parameters 2 , select dropout as 0.0;

[0096] Select the number of encoder layers depth, and get the R corresponding to the number of encoder layers 2 , select depth as 4;

[0097] Select the size feature-dimension of the hidden layer and get the R corresponding to the size of different hidden layers 2 , select feature-dimension as 64;

[0098] Select the number of slices n-slices and get the R corresponding to different numbers of slices 2 , select n-slices as 2;

[0099] Select the attenuation rate r and get the R corresponding to different attenuation rates 2 , select r as 1.

[0100] (3) Prediction of DNA gyrase inhibitory activity of the chemical to be evaluated.

[0101] BDBM50423645 was selected from the test set as the chemical to be evaluated to predict its inhibitory activity against DNA gyrase. The chemical structure of BDBM50423645 is shown in the following formula:

[0102]

[0103] BDBM50423645 is a fluorinated organic compound that is widely present in the environment. As the chemical to be tested in this embodiment, the SMILES code of BDBM50423645 was queried through the PubChem molecular database, specifically "CO[C@@H]1[C@@H](OC(N)=O)[C@H](O)[C@@H](Oc2ccc3c(=O)c(NC(=O)c4ccc(O)c(CC=C(C)C)c4)c(O)oc3c2C)OC1(C)C", and standardization, desolvation, charge correction, deionization and other chemical molecular structure cleaning operations were performed to avoid interference in subsequent calculations. These SMILES code cleaning operations can perform molecular structure inspection and processing to ensure that the structure of the molecule is chemically reasonable; correct common drawing errors and standardize functional groups; hide hydrogen atoms; recalculate stereochemistry; remove covalently bonded impurities; ensure that the strongest acid group is ionized first; produce neutral molecules as much as possible; standardize tautomers; and remove salts and covalently bound metals contained in the compound. The processed SMILES code was used as the input information of the model to calculate the p(IC 50 ) are predicted to be 6.23 and 6.85 respectively.

[0104] The p(IC 50 ) were 7.00 and 6.70 respectively, and the predicted results were basically consistent with the facts.

[0105] In summary, the present invention establishes a multi-task graph neural network model, which can predict the inhibitory activity of chemicals on DNA gyrase based only on the SMILES codes of the chemicals.

[0106] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for predicting the inhibitory activity of a chemical on DNA gyrase, characterized in that: The DNA gyrase comprises GyrA and GyrB, and the method comprises the following steps: S1. Obtain chemical structures and inhibitory activity data of the chemicals on GyrA and GyrB, perform a structure cleaning operation on the chemical structures, perform a scaling operation on the inhibitory activity data of GyrA and GyrB to construct a data set, and randomly divide the data set into a training set and a test set according to the proportion; S2. Construct a multi-task graph neural network model with chemical SMILES codes as input, whose output is the predicted value of the inhibitory activity data scaling value of the chemical on GyrA and GyrB; the multi-task graph neural network model includes a task-specific input layer, a molecular graph embedding layer, a molecular fragment graph embedding layer, multiple molecular graph neural network encoders, multiple molecular fragment graph neural network encoders, a dual-path combiner and a task-specific predictor, wherein the task-specific input layer includes a parallel GyrA task input layer and a GyrB task input layer, the molecular graph neural network encoder and the molecular fragment graph neural network encoder are both composed of a neighborhood attention module, an information aggregation module, a gating mechanism to update node feature module and a feature attention module, the feature attention module includes a maximum pooling layer, a sum pooling layer, a multi-layer perceptron global pooling layer and an activation function; the dual-path combiner selects a graph convolutional attention layer, the graph convolutional attention layer includes multiple attention heads and a ReLU activation function; the task-specific predictor includes 2 prediction heads, each of which includes n l Layer fully connected layer; The task-specific input layer extracts GyrA and GyrB data respectively through the parallel GyrA task input layer and the GyrB task input layer and inputs the data into the shared network of the multi-task graph neural network model for subsequent reasoning, wherein the shared network of the multi-task graph neural network model includes a molecular graph embedding layer, a molecular fragment graph embedding layer, multiple molecular graph neural network encoders, multiple molecular fragment graph neural network encoders and a dual-pathway combiner; The molecular graph embedding layer calculates the molecular graph of the chemical according to the chemical SMILES code using RDKit to extract molecular graph features from the molecular graph; wherein the molecular graph features include a×d at The atomic feature matrix, the adjacency matrix of length b and the b×d bt The chemical bond characteristic matrix is ​​, where a is the number of atoms in the chemical, b is twice the number of chemical bonds in the chemical, and d is at is the number of atomic features, d bt is the number of chemical bond features, and the adjacency matrix is ​​used to represent the atomic numbers of the bonding atom pairs in the molecular graph; The molecular fragmentation graph embedding layer calculates the molecular graph of the chemical according to the chemical SMILES code using RDKit, and extracts the molecular fragmentation features from the molecular graph based on the BRICS principle; wherein the molecular fragmentation features include a×d at The atomic feature matrix of length b f The molecular fragment adjacency matrix, b f ×d bt The chemical bond feature matrix of the molecular fragments and the molecular fragment index list of length a, where b f It is twice the number of chemical bonds remaining after the molecule is broken into molecular fragments. The molecular fragment index list is used to indicate the index of the molecular fragment where each atom in the chemical is located; The various molecular graph features and molecular fragment features extracted from the molecular graph embedding layer and the molecular fragment graph embedding layer are first transformed into features of length d h Then, each type of feature vector is input into the molecular graph neural network encoder and the molecular fragment graph neural network encoder respectively to obtain the molecular graph feature sequence and the molecular fragment feature sequence respectively, where d h The length of is the same as the size of the hidden layer; wherein, in the molecular graph neural network encoder and the molecular fragment graph neural network encoder, the molecular graph neural network encoder and the molecular fragment graph neural network encoder update the node features of the molecular graph and the molecular fragment graph through their respective neighborhood attention modules and information aggregation modules, and then update the node feature module through a gating mechanism to balance between retaining the original information and using new information. In the feature attention module, the global features of the molecular graph are extracted through the maximum pooling layer and the sum pooling layer, and the attention weights are generated by the multi-layer perceptron to perform weighted adjustment on the node features. Finally, the features of all nodes are aggregated into a global feature vector through the global pooling layer and the activation function, thereby obtaining the molecular graph feature sequence and the molecular fragment feature sequence. The molecular graph feature sequence and the molecular fragment feature sequence are input into the dual-pathway combiner. The dual-pathway combiner uses the graph convolution attention layer to learn the relationship between the molecular graph feature sequence and the molecular fragment feature sequence. Specifically, it adopts n h Each attention head independently calculates the attention weights of different molecular fragments, and performs nonlinear transformation through the ReLU activation function to obtain the molecule-molecular fragment interaction feature sequence; Then, the molecule-molecule fragment interaction feature sequence and the molecular graph feature sequence output by the molecular graph neural network encoder are spliced ​​and input into the task-specific predictor, which passes through the n prediction heads of the two prediction heads. l After the fully connected layer, the predicted values ​​of the inhibitory activity data scaling values ​​of the chemicals on GyrA and GyrB are output respectively; S3. Use cross-validation to determine the optimal hyperparameters of the multi-task graph neural network model, and train it using the training set to obtain the trained optimal multi-task graph neural network model, which is used as the DNA gyrase inhibitory activity prediction model to predict the inhibitory activity data scaling values ​​of chemicals on GyrA and GyrB; S4. Input the SMILES code of the chemical to be tested in the test set or the SMILES code of the chemical to be evaluated into the DNA gyrase inhibitory activity prediction model to obtain the predicted value of the inhibitory activity data scaling value of the chemical on GyrA and GyrB.

2. The method for predicting the inhibitory activity of a chemical on DNA gyrase according to claim 1, wherein: The step S1 includes the following sub-steps: S1.

1. Obtain chemical structures and inhibitory activity data on GyrA and GyrB from public databases. The chemical structures are represented by SMILES codes, and the inhibitory activity data are used to indicate the inhibitory activity of the chemical on GyrA and GyrB. S1.

2. Perform structure cleaning operations on chemical structures, specifically including: standardizing, removing solvents, charge correction, and deionizing chemical molecular structure cleaning operations on chemical SMILES codes to standardize and unify the form of chemical SMILES codes; wherein, these structure cleaning operations can perform molecular structure inspection to ensure that the structure of the molecule is chemically reasonable; correct drawing errors and standardize functional groups; hide hydrogen atoms; recalculate stereochemistry; remove covalently bonded impurities; ensure that the strongest acid group is ionized first; generate neutral molecules; standardize tautomers; and remove salts and covalently bound metals contained in the compound; S1.

3. performing a scaling operation on the numerical range of the inhibitory activity data of the chemical on GyrA and GyrB by taking the negative logarithm of the original numerical value of the inhibitory activity data of the chemical on GyrA and GyrB; S1.4, constructing GyrA and GyrB datasets based on the chemical SMILES codes after the structure cleaning operation in step S1.2 and the inhibitory activity data of GyrA and GyrB after the scaling operation in step S1.3; S1.

5. Randomly divide the GyrA and GyrB datasets into training and test sets in proportion.

3. The method for predicting the inhibitory activity of a chemical on DNA gyrase according to claim 1, wherein: The step S3 includes the following sub-steps: S3.

1. Preset multiple sets of different model hyperparameters as {lr, weight-decay, batchsize_a, batchsize_b, dropout, depth, feature-dimension, n-slices, r}, where lr is the learning rate, weight-decay is the weight decay parameter, batchsize_a and batchsize_b are the number of GyrA and GyrB samples selected for one training, dropout is the random dropout regularization parameter, depth is the number of encoder layers, feature-dimension is the size of the hidden layer, n-slices is the number of slices, r is the decay rate, lr is used to control the progress of convergence to the local minimum, weight-decay and dropout are used to prevent model overfitting, depth, feature-dimension, n-slices and r are used to adjust model complexity, and the size of batchsize_a and batchsize_b affects the degree and speed of model optimization; S3.

2. For each set of model hyperparameters, a multi-task graph neural network model is trained using the GyrA and GyrB data in the training set. The multi-task graph neural network model corresponding to this set of model hyperparameters is evaluated based on statistical parameters through N-fold cross-validation to determine the optimal model hyperparameters. The multi-task graph neural network model corresponding to the trained optimal model hyperparameters is obtained and used as a DNA gyrase inhibitory activity prediction model to predict the inhibitory activity data scaling values ​​of chemicals on GyrA and GyrB.

4. The method for predicting the inhibitory activity of a chemical on DNA gyrase according to claim 3, wherein: The statistical parameter is one of the coefficient of determination, root mean square error, mean square error, and mean absolute error.

5. The method for predicting the inhibitory activity of a chemical on DNA gyrase according to claim 3, wherein: For each set of model hyperparameters, the multi-task graph neural network model is trained using the data in the training set. The multi-task graph neural network model corresponding to the set of model hyperparameters is evaluated based on statistical parameters through N-fold cross-validation to determine the optimal model hyperparameters, specifically including: S3.2.

1. Train the multi-task graph neural network model based on a preset set of model hyperparameters using N-fold cross-validation. Randomly divide the data in the training set into N subsets, use one of the subsets as the validation set each time, and use the data in the remaining N-1 subsets to train the multi-task graph neural network model. During the training process, calculate statistical parameters based on the predicted values ​​of the scaled values ​​of the inhibitory activity data of the chemicals on GyrA and GyrB output by the multi-task graph neural network model and the corresponding true values ​​in the training set. Adjust the parameters of the multi-task graph neural network model with the optimal statistical parameters as the optimization goal, and then use the data in the validation set to validate the multi-task graph neural network model and calculate the corresponding statistical parameters. Repeat the above training process N times to finally obtain N statistical parameters, and calculate the average value of the N statistical parameters. S3.2.

2. Repeat step S3.2.1 to obtain the average statistical parameters corresponding to all groups of model hyperparameters, and select a group of model hyperparameters corresponding to the optimal statistical parameter average as the optimal model hyperparameters.

Citation Information

Patent Citations

  • Chemical estrogen receptor activation activity prediction model and a screening method

    CN112634993A

  • Method and apparatus for molecular toxicity prediction based on multi-task graph neural network

    CN113257369A