A virtual screening method for quorum sensing lead compounds and its application

Through a virtual screening method based on reinforcement learning, the features of PhcA and PhcR protein structures are extracted using GNN network and cross network, and the problem of difficulty in blocking rhizosphere invasion and predicting protein binding affinity in the prior art is solved, and efficient screening of population sensing lead compounds and control of cholesterol is achieved.

CN116189759BActive Publication Date: 2025-05-06NANJING AGRICULTURAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310234744.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-13
Publication Date
2025-05-06
Estimated Expiration
2043-03-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively block the rhizosphere invasion process of soil-borne fuzzy, and in the absence of protein structure information, it is difficult to predict the binding affinity of proteins and drug molecules.

Method used

A virtual screening method based on reinforcement learning is adopted to improve the performance of the virtual screening model of the group sensing pilot compound of the graph network by automatically searching the best graph structure. This method is based on the PhcA and PhcR protein structures, uses the GNN network and cross network to extract the characteristics of molecules and proteins, and combines with the LSTM controller to optimize the model structure to achieve efficient screening of population sensing pilot compounds.

Benefits of technology

This method can effectively extract the multi-dimensional characteristics of molecules and proteins, improve the accuracy and efficiency of virtual screening, provide the possibility of discovering new compounds with population induction activity, and provides new ideas and means for the control and prevention of bacteria such as Cyperus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189759B_ABST
    Figure CN116189759B_ABST
Patent Text Reader

Abstract

The present invention discloses a virtual screening method for quorum sensing lead compounds, and the main process includes: the input molecular compound structure is pre-processed to construct a molecular adjacency matrix, and is sent to the GNN1 network to generate compound features; the input protein sequence is extracted from its protein amino acid composition and dipeptide frequency to form a preliminary protein feature vector, which is sent to the cross network to generate cross-fusion features; at the same time, the protein sequence is generated into a corresponding contact map, which is then sent to the GNN2 network to generate protein sequence features; finally, the three feature combinations are sent to the fully connected layer to predict the affinity value. The present invention can be used to discover new compounds with quorum sensing activity, and provide new ideas and means for the control and prevention of bacteria such as Ralstonia solanacearum; at the same time, the method can efficiently screen out compounds that bind to PhcA and PhcR proteins, thereby discovering compounds with quorum sensing activity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medicinal chemistry, and in particular to a targeted virtual screening method and application of a Ralstonia solanacearum quorum sensing lead compound. Background Art

[0002] Ralstonia solanacearum is one of the most destructive soil-borne pathogens in the world. It is widely distributed in tropical, subtropical and temperate climate regions around the world, and is gradually spreading to high-latitude and high-altitude areas. The virulence behavior of soil-borne Ralstonia solanacearum during its invasion of the crop rhizosphere is regulated by quorum sensing. Ralstonia solanacearum has two quorum sensing systems: the AHL system and the trihydroxypalmitic acid methyl ester (3-OH PAME) system, of which the AHL system does not affect virulence. Ralstonia solanacearum uses the 3-OH PAME quorum sensing system to globally regulate metabolism and virulence behavior, coordinate the assembly of various secretion systems, and direct the temporal expression and secretion of various toxic factors, thereby successfully completing the rhizosphere invasion process. The system is composed of the PhcBSR synthesis component and the regulatory factor PhcA. Among them, PhcB is responsible for synthesizing the signal molecule 3-OH PAME, and PhcS is responsible for receiving the sensory signal molecule. When the concentration of 3-OH PAME exceeds a certain threshold, PhcS activates PhcR, thereby relieving PhcR's inhibition of PhcA. PhcA not only regulates the primary metabolism and AHL quorum sensing system of Ralstonia solanacearum, but also regulates the motility, iron carriers, biofilms, cell wall degrading enzymes, type III virulence factors that destroy the plant immune system, and extracellular polysaccharides, and other virulence behaviors closely related to the rhizosphere invasion process of Ralstonia solanacearum. Some studies have attempted to block the rhizosphere invasion process of soil-borne Ralstonia solanacearum by degrading quorum sensing molecules, but the blocking effect is not ideal.

[0003] At present, virtual screening is a very common strategy in computer-aided drug design and has been widely used. Drug-target affinity (DTA) prediction is an important step in virtual screening, which can quickly match targets and drugs and speed up the drug development process. DTA prediction provides information on the binding strength of drugs to target proteins and can be used to show whether small molecules bind to proteins. For proteins with known structure and site information, we can use molecular simulation and molecular docking for detailed simulation to obtain more accurate results, which is called structure-based virtual screening. However, there are still many proteins that do not have structural information. Even with homology models, it is still difficult to obtain structural information for many proteins. Therefore, using sequences (sequence-based virtual screening) to predict the binding affinity of proteins to drug molecules is an urgent issue, which is also the focus of the present invention.

[0004] Virtual screening based on molecular docking has become a core technology for computer-aided compound design and is widely used in the targeted development of new compounds. Therefore, using virtual screening of lead compounds for Ralstonia solanacearum quorum sensing to interfere with quorum sensing may be one of the important ways to control soil-borne bacterial wilt. Summary of the invention

[0005] The purpose of the present invention is to provide a targeted virtual screening method for quorum sensing lead compounds of Ralstonia solanacearum. The method can automatically search for the optimal graph structure based on reinforcement learning, improve the performance of the virtual screening model of quorum sensing lead compounds of the graph network, and is a screening method for quorum sensing lead compounds based on the PhcA and PhcR protein structures, which can be used to discover new compounds with quorum sensing activity, and provide new ideas and means for the control and prevention of bacteria such as Ralstonia solanacearum.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A virtual screening method for quorum sensing lead compounds, which takes the molecular compound structure and protein sequence as input and sends them to the preprocessing module to extract preliminary features, and then sends them to the prediction model network, wherein the structure and parameters of the prediction model are generated by LSTM controller training; the specific process is as follows:

[0008] The input molecular compound structure is preprocessed to construct a molecular adjacency matrix and sent to the GNN1 network to generate compound features;

[0009] The input protein sequence is used to extract its protein amino acid composition and dipeptide frequency to form a preliminary protein feature vector, which is then sent to the cross network to generate cross-fusion features. At the same time, the protein sequence is used to generate a corresponding contact map, which is then sent to the GNN2 network to generate protein sequence features.

[0010] Finally, the three features are combined and sent to the fully connected layer to predict the affinity value.

[0011] Furthermore, let the constructed molecular adjacency matrix be X1, the matrix element value of adjacent atoms in the molecular structure diagram is 1, and the non-adjacent atoms is 0. The size of the molecular adjacency matrix is ​​(n*n), where n is the number of nodes in the structure diagram, that is, the number of all atoms.

[0012] Furthermore, the protein sequence was processed using Pconsc4 software to output a probability matrix of whether the residue pairs were in contact or not, with a size of m*m. Values ​​greater than 0.5 in the matrix were retained, and other values ​​were set to 0. The filtered matrix was the protein contact map X2, where m is the number of residues.

[0013] Furthermore, the amino acid composition is the frequency of occurrence of each of the 20 amino acids constituting the sequence, and the frequency of the dipeptide is the frequency of occurrence of an amino acid pair consisting of any two amino acids.

[0014] Furthermore, the prediction model network is composed of a GNN1 network, a GNN2 network, and a cross network in parallel, which are merged and sent to a concatenation layer and a DROPOUT layer, and then connected to two fully connected layers.

[0015] Furthermore, the cross network is composed of 5 cross layers connected in series, and finally connected to a 128-dimensional fully connected layer. The output of the fully connected layer is f3. Each cross layer has the following formula:

[0016] C l+1 =C0C T l W c,l +b c,l +C l

[0017] Where: l = 1, 2, ..., 5, C l and C l+1 are the outputs of the lth layer and the l+1th layer cross layer, respectively. C0 is the combination of amino acid composition and dipeptide frequency X3, and W c,l and b c,l is the connection parameter between the two layers; all variables in the above formula are column vectors. The output of each layer is the output of the previous layer plus the feature cross.

[0018] Furthermore, the molecular adjacency matrix and protein contact map were respectively input into two different GNN1 and GNN2 networks, each network consisting of three GNN layers. The output features of the two GNN networks were f1 and f2, plus the cross-fusion features, which were concatenated into f1+f2+f3 to obtain the overall features of the corresponding small molecule-protein pairs for prediction; they were then sent to a fully connected layer with an output dimension of 128, and then to a second fully connected layer with an output dimension of 1, which was the affinity value predicted by the network.

[0019] Furthermore, the virtual screening network model is optimized through the LSTM controller. Model optimization is to use reinforcement learning in a certain parameter space to obtain the optimal structural parameters of two GNNs and the parameters of other neurons in the entire network; the GNN structure M needs to determine several parameters: sampling function (S), related measurement function (Att), aggregation function (Agg), multi-head attention number (K), output hidden embedding (Dim) and activation function (Act).

[0020] Furthermore, the optimization consists of two steps. First, LSTM predicts the corresponding operations of S, Att, Agg, Act, K, and Dim of a GNN1. Each prediction is performed by the softmax classifier of LSTM, and then the predicted value is input to the next time point to obtain the next parameter prediction; when the number of layers of GNN Layer reaches 3, the LSTM controller completes the generation of an architecture; the process is repeated to generate the parameters of GNN2; the entire prediction network is constructed and trained to obtain the weight parameters of the GNN network and other network layers; then, based on the accuracy obtained after network training, reinforcement learning is used to optimize the parameters of LSTM to obtain the optimized controller model; the two steps are executed alternately for a certain number of steps to obtain the final screening network model.

[0021] The above method can be applied in the virtual screening of lead compounds of Ralstonia solanacearum quorum sensing.

[0022] The beneficial effects of the present invention are as follows: first, the method extracts molecular adjacency matrix, protein contact map, and protein sequence cross-fusion features to form multi-dimensional features, which can better reflect the characteristics of molecules and proteins; second, the method uses reinforcement learning to optimize the model structure, avoiding the previous experience or a large number of manual selection of model parameters. The improvement and promotion of this method can effectively carry out virtual screening of lead compounds, which has broad prospects and extraordinary significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a virtual screening model diagram of the present invention;

[0024] Figure 2 It is a molecular graph representation. DETAILED DESCRIPTION

[0025] The present invention is described in detail below with reference to the accompanying drawings.

[0026] like Figure 1, a virtual screening method for quorum sensing lead compounds, which is a screening method for quorum sensing lead compounds based on PhcA and PhcR protein structures, and a screening method for Ralstonia solanacearum quorum sensing lead compounds established by computer-assisted virtual screening technology based on PhcA and PhcR protein structures, molecular biology technology, etc. The method is implemented through a virtual screening model. The input of the virtual screening model is the molecular compound structure and protein sequence, which are then sent to the preprocessing module to extract preliminary features, and then sent to the prediction model network. The structure and parameters of the prediction model are generated by LSTM controller training. Basic process: input the molecular formula of the compound, build the molecular adjacency matrix, and send it to the GNN1 subnetwork to generate compound features. Input the protein sequence, extract its protein amino acid composition and dipeptide frequency to form a preliminary protein feature vector, send it to the cross network, and generate cross-fusion features; at the same time, generate the corresponding contact map for the protein sequence, and then send it to the GNN2 network to generate protein sequence features. Finally, the three feature combinations are sent to the fully connected layer to predict the affinity value, and the output value is 0 without affinity and 1 with affinity. The following introduces the construction of molecular adjacency matrix, protein contact map, generation of protein sequence cross-fusion features, GNN structure optimization, and optimization of other parameters of the prediction network, which are achieved by the LSTM controller.

[0027] 1. Data preprocessing

[0028] (1) Construction of molecular adjacency matrix

[0029] The molecules are represented in the dataset in SMILES format. A molecular graph is constructed based on the drug SMILES string, which uses atoms as nodes and bonds as edges. The molecular graph structure construction process is shown as follows: Figure 2 As shown. Let the constructed molecular adjacency matrix be X1, the adjacent atomic matrix element value is 1, the non-adjacent atomic matrix element value is 0, and the size is (n*n), where n is the number of nodes in the graph, that is, the number of all atoms.

[0030] (2) Protein contact map

[0031] The protein contact map is a graphical representation method for describing the interaction between proteins. It shows the contact and interaction between proteins and is used to describe the protein structure and function. The protein sequence is processed using the Pconsc4 software, and the probability matrix of whether the residual pair is in contact is output. The size is m*m. The values ​​greater than 0.5 in the matrix are retained, and other values ​​are set to 0. The filtered matrix is ​​the protein contact map X2.

[0032] (3) Protein composition characteristics

[0033] The protein composition feature is a combination of amino acid composition and dipeptide frequency X3, with a size of 420 dimensions. The amino acid composition is the frequency of each of the 20 amino acids that make up the sequence. The dipeptide frequency is the frequency of occurrence of an amino acid pair consisting of any two amino acids. There are 20 amino acids that make up the protein sequence and 400 dipeptides.

[0034] 2. Prediction model architecture

[0035] (1) Cross-network

[0036] The cross network input is X3. The network consists of 5 cross layers connected in series, and finally a 128-dimensional fully connected layer. The output of the fully connected layer is f3. Each cross layer has the following formula:

[0037] C l+1 =C0C T l W c,l +b c,l +C l

[0038] Where: l = 1, 2, ..., 5. C l and C l+1 are the outputs of the crosslayer of the lth layer and the l+1th layer, C0 is X3, W c,l and b c,l is the connection parameter between the two layers. All variables in the above formula are column vectors. The output of each layer is the output of the previous layer plus the feature cross.

[0039] (2) Overall structure of affinity prediction network

[0040] The prediction model consists of two GNN networks (GNN1, GNN2) and a cross network in parallel. After merging, they are sent to a splicing layer and a DROPOUT layer, and then connected to two fully connected layers. The molecular adjacency matrix and protein contact map of drug molecules and proteins are input into two different GNN1 and GNN2 networks. Each network consists of three GNN layers. The output features of the two GNN networks are f1 and f2, and the cross-fusion features are added. After splicing, they are f1+f2+f3, and the overall features of the corresponding small molecule-protein pairs for prediction are obtained. Then it is sent to a fully connected layer with an output dimension of 128, and then to the second fully connected layer with an output dimension of 1, which is the affinity value predicted by the network. The specific structural parameters of the GNN layer and other network layers are obtained by the following network model optimization training process.

[0041] (3) LSTM controller realizes virtual screening network model optimization

[0042] Model optimization is to use reinforcement learning in a certain parameter space to obtain the best structural parameters of two GNNs and the parameters of other neurons in the entire network. The GNN structure M needs to determine several parameters: sampling function (S), related metric function (Att), aggregation function (Agg), multi-head attention number (K), output hidden embedding (Dim) and activation function (Act).

[0043] The specific description and corresponding parameter values ​​of each parameter are as follows:

[0044] 1. Output hidden embedding (Dim). Dim is the output dimension of each layer of GNN, which is an integer value.

[0045] 2. Sampling function (S). For each layer of GNN, a sampling function is required. Sampling is used in graph neural networks to select the receptive field for a given target node. Three types are used in this method: a. fixed number of neighbors sampling method, b. importance sampling method, and c. first-order neighbor sorting method.

[0046] 3. Related metric function (Att) and the number of multi-head attention K. For each layer of GNN, we select an Att method and the number of multi-head attention K. Att can choose two metric functions: GAT and GCN, corresponding to the two network structures of GAT and GCN respectively. The GCN network includes 2 graph convolution layers, 1 ReLU activation function layer and 1 Dropout layer. The GAT network includes K graph multi-head attention layers, 1 Softmax activation function layer and 1 Dropout layer. Among them, GAT assigns neighborhood importance by using the attention layer, and GCN assigns neighborhood importance according to the degree of the node.

[0047] 4. Aggregation function (Agg). For each layer of GNN, Agg aggregation is required. The optional aggregation functions Agg are: a. Sum aggregator, b. Mean aggregator, c. Pooling aggregator.

[0048] 5. Activation function (Act). For each layer of GNN, you need to use the Act activation function. The optional activation functions Act are: a.ReLU, b.LeakyReLU, c.ELU, d.Linear, e.Softmax. Increase the nonlinear fitting ability of the graph network and improve the expressiveness of the model.

[0049] The network model optimization uses an LSTM controller neural network to train the network. It consists of two steps. First, LSTM predicts the corresponding operation of [S, Att, Agg, Act, K, Dim] of a GNN1. Each prediction is performed by the softmax classifier of LSTM. Then the predicted value is input to the next time point to obtain the next parameter prediction. When the number of layers of GNN Layer reaches 3, the LSTM controller completes the generation of an architecture; repeat the process to generate the parameters of GNN2. Construct and train the entire prediction network to obtain the weight parameters of the GNN network and other network layers. Then, based on the accuracy obtained after network training, the parameters of LSTM are optimized by reinforcement learning to obtain the optimized controller model. The two steps are executed alternately for a certain number of steps to obtain the final screening network model.

[0050] The following is a further explanation of the model training:

[0051] The model training uses the public KIBA dataset. The dataset includes 229 unique proteins and 2,111 unique drugs, with 118,254 affinity pairs between proteins and drug molecules. In the training method, the dataset is divided into training set, validation set and test set in a ratio of 80%:10%:10%.

[0052] Training parameters: The controller is an LSTM network with 100 hidden units. It is trained using the ADAM optimizer with a learning rate of 0.0035. The controller samples the graph network structure, generates sub-models, and trains for 200 epochs. During training, L2 regularization with λ=0.0005 is applied. In addition, Dropout with p=0.5 is applied to the input of both layers as well as the normalized attention layer.

[0053] LSTM consists of three layers: input layer, hidden layer and output layer. The input dimension is 6*1, the hidden layer neurons are 100, and the time step is 10.

[0054] The LSTM network parameters are set as follows:

[0055] Layer1: LSTM (input_size=8, hidden_size=100, num_layers=2)

[0056] Layer2: Dropout (p=0.5)

[0057] Layer3:Linear(hidden_size=500,n_class=1)

[0058] Among them, input_size represents the input data dimension; hidden_size represents the output dimension; num_layers represents the number of LSTM layers stacked, and the default is 1; n_class represents the output dimension of the LSTM network, and 1 represents the output regression value.

[0059] After the controller was trained 1000 times, we let the controller output the best model from the 200 sampled GNNs. The results show that this optimization method can design the best model for the original virtual screening model.

[0060] Experimental Results

[0061] After the controller was trained 1000 times, we let the controller output the best model from the 200 sampled GNNs. The results show that this optimization method can design the best model for the original virtual screening model.

[0062] The optimal virtual screening model structure after optimization and screening is as follows:

[0063] GNN1 structure:

[0064] Layer1: molecular graph attention layer GAT1 (input dimension = (n*n), output dimension = 128, number of hidden layer units = 128, number of head attention = 4, activation function = elu(), aggregation function = sum()).

[0065] Layer2: molecular graph convolution layer GCN2 (input dimension = 128, output dimension = 256, number of hidden layer units = 256, number of head attention = 4, activation function = relu(), aggregation function = max()).

[0066] Layer3: Molecular graph attention layer GAT3 (input dimension = 256, output dimension = 128, number of hidden layer units = 128, number of head attention = 8, activation function = elu(), aggregation function = avg()).

[0067] GNN2 structure:

[0068] Layer 1: protein graph convolutional layer GCN4 (input dimension = (m*m), output dimension = 64, number of hidden layer units = 64, number of head attention = 16, activation function = relu(), aggregation function = max()).

[0069] Layer2: Protein graph attention layer GCN5 (input dimension = 64, output dimension = 256, number of hidden layer units = 256, number of head attention = 4, activation function = elu(), aggregation function = pooling()).

[0070] Layer3: protein graph convolutional layer GCN6 (input dimension = 256, output dimension = 256, number of hidden layer units = 256, number of head attention = 16, activation function = relu(), aggregation function = max()).

[0071] Cross network: input dimension = 420, output dimension = 128, number of layers n = 5

[0072] Feature concatenation layer: Concat1 (input dimension = (128, 256, 128), output dimension = 512).

[0073] Dropout layer: (512, 512), ratio p=0.5.

[0074] Fully connected layer: Linear(512,128).

[0075] Fully connected layer: Linear(128,1).

[0076] After 300 iterations, the network reached a good state, and the corresponding parameters were saved for virtual screening of lead compounds for quorum sensing of Ralstonia solanacearum.

[0077] In summary, the present invention provides a method for screening quorum sensing lead compounds based on the structure of PhcA and PhcR proteins, which can be used to discover new compounds with quorum sensing activity, and provide new ideas and means for the control and prevention of bacteria such as Ralstonia solanacearum. The method combines computer-assisted virtual screening technology with molecular biology technology to efficiently screen out compounds that bind to PhcA and PhcR proteins, thereby discovering compounds with quorum sensing activity. The method has the advantages of simple operation, high efficiency and rapidity, high screening accuracy, etc., and can be widely used in the fields of biomedicine and agriculture.

[0078] It is worth noting that PhcA and PhcR in the present invention are two key proteins in Ralstonia solanacearum, and their structure and function play an important role in bacterial quorum sensing. However, in other bacteria, there may be different quorum sensing proteins, so it is necessary to screen and study according to different bacterial species in order to find quorum sensing lead compounds suitable for different bacteria.

[0079] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the protection scope of the present invention in any form, and all technical solutions obtained by equivalent replacement and the like fall within the protection scope of the present invention.

[0080] The parts not involved in the present invention are the same as the prior art or can be implemented by using the prior art.

Claims

1. A virtual screening method for quorum sensing lead compounds, characterized in that: The molecular compound structure and protein sequence are sent as input to the preprocessing module to extract preliminary features, and then sent to the prediction model network, where the structure and parameters of the prediction model are generated through LSTM controller training; the specific process is as follows: The input molecular compound structure is preprocessed to construct a molecular adjacency matrix and sent to the GNN1 network to generate compound features; The input protein sequence is used to extract its protein amino acid composition and dipeptide frequency to form a preliminary protein feature vector, which is then sent to the cross network to generate cross-fusion features. At the same time, the protein sequence is used to generate a corresponding contact map, which is then sent to the GNN2 network to generate protein sequence features. Finally, the three features are combined and sent to the fully connected layer to predict the affinity value; The prediction model network is composed of a GNN1 network, a GNN2 network, and a cross network in parallel, which are merged and sent to a concatenation layer and a DROPOUT layer, and then connected to two fully connected layers; The cross network consists of 5 cross layers in series, and finally a 128-dimensional fully connected layer. The output of the fully connected layer is f 3. Each cross layer has the following formula: C l+1 = C 0 C T l W c,l +b c,l +C l in: l =1, 2, …, 5 , C l and C l+1 They are l Layer and l +1 cross layer output, C 0 That is, the combination of amino acid composition and dipeptide frequency X 3 , X 3 As the crossover network input, W c,l and b c,l is the connection parameter between the two layers; all variables in the above formula are column vectors; the output of each layer is the output of the previous layer plus the feature cross; The molecular adjacency matrix and protein contact map are input into two different GNN1 and GNN2 networks respectively. Each network consists of three GNN layers. The output features of the two GNN networks are , plus the cross-fusion features, the splicing is , and obtain the overall characteristics of the corresponding small molecule-protein pair for prediction; then it is sent to a fully connected layer with an output dimension of 128, and then sent to a second fully connected layer with an output dimension of 1, which is the affinity value predicted by the network.

2. The virtual screening method for a quorum sensing lead compound according to claim 1, characterized in that The molecular adjacency matrix constructed is , the adjacent atomic matrix element value on the molecular structure diagram is 1, and the non-adjacent atomic matrix element value is 0. The size of the molecular adjacency matrix is , is the number of nodes in the structure graph, that is, the number of all atoms.

3. The virtual screening method for quorum sensing lead compounds according to claim 1, characterized in that: Use Pconsc4 software to process protein sequences and output the probability matrix of whether residue pairs are in contact , retain the values ​​greater than 0.5 in the matrix, set other values ​​to 0, and the filtered matrix is ​​the protein contact map X 2 ,in is the number of residuals.

4. The virtual screening method for quorum sensing lead compounds according to claim 1, characterized in that: The amino acid composition is the frequency of occurrence of each of the 20 amino acids that make up the sequence, and the frequency of dipeptides is the frequency of occurrence of amino acid pairs consisting of any two amino acids.

5. The virtual screening method for quorum sensing lead compounds according to claim 1, characterized in that: The virtual screening network model is optimized through the LSTM controller. Model optimization is to use reinforcement learning in a certain parameter space to obtain the optimal structural parameters of two GNNs and the parameters of other neurons in the entire network; GNN structure M Several parameters need to be determined: sampling function (S), relevant metric function (Att), aggregation function (Agg), number of multi-head attention (K), output hidden embedding (Dim) and activation function (Act).

6. The virtual screening method for quorum sensing lead compounds according to claim 5, characterized in that: The optimization consists of two steps. First, LSTM predicts the corresponding operations of S, Att, Agg, Act, K, and Dim of a GNN1. Each prediction is performed by the softmax classifier of LSTM, and then the predicted value is input to the next time point to obtain the next parameter prediction; when the number of layers of GNN Layer reaches 3, the LSTM controller completes the generation of an architecture; repeat the process to generate the parameters of GNN2; build and train the entire prediction network to obtain the weight parameters of the GNN network and other network layers; then use reinforcement learning to optimize the parameters of LSTM based on the accuracy obtained after network training to obtain the optimized controller model; the two steps are executed alternately for a certain number of steps to obtain the final screening network model.

7. Use of the virtual screening method for quorum sensing lead compounds according to any one of claims 1 to 6 in Ralstonia solanacearum.

Citation Information

Patent Citations

  • Intelligent prediction method for small molecule-protein binding affinity

    CN114333984A

  • Transform and graph neural network-based drug prediction algorithm

    CN115188412A

  • Drug and protein affinity prediction method and system based on graph neural network

    CN115631805A