A Drug Target Matching and Recommendation Method Based on Multi-View Graph Neural Network

By combining point and edge feature message passing and cross-attention mechanism with multi-view graph neural network, the problem of insufficient accuracy in drug target prediction in existing technology is solved, and more efficient drug target matching and recommendation is achieved.

CN119601078BActive Publication Date: 2025-11-14CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411665435.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-20
Publication Date
2025-11-14
Estimated Expiration
2044-11-20

AI Technical Summary

Technical Problem

Existing graph neural networks fail to fully utilize the graph structure and sequence information of drug molecules in predicting drug target interactions, resulting in insufficient prediction accuracy.

Method used

A multi-view graph neural network is used to extract the feature vectors of drug molecules through point and edge feature message passing, and the atomic, chemical bond and sequence information are fused by combining the cross attention mechanism to construct a drug target matching recommendation model.

Benefits of technology

It improves the accuracy of drug target interaction prediction by comprehensively considering local and global information of drug molecules, thus enhancing the model's predictive ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119601078B_ABST
    Figure CN119601078B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of computational biology, specifically relating to a drug target matching and recommendation method based on a multi-view graph neural network. The method includes: acquiring information about the drug target to be predicted, inputting it into a trained drug target recommendation model, and selecting drug targets with a prediction probability greater than an interaction threshold to recommend to the user. This invention utilizes a graph neural network based on point and edge message passing to extract drug molecule feature vectors, modeling and extracting features from both atomic and chemical bond perspectives. Simultaneously, considering the global nature of drug sequence information, a cross-attention mechanism is used to interact the atomic features, chemical bond features, and sequence features of the drug molecule, more accurately capturing the complex relationships between drug molecules and target proteins, and improving the accuracy of drug target interaction prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computational biology, specifically relating to a drug target matching and recommendation method based on multi-view graph neural networks. Background Technology

[0002] Drug development is a challenging process. Traditional biological experimental methods, such as high-throughput screening, can accurately screen for drug-target pairs that can react. However, this method is cumbersome, requires significant time and financial investment, and the existence of millions of compounds and hundreds of potential targets poses a considerable challenge. With the rapid development of bioinformatics, computational biology, and drug discovery technologies, bioinformatics databases have accumulated vast amounts of biological experimental data, and the focus of drug research has gradually shifted from traditional laboratory experiments to computational methods. In the drug development process, predicting drug-target interactions has become a crucial area.

[0003] Current mainstream methods utilize graph structure data of drug molecules and sequence data of protein molecules to predict drug-target interactions. However, existing graph neural networks often focus only on the information connections between different atoms in drug molecules, neglecting the equally important role of chemical bonds connecting atoms in the chemical properties of drug molecules. Feature extraction in graph neural networks should not only consider information transmission between atoms but also the information transmission between atoms and chemical bonds. Furthermore, drug sequence information is often discarded after conversion to a graph structure. However, graph structure data can only represent local information of drug molecules to a certain extent, while molecular sequences can reflect global information. Fully utilizing both graph and sequence structure data can further improve the accuracy of model predictions. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a drug target matching and recommendation method based on a multi-view graph neural network, comprising:

[0005] Obtain information on the drug targets to be predicted, input it into a trained drug target recommendation model, and select drug targets with a predicted probability greater than the interaction threshold to recommend to the user.

[0006] The training process for the drug target recommendation model includes:

[0007] S1: Obtain drug target information, including: drug SMILES sequence, molecular map structure and target amino acid sequence;

[0008] S2: Construct initial atomic feature vectors and initial chemical bond feature vectors based on the atomic and chemical bond properties of the drug, respectively. Extract the atomic feature vector h of the drug molecule graph using a graph neural network based on point feature message passing and a graph neural network based on edge feature message passing, respectively. atom and chemical bond eigenvector h bond ;

[0009] S3: Obtain the sequence characteristics h of the drug based on the drug SMILES sequence. seq The cross-attention mechanism is used to interact the atomic feature vector and chemical bond feature vector of the drug with the drug sequence feature vector, respectively. The resulting atomic and chemical bond feature vectors are then concatenated to form the final feature vector H of the drug molecule. d ;

[0010] S4: Obtain the target sequence feature vector H based on the target amino acid sequence. p The feature vectors of drug molecules and target sequences are interacted to obtain the matching probability of the drug target.

[0011] S5: Calculate the loss based on the probability output by the model and the acquired drug target matching information to update the weights of the recommendation model, and obtain the trained drug target recommendation model.

[0012] The beneficial effects of this invention are:

[0013] This invention utilizes a graph neural network based on point and edge message passing to extract feature vectors of drug molecules. It models and extracts features of drug molecules from both atomic and chemical bond perspectives. Considering the global nature of drug sequence information, it uses a cross-attention mechanism to interact the atomic features, chemical bond features, and sequence features of drug molecules, thereby more accurately capturing the complex relationship between drug molecules and target proteins and improving the accuracy of drug-target interaction prediction. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the overall process of a drug target matching and recommendation method based on a multi-view graph neural network according to the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] Example 1: In this embodiment, the present invention provides a drug target matching and recommendation method based on a multi-view graph neural network, such as... Figure 1 As shown, it includes:

[0017] The system obtains the SMILES sequences and molecular structures of the drugs to be predicted, as well as the amino acid sequence information of the targets to be predicted. It then inputs these information into a trained drug target recommendation model and selects drug targets with a prediction probability greater than the interaction threshold to recommend to the user.

[0018] The training process for the drug target recommendation model includes:

[0019] S1: Obtain drug target information, including drug SMILES sequences, molecular diagram structure, and target amino acid sequences;

[0020] S2: Construct initial atomic feature vectors and initial chemical bond feature vectors based on the atomic and chemical bond properties of the drug, respectively. Extract the atomic feature vector h of the drug molecule graph using a graph neural network based on point feature message passing and a graph neural network based on edge feature message passing, respectively. atom and chemical bond eigenvector h bond ;

[0021] S3: Obtain the sequence characteristics h of the drug based on the drug SMILES sequence. seq The cross-attention mechanism is used to interact the atomic feature vector and chemical bond feature vector of the drug with the drug sequence feature vector, respectively. The resulting atomic and chemical bond feature vectors are then concatenated to form the final feature vector H of the drug molecule. d ;

[0022] S4: Obtain the target sequence feature vector H based on the amino acid sequence of the target. p The feature vectors of drug molecules and target sequences are interacted to obtain the matching probability of the drug target.

[0023] S5: Calculate the loss based on the probability output by the model and the acquired drug target matching information to update the weights of the recommendation model, and obtain the trained drug target recommendation model.

[0024] Initial atomic feature vectors and initial chemical bond feature vectors are constructed based on the atomic and chemical bond properties of the drug, including:

[0025] S21: One-hot encoding is performed on the atomic number, number of chemical bonds, formal charge, chirality (unspecified, tetral-hedral CW / CCW, and other), and hybridization (sp, sp2, sp3, sp3d, and sp3d2) of each atom. The resulting 0-1 feature vector is used as the initial feature vector h of the atom. 0 (v), where v represents an atom and the feature dimension is d1.

[0026] S22: Perform one-hot encoding on the chemical bond type (single, double, triple, and aromatic), whether it is a conjugated bond, whether it is part of a ring, and its stereochemistry (none, any, E / Z, and cis / trans) of each chemical bond to obtain a 0-1 feature vector as the initial feature vector h of the chemical bond. 0 (e uv ), where e uv This represents the chemical bond connecting atoms u and v, with a characteristic dimension of d2.

[0027] Graph neural networks based on point feature message passing include:

[0028] A graph neural network based on point feature message passing contains T time steps. In each time step t, the atomic v is determined according to the message function. Vector m is obtained by aggregating information from atoms adjacent to atom v. t+1 (v), update function Using m t+1 (v), the initial eigenvector h of atom v. 0 (v) and the eigenvector h at time step t t (v) Obtain the eigenvector h of atom v at the next time step t+1. t+1 (v). m t+1 (v) and h t+1 The formula for calculating (v) is as follows.

[0029]

[0030]

[0031] Where N(v) represents the set of atoms adjacent to atom v, m t+1 (v) and h t+1 The feature dimension of (v) is d1.

[0032] The message function at time step t of a graph neural network based on point message passing. The calculation formula is:

[0033]

[0034] Where ⊙ denotes element-wise multiplication, k uv This represents the attention weight between atoms u and v. The message function first calculates the maximum value of each feature dimension in the set of neighboring atoms of atom v to obtain the information of the most important neighboring atoms. Then, by weighting and summing the feature vectors of all neighboring atoms, we obtain ∑ which contains information about all neighboring atoms. u∈N(v) k uv ht (u). Multiply the two feature vectors element-wise to obtain the aggregated information vector m of atom v in the next time step. t+1 (v).

[0035] The weight k between atom u and atom v uv The calculation formula is:

[0036]

[0037] Among them, W k Let be the third trainable parameter matrix, with dimensions 2d1×d1. || denotes concatenating the feature vectors. τ is the temperature coefficient hyperparameter used to control k. uv The steepness of the distribution, exp represents the exponential function, ReLU represents the activation function, h(u) represents the characteristic representation of atom u, h(v) represents the characteristic representation of atom v, N(u) represents the set of atoms adjacent to atom u, and h(l) represents the characteristic representation of atom l.

[0038] The update function of a graph neural network based on point message passing at time step t The calculation formula is:

[0039]

[0040] Update function Using the aggregated information vector m of atom v t+1 (v) and the initial feature vector h 0 (v) Update its feature vector h at the next time step. t+1 (v). ELU is the activation function for Exponential Linear Units; It is a trainable parameter matrix with dimensions 2d1×d1, where a is the number of atoms in the drug molecule; h t+1 (b) has a feature dimension of d1.

[0041] In the final time step, all atoms in the drug molecule are spliced ​​together to obtain the atomic feature vector h of the drug molecule. atom The dimension is a×d1, where a represents the number of atoms.

[0042] The graph neural network based on edge feature message passing contains T time steps. In each time step t, the feature vector of the chemical bond between atom u and atom v is determined according to the message function. Vector m is obtained by aggregation from neighboring chemical bonds. t+1 (e uv The update function uses vector m t+1 (e uv ) and chemical bonds e uvThe initial eigenvector h 0 (e uv ) to get e uv The eigenvector h at the next time step t+1 t+1 (e uv ). m t+1 (e uv ) and h t+1 (e uv The formula for calculating ) is:

[0043]

[0044]

[0045] Where N(v) represents the set of atoms adjacent to atom v but not containing u. The message function of the graph neural network based on edge message passing at time step t... The calculation formula is:

[0046]

[0047] The update function of a graph neural network based on edge message passing is calculated using the following formula at time step t:

[0048]

[0049] It is a trainable parameter matrix with dimensions 2d²×d². t+1 (e uv ) and h t+1 (e uv The feature vector dimension of ) is d2. In the last time step T of the graph neural network based on edge feature message passing, the chemical bond e uv The feature vectors of atoms u and v are transformed into feature vectors of atoms u and v using the following formula:

[0050]

[0051]

[0052] Among them, W e It is the fourth trainable parameter matrix with dimensions d2×d1.

[0053] In the final time step, all atoms in the drug molecule are spliced ​​together to obtain the chemical bond feature vector h of the drug molecule. bond The dimension is a×d1, where a represents the number of atoms.

[0054] The sequence characteristics h of the drug were obtained based on the drug's SMILES sequence. seqThe cross-attention mechanism is used to interact the atomic feature vector and chemical bond feature vector of the drug with the drug sequence feature vector, respectively. The resulting atomic and chemical bond feature vectors are then concatenated to form the final feature vector H of the drug molecule. d ,include:

[0055] S31: Assign a randomly initialized feature vector of dimension d1 to each character in the SMILES sequence, and use a one-dimensional gated convolutional neural network to extract the features h of the SMILES sequence. seq The dimension is c×d1.

[0056] S32: The atomic characteristic vector h of the drug molecule atom As the query in the attention mechanism, the sequence feature h seq Using the attention mechanism as the key and value, the atomic feature vector H after the interaction is calculated. atom ,include:

[0057]

[0058] Among them, H atom The feature dimension is d1, T1 represents matrix transpose, and Softmax represents the activation function.

[0059] S33: The chemical bond feature vector h of the drug molecule bond As the query in the attention mechanism, the sequence feature h sep Using these as Key and Value, the atomic feature vector H after interaction is calculated based on the attention mechanism. bond ,include:

[0060]

[0061] Among them, H bond The feature dimension is d1.

[0062] S34: H atom and H bond After splicing, the final feature vector H of the drug molecule is obtained. d .

[0063] The target sequence feature vector H is obtained based on the amino acid sequence of the target. p The feature vectors of the drug molecule and the feature vector of the target sequence are interacted to obtain the matching probability of the drug target, including:

[0064] S41: Divide the target amino acid sequence into multiple overlapping ternary amino acid subsequences, and assign a randomly initialized feature vector of dimension d1 to each ternary amino acid subsequence. Extract features using a one-dimensional gated convolutional neural network to obtain the target sequence feature vector H. p The dimension is b×d1;

[0065] The formula for calculating the i-th layer of a one-dimensional gated convolutional neural network is:

[0066]

[0067] Among them, X i+1 It is based on the input X of the i-th layer i The obtained feature map vector, and X 0 A randomly initialized feature vector is assigned to each ternary amino acid subsequence. and The fifth and sixth trainable parameter matrices are f×m1×m2, where f represents the kernel size and m1 and m2 represent the dimensions of the input and output, respectively. The output of the last layer is H. p .

[0068] S42: The feature vector H of the drug molecule d As the query in the attention mechanism, the target sequence feature H p Using the key and value, the interaction feature vector is calculated based on the attention mechanism. The formula is:

[0069]

[0070] S43: The feature vector H of the drug molecule p As the query in the attention mechanism, the target sequence feature H d Using the key and value, the interaction feature vector is calculated based on the attention mechanism. The formula is:

[0071]

[0072] S44: Transforming the eigenvector through linear transformation and Mapping to the same dimension d, concatenating the two data points and feeding them into a binary classification linear layer yields the probability of drug-target matching. The formula is:

[0073]

[0074] Among them, W outLet represent the trainable parameter matrix of the binary classification linear layer, and σ represent the Sigmoid activation function.

[0075] The weights of the recommendation model are updated by calculating the loss based on the probability output by the model and the acquired drug target matching information, resulting in a trained drug target recommendation model. The loss function is calculated as follows:

[0076]

[0077] in Let y represent the total loss of the model, N represent the number of samples, and y represent the total loss of the model. i Let represent the label of the i-th sample. This represents the model's prediction result for the i-th sample.

[0078] The system acquires information on the drug molecules and target proteins to be predicted, inputs them into the trained model, obtains the interaction prediction probability of the drug targets, and selects drug targets with interaction probabilities greater than the interaction threshold to recommend to the user; preferentially, the interaction threshold is set to 0.65.

[0079] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A drug target matching and recommendation method based on multi-view graph neural networks, characterized in that, include: Obtain information on the drug targets to be predicted, input it into a trained drug target recommendation model, and select drug targets with a predicted probability greater than the interaction threshold to recommend to the user. The training process for the drug target recommendation model includes: S1: Obtain drug target information, including: drug SMILES sequence, molecular map structure and target amino acid sequence; S2: Construct initial atomic feature vectors and initial chemical bond feature vectors based on the atomic and chemical bond properties of the drug, respectively. Extract the atomic feature vector h of the drug molecule graph using a graph neural network based on point feature message passing and a graph neural network based on edge feature message passing, respectively. atom and chemical bond eigenvector h bond ; Initial atomic feature vectors and initial chemical bond feature vectors are constructed based on the atomic and chemical bond properties of the drug, including: S21: One-hot encoding is performed on the atomic number, number of chemical bonds, formal charge, chirality, and hybridization of each atom. The resulting 0-1 feature vector is used as the initial feature vector h of the atom. 0 (v), where v represents an atom and the feature dimension is d. 1; S22: One-hot encoding is performed on the chemical bond type, whether it is a conjugated bond, whether it is part of a ring, and its stereochemistry for each chemical bond. The resulting 0-1 feature vector is used as the initial feature vector h of the chemical bond. 0 (e uv ), where e uv The chemical bond connecting atoms u and v has a characteristic dimension of d2. Graph neural networks based on point feature message passing include: The graph neural network based on point feature message passing contains T time steps. In each time step t, atom v follows the message function of the graph neural network based on point feature message passing. The vector m at the next time step t+1 is obtained by aggregating information from the atoms adjacent to atom v. t+1 (v), Update function of graph neural network based on point feature message passing Using vector m t+1 (v), the initial eigenvector h of atom v 0 (v) and the eigenvector h at time step t t (v) Obtain the eigenvector h of atom v at the next time step t+1. t+1 (v); In the last time step T, all atomic features are integrated to obtain the atomic feature vector h. atom ; m t+1 (v) and h t+1 The formula for calculating (v) is as follows: Where N(v) represents the set of atoms adjacent to atom v, m t+1 (v) and h t+1 (v) has a feature dimension of d1, h t (u) represents the eigenvector of the atom u adjacent to atom v at time step t; Graph neural networks based on edge feature message passing include: The graph neural network based on edge feature message passing contains T time steps. In each time step t, the feature vector of the chemical bond between atom u and atom v is determined by the message function of the graph neural network based on edge feature message passing. The vector m at the next time step t+1 is obtained by polymerizing from neighboring chemical bonds. t+1 (e uv The update function of a graph neural network based on edge feature message passing. Using vector m t+1 (e uv ), chemical bond e uv The initial eigenvector h 0 (e uv ) and the eigenvector h at time step t t (e uv ) to get e uv The eigenvector h at the next time step t+1 t+1 (e uv In the last time step T, all atomic features are integrated to obtain the atomic feature vector h. bond ; m t+1 (e uv ) and h t+1 (e uv The formula for calculating ) is: Where N(v) represents the set of atoms adjacent to atom v but not containing atom u, h t (e vk (e) represents the chemical bond between atom u and atom k. vk The eigenvector at time step t; S3: Obtain the sequence characteristics h of the drug based on the drug SMILES sequence. seq The cross-attention mechanism is used to interact the atomic feature vector and chemical bond feature vector of the drug with the drug sequence feature vector, respectively. The resulting atomic and chemical bond feature vectors are then concatenated to form the final feature vector H of the drug molecule. d ; The sequence characteristics h of the drug were obtained based on the drug's SMILES sequence. seq The cross-attention mechanism is used to interact the atomic feature vector and chemical bond feature vector of the drug with the drug sequence feature vector, respectively. The resulting atomic and chemical bond feature vectors are then concatenated to form the final feature vector H of the drug molecule. d ,include: S31: Assign a randomly initialized feature vector of dimension d1 to each character in the SMILES sequence, and use a one-dimensional gated convolutional neural network to extract the features h of the SMILES sequence. seq The dimension is c×d1, where c represents the length of the SMILES sequence; S32: The atomic characteristic vector h of the drug molecule atom As the query in the attention mechanism, the sequence feature h seq As the Key and Value in the attention mechanism, the atomic feature vector H after interaction is calculated using the attention mechanism. atom ; S33: The chemical bond feature vector h of the drug molecule bond As the query in the attention mechanism, the sequence feature h seq As the Key and Value in the attention mechanism, the atomic feature vector H after interaction is calculated according to the attention mechanism. bond ; S34: H atom and H bond After splicing, the final feature vector H of the drug molecule is obtained. d ; S4: Obtain the target sequence feature vector H based on the target amino acid sequence. p The feature vectors of drug molecules and target sequences are interacted to obtain the matching probability of the drug target. The target sequence feature vector H is obtained based on the amino acid sequence of the target. p By interacting the feature vectors of drug molecules and target sequences, the matching probability of the drug target is obtained, including... S41: Divide the target amino acid sequence into multiple overlapping ternary amino acid subsequences, assign a randomly initialized feature vector of dimension d1 to each ternary amino acid subsequence, and extract the feature vector H of the target sequence using a one-dimensional gated convolutional neural network. p The dimension is b×d1, where b represents the length of the ternary amino acid subsequence; S42: The feature vector H of the drug molecule d As the query in the attention mechanism, the target sequence feature H p Using the key and value, the interaction feature vector is calculated based on the attention mechanism. S43: The feature vector H of the drug molecule p As the query in the attention mechanism, the target sequence feature H d Using the key and value, the interaction feature vector is calculated based on the attention mechanism. S44: Transforming the eigenvector through linear transformation and Mapping to the same dimension d, concatenating the two and feeding them into a binary classification linear layer, we obtain the probability of drug and target matching; S5: Calculate the loss based on the probability output by the model and the acquired drug target matching information to update the weights of the recommendation model, and obtain the trained drug target recommendation model.

2. The drug target matching and recommendation method based on a multi-view graph neural network according to claim 1, characterized in that, The message function of the graph neural network based on point feature message passing include: Where N(v) represents the set of atoms adjacent to atom v, ⊙ represents element-wise multiplication, and k uv h represents the attention weight between atoms u and v. t (u) represents the eigenvector of the atom u adjacent to atom v at time step t; The update function of the graph neural network based on point feature message passing include: Where ELU stands for Exponential Linear Units activation function. H represents the first trainable parameter matrix, with dimensions 2d1×d1; t+1 (v) has a feature dimension of d1, h t (v) represents the eigenvector of atom v at time step t, m t+1 (v) represents the vector of atom v at the next time step t+1, h 0 (v) represents the initial eigenvector of atom v.

3. The drug target matching and recommendation method based on a multi-view graph neural network according to claim 1, characterized in that, The message function of the graph neural network based on edge feature message passing include: The update function of the graph neural network based on edge feature message passing include: Where ELU represents the Exponential Linear Units activation function, N(v) represents the set of atoms adjacent to atom v but not containing u, and ⊙ represents element-wise multiplication. Denotes the second trainable parameter matrix, with dimensions 2d²×d²; m t+1 (e uv ) and h t+1 (e uv The feature vector dimension of h is d2. t (e vk (e) represents the chemical bond between atom u and atom k. vk At time step t, the eigenvector h t (e uv (e) represents the chemical bond between atoms u and v. uv At time step t, the eigenvector h 0 (e uv (e) represents the chemical bond between atoms u and v. uv The initial eigenvector, m t+1 (e uv (e) represents the chemical bond between atoms u and v. uv The feature vector at the next time step t+1.

Citation Information

Patent Citations

  • Drug and target interaction prediction method and device, equipment and storage medium

    CN113160894A

  • Transform and graph neural network-combined drug target interaction prediction method

    CN116417093A