Drug target prediction method and device, electronic equipment and storage medium

By dynamically gated fusion drugs with graph structural features and sequence features, combined with dynamic feature enhancement and context-aware optimization, the problems of feature characterization incompleteness and lack of cross-modal association mechanisms in existing drug target prediction methods are solved, and high-precision prediction of drug targets is achieved.

CN120496668APending Publication Date: 2025-08-15GUANGDONG INST OF INTELLIGENT SCI & TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510491718.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing drug target prediction methods have problems with incompleteness of characteristic characterization and lack of cross-modal association mechanisms, which makes the interpretability of the model difficult to meet the actual needs of medicinal chemists.

Method used

By obtaining the graph structural features and sequence features of the target drug for dynamic gating, combining dynamic feature enhancement and context-aware optimization, neural network training prediction models are used to achieve target affinity score prediction between drugs and proteins.

Benefits of technology

It realizes high-precision prediction of drug targets, breaks through the accuracy bottleneck of traditional single-modal methods, and improves the efficiency of feature fusion and the interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496668A_ABST
    Figure CN120496668A_ABST
Patent Text Reader

Abstract

The invention discloses a drug target prediction method and device, electronic equipment and a storage medium. The method comprises the following steps: performing dynamic gating fusion on graph structure features and sequence features of a target drug to obtain drug features; performing dynamic feature enhancement on the digitized sequence of the target protein, and further obtaining protein features through context sensing optimization; obtaining a target affinity score of the target drug and the target protein by using a prediction model based on the drug characteristics and the protein characteristics; wherein the prediction model is obtained through feature representation training marked with actual affinity scores on the basis of a neural network. According to the method, the prediction precision of drug target interaction is improved through deep fusion of drug multi-modal features and protein sequence context perception optimization. The method can realize high-precision prediction of the drug target, and can be widely applied to the technical field of drug target prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug target prediction, and in particular to a drug target prediction method, device, electronic equipment and storage medium. Background Art

[0002] Predicting drug-target interaction affinity involves a system of bioinformatics and cheminformatics technologies. Its core approach is to establish interpretable prediction models by quantifying the binding strength between drug molecules and biological targets. The current mainstream approach in this field relies on deep mining of databases of known drug-target interactions and employing machine learning algorithms to build prediction systems, effectively assessing the binding ability of novel compounds to specific targets. This technology plays a key supporting role in the new drug development process, significantly shortening the lead compound screening cycle and reducing preclinical research costs.

[0003] Existing technologies suffer from the following inherent flaws: Single-modality analysis methods suffer from incomplete feature representation: ① Sequence-based methods struggle to capture the dynamic influence of conformation on molecular recognition, such as the contribution of the spatial arrangement of drug molecules to drug-target binding; ② Structural-based methods tend to overlook the functional relevance of conserved sequence regions, such as the biological properties of key amino acid residues. More significantly, existing algorithms lack effective cross-modal correlation mechanisms at the feature fusion level, making the model's interpretability difficult to meet the practical needs of medicinal chemists. Summary of the Invention

[0004] The main purpose of the embodiments of the present invention is to propose a drug target prediction method, device, electronic device and storage medium in order to solve at least one problem of the prior art. The present invention can achieve high-precision prediction of drug targets.

[0005] To achieve the above objectives, one aspect of an embodiment of the present invention provides a drug target prediction method, the method comprising:

[0006] Obtain the graph structure features and sequence features of the target drug, and perform dynamic gating fusion on the graph structure features and sequence features to obtain drug features;

[0007] Dynamic feature enhancement is performed on the digital sequence of the target protein, and protein features are obtained through context-aware optimization;

[0008] Based on drug characteristics and protein characteristics, the prediction model is used to predict the target affinity score between the target drug and the target protein;

[0009] The prediction model is trained based on a neural network using feature representations labeled with actual affinity scores.

[0010] In some embodiments, obtaining the graph structure features of the target drug includes the following steps:

[0011] Get the SMILES of the target drug and convert the SMILES into a mol object;

[0012] Extract the atomic parameter features of each atom from the mol object, use the atomic parameter features as node features, and construct an edge set based on the chemical bond relationship between each atom;

[0013] Obtain the molecular graph structure of the target drug based on node features and edge sets;

[0014] Based on the molecular graph structure, a graph convolutional network is used to extract the feature vector of each atomic node, and the attention score between each atomic node is obtained through the graph self-attention network;

[0015] Normalize the attention scores to get the attention weights;

[0016] The feature vector is updated using the attention weight, and then the updated feature vectors of all atomic nodes are combined to obtain the graph structure features.

[0017] In some embodiments, obtaining the sequence characteristics of the target drug includes the following steps:

[0018] Represent the substructure of target drugs through MACCS molecular fingerprints;

[0019] Based on the substructure of the target drug, the corresponding query matrix, key matrix and value matrix are obtained;

[0020] Based on the query matrix, key matrix and value matrix, the association between substructures is captured through the attention mechanism to obtain the sequence features of the target drug.

[0021] In some embodiments, dynamically gating and fusing graph structure features and sequence features to obtain drug features includes the following steps:

[0022] Perform dimension splicing on graph structure features and sequence features to obtain splicing features;

[0023] Based on the concatenated features, dynamic weights at the feature dimension level are generated through learnable nonlinear transformations.

[0024] Perform linear transformation on the splicing features to obtain the target dimension features;

[0025] According to the target dimension features, feature dimension-level dynamic weights and graph structure features, residual feature reorganization based on gated weights is performed to obtain drug features.

[0026] In some embodiments, dynamic feature enhancement is performed on a digitized sequence of a target protein to obtain protein features through context-aware optimization, comprising the following steps:

[0027] Perform sequence digitization mapping on the target protein based on a pre-built amino acid vocabulary to obtain a digitized sequence as the initial representation;

[0028] An interval-constrained random initialization strategy is used to generate learnable masks;

[0029] Use learnable masks to perform dimension-level feature selection on the initial representation to obtain encoded features;

[0030] Based on the encoding features, the sequence representation of the target protein is obtained through collaborative training through the attention mechanism;

[0031] Multi-layer convolution operations are performed on the sequence representation to obtain protein features.

[0032] In some embodiments, the feature representation includes multiple pairs of feature representation combinations, each pair of feature representation combinations includes a pre-collected drug feature and a protein feature corresponding to a drug and a protein with known actual affinity scores; the method further includes the following steps:

[0033] The features in the feature representation combination are concatenated and input into the neural network for processing to obtain the predicted affinity score;

[0034] Based on the predicted affinity score and the actual affinity score corresponding to each pair of feature representation combinations, the mean square error is used to construct the loss function, and then the network parameters of the neural network are feedback-adjusted through the error result of the loss function to obtain the prediction model.

[0035] In some embodiments, based on drug characteristics and protein characteristics, a prediction model is used to predict a target affinity score between a target drug and a target protein, comprising the following steps:

[0036] Perform feature splicing on drug features and protein features to obtain a splicing vector;

[0037] The splicing vector is input into the prediction model to predict the target affinity score between the target drug and the target protein.

[0038] To achieve the above-mentioned purpose, another aspect of the present invention provides a drug target prediction device, comprising:

[0039] The first module is used to obtain the graph structure features and sequence features of the target drug, and dynamically gate the graph structure features and sequence features to obtain drug features;

[0040] The second module is used to dynamically enhance the digital sequence of the target protein and then obtain protein features through context-aware optimization;

[0041] The third module is used to predict the target affinity score between the target drug and the target protein using a prediction model based on drug characteristics and protein characteristics;

[0042] The prediction model is trained based on a neural network using feature representations labeled with actual affinity scores.

[0043] In some embodiments, the feature representation includes multiple pairs of feature representation combinations, each pair of feature representation combinations includes pre-collected drug features and protein features corresponding to drugs and proteins with known actual affinity scores; the apparatus further includes a fourth module specifically configured to perform the following operations:

[0044] The features in the feature representation combination are concatenated and input into the neural network for processing to obtain the predicted affinity score;

[0045] Based on the predicted affinity score and the actual affinity score corresponding to each pair of feature representation combinations, the mean square error is used to construct the loss function, and then the network parameters of the neural network are feedback-adjusted through the error result of the loss function to obtain the prediction model.

[0046] To achieve the above object, another aspect of an embodiment of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the above method when executing the computer program.

[0047] To achieve the above object, another aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above method is implemented.

[0048] The embodiments of the present invention include at least the following beneficial effects: The present invention provides a drug target prediction method, device, electronic device and storage medium, which obtains the graph structure features and sequence features of the target drug, and dynamically gates and fuses the graph structure features and sequence features to obtain drug features; dynamically enhances the digital sequence of the target protein, and then obtains protein features through context-aware optimization; based on the drug features and protein features, a prediction model is used to predict the target affinity score of the target drug and the target protein; wherein the prediction model is obtained by training a feature representation based on a neural network labeled with an actual affinity score. The present invention can break through the accuracy bottleneck of traditional single-modal methods by dynamically gating and fusing the graph structure features and sequence features of the drug. Furthermore, the present invention can achieve deep coupling of the feature space of drugs and proteins by extracting protein features through dynamic feature enhancement combined with context-aware optimization. The present invention can achieve high-precision prediction of drug targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a flow chart of a drug target prediction method provided by an embodiment of the present invention;

[0050] Figure 2 Schematic diagram of the process of obtaining drug characteristics according to an embodiment of the present invention;

[0051] Figure 3 Schematic diagram of the process of obtaining protein features according to an embodiment of the present invention;

[0052] Figure 4 Schematic diagram of the affinity prediction structure provided by an embodiment of the present invention;

[0053] Figure 5 Schematic diagram of the structure of the drug target prediction device provided by an embodiment of the present invention;

[0054] Figure 6 It is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to make the objects, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the appended claims.

[0056] It will be understood that the terms "first," "second," and the like used in the present invention may be used to describe various concepts in the present invention, but unless otherwise specified, these concepts are not limited by these terms. These terms are merely used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the terms "if" and "if" as used herein may be interpreted as "at the time of," "when," or "in response to a determination."

[0057] The terms "at least one", "plurality", "each", "any", etc. used in the present invention include at least one, two or more, multiple, two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0058] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those commonly understood by those skilled in the art to which the present invention pertains. The terms used in the present invention are for the purpose of describing the embodiments of the present invention only and are not intended to limit the present invention.

[0059] The drug target prediction method provided by the embodiment of the present invention relates to the technical field of drug target prediction. The drug target prediction method provided by the embodiment of the present invention can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the drug target prediction method, etc., but is not limited to the above forms.

[0060] The present invention can be used in a wide variety of general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0061] Figure 1 This is an optional flowchart of the drug target prediction method provided by an embodiment of the present invention. Figure 1 The method may include but is not limited to steps S100 to S300.

[0062] S100, obtaining graph structure features and sequence features of the target drug, and performing dynamic gating fusion on the graph structure features and sequence features to obtain drug features;

[0063] It should be noted that, in some embodiments, obtaining the graph structure features of the target drug may include the following steps: obtaining the SMILES of the target drug and converting the SMILES into a mol object; extracting the atomic parameter features of each atom from the mol object, using the atomic parameter features as node features, and constructing an edge set based on the chemical bond relationship between each atom; obtaining the molecular graph structure of the target drug based on the node features and the edge set; based on the molecular graph structure, using a graph convolutional network to extract the feature vector of each atomic node, and obtaining the attention score between each atomic node through a graph self-attention network; normalizing the attention score to obtain the attention weight; using the attention weight to update the feature vector, and then combining the updated feature vectors of all atomic nodes to obtain the graph structure features. Specifically, SMILES can be converted into a mol object using the RDKit tool.

[0064] For example, in some specific implementations, graph structure feature extraction can be implemented as follows:

[0065] The RDKit toolkit is used to convert the drug's SMILES (Simplified Molecular Input Lineentry System) data representation into a mol object, which is then converted into a molecular graph structure. Nodes are atoms, node features are represented by atomic properties, and edges are chemical bonds. Each atom's features are extracted, including atomic number, degree, formal charge, radical electron count, and aromaticity. Node features are defined as:

[0066] x i =[AtomicNum i ,Degree i ,FormalCharge i ,RadicalElectrons i ,IsAromatic i ]

[0067] Where AtomicNum i ,Degree i ,FormalCharge i ,RadicalElectrons i ,IsAromatic i Represents the i-th atom x i Atomic number, degree, formal charge, radical electron number and aromaticity.

[0068] Edges are chemical bonds. If atoms i and j form a bond, they are recorded in the edge set E. The edge set is constructed by the chemical bond relationship between atoms:

[0069] E={(x i ,x j )∣bond exists between atom i and atom j}

[0070] GCN extracts the influence of adjacent nodes and updates the node features. The following are the changes in the features of atoms in each layer:

[0071]

[0072] Among them, H (l+1) represents the feature matrix of the (l+1) layer nodes, l represents the number of convolutional network layers, N is the set of neighbor nodes of node i, is the sum of the adjacency matrix and the identity matrix, yes The degree matrix of is a diagonal matrix, and its diagonal elements are is the weight matrix of layer l, and σ is the activation function.

[0073] Capturing global structural dependencies through graph attention networks:

[0074] e ij =LeakyReLU(a T [Wh i / / Wh j ])

[0075] where e ij represents the unnormalized attention score between node i and node j, a is the learnable attention parameter vector, T is the matrix transpose, W is the learnable weight matrix, and h i 、h j is the feature representation of node i and node j after passing through the GCN layer, / / It is a vector splicing operation.

[0076] Normalize the attention coefficient by Softmax:

[0077]

[0078] α ij Represents the normalized attention weight, which is the exp exponential function, and N(i) represents all neighbor nodes of i (including itself).

[0079] Update the node features again using the attention weights:

[0080]

[0081] h i ' represents the updated feature vector of node i, N(i) represents all neighbor nodes of i (including itself), and the updated drug structure feature representation D1 (i.e., graph structure feature) is obtained.

[0082] It should be noted that, in some embodiments, obtaining the sequence characteristics of the target drug may include the following steps: representing the substructure of the target drug through the MACCS molecular fingerprint; obtaining the corresponding query matrix, bond matrix and value matrix based on the substructure of the target drug; based on the query matrix, bond matrix and value matrix, capturing the association between the substructures through the attention mechanism to obtain the sequence characteristics of the target drug.

[0083] For example, in some specific implementations, sequence feature extraction can be implemented as follows:

[0084] In addition, MACCS molecular fingerprints (166-bit binary vectors) are used to represent the sequence characteristics of molecules, where each bit represents the presence or absence of a specific substructure in the molecule:

[0085]

[0086] MS(i) represents the composition of the MACCS molecular fingerprint, and the Transformer encoder is used to capture the association between substructures:

[0087]

[0088] where Q D ,K D ,V D are the query matrix, key matrix and value matrix of the drug sequence respectively, d k is the dimension of the key.

[0089] Output sequence feature vector

[0090] It should be noted that, in some embodiments, the dynamic gating fusion of graph structure features and sequence features to obtain drug features may include the following steps: dimensionally splicing the graph structure features and sequence features to obtain spliced features; based on the spliced features, generating feature dimension-level dynamic weights through learnable nonlinear transformations; performing linear transformations on the spliced features to obtain target dimension features; and performing residual feature reorganization based on gating weights according to the target dimension features, feature dimension-level dynamic weights, and graph structure features to obtain drug features.

[0091] For example, in some specific implementations, the dynamic gating fusion mechanism can be implemented as follows:

[0092] Traditional feature fusion methods suffer from two major drawbacks: static fusion limitations: fixed weight assignment cannot adapt to different molecular properties. For example, for cyclic compounds, structural features must be emphasized, while for long-chain molecules, sequence features must be enhanced. Information interference risk: simple concatenation can lead to conflicts in feature space (for example, the linear combination of 166-dimensional graph features and 166-dimensional sequence features can cause orthogonal interference). To address these two issues, a dynamic gated fusion mechanism is implemented to achieve context-aware adaptive weight assignment and automatically learn the optimal fusion strategy through a differentiable function. Finally, the Hadamard product is used to achieve refined feature reconstruction, as described in detail below.

[0093] The feature D1 describing the drug graph structure and the feature D2 describing the drug sequence are concatenated in one dimension, and the residual connection retains the original graph structure features to avoid gradient disappearance or explosion.

[0094] V=[D1 / / D2]

[0095] Generate feature dimension-level dynamic weights through learnable nonlinear transformations:

[0096]

[0097] Wg Represents the weight matrix corresponding to the gate layer, b g Indicates bias.

[0098] Perform a linear transformation on the concatenated features to fit the target dimension:

[0099]

[0100] W t represents the weight matrix corresponding to the linear transpose, b t represents the transposed bias.

[0101] Dynamically generate weight coefficients, adjust the modal contribution ratio according to molecular characteristics, and reorganize residual features based on gated weights:

[0102] D=gate⊙T+(1-gate)⊙W g D1

[0103] Where ⊙ represents the Hadamard product (element-wise multiplication). The Hadamard product is used to achieve independent adjustment of feature dimensions, suppress orthogonal conflicting dimensions, and strengthen complementary dimensions to obtain the drug feature representation D (i.e., drug features).

[0104] S200, dynamically enhancing the digital sequence of the target protein, and then obtaining protein features through context-aware optimization;

[0105] It should be noted that, in some embodiments, step S200 may include the following steps: performing sequence digital mapping processing on the target protein based on a pre-constructed amino acid vocabulary to obtain a digital sequence as an initial representation; generating a learnable mask using an interval-constrained random initialization strategy; performing dimension-level feature selection on the initial representation using the learnable mask to obtain encoding features; based on the encoding features, performing collaborative training through an attention mechanism to obtain a sequence representation of the target protein; and performing multi-layer convolution operations on the sequence representation to obtain protein features.

[0106] For example, in some embodiments, protein feature extraction can be achieved as follows:

[0107] Sequence digital processing:

[0108] Construct an amino acid vocabulary to implement sequence mapping (e.g., A→1, C→2, E→4), generating integer sequences of length 1000. Obtain the initial representation through a learnable embedding layer:

[0109]

[0110] Adaptive Enhancement Coding:

[0111] Traditional protein encoders have two major bottlenecks: fixed-pattern feature extraction cannot adapt to the action modes of different drug molecules; irrelevant amino acid residues (such as flexible loop regions) interfere with key sites (such as active centers).

[0112] A dynamic feature enhancement module is proposed, which adopts an interval-constrained random initialization strategy to learn mask initialization:

[0113] M init =0.5·rand(1,1000,128)+0.5,(M init ∈[0.5,1.0])

[0114] Where rand() represents a random floating point number between 0 and 1.

[0115] Dynamic feature enhancement:

[0116] Perform dimension-wise feature selection on the original embeddings:

[0117] p′=p⊙M

[0118] Where M is a learnable mask. When M[i,j]→1.0 strengthens the features from the i-th position to the j-th dimension, M[i,j]→0.5 weakens the non-critical features.

[0119] Context-aware optimization:

[0120] Co-training of mask matrix and attention mechanism:

[0121] Q P =p′W Q ,K P =p′W k ,V P =p′W V

[0122] where Q p ,K p ,V p are the query matrix, protein sequence key matrix, and protein sequence value matrix obtained after the protein sequence is masked. Q , W K , W V , weight matrices corresponding to query, key and value matrices.

[0123] The attention weight calculation is based on the enhanced features, forming a two-stage optimization from the mask screening key feature dimension stage to the attention focusing key sequence positions.

[0124]

[0125] Get the sequence representation P with added attention focus T, retain the most significant feature P in each dimension through global maximum pooling T ′:

[0126]

[0127] L is the length of the sequence, i is the index of the sequence length, and j is the index of the feature dimension. In order to further extract the local spatial features of adjacent amino acids in the protein sequence, the updated P T 'A one-dimensional convolution operation is used (taking three layers of convolution as an example, the specific number of convolution layers can be adjusted according to actual needs, and the specific application of the embodiment of the present invention is only for illustration):

[0128] C1=ReLU(Conv1D(P T ′,W1))

[0129] C2=ReLU(Conv1D(C1,W2))

[0130] P = ReLU(Conv1D(C2,W3))

[0131] C1, C2, P represent the protein feature vectors after the first, second and third convolution operations, respectively, and W1, W2, W3 represent the weight matrices of the first, second and third convolution operations.

[0132] Each convolution uses a different weight matrix and convolution kernel size to capture local features at different scales. Finally, a maximum pooling operation is used to obtain the maximum value of each convolution channel to obtain the final feature vector representation P (i.e., protein feature) of the protein.

[0133] S300: Based on the drug characteristics and protein characteristics, a prediction model is used to predict the target affinity score between the target drug and the target protein.

[0134] Among them, the prediction model is trained based on a neural network through feature representations labeled with actual affinity scores;

[0135] It should be noted that the feature representation includes multiple pairs of feature representation combinations, each pair of feature representation combinations includes pre-collected drug features and protein features corresponding to drugs and proteins with known actual affinity scores; in some embodiments, the method may further include the following steps: concatenating the features in the feature representation combination and inputting them into a neural network for processing to obtain a predicted affinity score; based on the predicted affinity score and the actual affinity score corresponding to each pair of feature representation combinations, a loss function is constructed using the mean square error, and the network parameters of the neural network are then adjusted through feedback using the error result of the loss function to obtain a prediction model. In some optional implementations, the neural network used in the prediction model may adopt, but is not limited to, network structures such as a fully connected network, a multi-layer fully connected network, and a residual network.

[0136] For example, in some specific implementations, taking a neural network structure formed by three fully connected layers plus an activation function as an example (the specific number of fully connected layers can be adjusted according to actual needs, and the specific application of the embodiment of the present invention is only for illustration), affinity prediction training can be implemented through the following process:

[0137] The drug feature representation D and the protein feature representation P are concatenated into a vector v to describe the characteristics of the drug-protein interaction.

[0138] v=[D / / P]

[0139] The concatenated features v are used as the input of the fully connected network to extract deep features layer by layer and predict drug target affinity. The fully connected layer is calculated, and each layer performs the following operations:

[0140] v (l) =σ(W (l) v (l-1) +b (l) )

[0141] v (l) W: Output feature of the lth layer. (l) : The weight matrix of the first layer. b (l) : bias of the lth layer. σ: activation function LeakyReLU.

[0142] Final output layer:

[0143] y=Linear(v (3) )=W (4) v (3) +b (4)

[0144] W (4) is the weight matrix of the fourth layer, b (4) is the bias score of the fourth layer.

[0145] The affinity score y between the drug and the protein is obtained through a four-layer network.

[0146] The loss function uses the mean square error (MSE):

[0147]

[0148] where y i is the predicted affinity score, is the actual affinity score, and N represents the total number of predictions. The fully connected network is then adjusted based on the error results. The training process can be adjusted iteratively for convergence, and specific training parameters (such as learning rate and number of iterations) can be adaptively adjusted based on actual training requirements.

[0149] It should be noted that in some embodiments, based on drug features and protein features, using a prediction model to predict the target affinity score between the target drug and the target protein can include the following steps: feature splicing of the drug features and the protein features to obtain a splicing vector; inputting the splicing vector into the prediction model to predict and output the target affinity score between the target drug and the target protein.

[0150] For example, in some specific embodiments, the logical process of predicting the target affinity score between the target drug and the target protein based on the pre-trained prediction model is the same as the logical process of predicting the affinity score in the aforementioned embodiment, and will not be repeated here.

[0151] In order to explain the principle of the technical solution of the present invention in detail, the overall process of the present invention is described below in combination with some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and cannot be regarded as a limitation of the present invention.

[0152] First, it's important to note that current drug-target interaction prediction methods suffer from the following key flaws: Rigid multimodal feature fusion: Existing methods often employ simple concatenation or weighted averaging of sequence (SMILES) and graph structure (molecular graph) features, lacking a dynamic feature selection mechanism. This results in inefficient information fusion and an inability to adaptively distinguish the importance of key chemical bonds and sequence patterns. Protein representation lacks adaptive regulation: Traditional protein encoders employ fixed-mode feature extraction, failing to dynamically adjust feature focus areas based on specific drug-target pairings, limiting the ability to learn context-sensitive features.

[0153] In view of this, the present invention specifically proposes a multimodal fusion prediction framework by constructing a dual characterization system for drug molecules: on the one hand, analyzing its sequence structure information (semantic features of SMILES sequences), and on the other hand, extracting its graph structure information (atomic-level topological connection relationships), and innovatively introducing a dynamic gating fusion mechanism and a protein coding adaptive enhancement module to achieve deep coupling of feature space. This technical solution breaks through the accuracy bottleneck of traditional single-modal methods and provides reliable engineering support for the structure-based prediction and design of drug-target interaction affinity. Specifically, the method flow of the present invention can be implemented as follows:

[0154] Step 1. Drug feature extraction:

[0155] The drug feature extraction system of the present invention introduces a breakthrough dynamic gating fusion mechanism through collaborative learning of graph structure modeling and sequence feature analysis to achieve intelligent fusion of cross-modal features. Figure 2 As shown, the specific process can be implemented as follows:

[0156] Graph structure feature extraction:

[0157] The RDKit toolkit is used to convert the drug's SMILES (Simplified Molecular Input Lineentry System) data representation into a mol object, which is then converted into a molecular graph structure. Nodes are atoms, node features are represented by atomic properties, and edges are chemical bonds. Each atom's features are extracted, including atomic number, degree, formal charge, radical electron count, and aromaticity. Node features are defined as:

[0158] x i =[AtomicNum i ,Degree i ,FormalCharge i ,RadicalElectrons i ,IsAromatic i ]

[0159] Where AtomicNum i ,Degree i ,FormalCharge i ,RadicalElectrons i ,IsAromatic i Represents the i-th atom x i Atomic number, degree, formal charge, radical electron number and aromaticity.

[0160] Edges are chemical bonds. If atoms i and j form a bond, they are recorded in the edge set E. The edge set is constructed by the chemical bond relationship between atoms:

[0161] E={(x i ,x j )∣bond exists between atom i and atom j}

[0162] GCN extracts the influence of adjacent nodes and updates the node features. The following are the changes in the features of atoms in each layer:

[0163]

[0164]

[0165] Among them, H (l+1) represents the feature matrix of the (l+1) layer nodes, l represents the number of convolutional network layers, N is the set of neighbor nodes of node i, is the sum of the adjacency matrix and the identity matrix, yes The degree matrix of is a diagonal matrix, and its diagonal elements are is the weight matrix of layer l, and σ is the activation function.

[0166] Capturing global structural dependencies through graph attention networks:

[0167] e ij =LeakyReLU(a T [Wh i / / Wh j ])

[0168] where e ij represents the unnormalized attention score between node i and node j, a is the learnable attention parameter vector, T is the matrix transpose, W is the learnable weight matrix, and h i 、h j is the feature representation of node i and node j after passing through the GCN layer, / / It is a vector splicing operation.

[0169] Normalize the attention coefficient by Softmax:

[0170]

[0171] α ij Represents the normalized attention weight, which is the exp exponential function, and N(i) represents all neighbor nodes of i (including itself).

[0172] Update the node features again using the attention weights:

[0173]

[0174] h i ' represents the updated feature vector of node i, N(i) represents all neighbor nodes of i (including itself), and the updated drug structure feature representation D1 (i.e., graph structure feature) is obtained.

[0175] Sequence feature extraction:

[0176] In addition, MACCS molecular fingerprints (166-bit binary vectors) are used to represent the sequence characteristics of molecules, where each bit represents the presence or absence of a specific substructure in the molecule:

[0177]

[0178] MS(i) represents the composition of the MACCS molecular fingerprint, and the Transformer encoder is used to capture the association between substructures:

[0179]

[0180] where Q D ,K D ,V D are the query matrix, key matrix and value matrix of the drug sequence respectively, d k is the dimension of the key.

[0181] Output sequence feature vector

[0182] Dynamic gated fusion mechanism:

[0183] Traditional feature fusion methods have two major drawbacks: static fusion limitations: fixed weight allocation cannot adapt to different molecular characteristics, such as the need to strengthen structural features for cyclic compounds and sequence features for long-chain molecules; and information interference risk: simple splicing may lead to feature space conflicts (such as orthogonal interference between the linear combination of 166-dimensional graph features and 166-dimensional sequence features). To address these two issues, this paper implements context-aware adaptive weight allocation through a dynamic gated fusion mechanism and automatically learns the optimal fusion strategy through differentiable functions. Ultimately, the Hadamard product is used to achieve refined feature reconstruction, as described in detail below.

[0184] The feature D1 describing the drug graph structure and the feature D2 describing the drug sequence are concatenated in one dimension, and the residual connection retains the original graph structure features to avoid gradient disappearance or explosion.

[0185] V=[D1 / / D2]

[0186] Generate feature dimension-level dynamic weights through learnable nonlinear transformations:

[0187]

[0188] W g Represents the weight matrix corresponding to the gate layer, b g Indicates bias.

[0189] Perform a linear transformation on the concatenated features to fit the target dimension:

[0190]

[0191] W t represents the weight matrix corresponding to the linear transpose, b t represents the transposed bias.

[0192] Dynamically generate weight coefficients, adjust the modal contribution ratio according to molecular characteristics, and reorganize residual features based on gated weights:

[0193] D=gate⊙T+(1-gate)⊙W g D1

[0194] Where ⊙ represents the Hadamard product (element-wise multiplication). The Hadamard product is used to achieve independent adjustment of feature dimensions, suppress orthogonal conflicting dimensions, and strengthen complementary dimensions to obtain the drug feature representation D (i.e., drug features).

[0195] Step 2. Protein feature extraction:

[0196] like Figure 3 As shown in Figure 2, protein feature extraction includes the following core steps:

[0197] Sequence digital processing:

[0198] Construct an amino acid vocabulary to implement sequence mapping (e.g., A→1, C→2, E→4), generating integer sequences of length 1000. Obtain the initial representation through a learnable embedding layer:

[0199]

[0200] Adaptive Enhancement Coding:

[0201] Traditional protein encoders have two major bottlenecks: fixed-pattern feature extraction cannot adapt to the action modes of different drug molecules; irrelevant amino acid residues (such as flexible loop regions) interfere with key sites (such as active centers).

[0202] A dynamic feature enhancement module is proposed, which adopts an interval-constrained random initialization strategy to learn mask initialization:

[0203] M init =0.5·rand(1,1000,128)+0.5,(M init ∈[0.5,1.0])

[0204] Where rand() represents a random floating point number between 0 and 1.

[0205] Dynamic feature enhancement:

[0206] Perform dimension-wise feature selection on the original embeddings:

[0207] p′=p⊙M

[0208] Where M is a learnable mask. When M[i,j]→1.0 strengthens the features from the i-th position to the j-th dimension, M[i,j]→0.5 weakens the non-critical features.

[0209] Context-aware optimization:

[0210] Co-training of mask matrix and attention mechanism:

[0211] Q P =p′W Q ,K P=p′W K ,V P =p′W V

[0212] where Q p ,K p ,V p are the query matrix, protein sequence key matrix, and protein sequence value matrix obtained after the protein sequence is masked. Q , W K , W v , weight matrices corresponding to query, key and value matrices.

[0213] The attention weight calculation is based on the enhanced features, forming a two-stage optimization from the mask screening key feature dimension stage to the attention focusing key sequence positions.

[0214]

[0215] Get the sequence representation P with added attention focus T , retain the most significant feature P in each dimension through global maximum pooling T ′:

[0216]

[0217] L is the length of the sequence, i is the index of the sequence length, and j is the index of the feature dimension. In order to further extract the local spatial features of adjacent amino acids in the protein sequence, the updated P T 'A one-dimensional convolution operation is used (taking three layers of convolution as an example, the specific number of convolution layers can be adjusted according to actual needs, and the specific application of the embodiment of the present invention is only for illustration):

[0218] C1=ReLU(Conv1D(P T ′,W1))

[0219] C2=ReLU(Conv1D(C1,W2))

[0220] P = ReLU(Conv1D(C2,W3))

[0221] C1, C2, P represent the protein feature vectors after the first, second and third convolution operations, respectively, and W1, W2, W3 represent the weight matrices of the first, second and third convolution operations.

[0222] Each convolution uses a different weight matrix and convolution kernel size to capture local features at different scales. Finally, a maximum pooling operation is used to obtain the maximum value of each convolution channel to obtain the final feature vector representation P (i.e., protein feature) of the protein.

[0223] Step 3. Affinity prediction:

[0224] like Figure 4 As shown, the drug feature representation D and the protein feature representation P are concatenated into a vector v, which is used to describe the characteristics of the drug-protein interaction.

[0225] v=[D / / P]

[0226] The concatenated features v are used as the input of the fully connected network to extract deep features layer by layer and predict drug target affinity. The fully connected layer is calculated, and each layer performs the following operations:

[0227] v (l) =σ(W (l) v (l-1) +b (l) )

[0228] v (l) W: Output feature of the lth layer. (l) : The weight matrix of the first layer. b (l) : bias of the lth layer. σ: activation function LeakyReLU.

[0229] Final output layer:

[0230] y=Linear(v (3) )=W (4) v (3) +b (4)

[0231] W (4) is the weight matrix of the fourth layer, b (4) is the bias score of the fourth layer.

[0232] The affinity score y between the drug and the protein is obtained through a four-layer network.

[0233] The loss function uses the mean square error (MSE):

[0234]

[0235] where y i is the predicted affinity score, is the actual affinity score, and N represents the total number of predictions. The fully connected network is then adjusted based on the error results. The training process can be adjusted iteratively for convergence, and specific training parameters (such as learning rate and number of iterations) can be adaptively adjusted based on actual training requirements.

[0236] In some specific application scenarios, in order to verify the effectiveness and advantages of the drug-target interaction affinity prediction method proposed in the present invention, the present invention conducted an ablation experiment, focusing on comparing the performance of two input methods: drug sequence and graph structure sequence. The experiment compared and analyzed from multiple dimensions such as accuracy, stability and generalization ability. The results showed that the method of the present invention is superior to other methods in all indicators. It can better capture key information in terms of accuracy, can adapt to different environments and data sets in terms of stability, and has better adaptability in terms of generalization ability. In order to further verify its competitiveness, the present invention conducted an in-depth comparison with a variety of the latest methods. The experimental results show that the method of the present invention performs well in prediction accuracy, and can more accurately predict drug-target interaction affinity in multiple data set application scenarios, providing more reliable technical support for drug research and development.

[0237] In summary, the present invention aims to overcome the above limitations through the following innovations: (1) Designing a dynamic gated fusion mechanism to automatically generate modal weight coefficients through learnable nonlinear transformations, thereby achieving dynamic adaptive fusion of drug sequence features and graph structure features. (2) Proposing a protein coding adaptive enhancement module to reshape the feature space of the original embedding through a trainable random weight mask, enabling the model to autonomously adjust the feature focus dimension according to different drug-target pairing scenarios.

[0238] Specifically, the purpose of the present invention is to address the limitations of existing methods in drug and protein feature extraction. In terms of drug feature representation, the molecular information of the drug is converted through graph structure and sequence structure respectively, and then the two are fused through a dynamic feature selection mechanism to comprehensively capture the multidimensional characteristics of the drug and prevent the information loss caused by single modality representation in traditional methods. In terms of protein feature representation, the protein coding adaptive enhancement module and sequence and spatial structure are combined to extract key features, so as to more comprehensively predict the interaction affinity between protein and drug. Finally, by fusing the multimodal features of drugs and proteins, the present invention constructs an efficient prediction model to accurately predict drug-target interaction affinity.

[0239] Compared with the prior art, the present invention has at least the following beneficial effects:

[0240] At the drug characterization level, the dynamic gating fusion mechanism transcends the limitations of traditional static fusion: based on the chemical properties of drug molecules, this mechanism automatically adjusts the fusion weights of graph structural features and sequence features—focusing on strengthening the graph feature representation of interatomic bonding relationships and spatial conformations, while for linear molecules, it emphasizes the sequence feature analysis of substructure combination patterns. This achieves fine-grained fusion at the feature dimension level, enabling the model to adapt to the multimodal information complementarity of different drug types. At the protein representation level, a trainable mask matrix and attention mechanism form a two-level dynamic optimization: the mask layer first reshapes the feature space of the original embedding, weakening the interference information of surface disordered regions; the multi-head attention layer then further focuses on the spatial dependencies of key amino acid sites based on the masked features. This synergistic "feature enhancement-spatial focusing" mechanism significantly improves the context-awareness of target feature extraction. The two innovative technologies form an organic synergy through gradient back propagation - the gating weight distribution on the drug side guides the protein mask to focus on the ligand binding sensitive area, while the dynamic feature enhancement on the target side reversely optimizes the drug modality fusion strategy, and finally constructs a bidirectional adaptive interaction analysis framework, providing a more robust feature representation basis for the prediction of interactions between complex structure drugs and new targets.

[0241] On the other hand, Figure 5 As shown, the embodiment of the present invention further provides a drug target prediction device 900, which may include:

[0242] The first module 901 is used to obtain the graph structure features and sequence features of the target drug, and perform dynamic gating fusion on the graph structure features and sequence features to obtain drug features;

[0243] The second module 902 is used to perform dynamic feature enhancement on the digitized sequence of the target protein, and then obtain protein features through context-aware optimization;

[0244] The third module 903 is used to predict the target affinity score between the target drug and the target protein using a prediction model based on the drug characteristics and the protein characteristics;

[0245] The prediction model is trained based on a neural network using feature representations labeled with actual affinity scores.

[0246] In some embodiments, the feature representation includes multiple pairs of feature representation combinations, each pair of feature representation combinations includes pre-collected drug features and protein features corresponding to drugs and proteins with known actual affinity scores; the apparatus may further include a fourth module specifically configured to perform the following operations:

[0247] The features in the feature representation combination are concatenated and input into the neural network for processing to obtain the predicted affinity score;

[0248] Among them, the fully connected network includes multiple layers of fully connected layers and activation functions connected in sequence;

[0249] Based on the predicted affinity score and the actual affinity score corresponding to each pair of feature representation combinations, the mean square error is used to construct the loss function, and then the network parameters of the neural network are feedback-adjusted through the error result of the loss function to obtain the prediction model.

[0250] The contents of the method embodiments of the present invention are all applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0251] An embodiment of the present invention further provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-described drug target prediction method when executing the computer program. The electronic device can be any intelligent terminal, including a tablet computer and an in-vehicle computer.

[0252] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0253] See also Figure 6 , Figure 6 The hardware structure of an electronic device 1000 according to another embodiment is shown. The electronic device 1000 includes:

[0254] The processor 1001 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention.

[0255] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called by the processor 1001 to execute the drug target prediction method of the embodiment of the present invention.

[0256] Input / output interface 1003, used to implement information input and output;

[0257] Communication interface 1004, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0258] Bus 1005 , which transmits information between various components of the device (e.g., processor 1001 , memory 1002 , input / output interface 1003 , and communication interface 1004 );

[0259] The processor 1001 , the memory 1002 , the input / output interface 1003 and the communication interface 1004 are connected to each other in communication within the device via a bus 1005 .

[0260] An embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned drug target prediction method is implemented.

[0261] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0262] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0263] The drug target prediction method, drug target prediction device, electronic device and storage medium provided by the embodiments of the present invention obtain the graph structure features and sequence features of the target drug, perform dynamic gating fusion on the graph structure features and sequence features to obtain drug features; perform dynamic feature enhancement on the digitized sequence of the target protein, and then obtain protein features through context-aware optimization; based on the drug features and protein features, use a prediction model to predict the target affinity score of the target drug and the target protein; wherein the prediction model is obtained by training a feature representation based on a neural network labeled with the actual affinity score. The present invention can break through the accuracy bottleneck of traditional single-modal methods by dynamically gating the graph structure features and sequence features of the drug. Furthermore, the present invention can achieve deep coupling of the feature space of drugs and proteins by extracting protein features through dynamic feature enhancement combined with context-aware optimization. The present invention can achieve high-precision prediction of drug targets.

[0264] The embodiments described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.

[0265] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0266] The system embodiment described above is merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0267] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0268] The terms "first," "second," "third," "fourth," and the like (if any) in the description of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in orders other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products, or apparatus.

[0269] It should be understood that in the present invention, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can represent: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0270] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, and can be electrical, mechanical or other forms.

[0271] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0272] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0273] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store programs.

[0274] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the invention is not limited thereby. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the invention should be within the scope of the invention.

Claims

1. A drug target prediction method, characterized in that: The method comprises the following steps: Obtaining graph structural features and sequence features of the target drug, and performing dynamic gating fusion on the graph structural features and the sequence features to obtain drug features; Dynamic feature enhancement is performed on the digital sequence of the target protein, and protein features are obtained through context-aware optimization; Based on the drug characteristics and the protein characteristics, using a prediction model to predict a target affinity score between the target drug and the target protein; The prediction model is obtained by training a neural network through feature representations marked with actual affinity scores.

2. The drug target prediction method according to claim 1, wherein The step of obtaining the graph structure features of the target drug includes the following steps: Obtaining the SMILES of the target drug, and converting the SMILES into a mol object; Extracting atomic parameter features of each atom from the mol object, using the atomic parameter features as node features, and constructing an edge set based on the chemical bond relationship between each atom; Obtaining a molecular graph structure of the target drug according to the node features and the edge set; Based on the molecular graph structure, a graph convolutional network is used to extract the feature vector of each atomic node, and the attention score between each atomic node is obtained through a graph self-attention network; Normalizing the attention score to obtain an attention weight; The feature vector is updated using the attention weight, and the updated feature vectors of all atomic nodes are combined to obtain the graph structure feature.

3. The drug target prediction method according to claim 1, wherein The method of obtaining the sequence characteristics of the target drug comprises the following steps: The substructure of the target drug is represented by MACCS molecular fingerprint; Obtaining a corresponding query matrix, a key matrix, and a value matrix based on the substructure of the target drug; Based on the query matrix, the key matrix and the value matrix, the association between the substructures is captured through an attention mechanism to obtain the sequence features of the target drug.

4. The drug target prediction method according to claim 1, wherein The method of performing dynamic gating fusion on the graph structure feature and the sequence feature to obtain drug features comprises the following steps: Performing dimension splicing on the graph structure feature and the sequence feature to obtain a splicing feature; Based on the concatenated features, generating feature dimension-level dynamic weights through a learnable nonlinear transformation; Performing a linear transformation on the splicing features to obtain target dimension features; The drug features are obtained by performing residual feature reorganization based on gated weights according to the target dimension features, the feature dimension-level dynamic weights and the graph structure features.

5. The drug target prediction method according to claim 1, wherein The method of dynamically enhancing the digitized sequence of the target protein and then obtaining protein features through context-aware optimization includes the following steps: Performing sequence digitization mapping processing on the target protein based on a pre-constructed amino acid vocabulary to obtain the digitized sequence as an initial representation; An interval-constrained random initialization strategy is used to generate learnable masks; Performing dimension-level feature selection on the initial representation using the learnable mask to obtain encoding features; Based on the encoding features, collaborative training is performed through an attention mechanism to obtain a sequence representation of the target protein; A multi-layer convolution operation is performed on the sequence representation to obtain the protein feature.

6. The drug target prediction method according to claim 1, wherein The feature representation includes multiple pairs of feature representation combinations, each pair of the feature representation combination includes pre-collected drug features and protein features corresponding to the drug and protein with known actual affinity scores; the method further includes the following steps: Concatenating the features in the feature representation combination and inputting them into the neural network for processing to obtain a predicted affinity score; Based on the predicted affinity score and the actual affinity score corresponding to each pair of the feature representation combination, a loss function is constructed using the mean square error, and then the network parameters of the neural network are feedback-adjusted through the error result of the loss function to obtain the prediction model.

7. The drug target prediction method according to claim 1, wherein The method of predicting the target affinity score between the target drug and the target protein using a prediction model based on the drug characteristics and the protein characteristics comprises the following steps: Performing feature splicing on the drug feature and the protein feature to obtain a splicing vector; The splicing vector is input into the prediction model to predict and output the target affinity score between the target drug and the target protein.

8. A drug target prediction device, characterized in that: The device comprises: The first module is used to obtain the graph structure features and sequence features of the target drug, and dynamically gate the graph structure features and the sequence features to obtain drug features; The second module is used to dynamically enhance the digital sequence of the target protein and then obtain protein features through context-aware optimization; The third module is configured to predict the target affinity score between the target drug and the target protein using a prediction model based on the drug characteristics and the protein characteristics; The prediction model is obtained by training a neural network through feature representations marked with actual affinity scores.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Drug target intelligent prediction method and system

    CN120954490A

  • Molecular optimization method, system and device based on neural network and storage medium

    CN121237260A

  • Drug target prediction method based on cross-modal attention and uncertainty evaluation

    CN122050487A

  • Drug target affinity prediction method fusing PPI quality and uncertainty

    CN122290687A