A method for predicting drug-target interactions based on Freud algorithm network
Patent Information
- Application Number
- CN202510858302.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-25
AI Technical Summary
现有技术在药物分子与靶点蛋白结合识别中存在实验成本高昂、周期漫长的问题,且难以有效融合药物与蛋白质的复杂交互特征,模型特征粒度不足,尤其在融合蛋白残基层级信息与分子图结构特征方面存在缺失。
The Freud algorithm network is adopted, combining molecular graph structure modeling, sequence embedding coding, Freud algorithm neural network, the nonlinear transformation mechanism of the Kolmogolov-Arnold network and the multi-head self-attention mechanism, to capture the complex multi-scale interaction characteristics between the drug and the target protein, and improve the prediction accuracy through the dual-channel feature coding strategy and the cross-modal attention fusion module.
显著提升了药物筛选速度与靶点识别能力,提高了新药研发的效率和准确性,模型在药物-靶点亲和度预测方面表现优越。
Smart Images

Figure CN120375913B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics and relates to a method for predicting drug-target interactions based on a Freud algorithm network, which includes technologies such as the Freud algorithm, deep learning, the Kolmogorov-Arnold network, and the attention mechanism. Background Art
[0002] In the process of new drug research and development, the binding recognition between drug molecules and target proteins is a key step. However, the experimental cost of this process is high and the cycle is long. Therefore, the development of efficient and accurate computational prediction methods has become an important direction of current research. In existing technologies, deep learning methods are usually used to convert drug and protein codes into vectors and then calculate their affinity. However, the existing methods have the following shortcomings: the complex interaction characteristics of drugs and proteins are not effectively integrated, and it is difficult to characterize the deep binding rules; the model feature granularity is insufficient, especially in terms of the residue level information of the fusion protein and the molecular graph structure characteristics; the feature interaction process mainly relies on simple vector splicing or linear models, and the expression ability is limited. Therefore, how to improve the model's learning ability of complex drug-protein interactions by introducing a new deep feature modeling structure has become a technical problem that needs to be solved urgently. Summary of the Invention
[0003] This paper proposes a method for predicting drug-target interactions based on a Freud algorithm network, aiming to improve the accuracy of drug-target affinity predictions, optimize the efficiency of new target discovery, and assist in drug combination optimization. This method integrates molecular graph structure modeling, sequence embedding coding, Freud algorithm neural networks, the nonlinear transformation mechanism of the Kolmogorov-Arnold network, and a multi-head self-attention mechanism to effectively capture the complex multi-scale interaction characteristics between drugs and target proteins. This method can significantly improve drug screening speed and target identification capabilities in new drug development. The specific technical process is as follows:
[0004] Step 1: First, the drug molecule is structurally modeled. Based on the simplified molecular linear input system sequence, it is parsed and converted into a two-dimensional molecular graph structure through the chemical information library. For the target protein data, a residue-level graph structure model that combines the protein's three-dimensional structure and sequence information is used for encoding. The edge features include the residue connection relationship and its spatial distance characteristics.
[0005] Step 2: Based on the drug molecule and protein standardized data preprocessed in step 1, a dual-channel feature encoding strategy is used to extract their structural and sequential information respectively, and finally a mathematical feature expression of drug topological fingerprint and protein sequence embedding is formed as the input for subsequent neural network fusion modeling.
[0006] Step 3: The molecular fingerprint tensor and protein embedding tensor are separately input into the Floyd feature fusion neural network architecture. This parallel architecture consists of two branch channels, each independently processing and fusion mapping the drug and protein features. Each branch integrates multiple Floyd algorithm neural network units. Within each unit, a shortest-distance path search principle and a recursive feature iteration mechanism are used to achieve multi-scale dynamic feature extraction and optimization. Furthermore, to enhance the model's ability to fit high-order, complex features nonlinearly, a Kolmogorov-Arnold network module is integrated within each Floyd network unit.
[0007] Step 4: Input the three-dimensional tensor of molecules and proteins output in step 3 into the cross-modal attention fusion module, perform nonlinear normalization on the calculated scores through the S-shaped growth curve, and convert them into standardized probability values in the interval [0,1] as the final prediction result of the binding affinity between the drug and the target protein for binding affinity prediction.
[0008] A method for predicting drug-target interactions based on a Freudian algorithm network is implemented as follows: The present invention preprocesses drug molecules by representing them as molecular graph structures to capture rich chemical topological information. Drug molecules are modeled as two-dimensional undirected graphs. A simplified linear input system sequence molecular representation is converted into a molecular graph model using cheminformatics tools. During feature encoding, each common atom type is first mapped to a unique integer index based on a predefined dictionary, and a one-hot encoding is used to generate an atom feature vector. Aromaticity information is then extracted, and chemical bond types are encoded using one-hot encoding, with single, double, triple, and aromatic bonds categorized into edge feature vectors. Finally, by integrating node and edge features, a standardized feature representation of the drug molecule graph is achieved, providing a unified format for input to subsequent neural network models. For target protein data, a residue-level graph Transformer model is used based on graph structure modeling to encode protein spatial structure information, extracting three-dimensional conformational features between residues and potential binding pockets. A binding site prediction algorithm is used to locate potential drug binding regions within the complete three-dimensional protein structure using the calculated bounding box coordinates.
[0009] A method for predicting drug-target interactions based on a Freud algorithm network. The implementation process of step 2 is as follows:
[0010] Based on the standardized data of drug molecules and proteins preprocessed in step 1, a dual-channel feature encoding strategy is used to extract their structural and sequence information respectively: for drug molecule data, the extended connectivity fingerprint algorithm is applied to recursively extract molecular topology and functional group environment features in a multi-order neighborhood centered on atoms, and the local structure is encoded into a binary vector of fixed length using hash mapping rules to form a fingerprint representation describing the molecular topological structural properties; for target protein sequence data, the second-generation protein sequence pre-training model of the evolutionary scale modeling framework is introduced, and the amino acid sequence is feature extracted at the residue level through a deep neural network to generate a high-dimensional embedding vector containing comprehensive information such as spatial conformation, residue co-evolution association, and functional domain, ultimately forming a mathematical feature expression of drug topological fingerprint and protein sequence embedding as the input for subsequent neural network fusion modeling.
[0011] A method for predicting drug-target interactions based on a Freud algorithm network. The implementation process of step 3 is as follows:
[0012] The molecular fingerprint tensor and protein embedding tensor are separately input into a parallel feature fusion neural network architecture. This parallel architecture consists of two branch channels, each independently processing and fusion mapping the drug and protein features. Each branch integrates multiple Floyd algorithm neural network units, which implement multi-scale dynamic feature extraction and optimization based on the shortest distance path search principle and recursive feature iteration mechanism. Specifically, the Floyd algorithm utilizes a dynamic weight update strategy, combined with a Kolmogorov-Arnold network mapping module, and automatically combines multiple Fourier basis vectors through a learnable Fourier basis function expansion. During feature propagation, feature transfer weights are continuously adjusted to achieve long- and short-range interaction representation between different feature dimensions. The output three-dimensional molecule and protein tensors are input into a cross-modal attention fusion module, where a multi-channel self-attention mechanism is used to model the interaction between the multimodal high-order features of drugs and proteins. In the attention module, a multi-head parallel attention mechanism is used to capture potential molecule-protein interaction patterns and implicit spatial coupling relationships. The interactive features output by the attention module are row-normalized along the feature dimension to enhance feature stability and expression consistency; finally, the output normalized vector is subjected to feature fusion, and the cosine similarity aggregation operation is performed on the vector along the functional dimension to form a fused latent space embedding representation.
[0013] A method for predicting drug-target interactions based on a Freud algorithm network. The implementation process of step 4 is as follows:
[0014] The similarity score calculated in step 3 is nonlinearly normalized using an S-shaped growth curve and converted into a standardized probability value in the interval [0,1] as the final prediction result of the drug-target protein binding affinity for binding affinity prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a flow chart of the drug-target interaction prediction method based on the Freud algorithm network.
[0016] Figure 2 This is a flow chart for data preprocessing of molecular graphs and protein graphs.
[0017] Figure 3 This is a flow chart for feature extraction of molecules and proteins.
[0018] Figure 4 This is a flowchart of the network structure using the Floyd algorithm. DETAILED DESCRIPTION
[0019] The present invention is described in detail below with reference to the accompanying drawings and examples.
[0020] The present invention proposes a method for predicting drug-target interaction based on Freud algorithm network. The specific process is as follows: Figure 1 As shown in the figure, the four steps are as follows: data preprocessing of molecular graphs and protein graphs, extraction of molecular and protein features respectively through pre-training models, network training using the Floyd algorithm, and prediction of drug-target affinity.
[0021] Step 1: Data preprocessing flow chart for molecular graphs and protein graphs Figure 2 As shown, in order to achieve a unified graph structure input representation of drug molecules and target proteins, the present invention proposes a standardized graph construction method based on the fusion of node and edge features, which extracts features and formats the drug molecule and protein structure data respectively, and constructs a structured graph input adapted to the subsequent graph neural network model. In the molecular graph, both nodes and edges carry corresponding feature vectors, representing each drug molecule as a two-dimensional undirected graph structure. , where the node set Represents atoms in a molecule, edge sets Represents the covalent chemical bonds between atoms; the initial structural information of the drug molecule is in the form of a simplified molecular linear input system expression, which is parsed and converted into a standard molecular graph structure using chemical informatics tools. For the node features of the drug graph, 17 common molecular types are encoded using 17-dimensional one-hot encoding; the atomic form charge is 0 or 1, indicating whether it is charged; the number of free radical electrons is 0 or 1, indicating the number of unpaired electrons of the atom; for the edge features of the molecule, first including single bonds, double bonds, triple bonds, and aromatic bonds, represented by 4-dimensional one-hot encoding; judging whether it is a conjugated bond, indicating whether the bond participates in the conjugated structure, is represented by 0 or 1-dimensional one-hot encoding; judging whether it is in a ring, using 0 or 1-dimensional to indicate whether the bond constitutes part of the ring structure; judging the 6 types of hybrid orbitals sp, sp2, sp3, sp3d, sp3d2, and others, is represented by 6-dimensional one-hot encoding. The molecular feature encoding is shown in Table 1. In the protein graph, each target protein is represented as a protein pocket graph , where the node set represents protein residues, edge sets Represents the interaction or structural connection between residues; for protein graph node features, the standard 20 natural amino acids are covered, and the amino acid type of each residue is represented by 20-dimensional one-hot encoding. To reflect the compactness and conformational distribution of the internal structure of the residue, the maximum and minimum scaled distances between all atom pairs in the residue are extracted, and the following three groups of specific atomic distances are supplemented: and , and N, C and N; then further combined with the spatial conformation information of the protein, the main chain and side chain conformation angles of each residue were extracted, including the main chain torsion angle and the first Angle, a total of 4 dimensions, reflects the spatial rotational freedom and structural distribution pattern of the protein; the edge features of the protein graph include four types of structural edge features, with 1 dimension indicating whether two residues form a structural connection; with 1 dimension indicating The scaled distance between atoms; 1D is used to represent the scaled distance between residue center points; 2D is used to represent the maximum spatial distance between all atom pairs of residues. The protein feature encoding is shown in Table 2.
[0022] Table 1 Molecular feature coding
[0023]
[0024] Table 2 Protein feature codes
[0025]
[0026] Step 2: Extract the features of molecules and proteins through pre-training models. Figure 3As shown: For the extraction of drug molecule graph features, the present invention adopts a Morgan fingerprint generation method based on topological structure. First, the structural feature information of each atomic node in its local neighborhood is traversed through all atoms and their neighboring structures in the graph, and its topological structure information is encoded using a hash function, and finally a fixed-dimensional embedding representation vector is generated; the fingerprint can be a binary Boolean representation, or it can be converted into a continuous feature vector according to needs, where the dimension is fixed and includes atomic-level topological features such as functional groups, hybridization states, and ring structures. For the feature extraction of protein graphs, the present invention combines structural graph information with sequence embedding information for fusion modeling. On the one hand, the three-dimensional structural information of the protein is encoded by constructing a protein binding pocket graph, where the nodes are residues and the edges are the structural connection relationships between residues; on the other hand, the second-generation protein sequence pre-training model of the evolutionary scale modeling framework is introduced to embed the protein amino acid sequence. The second-generation protein sequence pre-training model of the evolutionary scale modeling framework can encode a one-dimensional protein sequence into a length of And the dimension is Contextual semantic representation matrix. Considering that the second-generation protein sequence pre-training model of the evolutionary scale modeling framework performs a global averaging operation on residue embeddings by default, which may lead to feature loss, the output of the residue-level embedding vector before average pooling is retained, and a customized length-wise average pooling is performed to generate the final protein representation vector.
[0027] Step 3: The flowchart of feature fusion of molecule and protein vectors through Freud network is as follows: Figure 4 As shown: First, for each input vector , each frequency component is respectively associated with a learnable weight coefficient and Perform element-by-element multiplication and accumulate the sum to form the first Feature expression of output dimensions:
[0028] (1)
[0029] and are the learnable weight parameters in the Floyd coefficient matrix, is the bias term, Is to construct the Fourier grid vector = [1,2,...,grid_size], the output is passed to the subsequent network layer after the nonlinear activation function ReLU. The features of drugs and targets are mapped and encoded through the Kolmogorov-Arnold network module, and then they are interacted and integrated through the joint attention mechanism. Drug Embedding and protein embedding The transformation is performed by a projector consisting of two Kolmogorov-Arnold network modules:
[0030] (2)
[0031] (3)
[0032] in, 、 Denote the embedding representation of drug and protein in nonlinear Fourier space respectively. To further characterize the complex dependency between drug atoms and protein residues, a joint attention mechanism is introduced to model the interaction features between embedding vectors. drug atoms and The interaction strength between protein residues The calculation is as follows:
[0033] (4)
[0034] in is a learnable weight matrix, Indicates the drug atoms and The interaction strength between protein residues. Then, the attention score is calculated :
[0035] (5)
[0036] Attention score Constructing the attention matrix , where each position represents the drug atoms and the Finally, for the unknown drug molecule and target protein pair to be predicted, feature fusion is adopted to embed through cosine similarity:
[0037] (6)
[0038] (7)
[0039] It is the final interaction representation, which serves as the prediction result of the drug's effect on the target protein and saves the trained model.
[0040] In step 4, after outputting the interaction vector between the drug and target protein in step 3, a sigmoid growth curve is introduced to perform nonlinear normalization, mapping the raw similarity score to a standardized probability value in the interval [0, 1], which serves as the final binding affinity prediction. This probability value intuitively reflects the binding probability between the drug molecule and the target protein, facilitating subsequent screening and decision analysis. For unknown drug-protein combinations, this method directly calls the trained model to obtain interaction features and generate predictions. Given that this task is a binary classification problem, binary cross-entropy is used as the loss function during training to optimize prediction accuracy and model robustness.
[0041] The proposed drug-target interaction prediction method was fully validated on the BIOSNAP public benchmark dataset. After 50 rounds of training, the model achieved an AUPR score of 0.933 on the test set, significantly outperforming the representative ConPlex method, which achieved an AUPR of 0.895 under the same conditions. This improvement in prediction accuracy demonstrates the superiority of the proposed method in terms of generalization performance and model expressiveness.
[0042] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A method for predicting drug-target interactions based on a Freud algorithm network, characterized in that: The Floyd algorithm is used to update weights for training, which includes four steps: preprocessing molecular graph and protein graph data, extracting features of molecular graph and protein graph, and using the Floyd algorithm network for training and prediction. The specific steps are as follows: Step 1: Drug molecules and target proteins are uniformly represented as graph structure inputs. First, a standardized graph construction method based on the fusion of node and edge features is used, where both nodes and edges carry corresponding feature vectors. Step 2: To extract drug molecule graph features, the topological structure information of each atomic node within its local neighborhood is encoded using a hash function. A protein binding pocket graph is then constructed to encode the protein's three-dimensional structural information. A second-generation protein sequence pre-training model based on the evolutionary scale modeling framework is introduced to embed the protein amino acid sequence. Step 3: Input the data into the neural network unit integrating multiple Floyd algorithms, use the Floyd dynamic weight update strategy, combine with the Kolmogorov-Arnold network mapping module, continuously adjust the feature transfer weight during the feature propagation process, and automatically combine multiple Fourier basis vectors through the learnable Fourier basis function expansion method to achieve long-range and short-range interactive expression between different feature dimensions. The implementation process of the Floyd algorithm is as follows: For each embedding vector , each frequency component is respectively associated with a learnable weight coefficient and Perform element-by-element multiplication and accumulate the sum to form the first Feature expression of output dimensions: ; and are the learnable weight parameters in the Floyd coefficient matrix, is the bias term, Is to construct the Fourier grid vector , passed to subsequent network layers; the features of drugs and targets are mapped and encoded through the Kolmogorov-Arnold network module, and then they are interacted and integrated through the joint attention mechanism; Step 4: The similarity score calculated in step 3 is nonlinearly normalized using the S-shaped growth curve and converted into a standardized probability value in the interval [0, 1] as the final prediction result.
2. The method for predicting drug-target interactions based on a Freud algorithm network according to claim 1, characterized in that: To achieve unified graph structure input of drug molecules and target proteins, the implementation process of molecular graph and protein graph preprocessing is as follows: Represent each drug molecule as a two-dimensional undirected graph structure , where the node set Represents atoms in a molecule, edge sets Represents covalent chemical bonds between atoms; in protein graphs, each target protein is represented as a protein pocket graph , where the node set represents protein residues, edge sets Represents the interaction or structural connection between residues; Drug graph node features include: 17 common molecular types, atomic formal charge, indicating whether it is charged; number of free radical electron pairs; molecular edge features, including single bonds, double bonds, triple bonds, and aromatic bonds; determination of whether it is a conjugated bond; whether the bond forms part of a ring structure; determination of 6 hybrid orbital types: sp, sp2, sp3, sp3d, sp3d2, and others; protein graph node features, including the standard 20 natural amino acids; to reflect the compactness and conformational distribution of the residue's internal structure, extract the maximum and minimum scaled distances between all atom pairs within the residue; and supplement the distances between three specific atoms. and , and N, C and N; then further combined with the spatial conformation information of the protein, the main chain and side chain conformation angles of each residue were extracted, including the main chain torsion angle and the first Angle; The edge features of protein graphs include four types of structural edge features, including whether two residues form a structural connection; Scaled distances between atoms; scaled distances between residue centers; maximum spatial distance between all pairs of atoms in a residue.
3. The method for predicting drug-target interaction based on Freud algorithm network according to claim 1, characterized in that: According to the protein graph structure, a special treatment is taken for the specific evolutionary scale modeling framework model. The feature extraction process of the protein graph is as follows: Encode the one-dimensional protein sequence as a And the dimension is contextual semantic representation matrix; considering that the second-generation protein sequence pre-training model of the evolutionary scale modeling framework performs a global average operation on the residue embedding by default, which may lead to feature loss, the output of the residue-level embedding vector before average pooling is retained, and a customized average pooling along the length direction is performed to generate the final protein representation vector.
Citation Information
Patent Citations
Comogolov GCN protein ligand affinity prediction method
CN118969062A
Cancer subtype classification method based on KAN network and multi-omics data
CN119252347A