Protein-ligand affinity prediction method based on ensemble learning
Through an integrated learning-based method, using the feature extraction and fusion of proteins and ligands, the dependence on three-dimensional data and computational resource consumption in the prior art is solved, and efficient and accurate protein-ligand affinity prediction is achieved.
Patent Information
- Application Number
- CN202510312872.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing protein-ligand affinity prediction methods rely on three-dimensional structural data, with high acquisition costs and limited quantity, insufficient feature extraction, large computing resources consumption, and insufficient generalization capabilities.
Using an integrated learning-based method, the physical and chemical characteristics and spatial conformational characteristics of proteins and ligands are extracted by obtaining the PDB files of proteins and the MOL2 files of ligands, and the feature extraction and fusion is performed using residue contact maps, attention convolution neural networks and isomerographic neural networks to reduce the input amount of model and improve prediction accuracy.
It significantly reduces the amount of model parameters, reduces the requirements for the training environment, improves the practicality and generalizability of the model, and improves the accuracy and generalization ability of predictions.
Smart Images

Figure CN120260668A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of protein ligands, and particularly relates to a method for predicting protein-ligand affinity based on ensemble learning. Background Art
[0002] Protein-Ligand Affinity refers to the binding strength between a protein and a ligand (such as a small molecule drug, a substrate, an inhibitor, etc.), usually expressed by a binding constant or a dissociation constant. Affinity is a key indicator in drug design and molecular docking research, directly affecting the biological activity and efficacy of the ligand.
[0003] Predicting protein-ligand interactions is a crucial step in drug research and development. Currently, the following two types of methods are mainly used in this field: (1) Methods based on physicochemical characteristics: such as Autodock Vina, which mainly uses the principles of molecular mechanics and statistical mechanics for prediction; (2) Methods based on deep learning: such as Pafnucy, which mainly relies on protein three-dimensional structure information for feature extraction and prediction. The above existing technologies have the following defects and deficiencies: 1. Strong data dependence: The existing methods heavily rely on protein three-dimensional structure data, which is costly to obtain and limited in quantity, restricting the wide application of the model; 2. Insufficient feature extraction: The existing methods often only focus on a single feature (such as physicochemical features or structural features), and cannot comprehensively capture the complex characteristics of protein-ligand interactions; 3. High consumption of computing resources: When using a three-dimensional convolutional neural network to process protein structures, the number of parameters is huge, and the requirements for computing resources are high, which is difficult for ordinary servers to bear; 4. Insufficient generalization ability: It performs well on the training set, but the prediction accuracy significantly decreases on the test set and the validation set, indicating that the model has limited prediction ability for unknown protein-ligand pairs. Summary of the Invention
[0004] To solve the problems existing in the above prior art, the present invention proposes a method for predicting protein-ligand affinity based on ensemble learning, which includes: obtaining the PDB file of the protein, the PDB file of the receptor, and the MOL2 file of the ligand; extracting features from the PDB file of the protein to obtain a protein feature vector; extracting the physicochemical features and spatial conformation features of the protein and the ligand from the PDB file and the MOL2 file; interacting the physicochemical features and spatial conformation features to obtain an interaction feature vector; processing the MOL2 file of the ligand to obtain a ligand feature vector; splicing the protein feature vector, the interaction feature vector, and the ligand feature vector, and inputting the spliced features into an affinity prediction model to obtain an affinity prediction result.
[0005] Advantages of the present invention:
[0006] The present invention converts the three-dimensional structure of proteins into residue contact maps, significantly reducing the model input while maximizing the retention of structural information. This processing method not only effectively reduces the number of model parameters but also greatly reduces the requirements for the training environment, making the model more practical and popularizable. The present invention characterizes ligands by combining molecular fingerprints and adjacency matrices, taking into account both the physicochemical properties of ligands and fully retaining the spatial information of their bonds and atoms. By treating ligands as graph structures, more representative and comprehensive feature information is obtained, thus significantly improving the performance of the model during training, testing, and validation. The present invention incorporates both the physicochemical properties and spatial structure information of the protein-ligand interaction region into the scope of feature extraction. By using advanced homogeneous graph neural networks to process these features, not only is the interaction mechanism between proteins and ligands better explained, but the accuracy of the final prediction results is also significantly improved. The present invention only requires the PDB file of the protein and the MOL2 file of the ligand. This feature effectively solves the problem of dependence on three-dimensional large datasets in traditional methods. The model simultaneously integrates physicochemical properties and spatial structure features, not only improving the interpretability of protein-ligand interactions but also significantly enhancing the accuracy of affinity prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 is the overall flowchart of the integrated model of the present invention;
[0008] Figure 2 is the diagram of the attention convolutional neural network model for processing proteins of the present invention;
[0009] Figure 3 is the diagram of the graph neural network for processing ligands of the present invention;
[0010] Figure 4 is the schematic diagram of the types of atoms and chemical bonds in the protein-ligand interaction part of the present invention;
[0011] Figure 5 is the diagram of the homogeneous graph neural network for processing the protein-ligand interaction part of the present invention;
[0012] Figure 6 is the schematic diagram of the multi-layer perceptron after the features are spliced and input of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0014] A protein-ligand affinity prediction method based on ensemble learning, such as Figure 1 shown. This method includes: obtaining the PDB file of the protein, the PDB file of the receptor, and the MOL2 file of the ligand; extracting features from the PDB file of the protein to obtain a protein feature vector; extracting the physical and chemical features and spatial conformation features of the protein and the ligand from the PDB file and the MOL2 file; interacting the physical and chemical features and the spatial conformation features to obtain an interaction feature vector; processing the MOL2 file of the ligand to obtain a ligand feature vector; splicing the protein feature vector, the interaction feature vector, and the ligand feature vector, and inputting the spliced features into an affinity prediction model to obtain an affinity prediction result.
[0015] The ensemble model proposed by the present invention realizes accurate prediction of protein-ligand binding affinity through multi-level and multi-dimensional feature extraction and fusion. The overall process of the model is divided into three parts, which respectively process the features of the protein, the protein-ligand interaction, and the ligand, and finally output the predicted binding affinity (pKa) through a fully connected layer.
[0016] In this embodiment, the preprocessing of the PDB file of the protein includes: 1. PDB file parsing, reading the PDB file: 1. Use tools (such as Biopython, ProDy, or MDTraj) to parse the PDB file and extract the atomic coordinates and residue information of the protein. Retain the required main chain residue sequence. Select residues: Usually, the Cα atom (Alpha Carbon) is selected as the representative of the residue because each amino acid has only one Cα atom, and it is located on the main chain and can better reflect the overall structure of the protein. 2. Extract Cα atom coordinates, obtain Cα coordinates: Extract the three-dimensional coordinates (x, y, z) of the Cα atom of each residue from the PDB file. Process missing data: If the Cα coordinates of some residues are missing, interpolation or deletion of these residues can be selected. 3. Calculate the distance between residues, calculate the distance matrix: Calculate the Euclidean distance between all Cα atoms to generate a symmetric distance matrix. 4. Define a contact threshold, select a threshold: Define a distance threshold Used to determine whether two residues are in contact. Generate a contact map: Convert the distance matrix into a binary contact map, where 1 indicates that two residues are in contact and 0 indicates not in contact. Data augmentation, rotation and translation: Randomly rotate and translate the three-dimensional structure of the protein to generate different contact maps to increase data diversity. Noise addition: Add random noise to the Cα coordinates to improve the robustness of the model. Data storage, storage format: Store the generated contact maps in a format suitable for deep learning frameworks, such as NumPy arrays (.npy) for convenient reading during model training. File naming: Generate a unique file name for the contact map of each protein based on the protein file name for easy subsequent loading and processing.
[0017] In this embodiment, protein feature extraction includes: The PDB file of the protein is preprocessed to generate a residue contact map (distance map). The residue contact map is input into a self-designed attention convolutional neural network (Attention-based CNN). After processing by a two-dimensional convolutional layer (Conv2D), batch normalization (BN), ELU activation function, and residual block (ResBlock), gradient vanishing is prevented and deep features are extracted. Redundant information is removed through average pooling. A self-attention module is introduced, and the Softmax function is used to generate attention weights to enhance the representation ability of key features. Finally, a 128-dimensional protein feature vector is output.
[0018] The Attention Convolutional Neural Network (ACNN) includes the following core components: a two-dimensional convolutional layer, a batch normalization layer, an ELU activation function, a residual block, an attention module, and a key feature enhancement module. The processing flow of the attention convolutional neural network for the residue contact map specifically includes the following steps: Two-dimensional convolutional layer: Initially extract the spatial features in the residue contact map through two-dimensional convolutional operations to capture the local interaction patterns between protein residues. Batch normalization layer: Perform batch normalization on the features output by the convolution to accelerate model convergence and improve training stability. ELU activation function: Introduce non-linearity using the Exponential Linear Unit (ELU) activation function to enhance the model's expressive power. Residual block: Fuse shallow features and deep features through residual connections to alleviate the problem of gradient disappearance and promote the optimization of deep networks. Mean pooling layer: Perform mean pooling operations on the feature map to reduce the feature dimension and redundant parameters, and improve computational efficiency. Attention module: Use the attention mechanism to weight the features, automatically select the key features related to protein-ligand affinity, and enhance the model's focusing ability on important information. Key feature enhancement module: Further strengthen the representation ability of key features through specific feature enhancement strategies to improve the model's prediction accuracy. Softmax activation function: Use the Softmax activation function in the output layer to ensure the probability distribution characteristics of the output results, thereby obtaining the required protein features. Through the above steps, the attention convolutional neural network can efficiently process the residue contact map.
[0019] In this embodiment, the extraction of protein-ligand interaction features includes: extracting the physicochemical features and spatial conformation (chemical bond) features of proteins and ligands from PDB and MOL2 files. Through processing by an Isomorphic Graph Neural Network (GNN), the interaction information of nodes (such as atoms in molecules) and edges (such as chemical bonds in molecules) is captured. An interaction feature vector of 128 dimensions is output.
[0020] In this embodiment, the ligand feature extraction includes: converting the MOL2 file into a standardized SMILES sequence. Further processing generates a molecular fingerprint and an adjacency matrix. Input into a graph neural network (GNN), and through embedding, linear layer, ReLU activation function, and average pooling (Avg-Pool) processing, the features are iteratively updated. A ligand feature vector of 128 dimensions is output.
[0021] As Figure 4As shown, the extraction of the physicochemical and spatial conformation features of proteins and ligands includes the following steps: 1. Extraction of atomic features. Atomic features are used to describe the chemical and structural properties of each atom in proteins and ligands, specifically including: - Atomic type: Atoms are classified according to their atomic numbers, covering common elements (such as carbon (C), nitrogen (N), oxygen (O)) and special categories such as halogens and metals. - Atomic chirality: Describes the stereochemical properties of atoms, covering four possible chirality types, used to characterize the three-dimensional conformation of molecules. Formal charge: Represents the positive or negative charge that an atom may carry, used to understand the electron distribution and chemical properties of molecules. Hybridization type: Describes the hybridization state of atoms (such as sp 3 2, sp 2 3, sp etc.), reflecting the geometric shape of atoms and their impact on the molecular structure. Number of hydrogen atoms: Represents the number of hydrogen atoms attached to an atom, used to characterize the chemical environment of the atom. Implicit valence state: Describes the number of chemical bonds that an atom can form, reflecting its chemical activity. Connectivity: Represents the number of other atoms connected to an atom, used to characterize the topological structure of the atom in the molecule. Whether it is an aromatic atom: A boolean value indicating whether the atom participates in an aromatic ring structure, and aromaticity has an important impact on the chemical properties of molecules. Whether it is an atom in a protein: A boolean value used to distinguish whether an atom belongs to a protein. Atomic mass: Expressed in atomic mass units (amu) and normalized by dividing by 100 to adjust the feature scale. 2. Extraction of chemical bond features. Chemical bond features are used to describe the types and spatial properties of chemical bonds in proteins and ligands, specifically including: Chemical bond type: Classified according to the type of chemical bond, such as single bond, double bond, triple bond, aromatic bond, etc. Chemical bond direction: Describes the directionality of chemical bonds, which is crucial for understanding the three-dimensional structure of molecules. Stereochemical properties of the bond: Describes the stereochemical attributes of chemical bonds (such as cis, trans, etc.), used to characterize the spatial conformation of molecules. Whether the chemical bond is in a ring: A boolean value indicating whether the chemical bond is located in the ring structure of the molecule. Whether the chemical bond belongs to a bond in a protein: A boolean value used to distinguish whether the chemical bond belongs to a protein. Construct an edge list and edge features. Based on the atomic and chemical bond information obtained through the above methods, construct an edge list and edge features, specifically including: Edge list: Represented as pairs of atomic indices, covering the chemical bonds within the ligand and the pairs of atoms with a distance less than the threshold in the ligand-protein interaction. Edge features: Include the type, direction, stereochemical attributes, whether in a ring, whether belonging to a protein of the chemical bond, and the length of the edge (i.e., the distance between atoms). Through the above steps, the extracted atomic features and chemical bond features can comprehensively describe the physicochemical properties and spatial conformation of proteins and ligands.
[0022] Processing the MOL2 file of the ligand includes the following steps: Reading the MOL2 file: Use the chemical informatics tool RDKit to parse the MOL2 file and extract the atomic and bond information of the ligand. Molecular structure preprocessing: Removing non-ligand parts: Remove solvent molecules or other non-ligand parts in the MOL2 file to ensure that only the target ligand is retained. Removing hydrogen atoms: Remove hydrogen atoms in the ligand to simplify the molecular structure and ensure chemical rationality. Standardizing formal charges: Standardize the formal charges of atoms in the ligand to ensure the correct molecular charge state. Stereochemistry processing: Check and correct the stereochemical information (such as chiral centers) of the molecule to ensure the accuracy of the molecular structure. Obtaining the SMILES sequence of the standardized ligand: Convert the preprocessed ligand into a standardized SMILES sequence for subsequent construction of the molecular graph. Constructing the MOL molecular graph: Convert the standardized SMILES sequence into a MOL molecular graph and extract atomic features, bond features, and edge features from it. Obtaining the molecular fingerprint: According to the extracted atomic features and edge features, set as the threshold to generate the molecular fingerprint. The molecular fingerprint is used to characterize the chemical structure and spatial information of the ligand. Converting to a NumPy array and storing: Convert the generated molecular fingerprint to a NumPy array and store it in a format suitable for processing by deep learning frameworks (such as a.npy file) for subsequent model training and prediction. Through the above steps, the MOL2 file of the ligand is efficiently processed, and the extracted molecular fingerprint can comprehensively characterize the chemical and structural properties of the ligand, providing high-quality feature inputs for protein-ligand affinity prediction.
[0023] As Figure 3 shown, processing the ligand by the graph neural network includes: Inputting the standardized SMILES sequence; Preprocessing to generate the molecular fingerprint and the adjacency matrix. Converting discrete features into continuous features through Embedding. Processing the features with a Linear layer and a ReLU activation function. Mean pooling to remove redundant information. Iteratively updating the features, and finally outputting the ligand features through a fully connected layer. Outputting the ligand feature vector.
[0024] In this embodiment, the input is the standard SMILES sequence of the ligand, which becomes the molecular fingerprint and the adjacency matrix as the input of the GNN after preprocessing. The schematic diagram of the GNN is as follows. The adjacency matrix and the molecular fingerprint of each ligand molecule are processed by embeddbing into continuous features, flattened after embedding to reduce the dimension, represented as continuous features for processing, and then, after passing through a linear layer, a relu activation function, and avg-pool mean pooling to remove redundancy, they are iteratively updated. Finally, the feature matrix after iterative update is output through the FC fully connected layer. After being processed by sum or mean, the features of each ligand molecule are integrated, and then the redundancy is removed through mean pooling to obtain the required features of the ligand molecule.
[0025] In this embodiment, the homogeneous graph neural network for processing protein-ligand interactions includes: inputting the atomic and chemical bond features of proteins and ligands into the network. Classifying and performing One-Hot encoding on the atomic and chemical bond features. Converting the features into low-dimensional continuous features through Embedding. Processing the features with a Linear layer and a ReLU activation function. Extracting the interaction information of nodes and edges by a homogeneous graph neural network (GINConv). Using batch normalization (BN) to alleviate internal covariate shift and accelerate model convergence. Iteratively updating the features, and finally outputting the interaction features through a fully connected layer. Outputting an interaction feature vector.
[0026] The affinity prediction model processes the concatenated features through the following steps: Feature concatenation: Concatenating the protein features, ligand features, and protein-ligand interaction features to form a unified feature vector as the input of the model. Multi-layer perceptron (MLP) processing: Inputting the concatenated feature vector into a multi-layer perceptron (MLP) for iterative update. As Figure 6 shown, the MLP consists of multiple fully connected layers, and each layer introduces a non-linear transformation through a non-linear activation function (such as ReLU or ELU) to gradually extract high-order feature representations. Final prediction value generation: Generating the final prediction value (such as the pKa value) through the output layer of the last layer of the MLP to characterize the affinity between the protein and the ligand. Model testing and verification: Comparing the generated prediction value with the true label, and evaluating the performance of the model through metrics such as mean squared error (MSE), mean absolute error (MAE), or correlation coefficient (R 2 ) etc. Using cross-validation or an independent test set to verify the model to ensure its generalization ability and robustness.
[0027] In Figure 1 , the first figure shows how to classify the features of each atom and bond during the interaction between a protein and a ligand, which are respectively the atom type, atom chirality, formal charge, hybridization type, number of hydrogen atoms, implicit valence state, connectivity, whether it is aromatic, and whether it is an atom in the protein. The bond features are the chemical bond type, bond direction, stereochemical property of the bond, whether the chemical bond is in a ring, and whether the chemical bond belongs to a bond in the protein. After one-hot encoding, it is returned as a list, and the physicochemical properties and spatial structure of the protein-ligand interaction part can be obtained. Finally, these obtained partial features are converted into tensors for use by the homogeneous graph neural network.
[0028] Figure 2It is a homogeneous graph neural network framework. After the input interaction information is processed through embedding, it is processed into low-dimensional continuous features. After iterative update through the linear layer, it is input into the ginconv homogeneous graph neural network for feature extraction. After being processed by the relu activation function, it becomes non-linear, enabling the model to express more complex relationships. After being processed by bn normalization, the internal covariates are adjusted, accelerating the convergence of the model training process and alleviating the problems of gradient disappearance and gradient explosion. After iterative update through the linear layer, the homogeneous graph neural network is used again to process and extract features, which are input into the linear layer for update, and the final features are output through the activation function.
[0029] The structure of the fully connected layer is as Figure 5 shown. The 128-dimensional feature vectors from proteins, interactions, and ligands are input into the fully connected layer, and feature fusion and prediction are performed through a multi-layer perceptron (MLP). The multi-layer perceptron consists of fully connected layers of sizes 128, 256, 128, 64, 32, and 1. Finally, the binding affinity (pKa) predicted by the model is output.
[0030] In terms of protein processing: This solution innovatively converts the three-dimensional structure of proteins into residue contact maps, significantly reducing the model input while maximizing the retention of structural information. This processing method not only effectively reduces the number of model parameters but also greatly reduces the requirements for the training environment, making the model more practical and popularizable.
[0031] In terms of ligand processing: This solution uses a combination of molecular fingerprints and adjacency matrices to represent ligands, considering both the physicochemical properties of ligands and fully retaining the spatial information of their bonding and atoms. By treating ligands as graph structures for processing, more representative and comprehensive feature information is obtained, thus significantly improving the performance of the model in the training, testing, and validation processes.
[0032] In terms of protein-ligand interaction processing: This solution innovatively proposes a new processing idea, incorporating both the physicochemical properties and spatial structure information of the protein-ligand interaction region into the feature extraction scope. By using an advanced homogeneous graph neural network to process these features, not only is the interaction mechanism between proteins and ligands better explained, but the accuracy of the final prediction results is also significantly improved.
[0033] In terms of data compatibility: This integrated model has loose requirements for data formats, only requiring the PDB file of the protein and the MOL2 file of the ligand. This feature effectively solves the problem of dependence on three-dimensional large datasets in traditional methods. The model simultaneously integrates physicochemical properties and spatial structure features, not only improving the interpretability of protein-ligand interactions but also significantly enhancing the accuracy of affinity prediction.
[0034] In the above-described embodiments, the object, technical solution, and advantages of the present invention have been further described in detail. It should be understood that the above-described embodiments are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for predicting protein-ligand affinity based on ensemble learning, characterized in that, Including: Obtaining the PDB file of the protein, the PDB file of the receptor, and the MOL2 file of the ligand; Performing feature extraction on the PDB file of the protein to obtain a protein feature vector; extracting the physicochemical features and spatial conformation features of the protein and the ligand from the PDB file and the MOL2 file; performing interactions on the physicochemical features and the spatial conformation features to obtain an interaction feature vector; processing the MOL2 file of the ligand to obtain a ligand feature vector; splicing the protein feature vector, the interaction feature vector, and the ligand feature vector, and inputting the spliced features into an affinity prediction model to obtain an affinity prediction result.
2. The method for predicting protein-ligand affinity based on ensemble learning according to claim 1, wherein Performing feature extraction on the PDB file of the protein includes: preprocessing the PDB file of the protein to generate a residue contact map; inputting the residue contact map into an attention convolutional neural network to obtain a protein feature vector.
3. The method for predicting protein-ligand affinity based on ensemble learning according to claim 2, wherein Preprocessing the PDB file of the protein includes: parsing the protein PDB file to obtain the atomic coordinates and residue information of the protein; selecting the Cα atom as the residue; extracting the three-dimensional coordinates (x, y, z) of the Cα atom of each residue; deleting the Cα atoms of the residues with missing coordinates; calculating the Euclidean distance between all Cα atoms to generate a symmetric distance matrix; setting a distance threshold to convert the distance matrix into a binary contact map; randomly rotating and translating the three-dimensional structure of the protein to generate different contact maps; adding random noise to the Cα coordinates; storing the generated contact maps.
4. The method for predicting protein-ligand affinity based on ensemble learning according to claim 2, wherein, The attention convolutional neural network includes: a two-dimensional convolutional layer, a batch normalization layer, an ELU activation function, a residual block, an attention module, and a key feature enhancement module; the processing flow of the attention convolutional neural network for the residue contact map includes: extracting the spatial shallow features in the residue contact map through two-dimensional convolution. Performing batch normalization on the shallow features; using the ELU activation function to process the normalized features to obtain deep features; fusing the shallow features and the deep features through the residual block to obtain a fused feature map. Performing an average pooling operation on the fused feature map; using the attention mechanism to weight the fused feature map to automatically select the key features of the protein; introducing noise into the data for feature enhancement; using the Softmax activation function to ensure the probability distribution characteristics of the output result and obtain protein features.
5. The method for predicting protein-ligand affinity based on ensemble learning according to claim 4, wherein Enhancing the key features includes: rotating the protein residue contact map; performing symmetric transformation on the protein residue contact map.
6. The method for predicting protein-ligand affinity based on ensemble learning according to claim 1, wherein Extracting the physicochemical characteristics and spatial conformation characteristics of proteins and ligands includes the following steps: extracting the atomic characteristics of proteins and ligands; specifically: atomic type, atomic chirality, formal charge, hybridization type, number of hydrogen atoms, implicit valence state, connectivity, whether it is an aromatic atom, whether it is an atom in a protein, and atomic mass; extracting the chemical bond type characteristics of proteins and ligands; specifically including chemical bond type, chemical bond direction, stereochemical properties of the bond, whether the chemical bond is in a ring, and whether the chemical bond belongs to a bond in a protein; constructing an edge list and edge features based on the atomic characteristics and chemical bond type characteristics; obtaining the physicochemical properties and spatial conformation of proteins and ligands according to the edge list and edge features.
7. A method for predicting protein-ligand affinity based on ensemble learning according to claim 1, characterized in that, Processing the MOL2 file of the ligand includes: parsing the MOL2 file to extract the atomic and chemical bond information of the ligand; preprocessing the ligand; converting the preprocessed ligand into a standardized SMILES sequence; converting the standardized SMILES sequence into a MOL molecular graph, and extracting atomic features, chemical bond features, and edge features from it. According to the extracted atomic features and edge features, set as the threshold to generate a molecular fingerprint. Convert the generated molecular fingerprint into a NumPy array and store it in the Numpy file format.
8. A method for predicting protein-ligand affinity based on ensemble learning according to claim 7, wherein, The preprocessing of ligands includes: removing non-ligand parts and solvent molecules to obtain the target ligand; removing hydrogen atoms in the target ligand; standardizing the formal charges of atoms in the ligand; checking and correcting the stereochemical information of the molecule.
9. The method for predicting protein-ligand affinity based on ensemble learning according to claim 1, wherein, The affinity prediction model processes the spliced features as follows: splicing protein features, ligand features, and protein-ligand interaction features, and inputting the spliced feature vector into a multi-layer perceptron for iterative update; the multi-layer perceptron consists of multiple fully connected layers, and each layer introduces a non-linear transformation through a non-linear activation function to gradually extract high-order feature representations; through the output layer of the last layer of the multi-layer perceptron, a final predicted value is generated, and the final predicted value represents the affinity of the protein-ligand.
Citation Information
Patent Citations
Drug-target interaction prediction method based on graph convolution and word vector
CN110289050A
Drug target binding affinity prediction method based on three-branch CNN
CN116189795A
Protein and ligand affinity prediction method based on graph neural network decoupling
CN116312758A
Multi-modal protein-ligand binding affinity prediction method based on cross-channel fusion
CN117912545A
Application of multi-modal feature fusion model in drug target binding affinity prediction
CN119479783A