Protein-ligand binding affinity prediction method, system, equipment and medium

By extracting and strengthening the characteristics of protein sequences and ligand SMILES sequences, and performing fusion prediction, the problem of predicting protein-ligand binding affinity in the prior art is solved, and higher prediction accuracy is achieved.

CN120108490APending Publication Date: 2025-06-06JIANGNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510294948.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art cannot better capture the in-depth characteristics of proteins and small molecule ligands, resulting in the insignificant effect of predicting protein-ligand binding affinity.

Method used

By collecting the protein-ligand binding affinity dataset, the global and local features of the protein sequence and ligand SMILES sequence were extracted, and feature strengthening and fusion were performed, and feature prediction was performed using deep learning models such as Transformer and convolutional neural networks.

Benefits of technology

The accuracy of protein-ligand binding affinity prediction is improved, and the problems of high cost and low accuracy of traditional methods are overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108490A_ABST
    Figure CN120108490A_ABST
Patent Text Reader

Abstract

The invention relates to a protein-ligand binding affinity prediction method, system, equipment and medium, and relates to the technical field of biological simulation prediction.The method comprises the steps that a protein-ligand binding affinity data set is collected and comprises a protein sequence and a ligand SMILES sequence; extracting global features and local features of the protein sequence and the ligand SMILES sequence; respectively performing feature enhancement on the global features and the local features of the protein sequence and the ligand SMILES sequence to obtain the global features and the local features of the enhanced protein sequence and the ligand SMILES sequence; performing feature fusion on global features and local features of the enhanced protein sequence and the ligand SMILES sequence to obtain fused features; and performing feature prediction on the fused features to obtain the value of the protein-ligand binding affinity. According to the method, the value of the protein-ligand binding affinity can be effectively predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of biological simulation prediction, and in particular to a method, system, device and medium for predicting protein-ligand binding affinity. Background Art

[0002] The interaction between proteins and small molecule ligands mainly forms a stable complex through non-covalent binding such as hydrogen bonding, electrostatic interaction and hydrophobic interaction. During the binding process, proteins and small molecules may undergo conformational changes, making the binding tighter, which involves a variety of biological processes, such as intracellular signal transduction and metabolic processes. Understanding the mechanism of protein-ligand interaction (PLI) helps to design new drugs that exert therapeutic effects through interaction with proteins. Identifying the mechanism of protein-ligand interaction not only helps to accelerate the drug design process, but also can effectively present the pathogenesis of the disease and provide ideas for the discovery and design of new drugs. Protein-ligand binding affinity (PLA) is an indicator to measure the strength of the interaction between the target protein and the ligand drug. In addition, in metabolic engineering, the affinity between the enzyme (protein) and the substrate (ligand) is reflected by the enzyme kinetic parameter Michaelis constant, which is one of the most important parameters for measuring the catalytic efficiency of the enzyme. Therefore, the prediction of protein-ligand binding affinity has important scientific and application value.

[0003] At present, the main experimental methods for determining protein-ligand binding affinity are isothermal titration calorimetry and plasma resonance. Although traditional experimental methods have high reliability, they also bring problems such as high cost and long time consumption. The use of computer-assisted prediction of protein-ligand binding affinity can reduce experimental costs and shorten the measurement cycle, which is of great significance for accelerating drug development, enzyme kinetic parameter prediction and other application fields. Traditional computer-aided drug design methods, such as molecular docking and molecular dynamics simulation, usually evaluate protein-ligand binding affinity through scoring functions such as AutoDock, X-Score and ChemScore. Although the problems of high cost and long time consumption of traditional wet experiments have been solved to a certain extent. However, these scoring functions are based on incomplete physical models, approximated to simplify calculations, and heavily rely on manual planning and complex operations, making it difficult for traditional computational methods to accurately predict the binding affinity of a large number of different protein-ligand pairs.

[0004] In recent years, artificial intelligence technologies represented by deep learning have been successfully applied to the fields of bioinformatics and synthetic biology, such as the prediction and design of gene regulatory elements, the prediction and design of protease functions, and the analysis and optimization of synthetic pathways. For the prediction of protein-ligand binding affinity, there are many methods that use deep learning technology to predict protein-ligand binding affinity. Among them, according to the type of input data, it can be mainly divided into sequence-based methods and structure-based methods. Sequence-based methods mainly use the amino acid sequence of the protein and the SMILES sequence of the ligand as the input of the model. For example, the DeepDTA model only needs to use convolutional neural networks to extract the features of proteins and ligands based on the protein sequence and ligand SMILES sequence to predict protein-ligand binding affinity. On this basis, the DeepDTAF model uses dilated convolution and traditional convolution operations to extract the features of proteins and ligands and predict protein-ligand binding affinity based on the protein sequence, protein binding pocket and ligand SMILES sequence, as well as the secondary structure characteristics of proteins and pockets. The DataDTA model integrates and captures multi-scale interaction information based on four different inputs: protein sequence, ligand SMILES sequence, binding pocket descriptor, and ligand molecular fingerprint, and predicts protein-ligand binding affinity by fusing high-speed network modules and multi-head attention modules. So far, structure-based deep learning methods for predicting protein-ligand binding affinity have also received widespread attention. For example, the Pafnucy model uses a four-dimensional tensor to represent the protein-ligand complex, and uses a three-dimensional convolutional neural network to learn interaction information and predict protein-ligand binding affinity. The OnionNet model is based on the specific contact between ligands and protein atoms, groups these contact points into different distance ranges to cover local and non-local interaction information between ligands and proteins, and uses deep convolutional neural networks to predict protein-ligand binding affinity. In structure-based prediction methods, information such as protein structure is needed. At present, only about 100,000 protein structures have been resolved, while billions of proteins and peptides with known sequences are still structurally unknown, including the sequence and position of their active pockets. The reason for the limited structural coverage is that resolving the structure of a single protein may require months or even years of meticulous work. Although the static structures of a small number of proteins have been resolved, the function of a protein is mainly determined by its dynamic properties. Existing data sets usually only provide a single static structure of a protein, which may affect the evaluation of experimental results. In contrast, sequence-based prediction methods do not rely on static structures and can effectively capture the dynamic properties and functional information of proteins, overcoming the limitations of structural data to a certain extent and being able to predict protein-ligand binding affinity more conveniently and efficiently.

[0005] However, existing sequence-based prediction methods can effectively capture the dynamic properties and functional information of proteins, but cannot better capture the deeper characteristics of proteins and small molecule ligands, resulting in insignificant results in predicting protein-ligand binding affinity. Therefore, it is necessary to design new methods to better explore the characteristics of proteins and small molecule ligands, and thus effectively predict protein-ligand binding affinity. Summary of the invention

[0006] Therefore, the technical problem to be solved by the present invention is to overcome the problem that the prior art cannot better capture the deeper features of proteins and small molecule ligands, resulting in the insignificant effect in predicting protein-ligand binding affinity.

[0007] In order to solve the above technical problems, the present invention provides a method for predicting protein-ligand binding affinity, comprising: Step S1: collecting a protein-ligand binding affinity data set, wherein the protein-ligand binding affinity data set includes a protein sequence and a ligand SMILES sequence; Step S2: extracting global features and local features of protein sequences and ligand SMILES sequences; Step S3: enhancing the global features and local features of the protein sequence and the ligand SMILES sequence respectively, and obtaining the enhanced global features and local features of the protein sequence and the ligand SMILES sequence; Step S4: performing feature fusion on the global features and local features of the enhanced protein sequence and the ligand SMILES sequence to obtain fused features; Step S5: Perform feature prediction on the fused features to obtain the value of protein-ligand binding affinity.

[0008] In one embodiment of the present invention, the method for extracting the global features of the protein sequence and the ligand SMILES sequence in step S2 comprises: Use the pre-trained protein language model ESM-2 to extract protein feature representation from protein sequences; The pre-trained molecular language model MolFormer is used to extract the ligand molecular feature representation from the ligand SMILES sequence; The protein feature representation and ligand molecule feature representation are respectively input into multiple series-connected Transformer encoders to obtain the protein global feature matrix and the ligand molecule global feature matrix , and as a global feature of protein sequence and ligand SMILES sequence, where Enter the code length for the protein sequence, is the protein word vector dimension, is the length of the ligand SMILES sequence, is the ligand word vector dimension.

[0009] In one embodiment of the present invention, the method for extracting local features of the input protein sequence and the ligand SMILES sequence in step S2 comprises: The protein sequence and ligand SMILES sequence are integer encoded, and the integer encoded protein sequence and ligand SMILES sequence are input into the embedding layer. The embedding layer converts the integer encoded protein sequence and ligand SMILES sequence into a two-dimensional output matrix through the nn.Embedding module of Pytorch and ,in, Enter the code length for the protein sequence, is the length of the ligand SMILES sequence, is the embedding dimension; The matrix and Input the first and second dilated convolution submodules respectively to obtain the protein local feature matrix and the ligand local feature matrix , and as local features of protein sequences and ligand SMILES sequences, where The first dilated convolution submodule includes A series of first dilated convolution blocks, each of which includes Activation function, convolution operation , Activation function, convolution operation And residual connection, convolution operation And convolution operation The expansion rate is ,in , then The first dilated convolution block is expressed as: ; in, , ; The second dilated convolution submodule includes A series of second dilated convolution blocks, each of which includes Activation function, convolution operation , Activation function, Convolution operation and residual connection, convolution operation And convolution operation The expansion rate is ,in , then The second dilated convolution block is expressed as: ; in, , .

[0010] In one embodiment of the present invention, the method for enhancing the global features of the protein sequence and the ligand SMILES sequence in step S3 comprises: The selective kernel network SKNet is used to enhance the global features of protein sequences, including: Protein global feature matrix Through two parallel convolution operations with different dilation rates and , and obtain the feature matrix and , where the convolution operation The expansion rate is , convolution operation The expansion rate is ; The feature matrix and Add together to generate the feature matrix , the feature matrix Perform global average pooling Get channel characteristics , and the channel features Through the first fully connected layer and the second fully connected layer, we can get , where the first fully connected layer includes batch normalization BatchNorm and The activation function is expressed as: , ; in, and are the linear transformation matrices of the first and second fully connected layers, respectively, and and ; Will Input to Layer, for and Assign a weight based on The weights assigned to the layers are broadcasted using the feature matrix and Make Hadamard product with the corresponding weight and sum to get the final output feature ,Will As the global feature of the enhanced protein sequence, the formula is expressed as: ; in, yes The attention weight of the first channel of the layer, yes The attention weight of the second channel of the layer, The operation of the layer is expressed as: , ; At the same time, the selective kernel network SKNet is used to enhance the global characteristics of the ligand SMILES sequence, and the ,Will As the global feature of the enhanced ligand SMILES sequence, the formula is expressed as: .

[0011] In one embodiment of the present invention, the method for enhancing the local features of the protein sequence and the ligand SMILES sequence in step S3 comprises: The squeeze-excitation network SENet is used to enhance the local features of protein sequences, including: Protein local feature matrix Perform global average pooling , get channel characteristics , the formula is: ; Channel Features After two third fully connected layers, the importance weight of the channel is obtained , where each third fully connected layer includes a linear transformation matrix , Activation function, linear transformation matrix and The activation function is expressed as: ; The protein local feature matrix is ​​broadcasted using the broadcast mechanism and the corresponding channel weight Do the Hadamard product to get the final output feature ,Will As the local feature of the enhanced protein sequence, the formula is expressed as: ; At the same time, the selective kernel network SKNet is used to enhance the local features of the ligand SMILES sequence, and the ,Will As the local feature of the enhanced ligand SMILES sequence, the formula is expressed as: .

[0012] In one embodiment of the present invention, the method of fusing the global features and local features of the enhanced protein sequence and the ligand SMILES sequence in step S4 to obtain the fused features includes: The enhanced protein sequence global features Input three different linear layers to calculate the query matrix of global features of protein sequence , key matrix Sum Matrix ; The global features of the enhanced ligand SMILES sequence Input three different linear layers to calculate the query matrix of the global features of the ligand SMILES sequence , key matrix Sum Matrix ; The bond matrix and value matrix of the global features of the protein sequence and the ligand SMILES sequence are exchanged; The query matrix, key matrix, and value matrix of the global features of the updated protein sequence are expressed as: ; in, , , It is The linear transformation matrix of the attention head query matrix, key matrix and value matrix, is the dimension of each attention head key vector, is the number of attention heads; The query matrix, bond matrix, and value matrix of the global features of the updated ligand SMILES sequence are expressed as: ; in, , , It is The linear transformation matrix of the attention head query matrix, key matrix and value matrix; At the same time, the query matrix, key matrix and value matrix of the local features of protein sequence and ligand SMILES sequence are constructed: ; ; After obtaining the query matrix, key matrix, and value matrix for each feature, the attention features of all attention heads are concatenated along the channel dimension and used The function calculates the global and local attention matrices of protein sequences and ligand SMILES sequences, where The global attention matrix for protein sequences is: ; The local attention matrix of the protein sequence is also obtained , the global attention matrix of the ligand SMILES sequence , local attention matrix of ligand SMILES sequence ; The global and local attention matrices of the protein sequence and the ligand SMILES sequence are residually connected with the global features and local features of the corresponding enhanced protein sequence and the ligand SMILES sequence, and then input into the second feed-forward network after layer normalization. The second feed-forward network includes two fourth fully connected layers, and the two fourth fully connected layers include Activation function; The attention matrix after the second feedforward network is residually connected with the matrix before entering the second feedforward network, and finally the global feature vectors of the protein sequence containing interaction information are obtained through global maximum pooling. , protein sequence local feature vector , global feature vector of ligand SMILES sequence , local feature vector of ligand SMILES sequence ; Finally, the four features , , , Perform horizontal splicing to form fusion features ,in ; Use multi-head attention mechanism and the third feedforward network to integrate features Processing, where the multi-head attention mechanism is expressed as: ; in, Represents the features after the multi-head attention mechanism, represents the number of attention heads, Indicates The output of an attention head is represents the output transformation matrix, The output of each attention head is expressed as: ; in, , and They are The query, key, and value transformation matrices for each attention head, is the dimension of the key vector; The fusion features And the features after the multi-head attention mechanism Perform residual connection , and then the matrix Stretch to size dimensional column vector, normalized by layer Then the column vector is input by The third feed-forward network consisting of activation function and two fifth fully connected layers , and get the final output And as the fused feature, the formula is expressed as: ; Flatten represents a matrix stretching operation. It is expressed as: ; in, and The third feedforward network is The linear transformation matrices of the two fifth fully connected layers, and , .

[0013] In one embodiment of the present invention, the method of performing feature prediction on the fused features in step S5 to obtain the value of protein-ligand binding affinity includes: The fused features Input the multilayer perceptron to get the predicted value , to predict protein-ligand binding affinity, where the multilayer perceptron includes layer normalization , linear operation, The activation function is expressed as: ; in, , , are the linear transformation matrices of the three linear operations of the multilayer perceptron, and , , .

[0014] In order to solve the above technical problems, the present invention provides a prediction system for protein-ligand binding affinity, comprising: Acquisition module: used for acquiring a protein-ligand binding affinity data set, wherein the protein-ligand binding affinity data set includes a protein sequence and a ligand SMILES sequence; Feature extraction module: used to extract global features and local features of protein sequences and ligand SMILES sequences; Feature enhancement module: used to enhance the global features and local features of the protein sequence and the ligand SMILES sequence, respectively, to obtain the enhanced global features and local features of the protein sequence and the ligand SMILES sequence; Feature fusion module: used to fuse the global features and local features of the enhanced protein sequence and ligand SMILES sequence to obtain the fused features; Prediction module: used to predict the features after fusion and obtain the value of protein-ligand binding affinity.

[0015] To solve the above technical problems, the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method for predicting protein-ligand binding affinity as described above are implemented.

[0016] To solve the above technical problems, the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above protein-ligand binding affinity prediction method are implemented.

[0017] The above technical solution of the present invention has the following advantages compared with the prior art:

[0018] The method for predicting protein-ligand binding affinity described in the present invention solves the problems of high cost of traditional biological experimental methods and low accuracy of existing calculation models. The present invention can effectively improve the prediction accuracy of protein-ligand binding affinity by deeply exploring the characteristics of protein sequences and ligand SMILES sequences. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to make the contents of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings.

[0020] Figure 1 is a flow chart of the method of the present invention;

[0021] Figure 2 Schematic diagram of the structure of the PLMAM-PLA model in an embodiment of the present invention;

[0022] Figure 3 It is a schematic diagram of the structure of the protein language model ESM-2 and the molecular language model MolFormer in the PLMAM-PLA model of an embodiment of the present invention;

[0023] Figure 4Schematic diagram of the structure of the selective kernel network SKNet in the PLMAM-PLA model of an embodiment of the present invention;

[0024] Figure 5 Schematic diagram of the structure of the squeeze-excitation network SENet in the PLMAM-PLA model of an embodiment of the present invention;

[0025] Figure 6 It is a schematic diagram of the structure of the cross-attention module and the self-attention module in the PLMAM-PLA model of an embodiment of the present invention. DETAILED DESCRIPTION

[0026] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.

[0027] Embodiment 1

[0028] Reference Figure 1 The present invention relates to a method for predicting protein-ligand binding affinity, comprising: Step S1: collecting a protein-ligand binding affinity data set, wherein the protein-ligand binding affinity data set includes a protein sequence and a ligand SMILES sequence; Step S2: extracting global features and local features of protein sequences and ligand SMILES sequences; Step S3: enhancing the global features and local features of the protein sequence and the ligand SMILES sequence respectively, and obtaining the enhanced global features and local features of the protein sequence and the ligand SMILES sequence; Step S4: performing feature fusion on the global features and local features of the enhanced protein sequence and the ligand SMILES sequence to obtain fused features; Step S5: Perform feature prediction on the fused features to obtain the value of protein-ligand binding affinity.

[0029] The following is a detailed introduction to this embodiment: According to the technical solution of the present invention, a specific implementation case is given, such as Figure 2 As shown, the model in this implementation case is named PLMAM-PLA (parameter information is shown in Tables 1 and 2).

[0030] Step S1: Collecting protein-ligand binding affinity data set.

[0031] This example targets samples from the three data sets (universal set, refined set, and core set) in the 2016 version of PDBbind, selects the amino acid sequence of the protein (hereinafter referred to as the protein sequence) and the SMILES coding sequence of the small molecule ligand (hereinafter referred to as the ligand SMILES sequence) as input, and selects the protein-ligand binding affinity value as output. Among them, the number of samples in the universal set, refined set, and core set are 9221, 3685, and 290, respectively. Then, 1000 samples are randomly selected from the refined set as the validation set, the remaining samples in the refined set and the entire universal set are used as the training set, and the entire core set is used as the test data set. After deleting the duplicate samples of the three data sets, a total of 11906 training samples, 1000 validation samples, and 290 test samples are generated. Finally, the lengths of the protein sequence and the ligand SMILES sequence are set to be respectively. and , truncate the sequence part that exceeds the fixed length, otherwise fill it with 0.

[0032] Step S2: Extract global and local features of the input protein sequence and ligand SMILES sequence (see Figure 3 ).

[0033] This embodiment uses a pre-trained language model to extract global features of input protein sequences and ligand SMILES sequences. Specifically, a protein language model ESM-2 is constructed and trained, and the pre-trained protein language model ESM-2 is used to extract protein feature representations from protein sequences. ESM-2 is pre-trained on the basis of 250 million amino acid sequences using rotational position encoding, which can effectively predict and analyze the function and structure of proteins. A molecular language model MolFormer is constructed and trained, and the pre-trained molecular language model MolFormer is used to derive the molecular feature representation of the ligand from the ligand SMILES sequence. The model uses rotational position encoding and linear attention mechanism, and is combined with highly distributed pre-training of SMILES sequences of 1.1 billion unlabeled molecules from PubChem and ZINC data sets. In the pre-trained language models ESM-2 and MolFormer, protein sequences and ligand SMILES sequences are both segmented at the single-character level, where each amino acid in the protein sequence and each chemical symbol in the ligand SMILES sequence are used as a token. This word segmentation method enables the model to capture the local features of amino acids in proteins and the details of molecular structure in a more fine-grained manner, thereby providing richer information for understanding the structure and function of biological molecules. Global information needs to rely on the deep structure and self-attention mechanism of the pre-trained model to more accurately understand and characterize the structural and functional characteristics of biological molecules. The token sequence (i.e., the protein feature representation extracted by the protein language model ESM-2 and the ligand molecule feature representation extracted by the molecular language model MolFormer) is token encoded, rotated position encoded, and converted into a target input sequence that conforms to the model, and input into multiple series-connected Transformer encoders (from the Transformer model) to extract the long-distance dependencies of the sequence, wherein the Transformer encoder includes a self-attention sub-mechanism and a first feed-forward network (feed-forward network), the first feed-forward network includes two layers of linear transformation Linear and ReLU activation function, and the ReLU activation function is located between the two layers of linear transformation Linear, and the self-attention sub-mechanism and the first feed-forward network both include a residual connection and a layer normalization. Thus, the protein global feature matrix is ​​obtained. and the ligand molecule global feature matrix , and as a global feature of protein sequence and ligand SMILES sequence, where Enter the code length for the protein sequence, is the protein word vector dimension, is the length of the ligand SMILES sequence, is the ligand word vector dimension.

[0034] At the same time, this embodiment uses a convolutional neural network to extract local features of the input protein sequence and the ligand SMILES sequence, as follows: First, the protein sequence and ligand SMILES sequence are integer encoded according to the vocabulary, and the relevant information of the vocabulary can be obtained from the DeepDTA model. Then, the integer-encoded protein sequence and ligand SMILES sequence are passed through the embedding layer, which can use Pytorch's nn.Embedding module to convert the integer-encoded sequence into a dense two-dimensional output matrix and ,in, Enter the code length for the protein sequence, is the length of the ligand SMILES sequence, is the embedding dimension. Finally, the protein local feature matrix is ​​obtained by using the first and second dilated convolution submodules respectively. and the ligand local feature matrix , and serve as local features of protein sequences and ligand SMILES sequences.

[0035] The dilated convolution submodule (the first dilated convolution submodule) for extracting the protein local feature matrix is ​​composed of A series of first dilated convolution blocks, each of which includes Activation function, convolution operation , Activation function, convolution operation and residual connections, and The expansion rate is ,in , then The operation in the first dilated convolution block can be expressed as: ; in, , .

[0036] The dilated convolution submodule (second dilated convolution submodule) for extracting the local feature matrix of the ligand is composed of A series of second dilated convolution blocks, each of which includes Activation function, convolution operation , Activation function, Convolution operations and residual connections, and The expansion rate is ,in , then The operation in the second dilated convolution block can be expressed as: ; in, , .

[0037] Step S3: Strengthen the global and local features of protein sequence and ligand SMILES sequence.

[0038] This example uses the selective kernel network (SKNet) to enhance the global features of protein sequences and ligand SMILES sequences (see Figure 4 ). The following takes the global features of enhanced protein sequences as an example to discuss the implementation of SKNet.

[0039] First, the separation operation is performed: the protein global feature matrix obtained by the pre-trained model ESM-2 Through two parallel convolution operations with different dilation rates and , to capture multi-scale semantic information, where the convolution operation The expansion rate is , convolution operation The expansion rate is , and obtain the feature matrix and .

[0040] Secondly, perform the fusion operation: transform the feature matrix and Add together to generate the feature matrix Then the feature matrix Perform global average pooling The channel characteristics obtained , and then pass through the first fully connected layer and the second fully connected layer to obtain , where batch normalization BatchNorm and The activation function is expressed as: , ; in, and are the linear transformation matrices of the first and second fully connected layers, respectively, and and ; It should be noted that the linear transformation matrix is ​​a linear operation.

[0041] Finally, make a selection: Input to Layer, for and Assign a weight based on The weights obtained by the layer are broadcasted using the feature matrix and Make Hadamard product with the corresponding weight and sum to get the final output feature ,Will As the global feature of the enhanced protein sequence, the formula is expressed as: ; in, yes The attention weight of the first channel of the layer, yes The attention weight of the second channel of the layer; The operation of the layer can be expressed as: , .

[0042] In a similar way, this embodiment uses the selective kernel network SKNet to enhance the global features of the ligand SMILES sequence and obtains ,Will As the global feature of the enhanced ligand SMILES sequence, the formula is expressed as: .

[0043] At the same time, this example uses the squeeze-excitation network (SENet) to enhance the local features of the protein sequence and the ligand SMILES sequence (see Figure 5 ). Taking the enhancement of local features of protein sequences as an example, the implementation method of SENet is discussed.

[0044] First, perform the extrusion operation: convert the local feature matrix Perform global average pooling , get channel characteristics ,Right now: ; Secondly, perform the excitation operation: transform the channel features After two third fully connected layers, the importance weight of the channel can be obtained ; Among them, the third fully connected layer is composed of a linear transformation matrix , Activation function, linear transformation matrix and The activation function is composed of the following formula: ; Finally, a weighted operation is performed: the protein local feature matrix is ​​broadcasted using the broadcast mechanism and the corresponding channel weight Do the Hadamard product to get the final output feature ,Will As the local feature of the enhanced protein sequence, the formula is expressed as: .

[0045] In a similar way, the selective kernel network SKNet is used to enhance the local features of the ligand SMILES sequence and obtain ,Will As the local feature of the enhanced ligand SMILES sequence, the formula is expressed as: .

[0046] Step S4: feature fusion of the global features and local features of the enhanced protein sequence and ligand SMILES sequence.

[0047] First, we use the cross-attention module to learn the interaction features between proteins and ligands (see Figure 6 (A) in the figure), and then use the self-attention module for further feature fusion (see Figure 6 (B) in the figure.

[0048] See also Figure 6 In (A), the cross-attention module is used to learn the interaction features between proteins and ligands. Taking the cross-attention of the global features of protein sequence and ligand SMILES sequence as an example, the specific implementation method is as follows: The enhanced protein sequence global features Input three linear layers with different weight matrices to calculate the query matrix of global features of protein sequence , the bond matrix Sum Matrix ; The global features of the enhanced ligand SMILES sequence Input three linear layers with different weight matrices to calculate the query matrix of the global characteristics of the ligand SMILES sequence , the bond matrix Sum Matrix Then the key matrix and value matrix of the global features of the protein sequence and the ligand SMILES sequence are exchanged. The query matrix, key matrix and value matrix of the global features of the updated protein sequence are given by the following formula: ; in, , , It is The linear transformation matrix of the attention head query matrix, key matrix and value matrix, is the dimension of each attention head key vector, is the number of attention heads. Similarly, the query matrix, bond matrix, and value matrix of the global features of the updated ligand SMILES sequence are given by the following formulas: ; in, , , It is The linear transformation matrix of the attention head query matrix, key matrix and value matrix.

[0049] In a similar way, we can obtain the query matrix, key matrix, and value matrix of the local features of protein sequences and ligand SMILES sequences, where the linear transformation matrices of global and local features share the same set of parameters: ; .

[0050] After obtaining the query matrix, key matrix, and value matrix for each feature, the attention features of all attention heads are concatenated along the channel dimension and used The function calculates the global and local attention matrices of protein sequences and ligand SMILES sequences. For example, the global attention matrix of the protein sequence is: .

[0051] In a similar way, the local attention matrix of the protein sequence can be obtained , the global attention matrix of the ligand SMILES sequence , local attention matrix of ligand SMILES sequence Next, the global and local attention matrices of the protein sequence and ligand SMILES sequence are residually connected with the global features and local features of the corresponding enhanced protein sequence and ligand SMILES sequence, and input into the second feedforward network after layer normalization. Taking the global features of the protein sequence as an example, the above operation can be expressed as: ; The second feed-forward network consists of two fourth fully connected layers, and the two fourth fully connected layers include Then, the attention matrix after the second feedforward network is residually connected with the matrix before entering the second feedforward network, and finally the global feature vector of the protein sequence containing the interaction information is obtained through the global maximum pooling MaxPool. , protein sequence local feature vector , global feature vector of ligand SMILES sequence , local feature vector of ligand SMILES sequence For example, taking the global features of protein sequences as an example, the formula is expressed as: ; in, represents the global feature vector of proteins including interaction information, represents the second feed-forward network and It can be expressed as: ; in, and The second feedforward network is The linear transformation matrix of the two fully connected layers of , .

[0052] Finally, see Figure 6 In (B), the self-attention module is used to fuse the global and local features of proteins and ligands including interaction information. First, these four features , , , Horizontal splicing to form fusion features ,in Then, the self-attention mechanism is used to fusion features For processing, the self-attention module consists of a multi-head attention mechanism and a third feedforward network. The multi-head self-attention mechanism aggregates information in the vector space, while learning their weights and weighting them to capture higher-level semantic information of the input features. The multi-head attention mechanism is expressed as: ; in, Represents the features after the multi-head attention mechanism, represents the number of attention heads, Indicates The output of an attention head is represents the output transformation matrix, The output of each attention head is expressed as: ; in, , and They are The query, key and value transformation matrices of the attention heads, is the dimension of the key vector. The characteristics after the multi-head attention mechanism are , the fusion features And the features after multi-head attention Perform residual connection , and then the matrix Stretch to size dimensional column vector, and then normalized by layer Then the column vector is input by The third feed-forward network consisting of activation function and two fifth fully connected layers , and get the final output And as the fused feature, the formula is expressed as: ; Flatten represents a matrix stretching operation. It can be expressed as: ; in, and The third feedforward network is The linear transformation matrices of the two fifth fully connected layers, and , .

[0053] Step S5: Perform feature prediction on the fused features and output the protein-ligand binding affinity prediction results.

[0054] Input features Output the predicted value through the multi-layer perceptron , predicting protein-ligand binding affinity. Among them, the multi-layer perceptron includes layer normalization , Operation (retention probability is ), linear operations, Activation function. The formula is: ; in, , , are the linear transformation matrices of the three linear operations of the multilayer perceptron, and , , .

[0055] Step S6: Use the data set to train and test the model, and by observing the performance of the model on the validation set, determine when to stop training and evaluate the impact of hyperparameters such as learning rate, batch size, hidden layer size on model performance, and adjust the model's hyperparameters to optimize model performance. Input the training set in the protein-ligand binding affinity data set into the PLMAM-PLA model, calculate the mean square error loss function between the model's predicted value and the true value, and use the AdamW optimization method to update the model parameters. The mean square error loss function is: ; in, is the true value of protein-ligand binding affinity, is the predicted value output by the model, is the number of training batches. Finally, the trained model is tested on the test set and compared with other models.

[0056] The hyperparameters in the training process are shown in Table 2. After the training is completed, the model is tested using the test set in the dataset and compared with other representative models in terms of mean square error RMES, mean absolute error MAE, standard deviation SD, Pearson correlation coefficient R, and consistency coefficient CI evaluation indicators.

[0057] Table 1

[0058]

[0059] Table 2

[0060]

[0061] Representative models include: DeepDTA model that uses CNN to learn protein and ligand sequence features respectively.

[0062] The DeepDTAF model uses CNN to extract features using global protein features, local pocket features and ligand features as input; the Pafnucy model and OnionNet model use 3D-CNN to extract features using the cubic grid box in the atomic coordinates of the protein pocket-ligand complex as input; the DataDTA model uses a dual interaction aggregation neural network strategy to effectively learn multi-scale interaction features; the DPLA model integrates multi-level information such as sequence representation of proteins and ligands, structural property representation of amino acids in proteins and protein binding pockets, MACC key ligand molecular fingerprints and ligand molecular network characteristics; and the Happy model extracts information from ligand maps, pocket maps and interaction maps.

[0063] The performance comparison results of the PLMAM-PLA model and other models are shown in Table 3. It can be seen that the PLMAM-PLA model constructed according to the content of the present invention is superior to other benchmark models in most performances. Specifically, the RMSE, MAE, SD, R and CI values ​​of the PLMAM-PLA model are 1.223, 0.971, 1.185, 0.838 and 0.820, respectively. Among them, the mean square error RMSE, standard deviation SD, and Pearson correlation coefficient R are improved by 0.4%, 2.9%, and 1.3% respectively compared with the model ranked second in performance. In addition, the PLMAM-PLA model is a model based only on sequences, which is 9.7%, 9.5%, 11.3%, 6.2%, and 2.6% higher than the current best sequence model DeepDTAF in mean square error RMSE, mean absolute error MAE, standard deviation SD, Pearson correlation coefficient R and consistency coefficient CI. Among them, the Happy model and DeepDTA model are the protein-ligand binding affinity prediction models with the best and worst performance among the seven benchmark models, respectively. The results show that the PLMAM-PLA model fully exploits the protein sequence and ligand SMILES sequence information to improve the ability to predict protein-ligand affinity. In addition, the PLMAM-PLA model only extracts feature information from protein sequence and ligand SMILES sequence to predict protein-ligand affinity, but its prediction performance is better than that of structure-based and multimodal input-based models (DeepDTAF, OnionNet, Pafnucy, DataDTA, DPLA, Happy).

[0064] Table 3

[0065]

[0066] Embodiment 2

[0067] This embodiment provides a prediction system for protein-ligand binding affinity, comprising: Acquisition module: used for acquiring a protein-ligand binding affinity data set, wherein the protein-ligand binding affinity data set includes a protein sequence and a ligand SMILES sequence; Feature extraction module: used to extract global features and local features of protein sequences and ligand SMILES sequences; Feature enhancement module: used to enhance the global features and local features of the protein sequence and the ligand SMILES sequence, respectively, to obtain the enhanced global features and local features of the protein sequence and the ligand SMILES sequence; Feature fusion module: used to fuse the global features and local features of the enhanced protein sequence and ligand SMILES sequence to obtain the fused features; Prediction module: used to predict the features after fusion and obtain the value of protein-ligand binding affinity.

[0068] Embodiment 3

[0069] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for predicting protein-ligand binding affinity described in Embodiment 1 are implemented.

[0070] Embodiment 4

[0071] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method for predicting protein-ligand binding affinity described in the first embodiment are implemented.

[0072] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present application may be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.

[0073] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0074] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1A function specified in one or more boxes.

[0075] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0076] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0077] Obviously, the above embodiments are merely examples for clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived from these are still within the protection scope of the invention.

Claims

1. A method for predicting protein-ligand binding affinity, characterized in that: include: Step S1: collecting a protein-ligand binding affinity data set, wherein the protein-ligand binding affinity data set includes a protein sequence and a ligand SMILES sequence; Step S2: extracting global features and local features of protein sequences and ligand SMILES sequences; Step S3: enhancing the global features and local features of the protein sequence and the ligand SMILES sequence respectively, and obtaining the enhanced global features and local features of the protein sequence and the ligand SMILES sequence; Step S4: performing feature fusion on the global features and local features of the enhanced protein sequence and the ligand SMILES sequence to obtain fused features; Step S5: Perform feature prediction on the fused features to obtain the value of protein-ligand binding affinity.

2. The method for predicting protein-ligand binding affinity according to claim 1, characterized in that: The method for extracting the global features of the protein sequence and the ligand SMILES sequence in step S2 includes: Use the pre-trained protein language model ESM-2 to extract protein feature representation from protein sequences; The pre-trained molecular language model MolFormer is used to extract the ligand molecular feature representation from the ligand SMILES sequence; The protein feature representation and ligand molecule feature representation are respectively input into multiple series-connected Transformer encoders to obtain the protein global feature matrix and the ligand molecule global feature matrix , and as a global feature of protein sequence and ligand SMILES sequence, where Enter the code length for the protein sequence, is the protein word vector dimension, is the length of the ligand SMILES sequence, is the ligand word vector dimension.

3. The method for predicting protein-ligand binding affinity according to claim 1, characterized in that: The method for extracting local features of the input protein sequence and the ligand SMILES sequence in step S2 includes: The protein sequence and ligand SMILES sequence are integer encoded, and the integer encoded protein sequence and ligand SMILES sequence are input into the embedding layer. The embedding layer converts the integer encoded protein sequence and ligand SMILES sequence into a two-dimensional output matrix through the nn.Embedding module of Pytorch and ,in, Enter the code length for the protein sequence, is the length of the ligand SMILES sequence, is the embedding dimension; The matrix and Input the first and second dilated convolution submodules respectively to obtain the protein local feature matrix and the ligand local feature matrix , and as local features of protein sequences and ligand SMILES sequences, where The first dilated convolution submodule includes A series of first dilated convolution blocks, each of which includes Activation function, convolution operation , Activation function, convolution operation And residual connection, convolution operation And convolution operation The expansion rate is ,in , then The first dilated convolution block is expressed as: ; in, , ; The second dilated convolution submodule includes A series of second dilated convolution blocks, each of which includes Activation function, convolution operation , Activation function, Convolution operation and residual connection, convolution operation And convolution operation The expansion rate is ,in , then The second dilated convolution block is expressed as: ; in, , .

4. The method for predicting protein-ligand binding affinity according to claim 2, characterized in that: The method for enhancing the global features of the protein sequence and the ligand SMILES sequence in step S3 includes: The selective kernel network SKNet is used to enhance the global features of protein sequences, including: Protein global feature matrix Through two parallel convolution operations with different dilation rates and , and obtain the feature matrix and , where the convolution operation The expansion rate is , convolution operation The expansion rate is ; The feature matrix and Add together to generate the feature matrix , the feature matrix Perform global average pooling Get channel characteristics , and the channel features Through the first fully connected layer and the second fully connected layer, we can get , where the first fully connected layer includes batch normalization BatchNorm and The activation function is expressed as: , ; in, and are the linear transformation matrices of the first and second fully connected layers, respectively, and and ; Will Input to Layer, for and Assign a weight based on The weights assigned to the layers are broadcasted using the feature matrix and Make Hadamard product with the corresponding weight and sum to get the final output feature ,Will As the global feature of the enhanced protein sequence, the formula is expressed as: ; in, yes The attention weight of the first channel of the layer, yes The attention weight of the second channel of the layer, The operation of the layer is expressed as: , ; At the same time, the selective kernel network SKNet is used to enhance the global characteristics of the ligand SMILES sequence, and the ,Will As the global feature of the enhanced ligand SMILES sequence, the formula is expressed as: 。 5. The method for predicting protein-ligand binding affinity according to claim 3, characterized in that: The method for enhancing the local features of the protein sequence and the ligand SMILES sequence in step S3 includes: The squeeze-excitation network SENet is used to enhance the local features of protein sequences, including: Protein local feature matrix Perform global average pooling , get channel characteristics , the formula is: ; Channel Features After two third fully connected layers, the importance weight of the channel is obtained , where each third fully connected layer includes a linear transformation matrix , Activation function, linear transformation matrix and The activation function is expressed as: ; The protein local feature matrix is ​​broadcasted using the broadcast mechanism and the corresponding channel weight Do the Hadamard product to get the final output feature ,Will As the local feature of the enhanced protein sequence, the formula is expressed as: ; At the same time, the selective kernel network SKNet is used to enhance the local features of the ligand SMILES sequence, and the ,Will As the local feature of the enhanced ligand SMILES sequence, the formula is expressed as: 。 6. The method for predicting protein-ligand binding affinity according to claim 1, characterized in that: In step S4, the method of fusing the global features and local features of the enhanced protein sequence and the ligand SMILES sequence to obtain the fused features includes: The enhanced protein sequence global features Input three different linear layers to calculate the query matrix of global features of protein sequence , key matrix Sum Matrix ; The global features of the enhanced ligand SMILES sequence Input three different linear layers to calculate the query matrix of the global features of the ligand SMILES sequence , key matrix Sum Matrix ; The bond matrix and value matrix of the global features of the protein sequence and the ligand SMILES sequence are exchanged; The query matrix, key matrix, and value matrix of the global features of the updated protein sequence are expressed as: ; in, , , It is The linear transformation matrix of the attention head query matrix, key matrix and value matrix, is the dimension of each attention head key vector, is the number of attention heads; The query matrix, bond matrix, and value matrix of the global features of the updated ligand SMILES sequence are expressed as: ; in, , , It is The linear transformation matrix of the attention head query matrix, key matrix and value matrix; At the same time, the query matrix, key matrix and value matrix of the local features of protein sequence and ligand SMILES sequence are constructed: ; ; After obtaining the query matrix, key matrix, and value matrix for each feature, the attention features of all attention heads are concatenated along the channel dimension and used The function calculates the global and local attention matrices of protein sequences and ligand SMILES sequences, where The global attention matrix for protein sequences is: ; The local attention matrix of the protein sequence is also obtained , the global attention matrix of the ligand SMILES sequence , local attention matrix of ligand SMILES sequence ; The global and local attention matrices of the protein sequence and the ligand SMILES sequence are residually connected with the global features and local features of the corresponding enhanced protein sequence and the ligand SMILES sequence, and then input into the second feed-forward network after layer normalization. The second feed-forward network includes two fourth fully connected layers, and the two fourth fully connected layers include Activation function; The attention matrix after the second feedforward network is residually connected with the matrix before entering the second feedforward network, and finally the global feature vectors of the protein sequence containing interaction information are obtained through global maximum pooling. , protein sequence local feature vector , global feature vector of ligand SMILES sequence , local feature vector of ligand SMILES sequence ; Finally, the four features , , , Perform horizontal splicing to form fusion features ,in ; Use multi-head attention mechanism and the third feedforward network to integrate features Processing, where the multi-head attention mechanism is expressed as: ; in, Represents the features after the multi-head attention mechanism, represents the number of attention heads, Indicates The output of an attention head is represents the output transformation matrix, The output of each attention head is expressed as: ; in, , and They are The query, key, and value transformation matrices for each attention head, is the dimension of the key vector; The fusion features And the features after the multi-head attention mechanism Perform residual connection , and then the matrix Stretch to size dimensional column vector, normalized by layer Then the column vector is input by The third feed-forward network consisting of activation function and two fifth fully connected layers , and get the final output And as the fused feature, the formula is expressed as: ; Flatten represents a matrix stretching operation. It is expressed as: ; in, and The third feedforward network is The linear transformation matrices of the two fifth fully connected layers, and , .

7. The method for predicting protein-ligand binding affinity according to claim 1, characterized in that: The method of performing feature prediction on the fused features in step S5 to obtain the value of protein-ligand binding affinity includes: The fused features Input the multilayer perceptron to get the predicted value , to predict protein-ligand binding affinity, where the multilayer perceptron includes layer normalization , linear operation, The activation function is expressed as: ; in, , , are the linear transformation matrices of the three linear operations of the multilayer perceptron, and , , .

8. A protein-ligand binding affinity prediction system, characterized in that: include: Acquisition module: used for acquiring a protein-ligand binding affinity data set, wherein the protein-ligand binding affinity data set includes a protein sequence and a ligand SMILES sequence; Feature extraction module: used to extract global features and local features of protein sequences and ligand SMILES sequences; Feature enhancement module: used to enhance the global features and local features of the protein sequence and the ligand SMILES sequence, respectively, to obtain the enhanced global features and local features of the protein sequence and the ligand SMILES sequence; Feature fusion module: used to fuse the global features and local features of the enhanced protein sequence and ligand SMILES sequence to obtain the fused features; Prediction module: used to predict the features after fusion and obtain the value of protein-ligand binding affinity.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method for predicting protein-ligand binding affinity according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for predicting protein-ligand binding affinity according to any one of claims 1 to 7 are implemented.