Drug-target interaction prediction method based on large language model and image representation

By using large language model and image characterization technology, two-dimensional image representation of drugs and targets are generated and prediction models of two-channel convolutional neural networks are constructed, and the problem of insufficient prediction accuracy and generalization ability of drug-target interactions in the existing technology is solved, achieving more efficient and accurate prediction effects.

CN119993257AActive Publication Date: 2025-05-13ZHEJIANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510068501.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

Existing drug-target interaction prediction models rely on known protein structures and inaccurate characterization, resulting in insufficient prediction accuracy and generalization capabilities.

Method used

Using a method based on large language model and image representation, the target sequence and drug SMILES are encoded through ESM-2 and X-MOL, a feature matrix is ​​generated, and image characterization technology is used to convert it into two-dimensional image representation to construct a drug-target interaction prediction model MapCPI of a two-channel convolutional neural network.

Benefits of technology

It improves the accuracy and efficiency of drug-target interaction prediction, shows good generalization performance, and can effectively predict the interaction of unresolved structural proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993257A_ABST
    Figure CN119993257A_ABST
Patent Text Reader

Abstract

The invention discloses a drug-target interaction prediction method based on a large language model and image characterization. The method comprises the following steps: coding a target sequence by using a large language model ESM-2, coding a drug SMILES by using a large language model X-MOL, generating a feature matrix, and generating two-dimensional image characterization of a drug and a target through an image characterization technology; and constructing a drug-target interaction prediction model MapCPI, inputting the two-dimensional image characterization of the drug and the target into the MapCPI to obtain a drug-target interaction prediction probability, and judging whether the drug-target is combined or not. The model can effectively extract drug-target interaction characteristics by using the coding information without depending on target structure information, so that prediction of drug-target interaction is realized, and good generalization performance is shown. According to the method, the accuracy and efficiency of drug-target interaction prediction are improved by combining the large language model and the image characterization technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of drug-target affinity prediction, and in particular relates to a drug-target interaction prediction method based on a large language model and image representation. Background Art

[0002] Understanding and predicting the interaction between drugs and target proteins is a crucial research area in the process of drug discovery and development. Although traditional experimental methods can provide reliable data, they are expensive, time-consuming, and require huge resources, which limits the implementation of large-scale screening. Using computational methods to predict the interaction between drugs and target proteins in the early stages of drug development is expected to improve the success rate of drug discovery.

[0003] In recent years, with the rapid development of big data and computer technology, artificial intelligence methods have been widely used in the development of compound-protein interaction (CPI) prediction models. However, existing CPI models generally rely on known ligand information of proteins, have weak generalization capabilities and have several limitations. First, some existing methods rely on the three-dimensional structure of known proteins, and for proteins with unresolved structures, these methods cannot effectively predict their interactions. Secondly, the protein and compound representations in existing sequence-based CPI prediction methods are not accurate enough, and the correlation information between features is ignored, and the potential relationship between encoding features is not paid attention to. These problems have led to the lack of prediction accuracy and generalization ability of existing CPI models, limiting their application in real drug discovery. Summary of the invention

[0004] The present invention proposes a drug-target interaction prediction method based on a large language model and image representation, aiming to overcome the limitations of existing methods and provide a more accurate prediction tool for drug discovery.

[0005] A drug-target interaction prediction method based on a large language model and image representation comprises the following steps:

[0006] S1: Collect drug-target interaction benchmark datasets and obtain drug SMILES and target sequences;

[0007] S2: Use the large language model ESM-2 to encode the target sequence, use the large language model X-MOL to encode the drug SMILES, generate a feature matrix, and generate a two-dimensional image representation of the drug and target through image representation technology;

[0008] S3: Construct a drug-target interaction prediction model MapCPI, input the two-dimensional image representation of drugs and targets into MapCPI, and obtain the drug-target interaction prediction probability;

[0009] S4: Determine whether drug-target binding occurs based on the predicted probability of drug-target interaction.

[0010] In step S1, the drug-target interaction benchmark dataset includes a BindingDB dataset and a Human dataset.

[0011] In step S2, the target sequence is encoded using the large language model ESM-2, and the drug SMILES is encoded using the large language model X-MOL to generate a feature matrix, and a two-dimensional image representation of the drug and target is generated through image representation technology, specifically including:

[0012] 2.1) Use the large language model ESM-2 to encode the protein sequences in the protein database to generate a protein encoding matrix, and use the large language model X-MOL to encode all bioactive compounds in the compound database to generate a compound encoding matrix. The original feature matrices of targets and drugs are converted into vectors through average pooling operations;

[0013] 2.2) The pairwise distances between protein coding matrices are calculated based on cosine similarity to generate a protein feature distance matrix. Similarly, the pairwise distances between compound coding matrices are calculated based on cosine similarity to generate a compound feature distance matrix. The calculation formula is:

[0014]

[0015] Where a and b represent different features, f a and f b Represents the eigenvalue vector composed of all samples on this dimension feature;

[0016] 2.3) reducing the dimension of the protein feature distance matrix or the compound feature distance matrix by uniform manifold approximation and projection algorithm, and projecting it into a two-dimensional space to obtain the scatter distribution of the protein feature or the scatter distribution of the compound feature respectively;

[0017] 2.4) Using the Jonker-Volgenant algorithm, the scattered distribution of protein features or the scattered distribution of compound features is linearly distributed to obtain a regularized protein template image or a regularized compound template image. The regularized protein template image and the regularized compound template image constitute a two-dimensional image representation of the drug and the target.

[0018] In step 2.1), the protein database is TTD database and UniProt database, and the compound database is ChEMBL database.

[0019] In step S3, a dual-channel convolutional neural network is used to construct a drug-target interaction prediction model MapCPI.

[0020] In step S4, whether the drug binds to the target is evaluated using a benchmark data set, which is the BindingDB and Human databases.

[0021] Furthermore, a drug-target interaction prediction method based on a large language model and image representation comprises the following steps:

[0022] S1: Collect drug-target interaction benchmark datasets and obtain drug SMILES and target sequences;

[0023] S2: Use large language models ESM-2 and X-MOL to encode target sequences and drug SMILES respectively, generate feature matrices, and generate two-dimensional image representations through image representation technology;

[0024] S3: Construct a drug-target interaction prediction model MapCPI, input the two-dimensional image representation of drugs and targets into MapCPI, and obtain the drug-target interaction prediction probability.

[0025] S4: Determine whether drug-target binding occurs based on the predicted probability of drug-target interaction.

[0026] In step S1, the data set includes BindingDB and Human data sets. First, the two data sets are divided into training set, validation set and test set respectively, wherein the BindingDB data set sets random partitioning of interaction pairs and unseen target partitioning, and the Human data set sets random partitioning of interaction pairs and cold start partitioning.

[0027] Furthermore, in step S2, the large language models ESM-2 and X-MOL are used to encode the target sequence and drug SMILES respectively, generate feature matrices, and realize the two-dimensional image representation of targets and drugs through the two steps of template image generation and image conversion.

[0028] The template image generation technology includes:

[0029] Large language model encoding: Use ESM-2 to encode large-scale protein sequences in the TTD (Therapeutic Target Database) and UniProt databases to generate encoding matrices, use X-MOL to encode all bioactive compounds in the ChEMBL database to generate encoding matrices, and convert the original feature matrices of targets and drugs into vectors through average pooling operations.

[0030] Feature distance calculation: The pairwise distances between protein coding features are calculated based on cosine similarity to generate a feature distance matrix; similarly, the pairwise distances between compound coding features are calculated based on cosine similarity to generate a feature distance matrix. The calculation formula is:

[0031]

[0032] Where a and b represent different features, f a and f b Represents the eigenvalue vector composed of all samples on this dimension feature.

[0033] Data dimensionality reduction: High-dimensional features are projected into two-dimensional space through the uniform manifold approximation and projection algorithm (UMAP), and the coordinates of the scattered points are calculated to generate the corresponding scattered point distribution by searching for the equivalent fuzzy topological structure that is closest to the original distance relationship;

[0034] Optimal grid allocation: Using the Jonker-Volgenant (JV) algorithm, the scattered point distribution of all features is linearly allocated to obtain a regularized template image. The template image records the spatial position of each feature. This process minimizes the square distance between the scattered points and the grid coordinates. The calculation formula is as follows:

[0035]

[0036] where x scatter ,y grid Represent the scattered coordinate matrix and grid position matrix of all features respectively;

[0037] The image conversion technology first performs large language model encoding on the drug-target interaction pair under study to obtain encoding vectors, and then maps the eigenvalues ​​of the encoding vectors to corresponding positions based on the template images obtained above, and finally obtains the image representation ESM2image (size 51×51) of the target and the image representation XMOLimage (size 28×28) of the drug.

[0038] Furthermore, in step S3, a dual-channel convolutional neural network is used to extract features from ESM2image and XMOLimage simultaneously. The first layer is a convolutional layer, using multiple convolution kernels with a stride of 1 (proteins use 13×13 convolution kernels and compounds use 9×9 convolution kernels), and the second layer is a maximum pooling layer with a stride of 2 (proteins use 5×5 pooling kernels and compounds use 3×3 pooling kernels). The third layer is a convolutional base, which contains three parallel convolution operations, each using a different convolution kernel (proteins use 1×1, 5×5 and 9×9 convolution kernels, and compounds use 1×1, 3×3 and 5×5 convolution kernels).

[0039] The outputs of the convolutional bases are concatenated and fed into a max pooling layer, followed by a second convolutional base to further enhance feature extraction, followed by a global max pooling to extract the final embedding representation. The output of each convolutional layer is introduced with a ReLU activation function to introduce non-linearity.

[0040] The protein and compound embeddings generated by the two-channel convolutional neural network are concatenated to obtain the embedding vector z of the interaction pair. The classifier consists of four fully connected layers. The outputs of the first three layers are nonlinearly transformed by the ReLU activation function, and the output of the last layer is normalized by the softmax function to obtain the probability of drug-target interaction.

[0041] Furthermore, in step S4, the present invention and five advanced drug-target interaction prediction models are compared and evaluated on the BindingDB and Human benchmark data sets, and the performance of the present invention is comprehensively evaluated by comprehensively considering evaluation indicators such as Matthews Correlation Coefficient (MCC), Area Under the Receiver Operating Characteristic Curve (AUROC), and Area Under the Precision-Recall Curve (AUPRC).

[0042] MCC: A comprehensive model performance evaluation indicator used to evaluate the overall prediction performance of a binary classification model. The value ranges from -1 to 1, where 1 indicates a completely correct prediction, 0 indicates a random prediction, and -1 indicates a completely wrong prediction. The calculation formula is:

[0043]

[0044] AUROC: It is used to measure the ability of the classification model to distinguish between positive and negative classes. It represents the area under the ROC curve. The closer its value is to 1, the better the model performance. The horizontal axis of the ROC curve is the false positive rate FPR, and the vertical axis is the true positive rate TPR. The calculation formula is:

[0045] AUROC = ∫0 1 TPR(FPR)d(FPR)

[0046]

[0047] AUPRC: It is used to measure the balance performance of the precision and recall of the classification model at different thresholds. The closer the value is to 1, the better the performance. The horizontal axis of the curve is the recall rate, and the vertical axis is the precision rate. The calculation formula is:

[0048] AUPRC=∫0 1 Precision(Recall)d(Recall)

[0049]

[0050] In the above formula, TP represents the number of true positive samples (the number of positive samples correctly predicted), FP represents the number of false positive samples (the number of negative samples incorrectly predicted as positive), TN represents the number of true negative samples (the number of negative samples correctly predicted), and FN represents the number of false negative samples (the number of positive samples incorrectly predicted as negative).

[0051] The present invention is used to evaluate the drug-target interaction from the Human database, with a predicted probability value ranging from 0 to 1, with 0.5 as the threshold, a predicted probability higher than 0.5 indicates that the drug-target can interact, and a predicted probability lower than 0.5 indicates that the drug-target cannot interact.

[0052] Compared with the prior art, the present invention has the following advantages:

[0053] The present invention abandons protein structure information and directly extracts interaction features from protein sequences and compound SMILES for drug-target interaction prediction; secondly, the present invention combines biological large language model encoding with image-like dimension expansion technology to innovatively perform image representation of proteins and compounds based on feature correlation. This image representation strategy aggregates the features of proteins and compounds encoded by the large language model according to the potential relationship between the encoding features, and can accurately characterize interacting proteins and compounds. The present invention develops a drug-target interaction prediction model by combining a large language model and image representation technology, improves prediction accuracy and efficiency, and exhibits good generalization performance. The present invention belongs to the field of drug-target affinity prediction technology and provides a new tool for drug discovery. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic diagram of the process of the drug-target interaction prediction method based on a large language model and image representation in the present invention;

[0055] Figure 2 A schematic diagram of compound / protein data embedding based on a large language model and an image representation method in the present invention;

[0056] Figure 3 It is a schematic diagram of the structure of the drug-target interaction prediction method based on a large language model and image representation in the present invention;

[0057] Figure 4 This is a visualization result image of the drug-target interaction prediction performed in the present invention. DETAILED DESCRIPTION

[0058] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0059] In this implementation case, Figure 1 As shown, the present invention provides a drug-target interaction prediction method MapCPI based on a large language model and image representation, comprising the following steps:

[0060] Step S1: Collect drug-target interaction benchmark datasets and obtain drug SMILES and target sequences

[0061] First, two drug-target interaction benchmark datasets, BindingDB and Human, were collected from public literature, and the SMILES and target sequence data of the drugs were obtained. The BindingDB dataset contains 39,747 positive interaction pairs and 31,218 negative interaction pairs, which are randomly divided into training, validation, and test sets, and a subset of interaction pairs with targets not seen in the training set is extracted from the test set. For the Human dataset, it is first randomly divided into training, validation, and test sets according to a ratio of 8:1:1. Secondly, a cold start partition of the Human dataset is also set. First, 5% and 10% of the random interaction pairs are placed in the validation and test sets respectively, and all interaction pairs that are repeated with drugs and targets in the validation or test sets are removed from the remaining 85% of the interaction pairs to form a training set.

[0062] Step S2: Use the large language models ESM-2 and X-MOL to encode the target sequence and drug SMILES respectively, generate a feature matrix, and generate a two-dimensional image representation through image representation technology;

[0063] For all drug-target interaction pairs, the drug and target are encoded by the large language model to obtain feature vectors, and then converted into two-dimensional image representation using image representation technology, that is, the drug and target are represented as XMOLimage and ESM2image respectively. Figure 2 As shown in Figure 1, the specific image representation process mainly includes two stages: template image generation and image conversion. The detailed steps are as follows:

[0064] 1) Template image generation stage, including large language model encoding, feature distance calculation, data dimension reduction and optimal grid allocation, as described below:

[0065] 1.1 Large language model encoding: The whole proteome sequences of all species where drug targets are located were collected from the TTD (Therapeutic Target Database) and UniProt databases, and all drug-like SMILES with activity data were collected from the ChEMBL database. The protein large language model ESM-2 and compound large language model X-MOL based on Transformer with a parameter scale of 3 billion were used to encode protein sequences and compound SMILES, respectively, to generate original feature matrices. Through average pooling operations, the original feature matrices of proteins and compounds were converted into encoding vectors of length 2,560 and 768, respectively. The 1,300,449 protein sequences collected were encoded with ESM-2 to obtain an encoding matrix of 1,300,449×2560, and the 2,231,654 bioactive compound SMILES collected were encoded with X-MOL to obtain an encoding matrix of 2,231,654×768.

[0066] 1.2 Feature distance calculation: Based on the encoding results of 1.3 million protein sequences, the pairwise distances between protein coding features were calculated based on cosine similarity to generate a 2,560×2,560 feature distance matrix; based on the encoding results of 2.2 million compound SMILES, the pairwise distances between compound coding features were calculated based on cosine similarity to generate a 768×768 feature distance matrix. The calculation formula for this process is:

[0067]

[0068] Where a and b represent different features, f a and f b Represents the eigenvalue vector composed of all samples on this dimension feature.

[0069] 1.3 Data dimensionality reduction: Based on the above feature distance matrix, the uniform manifold approximation and projection algorithm (UMAP) is used to project the high-dimensional features into two-dimensional space, and the equivalent fuzzy topological structure closest to the original distance relationship is searched to calculate the coordinates of the scattered points to generate the corresponding scattered point distribution.

[0070] 1.4 Optimal grid allocation: Using the Jonker-Volgenant (JV) algorithm, the scattered point distribution of all features is linearly allocated to obtain a regularized template image. The template image records the spatial position of each feature. This process minimizes the square distance between the scattered points and the grid coordinates. The calculation formula is as follows:

[0071]

[0072] where x scatter ,y gridRepresent the scatter coordinate matrix and grid position matrix of all features respectively.

[0073] 2) Image conversion stage

[0074] Firstly, the drug-target interaction pairs under study were encoded by ESM-2 and X-MOL large language models to obtain encoding vectors respectively. Based on the template images obtained above, the eigenvalues ​​of the encoding vectors were mapped to the corresponding positions respectively. Finally, the image representation ESM2image (size 51×51) of the target and the image representation XMOLimage (size 28×28) of the drug were obtained.

[0075] Step S3: construct a drug-target interaction prediction model MapCPI, input the two-dimensional image representation of the drug and the target into MapCPI, and obtain the drug-target interaction prediction probability.

[0076] Use a two-channel convolutional neural network to extract features from ESM2image and XMOLimage simultaneously, such as Figure 3 As shown in the figure. The first layer is a convolutional layer, using a variety of convolutional kernels with a stride of 1 (72 convolutional kernels of size 13×13 are used for proteins, and 48 convolutional kernels of size 9×9 are used for compounds), followed by a maximum pooling layer with a stride of 2 to reduce the computational cost (5×5 pooling kernels are used for proteins, and 3×3 pooling kernels are used for compounds). The third layer is a convolutional base, which contains three parallel convolution operations, each using a different convolutional kernel (1×1, 5×5, and 9×9 convolutional kernels are used for proteins, and 1×1, 3×3, and 5×5 convolutional kernels are used for compounds). The outputs of the convolutional bases are concatenated and input into a maximum pooling layer. Another convolutional base is then used to further enhance the representation capability. After the outputs of the second convolutional base are concatenated, global maximum pooling is applied to extract the final embedded representation. The output of each convolutional layer introduces nonlinear characteristics through the ReLU activation function.

[0077] The protein and compound embeddings generated by the two-channel convolutional neural network are concatenated into a joint embedding vector z for downstream drug-target interaction prediction, such as Figure 3 As shown in the figure, the classifier consists of four fully connected layers with 512, 256, 64 and 2 neurons respectively. The outputs of the first three layers are transformed nonlinearly by the ReLU activation function. The output of the last layer is normalized by the softmax function to obtain the probability of drug-target interaction. The binary cross entropy loss function is then used to simultaneously optimize the parameters of the dual-channel convolutional neural network and the classifier.

[0078] S4: Determine whether drug-target binding occurs based on the predicted probability of drug-target interaction.

[0079] The present invention comprehensively considers three evaluation indicators: Matthews Correlation Coefficient (MCC), Area Under the Receiver Operating Characteristic Curve (AUROC), and Area Under the Precision-Recall Curve (AUPRC). The larger the values ​​of these indicators, the better the model prediction performance.

[0080] MCC: A comprehensive model performance evaluation indicator used to evaluate the overall prediction performance of a binary classification model. The value ranges from -1 to 1, where 1 indicates a completely correct prediction, 0 indicates a random prediction, and -1 indicates a completely wrong prediction. The calculation formula is:

[0081]

[0082] AUROC: It is used to measure the ability of the classification model to distinguish between positive and negative classes. It represents the area under the ROC curve. The closer its value is to 1, the better the model performance. The horizontal axis of the ROC curve is the false positive rate FPR, and the vertical axis is the true positive rate TPR. The calculation formula is:

[0083] AUROC = ∫0 1 TPR(FPR)d(FPR)

[0084]

[0085] AUPRC: It is used to measure the balance performance of the precision and recall of the classification model at different thresholds. The larger the value, the better the performance. The horizontal axis of the curve is the recall rate Recall, and the vertical axis is the precision Precision. The calculation formula is:

[0086] AUPRC=∫0 1 Preision(Recall)d(Recall)

[0087]

[0088] In the above formula, TP represents the number of true positive samples (the number of positive samples correctly predicted), FP represents the number of false positive samples (the number of negative samples incorrectly predicted as positive), TN represents the number of true negative samples (the number of negative samples correctly predicted), and FN represents the number of false negative samples (the number of positive samples incorrectly predicted as negative).

[0089] In this embodiment, as shown in Table 1, the developed MapCPI model shows excellent prediction performance in the random partitioning tasks of both BindingDB and Human datasets, achieving the highest AUROC, AUPRC and MCC indicators in the BindingDB dataset, and the highest AUROC and AUPRC indicators in the Human dataset. This result illustrates the excellence and feasibility of the model proposed in the present invention. Figure 4 As shown in Table 1, the developed MapCPI model achieved the best prediction performance in MCC, AUROC and AUPRC in the unseen target segmentation of BindingDB and the cold start segmentation of Human, further verifying its excellent generalization ability in drug-target interaction prediction. These visualization results intuitively support the conclusions in Table 1, further proving the superiority of the model of the present invention in drug-target interaction prediction.

[0090] Table 1 Experimental data comparison of MapCPI model and other methods

[0091]

[0092] In addition, in this embodiment, as shown in Table 2, the prediction results of the MapCPI model for 10 drug-target interactions from the Human database are completely consistent with the results recorded in the database. This further proves the excellent performance and reliability of the model of the present invention in predicting drug-target interactions.

[0093] Table 2 Results of MapCPI model in the prediction of 10 drug-target interactions

[0094]

[0095]

[0096]

[0097] Finally, it should be noted that although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention.

[0098] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.

[0099] The application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the protection scope of the claims attached to the present invention.

Claims

1. A drug-target interaction prediction method based on a large language model and image representation, characterized in that: The following steps are involved: S1: Collect drug-target interaction benchmark datasets and obtain drug SMILES and target sequences; S2: Use the large language model ESM-2 to encode the target sequence, use the large language model X-MOL to encode the drug SMILES, generate a feature matrix, and generate a two-dimensional image representation of the drug and target through image representation technology; S3: Construct a drug-target interaction prediction model MapCPI, input the two-dimensional image representation of drugs and targets into MapCPI, and obtain the drug-target interaction prediction probability; S4: Determine whether drug-target binding occurs based on the predicted probability of drug-target interaction.

2. The drug-target interaction prediction method based on large language model and image representation according to claim 1, characterized in that: In step S1, the drug-target interaction benchmark dataset includes a BindingDB dataset and a Human dataset.

3. The drug-target interaction prediction method based on large language model and image representation according to claim 1, characterized in that: In step S2, the target sequence is encoded using the large language model ESM-2, and the drug SMILES is encoded using the large language model X-MOL to generate a feature matrix, and a two-dimensional image representation of the drug and target is generated through image representation technology, specifically including: 2.1) Use the large language model ESM-2 to encode the protein sequences in the protein database to generate a protein encoding matrix, and use the large language model X-MOL to encode all bioactive compounds in the compound database to generate a compound encoding matrix. The original feature matrices of targets and drugs are converted into vectors through average pooling operations; 2.2) The pairwise distances between protein coding matrices are calculated based on cosine similarity to generate a protein feature distance matrix. Similarly, the pairwise distances between compound coding matrices are calculated based on cosine similarity to generate a compound feature distance matrix. The calculation formula is: Where a and b represent different features, f a and f b Represents the eigenvalue vector composed of all samples on this dimension feature; 2.3) reducing the dimension of the protein feature distance matrix or the compound feature distance matrix by uniform manifold approximation and projection algorithm, and projecting it into a two-dimensional space to obtain the scatter distribution of the protein feature or the scatter distribution of the compound feature respectively; 2.4) Using the Jonker-Volgenant algorithm, the scattered distribution of protein features or the scattered distribution of compound features is linearly distributed to obtain a regularized protein template image or a regularized compound template image. The regularized protein template image and the regularized compound template image constitute a two-dimensional image representation of the drug and the target.

4. The drug-target interaction prediction method based on large language model and image representation according to claim 3, characterized in that: In step 2.1), the protein database is the TTD database and the UniProt database.

5. The drug-target interaction prediction method based on large language model and image representation according to claim 3, characterized in that: In step 2.1), the compound database is the ChEMBL database.

6. The drug-target interaction prediction method based on large language model and image representation according to claim 1, characterized in that: In step S3, a dual-channel convolutional neural network is used to construct a drug-target interaction prediction model MapCPI.

Citation Information

Patent Citations

  • Drug target affinity prediction method based on deep learning

    CN110689965A

  • Multi-modal drug-protein target interaction prediction method and system

    CN115985386A

  • Drug target general prediction method and device based on self-supervised learning, and medium

    CN116013428A

  • Credible drug-target relevance prediction method

    CN116884474A

  • Drug target interaction relationship prediction method and system

    CN119068972A