Recognition method and device for drug-target corresponding relation based on comparative learning and storage medium

Through the drug-target correspondence relationship identification method based on contrast learning, the CBAN-Predictor model and bilinear attention mechanism are used to solve the problems of high cost and time-consuming in the traditional drug discovery process, and efficient drug screening and identification are achieved.

CN120220790APending Publication Date: 2025-06-27NORTHEAST FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510276531.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The traditional drug discovery process relies on laboratory screening, which makes it expensive and time-consuming.

Method used

Using a drug-target correspondence recognition method based on contrast learning, the CBAN-Predictor model is constructed, and the interaction between drugs and targets is predicted using contrast learning and bilinear attention mechanisms.

Benefits of technology

It significantly improves the efficiency of drug screening, reduces the workload and overhead of experimental screening, can quickly adapt to new data and new tasks, and has good scalability and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220790A_ABST
    Figure CN120220790A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of drug discovery, and particularly relates to a method and equipment for identifying a drug-target corresponding relation based on comparative learning, and a storage medium. The problems that the drug discovery technology in the prior art depends on laboratory screening, and the drug discovery cost is high are solved. The invention provides a drug-target corresponding relation identification method based on comparative learning. Comprising the following steps: step 1, constructing a CBAN-Predictor model; the method comprises the following steps: obtaining a trained CBAN-Predictor model; step 2, inputting the drug-target pair data with an unknown relationship into the trained CBAN-Predictor model, outputting a prediction probability, and completing drug-target corresponding relationship identification according to the prediction probability; the drug-target pair of which the prediction probability is higher than a threshold value considers that the drug and the target can interact with each other; the problems that the drug discovery technology in the prior art depends on laboratory screening, and the drug discovery cost is high are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of drug discovery, and particularly relates to a method, device, and storage medium for identifying drug-target correspondence relationships based on contrastive learning. Background Art

[0002] Drug target identification is one of the key steps in drug development. Drug targets are usually proteins, enzymes, or other molecules that play a key role in the occurrence and development of diseases. Identifying and validating these targets helps to understand the molecular mechanisms of diseases and develop therapeutic drugs targeting specific targets. The traditional drug discovery process relies on laboratory screening, which is costly and time-consuming. Summary of the Invention

[0003] The object of the present invention is to solve the problem that the existing drug discovery technology relies on laboratory screening and the cost of drug discovery is high.

[0004] A method for identifying drug-target correspondence relationships based on contrastive learning is proposed. It includes:

[0005] Step 1: Construct a CBAN-Predictor model; obtain a trained CBAN-Predictor model;

[0006] Step 2: Input the drug-target pair data with unknown relationships into the trained CBAN-Predictor model, output the prediction probability, and complete the identification of drug-target correspondence relationships according to the prediction probability;

[0007] Drug-target pairs with a prediction probability higher than the threshold are considered to be able to interact with each other;

[0008] In the above Step 1, constructing a CBAN-Predictor model; obtaining a trained CBAN-Predictor model; the specific process is as follows:

[0009] S1: Construct a first training set and a second training set according to public databases;

[0010] S2: Establish a CBAN-Predictor model, train the CBAN-Predictor model according to the first training set, and obtain a trained CBAN-Predictor model;

[0011] S3: Train the trained CBAN-Predictor model according to the second training set to obtain a well-trained CBAN-Predictor model.

[0012] The public databases in the above S1 include DAVIS, BindingDB, and Biosnap;

[0013] The specific process of constructing the first training set based on the public database is as follows:

[0014] Extract the drug data, target data, and dissociation constant data of drugs and targets from the DAVIS database and the BindingDB database;

[0015] The target data is protein, and the drug data is drug molecules;

[0016] Take the drug-target data pairs with a dissociation constant KD less than 30 as having interactions, and the label is 1;

[0017] Take the drug-target data pairs with a dissociation constant KD greater than 30 as having no interactions, and the label is 0;

[0018] Construct the first training set according to the drug-target data pairs and the labels of the drug-target data pairs;

[0019] The specific process of constructing the second training set based on the public database in S1 is as follows:

[0020] Extract one drug data and one target data with an interaction relationship from the BIOSNAP database as a positive sample data, and finally obtain X positive sample data; X is a positive integer;

[0021] Randomly sample one drug data and one target data from the BIOSNAP database as a negative sample data; finally obtain Y negative sample data, Y is a positive integer;

[0022] Construct Z triple data according to the X positive sample data and the Y negative sample data; finally obtain Z triple data, Z is a positive integer; take the Z triple data as the second training set;

[0023] One triple data includes: drug data, positive sample target data, negative sample target data,

[0024] Among them, the drug data is the drug data that is the same for the positive sample and the negative sample;

[0025] For example, one positive sample is drug A, target B, and one negative sample is drug A, target C; then the synthesized triple data is: (drug A, target B, target C)

[0026] BIOSNAP only has positive samples. We construct negative samples by random sampling, assuming that randomly paired drugs and proteins have no interactions;

[0027] The CBAN-Predictor model in S2 sequentially includes an input layer, a feature extraction module, a bilinear attention mechanism layer, a feature fusion layer, a similarity calculation layer, and an output layer;

[0028] The feature extraction module includes a protein feature extraction model and a drug feature extraction model;

[0029] The protein feature extraction model is ProtBert; the drug feature extraction model uses Morgan molecular fingerprints; the structure of ProtBERT is as Figure 4 shown. ProtBERT is an existing model based on the BERT architecture, which has been pre-trained on the BFD (Big Fantastic Database) database. This model learns the deep representation of protein data and can be used for tasks such as protein classification and function prediction;

[0030] Morgan molecular fingerprint is a fingerprint representation method used in chemoinformatics that can effectively describe the molecular structure. It is based on the ECFP (Extended Circular Fingerprint) algorithm and encodes molecules in a local environment manner.

[0031] A computer storage medium, characterized in that at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the described method for identifying the drug-target correspondence relationship based on contrast learning.

[0032] An apparatus for identifying the drug-target correspondence relationship based on contrast learning, characterized in that the apparatus includes a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the described method for identifying the drug-target correspondence relationship based on contrast learning.

[0033] The beneficial effects of the present invention are as follows:

[0034] By combining contrast learning and bilinear attention mechanism, the present invention solves the limitations of traditional drug target identification methods in processing complex biological data. Contrast learning effectively enhances the specificity and recognition ability of the model by distinguishing real interaction and non-interaction samples.

[0035] The bilinear attention mechanism further captures the fine interaction relationship between drugs and targets, improving the model's understanding and representation ability of complex non-linear relationships. Compared with traditional laboratory drug screening, the present invention overcomes the problems of long test time and high cost, and significantly improves the efficiency of drug screening.

[0036] By combining contrastive learning and bilinear attention mechanism, this method can automatically process and analyze large-scale biological data to achieve efficient prediction of drug-target interactions. Compared with traditional methods, the present invention can still accurately identify potential drug targets and candidate molecules without high-quality three-dimensional structure information, greatly reducing the workload and cost of experimental screening.

[0037] In addition, this method can quickly adapt to new data and new tasks, has good scalability and adaptability, further promotes the acceleration of the drug discovery process and the reduction of costs, and provides an efficient and economical alternative for new drug research and development. Brief Description of the Drawings

[0038] Figure 1 is the overall flowchart of the present invention;

[0039] Figure 2 is the schematic diagram of contrastive learning of the present invention;

[0040] Figure 3 is the flowchart of the bilinear attention mechanism of the present invention;

[0041] Figure 4 is the schematic diagram of the protbert model adopted by the pre-trained model in the present invention. Detailed Embodiments

[0042] Detailed Embodiment 1: Combine Figure 1 to illustrate the present invention, including:

[0043] Step 1: Construct a CBAN-Predictor model; obtain a trained CBAN-Predictor model;

[0044] Step 2: Input the drug-target pair data with unknown relationship into the trained CBAN-Predictor model, output the prediction probability, and complete the recognition of the drug-target correspondence relationship according to the prediction probability;

[0045] Drug-target pairs with a prediction probability higher than the threshold are considered to be able to interact with each other;

[0046] In the above Step 1 of constructing a CBAN-Predictor model; obtaining a trained CBAN-Predictor model; the specific process is as follows:

[0047] S1: Construct a first training set and a second training set according to a public database;

[0048] S2: Establish a CBAN-Predictor model, train the CBAN-Predictor model according to the first training set, and obtain a trained CBAN-Predictor model;

[0049] S3. Train the trained CBAN-Predictor model according to the second training set to obtain a trained CBAN-Predictor model.

[0050] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that,

[0051] The public database in S1 includes DAVIS, BindingDB, and Biosnap;

[0052] The specific process of constructing the first training set according to the public database is as follows:

[0053] Extract drug data, target data, and dissociation constant data of drugs and targets from the DAVIS database and the BindingDB database;

[0054] The target data is protein data, and the drug data is the chemical formula of the drug;

[0055] The drug molecule and the protein data are used as the two inputs of the CBAN-Predictor model, and the predicted probability is output.

[0056] Take the drug-target data pairs with dissociation constant K D less than 30 as having interactions, and the label is 1;

[0057] Take the drug-target data pairs with dissociation constant K D greater than 30 as having no interactions, and the label is 0;

[0058] Construct the first training set according to the drug-target data pairs and the labels of the drug-target data pairs;

[0059] Other steps and parameters are the same as those in Specific Embodiment 1.

[0060] Specific Embodiment 3: The difference between this embodiment and Specific Embodiment 1 is that,

[0061] The specific process of constructing the second training set according to the public database in S1 is as follows:

[0062] Extract one drug data and one target data with an interaction relationship from the BIOSNAP database as a positive sample data, and finally obtain X positive sample data; X is a positive integer;

[0063] Randomly sample one drug data and one target data from the BIOSNAP database as a negative sample data; finally obtain Y negative sample data, and Y is a positive integer;

[0064] Construct Z triple data based on X positive sample data and Y negative sample data; finally, Z triple data are obtained, where Z is a positive integer; use the Z triple data as the second training set;

[0065] One triple data includes: drug data, positive sample target data, negative sample target data,

[0066] Among them, the drug data is the same drug data for both positive and negative samples;

[0067] For example, one positive sample is drug A, target B, and one negative sample is drug A, target C; then the synthesized triple data is: (drug A, target B, target C)

[0068] BIOSNAP only has positive samples. We construct negative samples through random sampling, assuming that randomly paired drugs and proteins do not interact;

[0069] Other steps and parameters are the same as those in one of the specific implementation manners 1 to 2.

[0070] Specific implementation manner 4: The difference between this implementation manner and the specific implementation manners 1 to 4 is that,

[0071] In the S2, the CBAN-Predictor model successively includes an input layer, a feature extraction module, a bilinear attention mechanism layer, a feature fusion layer, a similarity calculation layer, and an output layer;

[0072] The feature extraction module includes a protein feature extraction model and a drug feature extraction model;

[0073] The protein feature extraction model is ProtBert; the drug feature extraction model uses Morgan molecular fingerprints; The structure of ProtBERT is as Figure 4 shown. ProtBERT is an existing model based on the BERT architecture, which has been pre-trained on the BFD (BigFantastic Database) database. This model learns the deep representation of protein data and can be used for tasks such as protein classification and function prediction;

[0074] Morgan molecular fingerprint is a fingerprint representation method for chemoinformatics, which can effectively describe the molecular structure. It is based on the ECFP (Extended Circular Fingerprint) algorithm and encodes molecules in a local environment manner.

[0075] Other steps and parameters are the same as those in one of the specific implementation manners 1 to 3.

[0076] Specific implementation manner 5: The difference between this implementation manner and the specific implementation manners 1 to 4 is that,

[0077] In S2, the CBAN-Predictor model is trained according to the first training set to obtain the trained CBAN-Predictor model. The specific process is as follows:

[0078] S2.1: Use the drug-target pair data in the first training set as the input of the input layer of the CBAN-Predictor model.

[0079] Then, the input drug-target pair data is subjected to feature extraction processing through the feature extraction module to obtain the feature representation of the drug-target pair data.

[0080] S2.2: Input the feature representation of the drug-target pair data obtained in S2.1 into the bilinear attention mechanism layer for bilinear transformation processing to obtain the attention weight matrix of the drug-target pair data.

[0081] S2.3: Input the feature representation of the drug-target pair data obtained in S2.1 and the attention weight matrix of the drug-target pair data obtained in S2.2 into the feature fusion layer for weighted fusion processing to obtain the fused feature representation of the drug-target pair data.

[0082] S2.4: Input the fused feature representation of the drug-target pair data obtained in S2.3 into the similarity calculation layer for similarity calculation processing to obtain the similarity score of the drug-target pair data, which is used as the output of the output layer of the CBAN-Predictor model.

[0083] S2.5: Calculate the first loss function according to the input and output of the CBAN-Predictor model, and perform iterative training on the CBAN-Predictor model according to the first loss function. Stop the iteration when the number of iterations reaches the set maximum number of iterations.

[0084] S2.6: Calculate the AUPRC performance metric value of the CBAN-Predictor model obtained in each iteration, and select the CBAN-Predictor model with the largest AUPRC performance metric value as the trained CBAN-Predictor model.

[0085] Other steps and parameters are the same as those in any one of the specific embodiments one to four.

[0086] Specific embodiment six: The difference between this embodiment and the specific embodiments one to five is that

[0087] In S2.1, the input pair of drug-target pair data is subjected to feature extraction processing through the feature extraction module to obtain the feature representation of a pair of drug-target pair data. The specific process is as follows:

[0088] S2.1.1: Input the drug data of a pair of drug-target data into the protein feature extraction model of the feature extraction module, and perform feature extraction processing to obtain a target feature representation.

[0089] The protein feature extraction model is an existing protein language model (ProtBert). The protein feature representation includes the global structure and functional information of the protein. Specifically, the protein feature representation includes the biophysical properties of amino acids to the structural, functional, and evolutionary features of the protein.

[0090] S2.1.2: Input the target data of a pair of drug-target data into the drug feature extraction model of the feature extraction module, and perform feature extraction processing to obtain a drug feature representation.

[0091] The drug feature extraction model is the Morgan molecular fingerprint, which extracts the feature representation in the drug molecular structure.

[0092] S2.1.3: Combine a target feature representation obtained in S2.1.1 and a drug feature representation obtained in S2.1.2 into the feature representation of a pair of drug-target data.

[0093] In S2.2, input the feature representation of the drug-target data obtained in S2.1 into the bilinear attention mechanism layer for bilinear transformation processing to obtain the attention weight matrix of the drug-target data, which is expressed by the formula:

[0094]

[0095] In the formula, R d is the drug small molecule feature representation, R d T is the transpose matrix of R d R p is the target protein feature representation; softmax() represents the softmax function; M0 represents the attention weight matrix of the drug-target data; 1 represents the fixed all-1 vector, P is a learnable parameter vector, U is the weight matrix of the drug feature representation; V is the weight matrix of the target feature representation, V T is the transpose matrix of V, and ⊙ represents the element-wise product; during the training process, learn and gradually update the weight matrix U of the drug molecular feature matrix, the weight matrix V of the target protein feature matrix, and the parameter vector P.

[0096] In S2.3, input the feature representation of the drug-target data obtained in S2.1 and the attention weight matrix of the drug-target data obtained in S2.2 into the feature fusion layer for weighted fusion processing to obtain the fused feature representation of the drug-target data; the specific process is as follows:

[0097] S2.3.1: Represent the drug feature as R d , the target feature as R p and the attention weight matrix M0 to obtain the intermediate feature representation f h , which is expressed by the formula as:

[0098] f h = (R d T U2) T ·M0·(R p T V2) (2)

[0099] where, · represents multiplication, U2 and V2 are another set of weight matrices, f h is the intermediate representation, R p T is the transposed matrix of R p , which captures the bilinear interaction between the input channels.

[0100] S2.3.2: Perform pooling on the intermediate feature representation f h to obtain a fused feature representation f, and use the fused feature representation f as the fused feature representation of the drug-target pair data, which is expressed by the formula as:

[0101] f = SumPooling(f h , s)

[0102] where, SumPooling represents sum pooling, s represents the step size, and the SumPooling function is a one-dimensional and non-overlapping sum pooling operation with a step size of s;

[0103] In S2.4, the fused feature representation of the drug-target pair data obtained in S2.3 is input into the similarity calculation layer for similarity calculation processing to obtain the similarity score of the drug-target pair data, which is expressed by the formula as:

[0104]

[0105] In the formula, MLP represents multi-layer perceptron processing; Sigmod represents non-linear transformation processing;

[0106] The fused feature representation f of the drug-target pair data obtained in S2.3 is input into the multi-layer perceptron (MLP), and finally through the non-linear transformation of the Sigmoid layer, the output is converted into a probability value between 0 and 1; the processing process is well-known to those skilled in the art;

[0107] In S2.5, the first loss function is the binary cross-entropy loss function, which is a well-known loss function in the art;

[0108] , which is expressed by the formula as:

[0109]

[0110] In the formula, L1 represents the first loss function, the binary cross-entropy loss function, which is widely used in binary classification problems. Its role is to measure the difference between the model's prediction result and the true label (y). The goal during model training is to minimize this loss function, so that the prediction probability is as close as possible to the true label y. y represents the true label, usually 0 or 1 (the label in the classification task), and log represents the logarithmic function, which is used to calculate the cross-entropy loss; represents the probability value predicted by the model, indicating the probability that the sample belongs to the positive class.

[0111] The maximum number of iterations in S2.6 is 50 times;

[0112] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.

[0113] Specific embodiment seven: The difference between this embodiment and the first to sixth specific embodiments is that

[0114] in S3, the trained CBAN-Predictor model is trained according to the second training set to obtain a trained CBAN-Predictor model. The specific process is as follows:

[0115] S3.1: Use the triple data in the second training set as the input to the input layer of the trained CBAN-Predictor model;

[0116] Then, the input triple data is subjected to feature extraction processing through the feature extraction module to obtain the feature representation of the triple data;

[0117] S3.2: Input the feature representation of the triple data obtained in S3.1 into the bilinear attention mechanism layer for bilinear transformation processing to obtain the attention weight matrix of the triple data;

[0118] S3.3: Input the feature representation of the triple data obtained in S3.1 and the attention weight matrix of the triple data obtained in S3.2 into the feature fusion layer for weighted fusion processing to obtain the fused feature representation of the triple data;

[0119] S3.4: Input the fused feature representation of the triple data obtained in S3.3 into the similarity calculation layer for similarity calculation processing to obtain the similarity score of the triple data, which is used as the output of the output layer of the trained CBAN-Predictor model;

[0120] S3.5: Calculate the second loss function based on the input and output of the trained CBAN-Predictor model, and iteratively train the trained CBAN-Predictor model according to the second loss function; stop the iteration when the number of iterations reaches the set maximum number of iterations;

[0121] S3.6: Calculate the AUPRC performance metric value of the trained CBAN-Predictor model obtained in each iteration, and select the trained CBAN-Predictor model with the largest AUPRC performance metric value as the trained CBAN-Predictor model

[0122] Other steps and parameters are the same as those in any one of the specific embodiments 1 to 6.

[0123] Specific embodiment 8: The difference between this embodiment and the specific embodiments 1 to 7 is that

[0124] In S3.1, a set of input triple data is subjected to feature extraction processing by a feature extraction module to obtain a feature representation of the set of triple data. The specific process is as follows:

[0125] S3.1.1: Input the positive sample target data of a set of input triple data into the protein feature extraction model of the feature extraction module, and perform feature extraction processing to obtain a positive sample target feature representation;

[0126] The protein feature extraction model is an existing protein language model (ProtBert). The protein feature representation includes the global structure and function information of the protein. Specifically, the protein feature representation includes the biophysical properties of amino acids to the structural, functional, and evolutionary features of the protein;

[0127] S3.1.1: Input the negative sample target data of a set of input triple data into the protein feature extraction model of the feature extraction module, and perform feature extraction processing to obtain a negative sample target feature representation;

[0128] S3.1.3: Input the drug data of a set of input triple data into the drug feature extraction model of the feature extraction module, and perform feature extraction processing to obtain a drug feature representation;

[0129] The drug feature extraction model is Morgan molecular fingerprint, which extracts the feature representation in the drug molecular structure;

[0130] S3.1.4: Combine a positive sample target feature representation obtained in S3.1.1 and a drug feature obtained in S3.1.3 into a positive sample feature representation;

[0131] Combine a negative sample target feature representation obtained in S3.1.2 and a drug feature representation obtained in S3.1.3 into a negative sample feature representation;

[0132] Combine the positive sample feature representation and the negative sample feature representation into the feature representation of a group of triple data;

[0133] In S3.2, input the feature representation of the triple data obtained in S2.1 into the bilinear attention mechanism layer for bilinear transformation processing to obtain the attention weight matrix of the triple data, which is expressed by the formula:

[0134]

[0135] In the formula, R d is the drug small molecule feature representation, R d T is the transpose matrix of R d ; R p1 is the positive sample target protein feature representation; softmax() represents the softmax function; M1 represents the attention weight matrix of the positive sample; 1 represents the fixed all-1 vector, P is a learnable parameter vector, U is the weight matrix of the drug feature representation; V is the weight matrix of the target feature representation, V T is the transpose matrix of V, and ⊙ represents the element-wise product; during the training process, learn and gradually update the weight matrix U of the drug molecule feature matrix, the weight matrix V of the target protein feature matrix, and the parameter vector P;

[0136]

[0137] In the formula, R d is the drug small molecule feature representation, R d T is the transpose matrix of R d ; R p2 is the negative sample target protein feature representation; softmax() represents the softmax function; M2 represents the attention weight matrix of the negative sample; 1 represents the fixed all-1 vector, P is a learnable parameter vector, U is the weight matrix of the drug feature representation; V is the weight matrix of the target feature representation, V T is the transpose matrix of V, and ⊙ represents the element-wise product; during the training process, learn and gradually update the weight matrix U of the drug molecule feature matrix, the weight matrix V of the target protein feature matrix, and the parameter vector P;

[0138] In S3.3, input the feature representation of the triple data obtained in S3.1 and the attention weight matrix of the triple data obtained in S2.2 into the feature fusion layer for weighted fusion processing to obtain the fused feature representation of the triple data; the specific process is:

[0139] S3.3.1: Represent the drug feature as R d , the positive sample target feature representation R p1 and the positive sample attention weight matrix M1 to obtain the positive sample intermediate feature representation f h1 , which is expressed by the formula as:

[0140] f h1 =(R d T U3) T ·M1·(R p1 T V3) (6)

[0141] where U3 and V3 are another set of weight matrices, and R p1 T is the transpose matrix of R p1 , which captures the bilinear interaction between input channels.

[0142] S3.3.2: Perform pooling on the intermediate feature representation f h1 to obtain a fused feature representation f1, and use the fused feature representation f1 as the fused feature representation of the triple data, which is expressed by the formula as:

[0143] f1 = SumPooling(f h1 , s)

[0144] where, SumPooling represents the sum pooling process, and s represents the step size,

[0145] In S3.4, the fused feature representation of the triple data obtained in S3.3 is input into the similarity calculation layer for similarity calculation processing to obtain the similarity score of the triple data, which is expressed by the formula as:

[0146]

[0147] In S3.5, the second loss function is the contrastive loss function; it is a well-known loss function in the art;

[0148] The calculation formula of the contrastive loss function is:

[0149] L2 = max(d(a, p) - d(a, n) + margin, 0) (8)

[0150] In the formula, L2 represents the contrast loss function, d(a, p) represents the distance between the anchor sample and the positive class sample, and theoretically, the lower this value is, the better. d(a, n) represents the distance between the anchor sample and the negative class sample, and theoretically, the higher this value is, the better; margin represents the expected distance interval. The function of this value is that if the distance difference between the positive class sample and the negative class sample is less than margin, the loss will increase; if their difference exceeds margin, the loss is 0

[0151] The maximum number of iterations in S3.6 is 50 times

[0152] Other steps and parameters are the same as those in any one of the specific embodiments one to seven

[0153] Specific embodiment nine: This embodiment is a computer storage medium, and at least one instruction is stored in the storage medium. The at least one instruction is loaded and executed by a processor to implement the method for identifying the drug-target correspondence relationship based on contrast learning

[0154] It should be understood that the instruction includes a computer program product, software, or computerized method corresponding to any method described in the present invention; the instruction can be used to program a computer system or other electronic devices. The computer storage medium may include a readable medium on which the instruction is stored, and may include, but is not limited to, a magnetic storage medium and an optical storage medium; the magneto-optical storage medium includes a read-only memory ROM, a random access memory RAM, an erasable programmable memory (e.g., EPROM and EEPROM), and a flash memory layer, or other types of media suitable for storing electronic instructions

[0155] Specific embodiment ten: This embodiment is a device for identifying the drug-target correspondence relationship based on contrast learning. The device includes a processor and a memory. It should be understood that it includes any device including a processor and a memory described in the present invention. The device may further include other units and modules for displaying, interacting, processing, controlling, etc. through signals or instructions, and other functions

[0156] At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the method for identifying the drug-target correspondence relationship based on contrast learning

[0157] Those skilled in the art should understand that at least one stored instruction is a computer program product corresponding to a method or a system. Therefore, the present application can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented using various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript, etc.

[0158] The present application is described with reference to the flowcharts and / or block diagrams of methods, systems, and computer program products according to the embodiments of the present application, and can also be used for corresponding devices. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0159] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0161] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0162] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

[0163] Summarize the beneficial effects of the present invention in combination with Specific Embodiments 1 to 10:

[0164] Drug-Target Interaction (DTI) prediction based on contrastive learning, specifically involving using protein data representations generated by a pre-trained protein language model and small molecule drug features generated by Morgan fingerprints to map drugs and proteins into a shared latent space, combined with a bilinear attention mechanism to more precisely capture the complex interactions between the two. At the same time, based on the distances between the learned protein data representations for binding prediction, it can quickly and efficiently screen potential drug molecules.

[0165] Since experimental screening of potential drug molecules is a crucial and time-consuming step in the drug discovery process, and existing computational methods have limitations, this study aims to develop a contrastive learning-based computational method that can accurately predict DTI on large-scale data, thereby accelerating the drug discovery process.

[0166] Two different stages are adopted when training the model, each stage having different training objectives and loss functions to balance generalization ability and specificity.

[0167] In the first stage, the model uses a pre-trained protein language model (ProtBert) to generate protein features and generates drug features through Morgan fingerprints. Then, through the bilinear attention mechanism, it calculates attention weights for each combination of protein and drug pairs to generate a weighted interaction feature representation. This structure allows the model to learn more complex relationships in all interactions between the two features. The obtained weighted features are input into a non-linear transformation to be mapped into a shared latent space, and the interaction is predicted based on the cosine distance. The non-linear transformation parameters of the model are optimized on a low-coverage dataset through binary cross-entropy loss.

[0168] In the second stage, the target protein, drug, and decoy triple are also mapped into the shared latent space and optimized on a high-coverage dataset through a triple distance loss function to minimize the target-drug distance and maximize the target-decoy distance to ensure discriminative ability.

[0169] During the training process, the model alternates between two stages to optimize both objectives simultaneously. By adjusting the learning rate and contrastive loss, it finally achieves good generalization performance on unseen data and high-specificity recognition of real interactions. It belongs to the field of drug discovery.

[0170] The above are only descriptions of the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention and is based on the technical essence of the present invention, any simple modification, equivalent replacement, and improvement of the above embodiments still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for identifying drug-target correspondence based on contrastive learning, characterized in that: The following steps are involved: Step 1: Build a CBAN-Predictor model; obtain the trained CBAN-Predictor model; Step 2: Input the drug-target pair data with unknown relationship into the trained CBAN-Predictor model, output the predicted probability, and complete the drug-target correspondence identification based on the predicted probability; In the step 1, a CBAN-Predictor model is constructed; a trained CBAN-Predictor model is obtained; the specific process is as follows: S1. Construct a first training set and a second training set based on a public database; S2, establishing a CBAN-Predictor model, and training the CBAN-Predictor model according to the first training set to obtain a trained CBAN-Predictor model; S3. Train the trained CBAN-Predictor model according to the second training set to obtain a trained CBAN-Predictor model.

2. The method for identifying drug-target correspondence based on contrastive learning according to claim 1, characterized in that: The public databases in S1 include DAVIS, BindingDB and Biosnap; The specific process of constructing the first training set according to the public database is: Extract drug data, target data, and dissociation constant data of drugs and targets from the DAVIS database and the BindingDB database; The dissociation constant K D Drug-target data pairs with less than 30 are considered to have interactions and are labeled as 1; The dissociation constant K D Drug-target data pairs greater than 30 are considered to have no interaction and are labeled 0; Construct a first training set based on the drug-target data pairs and the labels of the drug-target data pairs;.

3. The method for identifying drug-target correspondence based on contrastive learning according to claim 2, characterized in that: The specific process of constructing the second training set according to the public database in S1 is: Extract a drug data and a target data with interaction relationship in the BIOSNAP database as a positive sample data, and finally obtain X positive sample data; X is a positive integer; Randomly sample a drug data and a target data in the BIOSNAP database as a negative sample data; finally, obtain Y negative sample data, where Y is a positive integer; Construct Z triplet data according to X positive sample data and Y negative sample data; finally obtain Z triplet data, where Z is a positive integer; Take Z triplet data as the second training set; The triplet data includes: drug data, positive sample target data, and negative sample target data.

4. The method for identifying drug-target correspondence based on contrastive learning according to claim 3, characterized in that: The CBAN-Predictor model in S2 includes an input layer, a feature extraction module, a bilinear attention mechanism layer, a feature fusion layer, a similarity calculation layer and an output layer in sequence; The feature extraction module includes a protein feature extraction model and a drug feature extraction model.

5. The method for identifying drug-target correspondence based on contrastive learning according to claim 4, characterized in that: In S2, the CBAN-Predictor model is trained according to the first training set to obtain a trained CBAN-Predictor model; the specific process is: S2.1: The drug-target pair data in the first training set is used as the input layer of the CBAN-Predictor model; Then the input drug-target pair data is processed by the feature extraction module to obtain the feature representation of the drug-target pair data; S2.2: Input the feature representation of the drug-target pair data obtained in S2.1 into the bilinear attention mechanism layer for bilinear transformation to obtain the attention weight matrix of the drug-target pair data; S2.3: The feature representation of the drug-target pair data obtained in S2.1 and the attention weight matrix of the drug-target pair data obtained in S2.2 are input into the feature fusion layer for weighted fusion processing to obtain a fused feature representation of the drug-target pair data; S2.4: Input the fusion feature representation of the drug-target pair data obtained in S2.3 into the similarity calculation layer for similarity calculation processing to obtain the similarity score of the drug-target pair data as the output of the CBAN-Predictor model output layer; S2.5: Calculate a first loss function according to the input and output of the CBAN-Predictor model, and iteratively train the CBAN-Predictor model according to the first loss function; stop the iteration when the number of iterations reaches the set maximum number of iterations; S2.6: Calculate the AUPRC performance index value of the CBAN-Predictor model obtained in each iteration, and select the CBAN-Predictor model with the largest AUPRC performance index value as the trained CBAN-Predictor model.

6. The method for identifying drug-target correspondence based on contrastive learning according to claim 5, characterized in that: In S2.1, the input pair of drug-target pair data is subjected to feature extraction processing by the feature extraction module to obtain a feature representation of the pair of drug-target pair data. The specific process is as follows: S2.1.1: Input the drug data of a pair of drug-target pair data into the protein feature extraction model of the feature extraction module, perform feature extraction processing to obtain a target feature representation; S2.1.2: Input the target data of a pair of drug-target pair data into the drug feature extraction model of the feature extraction module, perform feature extraction processing to obtain a drug feature representation; S2.1.3: Combine a target feature representation obtained in S2.1.1 and a drug feature representation obtained in S2.1.2 into a feature representation of a drug-target pair data; In S2.2, the feature representation of the drug-target pair data obtained in S2.1 is input into the bilinear attention mechanism layer for bilinear transformation processing to obtain the attention weight matrix of the drug-target pair data, which is expressed as follows: In the formula, R d is the drug characteristic representation, R d T YesR d The transposed matrix, R p It is the target feature representation; softmax() represents the softmax function; M0 represents the attention weight matrix of the drug-target pair data; 1 represents the all-1 vector, P is the parameter vector, U is the weight matrix of the drug feature representation; V is the weight matrix of the target feature representation, V T is the transposed matrix of V, ⊙ represents the element-by-element product; In S2.3, the feature representation of the drug-target pair data obtained in S2.1 and the attention weight matrix of the drug-target pair data obtained in S2.2 are input into the feature fusion layer for weighted fusion processing to obtain the fused feature representation of the drug-target pair data; the specific process is: S2.3.1: Express R according to drug characteristics d , target feature representation R p And the attention weight matrix M0 obtains the intermediate feature representation f h , expressed as: f h =(R d T U2) T ·M0·(R p T V2) (2) Among them, · represents multiplication, U2 and V2 are another set of weight matrices, and f h is the intermediate representation, R p T YesR p The transposed matrix of S2.3.2: Representing the intermediate features f h Pooling is performed to obtain a fused feature representation f, which is used as the fused feature representation of the drug-target pair data and is expressed as follows: f=SumPooling(f h ,s) Among them, SumPooling represents summing pooling processing, s represents the step size, In S2.4, the fusion feature representation of the drug-target pair data obtained in S2.3 is input into the similarity calculation layer for similarity calculation processing to obtain the similarity score of the drug-target pair data, which is expressed by the formula: In the formula, MLP represents multi-layer perceptron processing; Sigmod represents nonlinear transformation processing; The first loss function in S2.5 is a binary cross entropy loss function. The maximum number of iterations in S2.6 is 50.

7. The method for identifying drug-target correspondence based on contrastive learning according to claim 6, characterized in that: In S3, the trained CBAN-Predictor model is trained according to the second training set to obtain a trained CBAN-Predictor model. The specific process is as follows: S3.1: Use the triplet data in the second training set as the input layer of the trained CBAN-Predictor model; Then the input triple data is processed by the feature extraction module to obtain the feature representation of the triple data; S3.2: Input the feature representation of the triple data obtained in S3.1 into the bilinear attention mechanism layer for bilinear transformation to obtain the attention weight matrix of the triple data; S3.3: The feature representation of the triple data obtained in S3.1 and the attention weight matrix of the triple data obtained in S3.2 are input into the feature fusion layer for weighted fusion processing to obtain the fused feature representation of the triple data; S3.4: Input the fusion feature representation of the triple data obtained in S3.3 into the similarity calculation layer for similarity calculation processing to obtain the similarity score of the triple data as the output of the output layer of the trained CBAN-Predictor model; S3.5: Calculate a second loss function according to the input and output of the trained CBAN-Predictor model, and iteratively train the trained CBAN-Predictor model according to the second loss function; stop the iteration when the number of iterations reaches the set maximum number of iterations; S3.6: Calculate the AUPRC performance index value of the trained CBAN-Predictor model obtained in each iteration, and select the trained CBAN-Predictor model with the largest AUPRC performance index value as the trained CBAN-Predictor model.

8. The method for identifying drug-target correspondence based on contrastive learning according to claim 7, characterized in that: In S3.1, a set of triplet data input is processed by a feature extraction module to obtain a feature representation of a set of triplet data. The specific process is as follows: S3.1.1: Input a set of positive sample target data of the input triplet data into the protein feature extraction model of the feature extraction module, perform feature extraction processing to obtain a positive sample target feature representation; S3.1.1: Inputting a set of negative sample target data of triplet data into the protein feature extraction model of the feature extraction module, performing feature extraction processing to obtain a negative sample target feature representation; S3.1.3: Inputting a set of triplet data of drug data into a drug feature extraction model of a feature extraction module, performing feature extraction processing to obtain a drug feature representation; S3.1.4: Combine a positive sample target feature representation obtained in S3.1.1 and a drug feature obtained in S3.1.3 into a positive sample feature representation; Combine a negative sample target feature representation obtained in S3.1.2 and a drug feature representation obtained in S3.1.3 into a negative sample feature representation; The positive sample feature representation and the negative sample feature representation are combined into a set of feature representations of triple data; In S3.2, the feature representation of the triple data obtained in S2.1 is input into the bilinear attention mechanism layer for bilinear transformation processing to obtain the attention weight matrix of the triple data, which is expressed as follows: In the formula, R d is the drug characteristic representation, R d T YesR d The transposed matrix, R p1 is the positive sample target feature representation; softmax() represents the softmax function; M1 represents the attention weight matrix of the positive sample; 1 represents a fixed all-1 vector, P is the parameter vector, U is the weight matrix of the drug feature representation; V is the weight matrix of the target feature representation, V T is the transposed matrix of V, ⊙ represents the element-by-element product; In the formula, R d is the drug characteristic representation, R d T YesR d The transposed matrix, R p2 is the negative sample target feature representation; softmax() represents the softmax function; M2 represents the attention weight matrix of the negative sample; 1 represents a fixed all-1 vector, P is the parameter vector, U is the weight matrix of the drug feature representation; V is the weight matrix of the target feature representation, V T is the transposed matrix of V, ⊙ represents the element-by-element product; In S3.3, the feature representation of the triple data obtained in S3.1 and the attention weight matrix of the triple data obtained in S2.2 are input into the feature fusion layer for weighted fusion processing to obtain a fused feature representation of the triple data; The specific process is: S3.3.1: Expressing R according to drug characteristics d , positive sample target feature representation R p1 And the positive sample attention weight matrix M1 obtains the positive sample intermediate feature representation f h1 , expressed as: f h1 =(R d T U3) T ·M1·(R p1 T V3) (6) Among them, U3 and V3 are another set of weight matrices, R p1 T YesR p1 The transposed matrix of S3.3.2: Representing the intermediate features f h1 Pooling is performed to obtain a fused feature representation f1, which is used as the fused feature representation of the triple data, and is expressed as follows: f1=SumPooling(f h1 ,s) Among them, SumPooling represents summing pooling processing, s represents the step size, In S3.4, the fusion feature representation of the triple data obtained in S3.3 is input into the similarity calculation layer for similarity calculation processing to obtain the similarity score of the triple data, which is expressed by the formula: The second loss function in S3.5 is a contrast loss function; The maximum number of iterations in S3.6 is 50.

9. A computer storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement a method for identifying a drug-target correspondence relationship based on contrastive learning as described in any one of claims 1 to 8.

10. A device for identifying drug-target correspondence based on contrastive learning, characterized in that: The device includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement a method for identifying a drug-target correspondence relationship based on contrastive learning as described in any one of claims 1 to 8.