Drug target determination method and device, computer equipment and storage medium

Through unsupervised feature extraction and clustering processing of protein heat maps, the problem of inaccurate drug target identification in existing technologies is solved, and high-accuracy drug target identification under big data conditions is achieved.

CN120600099APending Publication Date: 2025-09-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410253548.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-05
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The statistical methods used in existing technologies to fit dissolution curves and group scoring methods are not very accurate in determining drug targets in the case of big data. The assumption that the data follows a specific distribution may not hold, resulting in inaccurate drug target determination.

Method used

By obtaining a protein heat map, an unsupervised feature extraction model is used to extract protein heat map features and perform clustering processing. Drug targets are determined based on the significance of protein concentration differences and the goodness of fit of the dissolution curve, avoiding the assumption that the data obeys a specific distribution.

Benefits of technology

It improves the accuracy of drug target identification, can handle large amounts of protein data, overcomes the problem of fewer positive samples, and ensures the feasibility of classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600099A_ABST
    Figure CN120600099A_ABST
Patent Text Reader

Abstract

The invention relates to a drug target determination method and device, computer equipment and a storage medium. The method comprises the following steps: obtaining a protein heat map corresponding to candidate proteins; based on a feature extraction model, extracting features of the protein heat map to obtain heat map features corresponding to the candidate proteins; based on the heat map features corresponding to the candidate proteins, clustering the candidate proteins to obtain a plurality of candidate protein sets; based on the temperature data and the protein concentration data of each candidate protein in a candidate protein set, determining the protein concentration difference significance corresponding to the candidate protein set, and based on the protein concentration difference significance and the protein concentration data of the candidate proteins, determining the protein concentration difference significance corresponding to the candidate protein set; and determining standard reference data of the candidate proteins, and determining a drug target based on the standard reference data of each candidate protein in the plurality of candidate protein sets. According to the embodiment of the invention, the drug target can be accurately determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer equipment, storage medium, and computer program product for determining drug targets. Background Art

[0002] Based on the principle that protein stability increases after binding to a ligand, thermoproteomics evaluates changes in protein thermal stability by heating cells (or protein extracts) with and without drugs and measuring the remaining soluble proteins at each temperature to determine the drug target. The discovery of drug targets can accelerate the drug development process and reduce R&D costs. In related technologies, a dissolution curve is fitted statistically. The dissolution curve is used to reflect the stability changes of proteins during temperature increase, and proteins are grouped and scored based on statistical analysis to determine drug targets.

[0003] However, statistical methods for fitting dissolution curves and group scoring usually assume that the data obey a certain distribution or model, and have certain assumptions about the distribution form and characteristics of the data. When the amount of data is huge, the assumptions may not hold, resulting in low accuracy in determining drug targets. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium and computer program product for determining drug targets that can improve the accuracy of drug target determination in order to address the above technical problems.

[0005] In a first aspect, the present application provides a method for determining a drug target. The method comprises:

[0006] Obtaining a protein heat map corresponding to the candidate protein; wherein the protein concentration in the protein heat map changes with changes in drug concentration and temperature;

[0007] Extracting features of the protein heat map based on a feature extraction model to obtain heat map features corresponding to the candidate protein; wherein the feature extraction model is obtained by unsupervised training based on protein heat map samples;

[0008] performing clustering processing on the candidate proteins based on the heat map features corresponding to the candidate proteins to obtain multiple candidate protein sets;

[0009] Based on the temperature data and protein concentration data of each candidate protein in the candidate protein set, the protein concentration difference significance corresponding to the candidate protein set is determined; based on the protein concentration difference significance and the protein concentration data of the candidate protein, the standard reference data of the candidate protein is determined; based on the standard reference data of each candidate protein in the multiple candidate protein sets, the drug target is determined.

[0010] In a second aspect, the present application further provides a device for determining a drug target, the device comprising:

[0011] An acquisition module is used to obtain a protein heat map corresponding to the candidate protein; wherein the protein concentration in the protein heat map changes with the change of drug concentration and temperature;

[0012] An extraction module, configured to extract features of the protein heat map based on a feature extraction model to obtain heat map features corresponding to the candidate protein; wherein the feature extraction model is obtained by unsupervised training based on protein heat map samples;

[0013] a clustering module, configured to perform clustering processing on the candidate proteins based on the heat map features corresponding to the candidate proteins to obtain a plurality of candidate protein sets;

[0014] A determination module is configured to determine the protein concentration difference significance corresponding to the candidate protein set based on the temperature data and protein concentration data of each candidate protein in the candidate protein set, determine the standard reference data of the candidate protein based on the protein concentration difference significance and the protein concentration data of the candidate protein, and determine the drug target based on the standard reference data of each candidate protein in the multiple candidate protein sets.

[0015] In one embodiment, the apparatus further includes a feature extraction model training module, wherein the feature extraction model training module is configured to:

[0016] Acquire a sample set; the sample set includes a plurality of protein heat map samples;

[0017] Extracting features of a first protein heat map sample in the sample set based on an initial feature extraction model to obtain a first heat map feature, and extracting features of a second protein heat map sample in the sample set to obtain a second heat map feature;

[0018] determining a contrast loss based on a similarity between the first heat map feature and the second heat map feature;

[0019] Based on the contrast loss, the initial feature extraction model is iteratively adjusted to obtain a feature extraction model.

[0020] In one embodiment, the training module of the feature extraction model is further used to:

[0021] Determine a transposed feature matrix corresponding to the second heat map feature based on the feature matrix corresponding to the second heat map feature;

[0022] Obtaining a cross-correlation matrix based on a feature matrix corresponding to the first heat map feature and a transposed feature matrix corresponding to the second heat map feature, wherein the cross-correlation matrix is ​​used to represent the similarity between the first heat map feature and the second heat map feature;

[0023] Based on the difference between the cross-correlation matrix and the identity matrix, a contrastive loss is determined.

[0024] In one embodiment, the training module of the feature extraction model is further used to:

[0025] Obtaining an original protein heat map sample from the sample set;

[0026] performing a first noise processing on the original protein heat map sample to obtain a processed first protein heat map sample, and performing a second noise processing on the original protein heat map sample to obtain a processed second protein heat map sample;

[0027] Based on the initial feature extraction model, the features of the processed first protein heat map sample are extracted to obtain a first heat map feature, and the features of the processed second protein heat map sample are extracted to obtain a second heat map feature.

[0028] In one embodiment, the feature extraction model includes a plurality of feature encoders, and the extraction module is further configured to:

[0029] Inputting the protein heat map into each feature encoder respectively and outputting corresponding sub-heat map features; wherein at least one of the following parameters: convolution kernel size, number of convolution kernels, and number of convolution layers in the two feature encoders is different;

[0030] The features of each sub-heat map are concatenated to obtain the features of the protein heat map.

[0031] In one embodiment, the acquisition module is further configured to:

[0032] Obtain the initial protein heat map corresponding to the candidate protein;

[0033] Data interpolation processing is performed on the initial protein heat map to obtain a protein heat map corresponding to the candidate protein.

[0034] In one embodiment, the acquisition module is further configured to:

[0035] Obtaining a dissolution curve and a vector machine model of the candidate protein; wherein the protein concentration in the dissolution curve changes with temperature;

[0036] Inputting the temperature data of the point to be inserted into the vector machine model, and outputting the protein concentration data of the point to be inserted;

[0037] performing interpolation processing on the melting curve based on the temperature data and protein concentration data of the point to be inserted to obtain a processed melting curve;

[0038] Based on the processed melting curves, a protein heat map corresponding to the candidate protein is obtained.

[0039] In one embodiment, the acquisition module is further configured to:

[0040] Acquire a sample set; wherein the sample set includes protein concentration data and corresponding temperature data;

[0041] inputting the temperature data of the sample set into an initial general vector machine model and outputting a predicted first protein concentration;

[0042] determining a first loss based on a difference between the predicted first protein concentration and the actual protein concentration;

[0043] Adjusting the network parameters of the initial general vector machine model, inputting the temperature data of the sample set into the adjusted general vector machine model, and outputting a predicted second protein concentration;

[0044] determining a second loss based on a difference between the predicted second protein concentration and the actual protein concentration;

[0045] When the second loss is less than or equal to the first loss, the adjusted general vector machine model is retained, and the network parameters of the adjusted general vector machine model are continuously adjusted until the target loss is less than a preset threshold.

[0046] In one embodiment, the acquisition module is further configured to:

[0047] Obtaining minimum temperature data and maximum temperature data corresponding to the melting curve;

[0048] Dividing the temperature interval between the minimum temperature data and the maximum temperature data into a preset number of sub-temperature intervals;

[0049] Based on the interval endpoints of the sub-temperature interval, temperature data of the point to be inserted is obtained, and the temperature data of the point to be inserted is input into the vector machine model.

[0050] In one embodiment, the determining module is further configured to:

[0051] determining a dissolution curve of the candidate protein based on the protein concentration data of the candidate protein;

[0052] Determining the average goodness of fit of the dissolution curves of the candidate protein based on the dissolution curves corresponding to each drug concentration of the candidate protein;

[0053] Based on the average goodness of fit and the significance of the protein concentration difference, the standard reference data of the target protein is determined.

[0054] In one embodiment, the determining module is further configured to:

[0055] determining the area under the curve of the melting curve corresponding to each candidate protein based on the temperature data and protein concentration data of each candidate protein in the candidate protein set;

[0056] The areas under the curves of the melting curves are fused to obtain the fused area under the curves;

[0057] Based on the area under the fusion curve and the distribution function, the significance of the difference in protein concentration of the candidate proteins is determined.

[0058] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the drug target determination method described in any embodiment of the present disclosure.

[0059] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for determining a drug target according to any embodiment of the present disclosure.

[0060] In a fifth aspect, the present application further provides a computer program product, comprising a computer program that, when executed by a processor, implements the method for determining a drug target as described in any embodiment of the present disclosure.

[0061] The above-mentioned method, apparatus, computer equipment, storage medium and computer program product for determining drug targets. The features of the protein heat map are extracted by a feature extraction model, and the extracted heat map features are clustered to obtain a plurality of candidate protein sets, wherein the protein heat map features contain information on protein concentration, drug concentration and temperature. In the process of grouping, the embodiment of the present disclosure does not need to assume that the data obeys a specific distribution, but clusters and groups according to the heat map features of the protein, so that it can cope with a large amount of protein data and accurately group them, thereby improving the accuracy of drug target determination. Furthermore, in the embodiment of the present disclosure, the feature extraction model can be trained in an unsupervised manner, which can overcome the technical problem of fewer positive samples in the field of drug targets and ensure the feasibility of classification based on the heat map features of the protein. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 Schematic diagram of a process for determining a drug target in one embodiment;

[0063] Figure 2 is a schematic diagram of a method for generating a protein heat map in one embodiment;

[0064] Figure 3 Schematic diagram of a process for determining a drug target in one embodiment;

[0065] Figure 4 Schematic diagram of a process for determining a drug target in one embodiment;

[0066] Figure 5 Schematic diagram of a method for training a feature extraction model in one embodiment;

[0067] Figure 6 Schematic diagram of the structure of a feature extraction model in one embodiment;

[0068] Figure 7 Schematic diagram of the structure of a vector machine model in one embodiment;

[0069] Figure 8 Schematic diagram of a process for determining a drug target in one embodiment;

[0070] Figure 9 Schematic diagram of a process for determining a drug target in one embodiment;

[0071] Figure 10 is a schematic diagram of a protein heat map in one embodiment;

[0072] Figure 11 is a schematic diagram of a protein heat map in another embodiment;

[0073] Figure 12A schematic diagram comparing an embodiment with the prior art;

[0074] Figure 13 A schematic diagram comparing an embodiment with the prior art;

[0075] Figure 14 A schematic diagram comparing an embodiment with the prior art;

[0076] Figure 15 is a structural block diagram of a device for determining a drug target in one embodiment;

[0077] Figure 16 is a diagram of the internal structure of a computer device in one embodiment;

[0078] Figure 17 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0079] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0080] In order to facilitate those skilled in the art to understand the technical solution provided by the embodiments of the present disclosure, the technical environment in which the technical solution is implemented is described below.

[0081] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0082] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0083] Computer vision (CV) is the science of making machines "see." Specifically, it refers to using cameras and computers to replace the human eye in identifying, locating, and measuring objects. Furthermore, it involves image processing, which transforms the computer's image into something more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems capable of extracting information from images or multidimensional data. Large model technology has revolutionized the development of computer vision. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and broadly apply to specific downstream tasks. Computer vision technologies generally include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. Common biometric recognition technologies include facial recognition and fingerprint recognition.

[0084] Key technologies in speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction. Large model technology is revolutionizing the development of speech technology. Pre-trained models such as WavLM and UniSpeech, which leverage the Transformer architecture, possess strong generalization and versatility, enabling them to effectively handle a wide range of speech processing tasks.

[0085] Natural language processing (NLP) is a key area of ​​research in computer science and artificial intelligence. It studies theories and methods that enable effective communication between humans and computers using natural language. Natural language processing involves natural language, the language we use daily, and is closely related to linguistics. Pretrained models are derived from large language models (LLMs) in the NLP field. After fine-tuning, LLMs can be widely applied to downstream tasks. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question-answering, and knowledge graphs.

[0086] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0087] Autonomous driving technology refers to the ability of a vehicle to drive itself without a driver. It typically includes technologies such as high-precision maps, environmental perception, computer vision, behavioral decision-making, path planning, and motion control. Autonomous driving encompasses multiple development paths, including single-vehicle intelligence, vehicle-road collaboration, and networked cloud control. Autonomous driving technology has broad application prospects, currently in logistics, public transportation, taxis, and smart transportation, and is expected to further develop in the future.

[0088] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The solution provided in the embodiments of this application involves computer vision technology, machine learning / deep learning technology, etc.

[0089] Two-dimensional proteomics (TPP, Thermal Proteome Profiling) is based on the principle that protein stability increases after binding to a ligand. Thermal proteomics evaluates changes in protein thermal stability by heating drug-treated and untreated cells (or protein extracts) and measuring the remaining soluble protein at each temperature, thereby determining the drug target. It is currently the only biophysical strategy that can identify drug targets in living cells or even tissues. Drug discovery mainly involves three aspects: drugs, targets, and disease phenotypes. Finding effective drugs and drug targets is one of the most widely used applications of proteomics. Proteomics can provide rich protein expression in cells or tissues. By comparing the differences in protein expression profiles between healthy or diseased tissues, cells or body fluids, differentially expressed proteins are discovered. These proteins may be potential biomarkers or drug targets.

[0090] Thermoproteomics (also known as cellular thermal shift assay (CETSA)) possesses powerful target discovery capabilities and is widely used in drug discovery and basic biology. Numerous different paradigms have emerged from these initial thermoproteomics approaches. For example, temperature-dependent thermoproteomic analysis (TPP-TR) and compound concentration-dependent thermoproteomic analysis (TPP-CCR) focus on analyzing protein stability changes under single-dimensional conditions (a single compound concentration / temperature profile corresponding to multiple temperature / compound concentration profiles), thereby identifying compound targets. Recently, two-dimensional thermoproteomic analysis (2D-TPP) has been improved upon this approach by integrating temperature-dependent and isothermal ligand concentration-dependent experiments into a single analytical scheme. 2D-TPP enables simultaneous analysis of the effects of temperature and compound concentration on protein thermal stability, enabling more precise quantification and direct estimation of compound-protein affinity. Furthermore, the proteomic global stability analysis (PISA) method has been developed by pooling samples from different temperature conditions. Hybrid thermoproteomic analysis (Hybrid-TPP) has been developed by expanding the range of temperature conditions by using the lowest temperature sample as a bridge.

[0091] The basic principle of thermoproteomics is that when temperatures rise, proteins denature, losing their function and stability. By measuring changes in protein concentration at different temperatures, its thermal stability can be assessed. When drugs or small molecules bind to proteins, they alter their thermal stability, causing changes in protein concentration at specific temperatures. By comparing the thermal stability of proteins before and after treatment, it is possible to identify which proteins the drug or small molecule interacts with.

[0092] In the related art, simple statistical analysis is used to group and score proteins based on the dissolution curve of the protein, for example, grouping based on the characteristic points or change trends of the dissolution curve. However, the method of grouping proteins based on the dissolution curve in the related art usually assumes that the data obeys a certain specific distribution or model. When the data scale is large, these assumptions may no longer hold. Therefore, the above method is difficult to deal with complex biological experimental data. As the amount of data increases, the prediction of the thermal stability of the protein becomes inaccurate, and it is difficult to accurately determine the drug target. In addition, the statistical algorithm is usually calculated based on the entire data set, and the algorithm calculation complexity is high, requiring a large amount of computing resources and time.

[0093] Based on practical technical requirements similar to those described above, the present application provides a method, apparatus, computer device, storage medium, and computer program product for determining a drug target.

[0094] In one embodiment, Figure 1As shown, a method for determining drug targets is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understandable that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0095] Step S101 , obtaining a protein heat map corresponding to the candidate protein; wherein the protein concentration in the protein heat map changes with changes in drug concentration and temperature.

[0096] The protein heat map can include protein concentration data, drug concentration data, and temperature data. Alternatively, protein concentration can be obtained using mass spectrometry data recorded by a mass spectrometer. For example, all peptides P1, P2, P3, etc. corresponding to protein A are found from the mass spectrometry data, and the peptides ranked by a preset number of digits are selected. The average intensity of the peptides ranked by the preset number of digits is taken as the protein concentration. In one exemplary embodiment, the abscissa represents drug concentration, the ordinate represents temperature, and the pixel value represents protein concentration, thereby obtaining a protein heat map.

[0097] In a specific implementation process, refer to Figure 2 As shown in the figure, the experimental process of two-dimensional thermoproteomics includes the following steps: First, cells are grown in the presence of n different drug concentrations, and each sample is treated at m different temperatures. Then, the remaining soluble proteins in the sample are extracted, and the proteins can be digested with trypsin and labeled with TMT (Tandem Mass Tag) for analysis by mass spectrometry. Ultimately, each protein is collected as an m×n matrix of data, called a protein heat map.

[0098] Step S103: extracting the features of the protein heat map based on a feature extraction model to obtain heat map features corresponding to the candidate protein; wherein the feature extraction model is obtained by unsupervised training based on protein heat map samples.

[0099] Among them, the feature extraction model may include a convolutional neural network (CNN) model, a residual neural network (ResNet) model, and a variant network of the convolutional neural network (such as VGG, Inception, MobileNet) model.

[0100] Specifically, the feature extraction model can be trained in an unsupervised manner based on protein heat map samples. In one exemplary embodiment, the network parameters of the feature extraction model can be adjusted by constructing a reconstruction loss. For example, the protein heat map samples are input into the feature extraction model to obtain encoded data, and reconstructed data is obtained based on the encoded data through reverse processing. The reconstruction loss is calculated based on the difference between the reconstructed data and the original data. In one exemplary embodiment, the network parameters of the feature extraction model can also be adjusted by constructing a contrast loss. Among them, the contrast loss can include contrast loss, triple loss, softmax-based loss function, angle-based loss function, etc.

[0101] In the disclosed embodiment, the protein heat map corresponding to the candidate protein is input into the feature extraction network, and the heat map features of the candidate protein are extracted based on the feature extraction network. The number of candidate proteins may include multiple, such as candidate protein A, candidate protein B and candidate protein C, etc. Then the above feature extraction process may include inputting candidate protein A into the feature extraction network to obtain the heat map features corresponding to candidate protein A; inputting candidate protein B into the feature extraction network to obtain the heat map features corresponding to candidate protein B; inputting candidate protein C into the feature extraction network to obtain the heat map features corresponding to candidate protein C. It should be noted that the types and numbers of candidate proteins are not limited to the above examples. In fact, candidate proteins may include all proteins discovered in the experiment.

[0102] Step S105 : clustering the candidate proteins based on the heat map features corresponding to the candidate proteins to obtain a plurality of candidate protein sets.

[0103] Specifically, a variety of clustering methods can be used to cluster candidate proteins, such as hierarchical clustering, density-based spatial clustering of applications with noise (DBSCAN), K-means clustering, and probability model-based clustering (GMM, Gaussian Mixture Model).

[0104] In the specific implementation process, taking the K-Means clustering method as an example, the heat map features of K candidate proteins can be randomly selected as the initial cluster centers. The heat map features of the remaining candidate proteins can be assigned to the cluster corresponding to the cluster center closest to them. For each cluster, the average of the heat map features of all candidate proteins is calculated and used as the new cluster center. This assignment and averaging process is repeated until the cluster center no longer changes or the preset number of iterations is reached.

[0105] Step S107: determining the significance of the protein concentration difference corresponding to the candidate protein set based on the temperature data and protein concentration data of each candidate protein in the candidate protein set; determining the standard reference data of the candidate protein based on the protein concentration difference significance and the protein concentration data of the candidate protein; and determining the drug target based on the standard reference data of each candidate protein in the multiple candidate protein sets.

[0106] The protein concentration difference significance is used to characterize the significance of the difference between proteins before and after drug administration. A higher protein concentration difference significance indicates a greater likelihood of a protein-drug interaction. Because candidate proteins within the same candidate protein set are similar in nature, in one exemplary embodiment, the protein concentration difference significance can be calculated using statistical methods such as the Student's t-test, Analysis of Variance (ANOVA), and Chi-square test to calculate the protein concentration difference significance (P value) for each candidate protein set.

[0107] In a specific implementation, using a t-test as an example, a dissolution curve for a candidate protein can be generated based on the temperature data and protein concentration data of the candidate protein. The protein concentration in the dissolution curve varies with temperature, and the area under the curve of the dissolution curve can be used to characterize the severity of the dissolution curve change. In an exemplary embodiment, the significance of the difference in protein concentration corresponding to the candidate protein set is determined based on the sum of the areas under the curves of the dissolution curves of each candidate protein in the candidate protein set, the mean of the sum of the areas under the curves, and the variance of the sum of the areas under the curves.

[0108] In one exemplary embodiment, a dissolution curve for a candidate protein can be generated based on protein concentration data for the candidate protein, such as data on changes in protein concentration with temperature. Furthermore, based on the actual protein concentration data and the predicted protein concentration data, a goodness of fit of the dissolution curve for the candidate protein can be obtained, such as R-squared (Coefficient of Determination), Root Mean Square Error (RMSE), and Mean Absolute Error (MAE).

[0109] In an exemplary embodiment, based on the protein concentration difference significance corresponding to the candidate protein set and the goodness of fit of the dissolution curve of the candidate protein, the protein concentration difference significance and the goodness of fit can be fused to obtain the standard reference data of the candidate protein. For example, the candidate protein set includes candidate protein set A, candidate protein set B, and candidate protein set C, wherein candidate protein set A includes candidate protein a1, candidate protein b1, and candidate protein c1; candidate protein set B includes candidate protein a2, candidate protein b2, and candidate protein c2. Then the standard reference data of candidate protein a1 can be expressed as the goodness of fit of the dissolution curve of a1 × the protein concentration difference significance of candidate protein set A; the standard reference data of candidate protein b1 can be expressed as the goodness of fit of the dissolution curve of b1 × the protein concentration difference significance of candidate protein set A; the standard reference data of candidate protein a2 can be expressed as the goodness of fit of the dissolution curve of a2 × the protein concentration difference significance of candidate protein set B. Similarly, the calculation method of the standard reference data of other candidate proteins is similar to that of the above embodiment, and this disclosure will not be repeated here.

[0110] In an exemplary embodiment, the candidate proteins in the plurality of candidate protein sets may be ranked based on standard reference data of each candidate protein, and a preset number of candidate proteins ranked high may be selected as drug targets.

[0111] In the above-mentioned method for determining drug targets, the features of the protein heat map are extracted by a feature extraction model, and the extracted heat map features are clustered to obtain a plurality of candidate protein sets, wherein the protein heat map features contain information on protein concentration, drug concentration and temperature. Compared with the traditional method of using statistical methods to simply group the dissolution curves of proteins based on feature points, the embodiment of the present disclosure does not need to assume that the data obeys a specific distribution during the grouping process, but clusters and groups according to the heat map features of the protein, thereby being able to cope with a large amount of protein data and accurately group them. Furthermore, in the embodiment of the present disclosure, the feature extraction model can be trained in an unsupervised manner, which can overcome the technical problem of fewer positive samples in the field of drug targets and ensure the feasibility of classification based on the heat map features of the protein.

[0112] In one embodiment, reference Figure 3 As shown, before extracting the features of the protein heat map based on the feature extraction model, the method further includes:

[0113] Step S301: Acquire a sample set; the sample set includes a plurality of protein heat map samples.

[0114] In the embodiment of the present disclosure, the sample set includes a plurality of protein heat map samples. In the protein heat map samples, the protein concentration changes with changes in drug concentration and temperature.

[0115] Step S303 : Based on the initial feature extraction model, extract the features of the first protein heat map sample in the sample set to obtain a first heat map feature, and extract the features of the second protein heat map sample in the sample set to obtain a second heat map feature.

[0116] Among them, the initial feature extraction model can include a convolutional neural network (CNN) model, a residual neural network (ResNet) model, and a variant network of the convolutional neural network (such as VGG, Inception, MobileNet) model.

[0117] Specifically, a protein heat map sample can be selected from the sample set as a first protein heat map sample, and a protein heat map sample can be selected from the sample set as a second protein heat map sample. The first protein heat map sample is input into an initial feature extraction model to output a first heat map feature, and the second protein heat map sample is input into the initial feature extraction model to output a second heat map feature.

[0118] Step S305 : determining a contrast loss based on the similarity between the first heat map feature and the second heat map feature.

[0119] Step S307: Iteratively adjust the initial feature extraction model based on the contrast loss to obtain a feature extraction model.

[0120] In an exemplary embodiment, a threshold value may be used to determine whether the first heat map feature and the second heat map feature belong to the same category. If they belong to the same category, the first heat map feature and the second heat map feature will be closer, otherwise they will be more dispersed. In a specific implementation process, for example, f(x i ) represents the first heat map feature, f(x j ) represents the second heat map feature, and the contrast loss can be expressed as follows:

[0121]

[0122] Among them, y ij is a label, when f(x i ) and f(x j ) is less than the above threshold, y ij The value is 1 if d is not set, otherwise it is 0. ij f(x i ) and f(xj ), optionally, the distance between them can be calculated using metrics such as Euclidean distance or cosine similarity, where α is the above threshold.

[0123] In another exemplary embodiment, the contrast loss may include a triplet loss. Specifically, the first protein heat map sample in the sample set is input into the initial feature extraction model to output a first heat map feature. A protein heat map sample is obtained from a protein heat map sample of the same category as the first protein heat map sample, as a second protein heat map sample of a positive sample, and the second protein heat map sample of the positive sample is input into the initial feature extraction model to obtain a second heat map feature. A protein heat map sample is obtained from a protein heat map sample of a different category from the first protein heat map sample, as a second protein heat map sample of a negative sample, and the second protein heat map sample of the negative sample is input into the initial feature extraction model to obtain a third heat map feature. The triplet loss can be expressed as follows:

[0124]

[0125] Among them, f(x i ) represents the first heat map feature, Represents the second heat map feature of the positive sample, Represents the second heat map feature of the negative sample, d() represents the distance metric function, and α represents the threshold, which is used to control the distance between the first heat map feature and the second heat map feature of the positive sample, the distance between the first heat map feature and the second heat map feature of the negative sample, and the difference between the two distances.

[0126] In another exemplary embodiment, the contrast loss may include an N-pair Loss loss. Specifically, for example, the first heat map feature is represented as: i , the second heat map feature is represented as: x j In an exemplary embodiment, the first heat map feature x i and the second heatmap feature x j Considered as a category, the first heatmap feature x i Considered as the center of the positive sample and category, the second heat map feature x j Considered as a negative sample, calculate the second heat map feature x j With category center x i The similarity between i,j Finally, the softmaz function is used to convert the similarity into a probability distribution. The goal of N-pair Loss is to maximize the probability of positive sample pairs and minimize the probability of negative sample pairs. It is specifically expressed as follows:

[0127]

[0128] N-pair Loss can effectively handle large-scale data sets and perform transfer learning in different tasks.

[0129] In another exemplary embodiment, the training method of the feature extraction model includes adjusting the model parameters based on Angular Loss (angle-based loss function). Specifically, the first heat map feature is represented by x i ; The second heat map feature is represented by x j . For the first heat map feature x i Perform normalization and get x i ′ ; For the second heat map feature x j Perform normalization and get x j ′ Calculate x i ′ and x j ′ The angle θ i,j For positive sample pairs, we hope that their angle is as small as possible, so we hope to maximize cosθ i,j ; For negative sample pairs, we hope that their angle is as large as possible, so we hope to minimize cosθ i,j The goal of Angular Loss is to maximize the similarity of positive samples, minimize the similarity of negative sample pairs, and ensure that the distance between sample pairs is within a certain range. Angular Loss loss is expressed as follows:

[0130]

[0131] in, represents the angle between the positive sample pairs, Represents the angle between negative sample pairs. Angular Loss can effectively process large-scale datasets and ensure that the distance between sample pairs is within a certain range.

[0132] In the embodiment of the present disclosure, according to the specific application scenario requirements, a contrast loss is determined from the above-mentioned multiple contrast losses, and the initial feature extraction model is iteratively adjusted to obtain a feature extraction model.

[0133] In the above embodiment, unsupervised training of the initial feature extraction model by constructing a contrastive loss eliminates the need for labeling sample data, thus enabling training using a large amount of unlabeled data. Furthermore, contrastive learning can capture subtle differences between samples, thereby providing a more robust and generalizable feature representation.

[0134] In one embodiment, determining the contrast loss based on the similarity between the first heat map feature and the second heat map feature includes:

[0135] Based on the feature matrix corresponding to the second heat map feature, a transposed feature matrix corresponding to the second heat map feature is determined.

[0136] Obtaining a cross-correlation matrix based on a feature matrix corresponding to the first heat map feature and a transposed feature matrix corresponding to the second heat map feature, wherein the cross-correlation matrix is ​​used to represent the similarity between the first heat map feature and the second heat map feature;

[0137] Based on the difference between the cross-correlation matrix and the identity matrix, a contrastive loss is determined.

[0138] Specifically, the feature matrix of the first heat map feature can be expressed as Z A , the feature matrix of the second heat map feature can be expressed as Z B , swap the rows and columns of the feature matrix of the second heat map feature to get Z B The transposed matrix is ​​expressed as (Z B ) T .

[0139] In an exemplary embodiment, the feature matrix of the first heat map feature and the transposed feature matrix of the second heat map feature are fused to obtain a cross-correlation matrix. For example, the cross-correlation matrix is ​​represented by C, then C = ZA × (ZB) T .

[0140] In an exemplary embodiment, the loss L is constructed based on the difference between the cross-correlation matrix and the identity matrix. BT as follows:

[0141]

[0142] Among them C ii and C ij Represents the elements in the cross-correlation matrix C, when L BT The smaller it is, the closer C is to the identity matrix.

[0143] In the above embodiment, a cross-correlation matrix is ​​obtained based on the feature matrix corresponding to the first heat map feature and the transposed feature matrix corresponding to the second heat map feature, wherein the cross-correlation matrix is ​​used to characterize the similarity between the first heat map feature and the second heat map feature. Based on the difference between the cross-correlation matrix and the identity matrix, the contrast loss is determined. When the contrast loss is smaller, the cross-correlation matrix is ​​closer to the identity matrix, indicating that the network expresses the first heat map feature and the second heat map feature more divergently and more independently, thereby optimizing the accuracy of feature extraction of the initial feature extraction model.

[0144] In one embodiment, reference Figure 4 As shown, based on the initial feature extraction model, the features of the first protein heat map sample in the sample set are extracted to obtain the first heat map features, and the features of the second protein heat map sample in the sample set are extracted to obtain the second heat map features, including:

[0145] Step S401: obtaining an original protein heat map sample from the sample set.

[0146] Step S403 : performing a first noise processing on the original protein heat map sample to obtain a processed first protein heat map sample, and performing a second noise processing on the original protein heat map sample to obtain a processed second protein heat map sample.

[0147] Step S405 : Based on the initial feature extraction model, the features of the processed first protein heat map sample are extracted to obtain a first heat map feature, and the features of the processed second protein heat map sample are extracted to obtain a second heat map feature.

[0148] The noise added by the noise processing may include Gaussian noise, salt and pepper noise, speckle noise, fuzzy noise, color noise, etc. Figure 5 As shown, an original first protein heat map sample is obtained from the sample set, for example, represented as X, and a first noise processing is performed on the original protein heat map sample to obtain a processed first protein heat map sample, for example, represented as Y A =X+NoiseA; perform a second noise processing on the original protein heat map sample to obtain a processed second protein heat map sample, for example, represented by Y B =X+NoiseB. In an exemplary embodiment, NoiseA and NoiseB may be Gaussian noises that obey the same distribution, but the noise values ​​are randomly generated random numbers.

[0149] In the above embodiment, by performing the first noise processing and the second noise processing on the original protein heat map sample, a processed first protein heat map sample and a processed second protein heat map sample are obtained, and features of the processed first protein heat map sample are extracted to obtain a first heat map feature, and features of the processed second protein heat map sample are extracted to obtain a second heat map feature. This can increase the richness of protein heat map samples in the sample set, facilitating the training of the feature extraction model.

[0150] In one embodiment, the feature extraction model includes a plurality of feature encoders, which extract features of the protein heat map based on a target feature extraction model, including:

[0151] Inputting the protein heat map into each feature encoder respectively and outputting corresponding sub-heat map features; wherein at least one of the following parameters: convolution kernel size, number of convolution kernels, and number of convolution layers in the two feature encoders is different;

[0152] The features of each sub-heat map are concatenated to obtain the features of the protein heat map.

[0153] In the embodiment of the present disclosure, the feature extraction model may include multiple feature encoders, such as two or more feature encoders. Different feature encoders have different network structures. Generally speaking, the number of feature encoders is not necessarily the greater the better. In an exemplary embodiment, reference Figure 6 As shown, the feature extraction model includes four feature encoders, which may include one or more of a convolution layer, a ReLU transfer function (Activation ReLU), and a pooling layer. In each encoder, at least one of the convolution kernel size, the number of convolution kernels, and the number of convolution layers is different.

[0154] In an exemplary embodiment, the protein heat map is input into each feature encoder respectively, for example, the protein heat map is input into Figure 6 Each of the four feature encoders in

[15] outputs a corresponding sub-heatmap feature. The sub-heatmap features output by the four feature encoders are further concatenated to obtain the protein heatmap feature.

[0155] In the above embodiment, the protein heat map is input into different feature encoders to obtain corresponding sub-heat map features. Different feature encoders can extract features of the protein heat map with different information dimensions, thereby extracting features with strong divergence, which facilitates subsequent accurate grouping based on the protein heat map features.

[0156] In one embodiment, obtaining a protein heat map corresponding to a candidate protein includes:

[0157] Obtain the initial protein heat map corresponding to the candidate protein;

[0158] Data interpolation processing is performed on the initial protein heat map to obtain a protein heat map corresponding to the candidate protein.

[0159] Specifically, the protein concentration, drug concentration and temperature in the initial protein heat map can be obtained through experiments. Figure 2In the example shown, cells are first grown in the presence of n different drug concentrations, and each sample is treated at m different temperatures. The remaining soluble proteins in the sample are then extracted, digested with trypsin, and labeled with TMT (Tandem Mass Tag). Mass spectrometry analysis is performed, ultimately yielding m×n matrix data for each protein, yielding an initial protein heatmap.

[0160] In the disclosed embodiments, at the same drug concentration, a dissolution curve can be plotted based on protein concentration data at different temperatures. This dissolution curve can reflect changes in protein stability as the temperature increases. In specific implementations, due to the large amount of protein data, the actual values ​​of each protein dissolution curve are typically only 5 to 10 points, and the data is relatively sparse. Interpolation can be used to increase the amount of data in the dissolution curve.

[0161] In an exemplary embodiment, the least squares method can be used to fit the melting curve. Based on the temperature data of the point to be inserted and the fitted melting curve, the corresponding protein data is predicted. In a specific implementation process, the fitting formula can be as follows:

[0162]

[0163] Where T represents temperature, and x, b, and plateau represent the parameters to be fitted.

[0164] In another exemplary embodiment, a vector machine model can be used to predict protein data of the insertion point. The vector machine model can include a support vector machine and a general vector machine. Optionally, the general vector machine is a single-layer neural network, a learning machine with an input layer, a hidden layer, and an output layer, which performs better in terms of overfitting restrictions. The input and output process of the vector machine model can be referred to Figure 7 As shown, it includes three steps:

[0165] (1) Middle layer input:

[0166]

[0167] (2) Nonlinear mapping:

[0168]

[0169] (3) Network output:

[0170]

[0171] Among them, W 1 and W 2 is the weight matrix of the network, b is the bias, β iControl parameters for the transfer function.

[0172] In the above embodiment, the protein heat map of the candidate protein is obtained by performing data interpolation processing on the initial protein heat map, so that the protein heat map of the candidate protein contains more data points, which facilitates the subsequent training of the feature extraction model and feature extraction of the protein heat map.

[0173] In one embodiment, reference Figure 8 As shown, data interpolation processing is performed on the initial protein heat map, including:

[0174] Step S801 , obtaining a dissolution curve and a vector machine model of the candidate protein; wherein the protein concentration in the dissolution curve varies with temperature.

[0175] In an exemplary embodiment, the dissolution curve of the candidate protein can be obtained experimentally, such as the above Figure 2 In another exemplary embodiment, the dissolution curve of the candidate protein can be obtained by an initial protein heat map of the candidate protein. For example, the initial protein heat map includes data corresponding to M drug concentrations. Then, the initial protein heat map can extract M dissolution curves, each of which represents protein concentration data and temperature data corresponding to the same drug concentration, and the protein concentration changes with temperature.

[0176] In the embodiment of the present disclosure, the vector machine model may include the model in the above embodiment that can predict protein concentration data based on temperature data.

[0177] Step S803: input the temperature data of the point to be inserted into the vector machine model, and output the protein concentration data of the point to be inserted.

[0178] Specifically, the temperature data for the point to be inserted can be set according to actual needs. Alternatively, interpolation can be performed based on where the experimental temperature data is relatively sparse. The temperature data for the point to be inserted is input into the vector machine model, and the corresponding protein concentration data is output.

[0179] Step S805 : performing interpolation processing on the melting curve based on the temperature data and protein concentration data of the point to be inserted to obtain a processed melting curve.

[0180] In an exemplary embodiment, the temperature data and protein concentration data of the point to be inserted are inserted into the melting curve to form a processed melting curve. For example, the candidate protein melting curve includes: temperature t1, protein concentration x1; temperature t5, protein concentration x5; temperature t6, protein concentration x6. The temperature data of the point to be inserted include t2, t3, and t4, and the protein concentration data predicted by the vector machine model include x2, x3, and x4. After the temperature data and protein concentration data of the point to be inserted are inserted into the melting curve, the data of the processed melting curve includes: temperature t1, protein concentration temperature x1; temperature t2, protein concentration temperature x2; temperature t3, protein concentration temperature x3; temperature t4, protein concentration temperature x4; temperature t5, protein concentration x5; temperature t6, protein concentration x6.

[0181] Step S807: obtaining a protein heat map corresponding to the candidate protein based on the processed melting curve.

[0182] Specifically, the processed dissolution curves can be inserted into the protein heat map according to the corresponding drug concentrations to generate a protein heat map corresponding to the candidate protein.

[0183] In the above embodiment, the protein concentration data of the point to be inserted in the dissolution curve is predicted by the vector machine model. Compared with fitting the dissolution curve based on statistics, the vector machine model does not require a large number of samples for learning and has higher accuracy when the dissolution curve sample points are small.

[0184] In one embodiment, reference Figure 9 As shown, the vector machine model includes a general vector machine, and before obtaining the dissolution curve and vector machine model of the candidate protein, it also includes:

[0185] Step S901: Acquire a sample set; wherein the sample set includes protein concentration data and corresponding temperature data.

[0186] Specifically, refer to Figure 7 As shown, in the embodiment of the present disclosure, the vector machine model includes a general vector machine. The sample set includes protein concentration data and corresponding temperature data.

[0187] Step S903: Input the temperature data of the sample set into the initial general vector machine model, and output the predicted first protein concentration.

[0188] Among them, the network parameters in the initial general vector machine model, such as Figure 7 W in 1 and W 2Initialization can be performed randomly. In an exemplary embodiment, the temperature data of a sample point can be randomly selected from the sample set and input into the initial general vector machine model to output the predicted first protein concentration.

[0189] Step S905 : determining a first loss based on the difference between the predicted first protein concentration and the actual protein concentration.

[0190] In one exemplary embodiment, the mean square error between the predicted first protein concentration and the actual protein concentration is used as the first loss. In another exemplary embodiment, the distance between the predicted first protein concentration and the actual protein concentration is used as the first loss, which is not limited in the present disclosure.

[0191] Step S907 , adjusting the network parameters of the initial general vector machine model, inputting the temperature data of the sample set into the adjusted general vector machine model, and outputting the predicted second protein concentration.

[0192] Specifically, the network parameters of the general vector machine model can be randomly selected for adjustment, for example, Figure 7 As shown, randomly select W 1 or W 2 Make adjustments, for example, in W 1 or W 2 Add a noise disturbance.

[0193] In an embodiment of the present disclosure, the temperature data of the sample set is input again into the adjusted general vector machine model, wherein the temperature data input again may be the same as the temperature data input previously, or may be different from the temperature data input previously, and the present disclosure does not impose any restrictions on this.

[0194] Step S909 : determining a second loss based on the difference between the predicted second protein concentration and the actual protein concentration.

[0195] In one exemplary embodiment, the mean square error between the predicted second protein concentration and the actual protein concentration is used as the second loss. In another exemplary embodiment, the distance between the predicted second protein concentration and the actual protein concentration is used as the second loss, which is not limited in the present disclosure.

[0196] Step S911: When the second loss is less than or equal to the first loss, retain the adjusted general vector machine model and continue to adjust the network parameters of the adjusted general vector machine model until the target loss is less than a preset threshold.

[0197] In one exemplary embodiment, when the second loss is less than or equal to the first loss, the adjusted general vector machine model is retained. In another exemplary embodiment, when the second loss is less than or equal to the first loss, the adjusted general vector machine model is discarded, and the pre-adjusted general vector machine model is re-predicted and adjusted until the target loss is less than a preset threshold.

[0198] In the above embodiment, the overfitting vector can be effectively controlled. In each round of training, the general vector machine only adjusts one parameter instead of modifying the entire system, as long as this change does not make the performance worse, it is acceptable. Specifically, during the training process of the vector machine model, there is no need to introduce gradient information when calculating the loss, which is suitable for scenarios where gradients are difficult to solve or the gradient cost is very high; at the same time, there is no need to rely on local gradient information, and it can be explored more flexibly in the search space. Furthermore, since no gradient information is introduced (gradient descent may fall into a local optimal solution), the embodiment of the present disclosure can better avoid local optimal solutions through a large amount of random sampling, which is helpful for global search. In addition, the embodiment of the present disclosure is applicable to random objective functions. If the objective function is random or noisy, gradient descent may be disturbed and produce unstable results, while the embodiment of the present disclosure can offset the influence of noise through multiple sampling to obtain more robust optimization results.

[0199] In summary, the disclosed embodiments are highly flexible and applicable to a variety of neuronal transfer functions and loss functions, as well as continuous, discontinuous, and even discrete system parameters. The above embodiments allow for direct search for appropriate projection vectors in a conventional vector space. Unlike conventional gradient descent algorithms, which require a large number of samples to learn, the general vector machine (GVM) performs better in application scenarios with fewer protein concentration samples.

[0200] In one embodiment, inputting the temperature data of the point to be inserted into the vector machine model includes:

[0201] Obtain the minimum temperature data and the maximum temperature data corresponding to the melting curve.

[0202] The temperature interval between the minimum temperature data and the maximum temperature data is divided into a preset number of sub-temperature intervals.

[0203] Based on the interval endpoints of the sub-temperature interval, temperature data of the point to be inserted is obtained, and the temperature data of the point to be inserted is input into the vector machine model.

[0204] In a specific implementation, for example, there are 10 temperature data points corresponding to the melting curve, and 100 points need to be expanded. Among these 10 points, the minimum temperature data t1 is 37, and the maximum temperature data t2 is 67. In an exemplary embodiment, the temperature range of 37 to 67 is divided. Optionally, the division can be performed at equal intervals, for example, into 101 parts, forming 101 sub-temperature ranges. Then, there are 102 endpoint values ​​of the sub-temperature ranges. Excluding the existing two data points of 37 and 67, the remaining 100 points can be used as the temperatures of the points to be inserted.

[0205] The above embodiment can conveniently and quickly obtain the temperature data of the point to be inserted into the melting curve, and the distribution of the temperature data is relatively balanced.

[0206] In one embodiment, based on the protein concentration difference significance and the protein concentration data of the candidate protein, determining the standard reference data of the candidate protein comprises:

[0207] Based on the protein concentration data for the candidate protein, a dissolution curve for the candidate protein is determined.

[0208] Based on the dissolution curves corresponding to each drug concentration of the candidate protein, the average goodness of fit of the dissolution curves of the candidate protein is determined.

[0209] Based on the average goodness of fit and the significance of the protein concentration difference, the standard reference data of the candidate protein is determined.

[0210] Specifically, protein concentration data and corresponding temperature data of the candidate protein are obtained, and based on the protein concentration data and corresponding temperature data at the same drug concentration, a dissolution curve of the candidate protein is obtained.

[0211] In an exemplary embodiment, for a dissolution curve corresponding to any drug concentration, the goodness of fit of the dissolution curve of the candidate protein at that drug concentration is determined based on the actual data and predicted data of the candidate protein at that drug concentration. For example, the root mean square error (RMS) is used to describe the goodness of fit of the dissolution curve, as expressed by the following formula:

[0212]

[0213] in, Indicates the goodness of fit of the dissolution curve before and after drug addition, y i represents the actual value of protein concentration, represents the predicted value of protein concentration, and n represents the number of data points of protein concentration in the melting curve.

[0214] For another example, the mean absolute error is used to describe the goodness of fit of the melting curve, and the formula is as follows:

[0215]

[0216] in, Indicates the goodness of fit of the dissolution curve before and after drug addition, y i represents the actual value of protein concentration, represents the predicted value of protein concentration, and n represents the number of data points of protein concentration in the melting curve.

[0217] In an exemplary embodiment, the goodness of fit of the dissolution curves corresponding to different drug concentrations may be averaged to obtain the goodness of fit of the protein dissolution curve.

[0218] In an exemplary embodiment, based on the average goodness of fit and the significance of the protein concentration difference, the standard reference data of the candidate protein is determined, which can be expressed as follows:

[0219]

[0220] Where p represents the significance of the difference in protein concentration. It represents the goodness of fit of the dissolution curve before and after drug addition, and M represents the number of drug concentration gradients.

[0221] In the above embodiment, the p-value represents the significance of the difference in protein concentration corresponding to the candidate protein set, and the p-value of each candidate protein in the same candidate protein set is the same. It represents the average goodness of fit of the candidate protein dissolution curve. The average goodness of fit of each candidate protein is different, so the standard reference data Score of each candidate protein is different.

[0222] Furthermore, in the standard reference data of the candidate protein, the p-value range is between 0 and 1, -log(p) is a number greater than 0, and the smaller the p-value, the more obvious the significant difference. At this time, the larger the -log(p), the larger the score. It represents the goodness of fit, ranging from 0 to 1. The larger the value, the better the curve fitting effect. Therefore, in the above embodiment, the candidate protein with a larger standard score is more likely to become a drug target.

[0223] In one embodiment, determining the average goodness of fit of the dissolution curves of the candidate protein based on the dissolution curves corresponding to each drug concentration of the candidate protein comprises:

[0224] Based on the dissolution curves of the candidate proteins in the candidate protein set at the drug concentration, actual data, predicted data and average data of the actual data of the protein concentration of the candidate proteins are determined.

[0225] The goodness of fit of the candidate protein at the drug concentration is determined based on the actual data of the protein concentration, the predicted data and the average data of the actual data.

[0226] The goodness of fit at each drug concentration is averaged to obtain the average goodness of fit of the dissolution curve of the candidate protein.

[0227] The actual data of the protein concentration of the candidate protein may include data from actual experiments, the predicted data may include fitting data of the dissolution curve or data predicted by a vector machine model, and the average data of the actual data may include the average value of the above actual data.

[0228] In a specific implementation, for example, the actual data of protein concentration in the melting curve X is represented by y i , the predicted data of protein concentration is expressed as The average data of the actual data is expressed as The goodness of fit corresponding to melting curve 1 can be expressed as follows:

[0229]

[0230] In an exemplary embodiment, the goodness of fit of each melting curve is averaged to obtain the average goodness of fit of the candidate protein, for example Where M represents the number of drug concentration gradients.

[0231] In the above embodiment, the goodness of fit of the dissolution curve is determined by using the actual data, predicted data and average data of the protein concentration of the candidate protein, so that the degree of fit between the dissolution curve and the actual data and the prediction accuracy can be accurately evaluated.

[0232] In one embodiment, determining the significance of the difference in protein concentration corresponding to the candidate protein set based on the temperature data and protein concentration data of each candidate protein in the candidate protein set includes:

[0233] determining the area under the curve of the melting curve corresponding to each candidate protein based on the temperature data and protein concentration data of each candidate protein in the candidate protein set;

[0234] The areas under the curves of the melting curves are fused to obtain the fused area under the curves;

[0235] Based on the area under the fusion curve and the distribution function, the significance of the difference in protein concentration of the candidate proteins is determined.

[0236] Specifically, for example, the candidate protein set includes candidate protein A, candidate protein B, and candidate protein C. In an exemplary embodiment, the area under the curve of the melting curve corresponding to the candidate protein can be calculated by integration. For example, the area under the curve of candidate protein A is expressed as S A , the area under the curve of candidate protein B is denoted as S B , the area under the curve of candidate protein C is denoted as S C In an exemplary embodiment, the fusion process of the areas under the curves of the melting curves may include summing the areas under the curves. For example, the fused area under the curve is represented by S AUC =S A +S B +S C In an exemplary embodiment, the mean of the area under the curve of the same candidate protein set fusion can be further calculated, which is expressed as and the variance of the area under the curve of the same candidate protein set fusion var(S AUC ).

[0237] In an exemplary embodiment, the significance of the difference in protein concentration of the candidate proteins is determined based on the fusion area under the curve, the mean of the fusion area under the curve, and the variance of the fusion area under the curve as follows:

[0238]

[0239] Where N represents the number of candidate proteins in the candidate protein set.

[0240] In the above embodiment, scipy.stats.t.sf represents a function in the scipy function library for calculating the p-value of the t distribution.

[0241] In the above embodiment, the temperature data and protein concentration data for each candidate protein in the candidate protein set are used to determine the area under the curve (AUC) of the dissolution curve corresponding to each candidate protein. Based on the combined AUC and distribution function, the significance of the protein concentration difference for each candidate protein is determined. Because the AUC reflects the severity of the dissolution curve change, the calculated significance of the difference can reflect the change in protein concentration before and after drug addition, thereby selecting candidate proteins with a high degree of change as drug targets.

[0242] In a specific embodiment, the method of the present application can be applied to the scenario of determining drug targets. In the related art, simple statistical analysis is used to group and score proteins based on the dissolution curve of the protein, for example, grouping is based on the characteristic points or change trends of the dissolution curve. However, the method of grouping proteins according to the dissolution curve in the related art usually assumes that the data obeys a certain specific distribution or model. When the data scale is large, these assumptions may no longer hold. Therefore, the above method is difficult to cope with complex biological experimental data. As the amount of data increases, the prediction of the thermal stability of the protein becomes inaccurate, and it is difficult to accurately determine the drug target. The present application provides a method for determining a drug target, which can accurately determine the target of a drug from a large number of candidate proteins.

[0243] The method for determining drug targets in the embodiments of the present disclosure may include the following steps: obtaining an initial protein heat map corresponding to the candidate protein, training a vector machine model, interpolation processing of the initial protein heat map, training a feature extraction model, feature extraction of the processed protein heat map, clustering of heat map features, and calculation of standard reference data for the candidate protein.

[0244] The initial protein heat map corresponding to the candidate protein is obtained. Specifically, the protein concentration, drug concentration and temperature in the initial protein heat map can be obtained through experiments. For example, Figure 2 In the example shown, cells are first grown in the presence of n different drug concentrations, and each sample is treated at m different temperatures. The remaining soluble proteins in the sample are then extracted, digested with trypsin, and labeled with TMT (Tandem Mass Tag), which can then be analyzed by mass spectrometry. Ultimately, m×n matrix data is collected for each protein, yielding an initial protein heatmap.

[0245] Training a general vector machine model, obtaining a sample set; wherein the sample set includes protein concentration data and corresponding temperature data. Inputting the temperature data of the sample set into an initial general vector machine model, outputting a predicted first protein concentration. Determining a first loss based on the difference between the predicted first protein concentration and the actual protein concentration. Adjusting the network parameters of the initial general vector machine model, and inputting the temperature data of the sample set into the adjusted general vector machine model, outputting a predicted second protein concentration. Determining a second loss based on the difference between the predicted second protein concentration and the actual protein concentration. If the second loss is less than or equal to the first loss, retaining the adjusted general vector machine model, and continuing to adjust the network parameters of the adjusted general vector machine model until the target loss is less than a preset threshold.

[0246] In one exemplary embodiment, when the second loss is less than or equal to the first loss, the adjusted general vector machine model is retained. In another exemplary embodiment, when the second loss is less than or equal to the first loss, the adjusted general vector machine model is discarded, and the pre-adjusted general vector machine model is re-predicted and adjusted until the target loss is less than a preset threshold.

[0247] Interpolation processing is performed on the initial protein heat map to obtain a melting curve and vector machine model for the candidate protein; the protein concentration in the melting curve varies with temperature. Temperature data for the point to be inserted is input into the vector machine model, which outputs protein concentration data for the point to be inserted. Based on the temperature data and protein concentration data for the point to be inserted, the melting curve is interpolated to obtain a processed melting curve. Based on the processed melting curve, a protein heat map corresponding to the candidate protein is obtained. In a specific implementation, for example, the melting curve contains 10 temperature data points, and 100 points need to be expanded. Among these 10 points, the minimum temperature data point t1 is 37, and the maximum temperature data point t2 is 67. In one exemplary embodiment, the temperature range from 37 to 67 is segmented. Alternatively, the segmentation can be performed at equal intervals, for example, into 101 sub-temperature ranges, resulting in 101 endpoint values ​​for the sub-temperature ranges. Excluding the existing data points of 37 and 67, the remaining 100 points can be used as the temperatures of the points to be inserted.

[0248] Training a feature extraction model involves obtaining a sample set; the sample set includes multiple protein heat map samples. Based on an initial feature extraction model, features of a first protein heat map sample in the sample set are extracted to obtain a first heat map feature, and features of a second protein heat map sample in the sample set are extracted to obtain a second heat map feature. Based on the similarity between the first heat map feature and the second heat map feature, a contrast loss is determined. Based on the contrast loss, the initial feature extraction model is iteratively adjusted to obtain a feature extraction model. In an exemplary embodiment, determining the contrast loss based on the similarity between the first heat map feature and the second heat map feature includes: determining a transposed feature matrix corresponding to the second heat map feature based on a feature matrix corresponding to the second heat map feature. Based on the feature matrix corresponding to the first heat map feature and the transposed feature matrix corresponding to the second heat map feature, a cross-correlation matrix is ​​obtained, wherein the cross-correlation matrix is ​​used to characterize the similarity between the first heat map feature and the second heat map feature. Based on the difference between the cross-correlation matrix and the identity matrix, a contrast loss is determined. In an exemplary embodiment, determining the contrast loss based on the similarity between the first heat map feature and the second heat map feature includes: determining a transposed feature matrix corresponding to the second heat map feature based on a feature matrix corresponding to the second heat map feature. Obtaining a cross-correlation matrix based on the feature matrix corresponding to the first heat map feature and the transposed feature matrix corresponding to the second heat map feature, wherein the cross-correlation matrix is ​​used to represent the similarity between the first heat map feature and the second heat map feature. Determining the contrast loss based on the difference between the cross-correlation matrix and a unit matrix.

[0249] Feature extraction of the processed protein heat map, specifically, inputting the processed protein heat map corresponding to the candidate protein into a feature extraction network, and extracting the heat map features of the candidate protein based on the feature extraction network. The number of candidate proteins may include multiple, such as candidate protein A, candidate protein B, and candidate protein C, etc., then the above feature extraction process may include inputting candidate protein A into the feature extraction network to obtain the heat map features corresponding to candidate protein A; inputting candidate protein B into the feature extraction network to obtain the heat map features corresponding to candidate protein B; and inputting candidate protein C into the feature extraction network to obtain the heat map features corresponding to candidate protein C. It should be noted that the types and number of candidate proteins are not limited to the above examples. In fact, candidate proteins can include all proteins discovered in the experiment.

[0250] In one embodiment, the feature extraction model includes multiple feature encoders, and based on the target feature extraction model, extracts the features of the protein heat map, including: inputting the protein heat map into each feature encoder respectively, and outputting corresponding sub-heat map features; wherein the following parameters in the two feature encoders are different: at least one of the convolution kernel size, the number of convolution kernels, and the number of convolution layers; splicing the features of each sub-heat map to obtain the features of the protein heat map.

[0251] In the embodiment of the present disclosure, the feature extraction model may include multiple feature encoders, such as two or more feature encoders. Different feature encoders have different network structures. Generally speaking, the number of feature encoders is not necessarily the greater the better. In an exemplary embodiment, reference Figure 6 As shown, the feature extraction model includes four feature encoders, which may include one or more of a convolution layer, a ReLU transfer function (Activation ReLU), and a pooling layer. In each encoder, at least one of the convolution kernel size, the number of convolution kernels, and the number of convolution layers is different.

[0252] In an exemplary embodiment, the protein heat map is input into each feature encoder respectively, for example, the protein heat map is input into Figure 6 Each of the four feature encoders in

[15] outputs a corresponding sub-heatmap feature. The sub-heatmap features output by the four feature encoders are further concatenated to obtain the protein heatmap feature.

[0253] Clustering of heat map features, specifically, a variety of clustering methods can be used to cluster candidate proteins, such as hierarchical clustering (Hierarchical Clustering), density-based spatial clustering of applications with noise (DBSCAN), K-means clustering, probability model-based clustering (GMM, Gaussian Mixture Model), etc.

[0254] In the specific implementation process, taking the K-Means clustering method as an example, the heat map features of K candidate proteins can be randomly selected as the initial cluster centers. The heat map features of the remaining candidate proteins can be assigned to the cluster corresponding to the cluster center closest to them. For each cluster, the average of the heat map features of all candidate proteins is calculated and used as the new cluster center. This assignment and averaging process is repeated until the cluster center no longer changes or the preset number of iterations is reached.

[0255] The standard reference data for the candidate proteins is calculated by determining the area under the curve (AUC) of the melting curve corresponding to each candidate protein in the candidate protein set based on the temperature data and protein concentration data for each candidate protein. The AUCs of the melting curves are fused to obtain a fused AUC. Based on the fused AUC and the distribution function, the significance of the protein concentration difference of the candidate proteins is determined.

[0256] Specifically, for example, the candidate protein set includes candidate protein A, candidate protein B, and candidate protein C. In an exemplary embodiment, the area under the curve of the melting curve corresponding to the candidate protein can be calculated by integration. For example, the area under the curve of candidate protein A is expressed as S A , the area under the curve of candidate protein B is denoted as S B , the area under the curve of candidate protein C is denoted as S C In an exemplary embodiment, the fusion process of the areas under the curves of the melting curves may include summing the areas under the curves. For example, the fused area under the curve is represented by S AUC =S A +S B +S C In an exemplary embodiment, the mean of the area under the curve of the same candidate protein set fusion can be further calculated, which is expressed as and the variance of the area under the curve of the same candidate protein set fusion var(S AUC Finally, the significance of the difference in protein concentration of the candidate proteins can be expressed as shown in Formula 14.

[0257] Based on the protein concentration data of the candidate protein, a dissolution curve of the candidate protein is determined. Based on the dissolution curves of the candidate protein in the candidate protein set at the drug concentration, actual data, predicted data, and average data of the actual data of the protein concentration of the candidate protein are determined. Based on the actual data, the predicted data, and the average data of the actual data of the protein concentration, a goodness of fit of the candidate protein at the drug concentration is determined.

[0258] The goodness of fit at each drug concentration is averaged to obtain an average goodness of fit for the dissolution curve of the candidate protein. Based on the average goodness of fit and the significance of the difference in protein concentration, the standard reference data for the candidate protein is determined. Finally, the standard reference data for the candidate protein can be expressed as shown in Formula 12.

[0259] Figure 10 and Figure 11The comparison of the fitting effects of the two-dimensional heat map of known targets in the embodiments of the present disclosure is shown. Among them, "raw data" represents the original experimental data, and "interpolation data" represents the fitted data. In an exemplary embodiment, the horizontal axis of the heat map represents the drug concentration, the vertical axis of the heat map represents the temperature, and the pixel value of the heat map represents the protein concentration quantified by mass spectrometry. Among them, the drug concentrations include 0μM, 0.02μM, 0.143μM, 1μM and 5μM. It can be seen that the result of using GVM fitting is closer to the distribution law of the experimental data.

[0260] In an exemplary embodiment, UMAP was used to visualize the protein distribution after clustering, and the results were as follows: Figure 12 As shown in the figure, it can be seen that after GVM fitting clustering, most of the targets can be retained, while Sigmoid only leaves a few targets due to fitting failure.

[0261] In an exemplary embodiment, the present disclosure uses a two-dimensional heat map after fitting and interpolation to score and sort proteins. Proteins with higher scores are likely to be potential targets. We also use the Panobinostat dataset for benchmarking. Plot the comparison of the results before and after sorting. Figure 13 shown.

[0262] In an exemplary embodiment, the ranking results are quantified into a unified indicator, the enrichment score, which is used to characterize the sensitivity of the algorithm in finding targets. The enrichment score calculation formula is as follows:

[0263]

[0264] refer to Figure 14 As shown, the GVM algorithm of the embodiment of the present disclosure is better than the existing Sigmoid algorithm.

[0265] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0266] Based on the same inventive concept, embodiments of the present application also provide a drug target determination device for implementing the aforementioned drug target determination method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of the one or more drug target determination device embodiments provided below can be found in the aforementioned limitations of the drug target determination method and will not be further elaborated here.

[0267] In one embodiment, Figure 15 As shown, a device for determining a drug target is provided, comprising:

[0268] Acquisition module 1501 is used to obtain a protein heat map corresponding to the candidate protein; wherein the protein concentration in the protein heat map changes with changes in drug concentration and temperature;

[0269] Extraction module 1503, configured to extract features of the protein heat map based on a feature extraction model to obtain heat map features corresponding to the candidate protein; wherein the feature extraction model is obtained by unsupervised training based on protein heat map samples;

[0270] A clustering module 1505 is configured to perform clustering processing on the candidate proteins based on the heat map features corresponding to the candidate proteins to obtain multiple candidate protein sets;

[0271] Determination module 1507 is used to determine the protein concentration difference significance corresponding to the candidate protein set based on the temperature data and protein concentration data of each candidate protein in the candidate protein set, determine the standard reference data of the candidate protein based on the protein concentration difference significance and the protein concentration data of the candidate protein, and determine the drug target based on the standard reference data of each candidate protein in the multiple candidate protein sets.

[0272] In one embodiment, the apparatus further includes a feature extraction model training module, wherein the feature extraction model training module is configured to:

[0273] Acquire a sample set; the sample set includes a plurality of protein heat map samples;

[0274] Extracting features of a first protein heat map sample in the sample set based on an initial feature extraction model to obtain a first heat map feature, and extracting features of a second protein heat map sample in the sample set to obtain a second heat map feature;

[0275] determining a contrast loss based on a similarity between the first heat map feature and the second heat map feature;

[0276] Based on the contrast loss, the initial feature extraction model is iteratively adjusted to obtain a feature extraction model.

[0277] In one embodiment, the training module of the feature extraction model is further used to:

[0278] Determine a transposed feature matrix corresponding to the second heat map feature based on the feature matrix corresponding to the second heat map feature;

[0279] Obtaining a cross-correlation matrix based on a feature matrix corresponding to the first heat map feature and a transposed feature matrix corresponding to the second heat map feature, wherein the cross-correlation matrix is ​​used to represent the similarity between the first heat map feature and the second heat map feature;

[0280] Based on the difference between the cross-correlation matrix and the identity matrix, a contrastive loss is determined.

[0281] In one embodiment, the training module of the feature extraction model is further used to:

[0282] Obtaining an original protein heat map sample from the sample set;

[0283] performing a first noise processing on the original protein heat map sample to obtain a processed first protein heat map sample, and performing a second noise processing on the original protein heat map sample to obtain a processed second protein heat map sample;

[0284] Based on the initial feature extraction model, the features of the processed first protein heat map sample are extracted to obtain a first heat map feature, and the features of the processed second protein heat map sample are extracted to obtain a second heat map feature.

[0285] In one embodiment, the feature extraction model includes a plurality of feature encoders, and the extraction module is further configured to:

[0286] Inputting the protein heat map into each feature encoder respectively and outputting corresponding sub-heat map features; wherein at least one of the following parameters: convolution kernel size, number of convolution kernels, and number of convolution layers in the two feature encoders is different;

[0287] The features of each sub-heat map are concatenated to obtain the features of the protein heat map.

[0288] In one embodiment, the acquisition module is further configured to:

[0289] Obtain the initial protein heat map corresponding to the candidate protein;

[0290] Data interpolation processing is performed on the initial protein heat map to obtain a protein heat map corresponding to the candidate protein.

[0291] In one embodiment, the acquisition module is further configured to:

[0292] Obtaining a dissolution curve and a vector machine model of the candidate protein; wherein the protein concentration in the dissolution curve changes with temperature;

[0293] Inputting the temperature data of the point to be inserted into the vector machine model, and outputting the protein concentration data of the point to be inserted;

[0294] performing interpolation processing on the melting curve based on the temperature data and protein concentration data of the point to be inserted to obtain a processed melting curve;

[0295] Based on the processed melting curves, a protein heat map corresponding to the candidate protein is obtained.

[0296] In one embodiment, the acquisition module is further configured to:

[0297] Acquire a sample set; wherein the sample set includes protein concentration data and corresponding temperature data;

[0298] inputting the temperature data of the sample set into an initial general vector machine model and outputting a predicted first protein concentration;

[0299] determining a first loss based on a difference between the predicted first protein concentration and the actual protein concentration;

[0300] Adjusting the network parameters of the initial general vector machine model, inputting the temperature data of the sample set into the adjusted general vector machine model, and outputting a predicted second protein concentration;

[0301] determining a second loss based on a difference between the predicted second protein concentration and the actual protein concentration;

[0302] When the second loss is less than or equal to the first loss, the adjusted general vector machine model is retained, and the network parameters of the adjusted general vector machine model are continuously adjusted until the target loss is less than a preset threshold.

[0303] In one embodiment, the acquisition module is further configured to:

[0304] Obtaining minimum temperature data and maximum temperature data corresponding to the melting curve;

[0305] Dividing the temperature interval between the minimum temperature data and the maximum temperature data into a preset number of sub-temperature intervals;

[0306] Based on the interval endpoints of the sub-temperature interval, temperature data of the point to be inserted is obtained, and the temperature data of the point to be inserted is input into the vector machine model.

[0307] In one embodiment, the determining module is further configured to:

[0308] determining a dissolution curve of the candidate protein based on the protein concentration data of the candidate protein;

[0309] Determining the average goodness of fit of the dissolution curves of the candidate protein based on the dissolution curves corresponding to each drug concentration of the candidate protein;

[0310] Based on the average goodness of fit and the significance of the protein concentration difference, the standard reference data of the target protein is determined.

[0311] In one embodiment, the determining module is further configured to:

[0312] determining the area under the curve of the melting curve corresponding to each candidate protein based on the temperature data and protein concentration data of each candidate protein in the candidate protein set;

[0313] The areas under the curves of the melting curves are fused to obtain the fused area under the curves;

[0314] Based on the area under the fusion curve and the distribution function, the significance of the difference in protein concentration of the candidate proteins is determined.

[0315] Each module in the aforementioned drug target determination device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0316] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 16As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store drug target determination data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for determining a drug target is implemented.

[0317] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 17 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a method for determining a drug target is implemented. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse, etc.

[0318] Those skilled in the art will understand that Figure 17The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0319] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0320] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0321] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for determining a drug target, characterized in that: The method comprises: Obtaining a protein heat map corresponding to the candidate protein; wherein the protein concentration in the protein heat map changes with changes in drug concentration and temperature; Extracting features of the protein heat map based on a feature extraction model to obtain heat map features corresponding to the candidate protein; wherein the feature extraction model is obtained by unsupervised training based on protein heat map samples; performing clustering processing on the candidate proteins based on the heat map features corresponding to the candidate proteins to obtain multiple candidate protein sets; Based on the temperature data and protein concentration data of each candidate protein in the candidate protein set, the protein concentration difference significance corresponding to the candidate protein set is determined; based on the protein concentration difference significance and the protein concentration data of the candidate protein, the standard reference data of the candidate protein is determined; based on the standard reference data of each candidate protein in the multiple candidate protein sets, the drug target is determined.

2. The method according to claim 1, characterized in that Before extracting the features of the protein heat map based on the feature extraction model, the method further includes: Acquire a sample set; the sample set includes a plurality of protein heat map samples; Extracting features of a first protein heat map sample in the sample set based on an initial feature extraction model to obtain a first heat map feature, and extracting features of a second protein heat map sample in the sample set to obtain a second heat map feature; determining a contrast loss based on a similarity between the first heat map feature and the second heat map feature; Based on the contrast loss, the initial feature extraction model is iteratively adjusted to obtain a feature extraction model.

3. The method according to claim 2, characterized in that The determining the contrast loss based on the similarity between the first heat map feature and the second heat map feature includes: Determine a transposed feature matrix corresponding to the second heat map feature based on the feature matrix corresponding to the second heat map feature; Obtaining a cross-correlation matrix based on a feature matrix corresponding to the first heat map feature and a transposed feature matrix corresponding to the second heat map feature, wherein the cross-correlation matrix is ​​used to represent the similarity between the first heat map feature and the second heat map feature; Based on the difference between the cross-correlation matrix and the identity matrix, a contrastive loss is determined.

4. The method according to claim 2, characterized in that Extracting features of a first protein heat map sample in the sample set based on an initial feature extraction model to obtain a first heat map feature, and extracting features of a second protein heat map sample in the sample set to obtain a second heat map feature, including: Obtaining an original protein heat map sample from the sample set; performing a first noise processing on the original protein heat map sample to obtain a processed first protein heat map sample, and performing a second noise processing on the original protein heat map sample to obtain a processed second protein heat map sample; Based on the initial feature extraction model, the features of the processed first protein heat map sample are extracted to obtain a first heat map feature, and the features of the processed second protein heat map sample are extracted to obtain a second heat map feature.

5. The method according to claim 1, wherein The feature extraction model includes a plurality of feature encoders, and extracts features of the protein heat map based on a target feature extraction model, including: Inputting the protein heat map into each feature encoder respectively and outputting corresponding sub-heat map features; wherein at least one of the following parameters: convolution kernel size, number of convolution kernels, and number of convolution layers in the two feature encoders is different; The features of each sub-heat map are concatenated to obtain the features of the protein heat map.

6. The method according to claim 1, characterized in that Get the protein heatmap corresponding to the candidate protein, including: Obtain the initial protein heat map corresponding to the candidate protein; Data interpolation processing is performed on the initial protein heat map to obtain a protein heat map corresponding to the candidate protein.

7. The method according to claim 6, characterized in that Performing data interpolation processing on the initial protein heat map, including: Obtaining a dissolution curve and a vector machine model of the candidate protein; wherein the protein concentration in the dissolution curve changes with temperature; Inputting the temperature data of the point to be inserted into the vector machine model, and outputting the protein concentration data of the point to be inserted; performing interpolation processing on the melting curve based on the temperature data and protein concentration data of the point to be inserted to obtain a processed melting curve; Based on the processed melting curves, a protein heat map corresponding to the candidate protein is obtained.

8. The method according to claim 7, characterized in that The vector machine model includes a general vector machine, and before obtaining the dissolution curve and vector machine model of the candidate protein, further includes: Acquire a sample set; wherein the sample set includes protein concentration data and corresponding temperature data; inputting the temperature data of the sample set into an initial general vector machine model and outputting a predicted first protein concentration; determining a first loss based on a difference between the predicted first protein concentration and the actual protein concentration; Adjusting the network parameters of the initial general vector machine model, inputting the temperature data of the sample set into the adjusted general vector machine model, and outputting a predicted second protein concentration; determining a second loss based on a difference between the predicted second protein concentration and the actual protein concentration; When the second loss is less than or equal to the first loss, the adjusted general vector machine model is retained, and the network parameters of the adjusted general vector machine model are continuously adjusted until the target loss is less than a preset threshold.

9. The method according to claim 7, characterized in that The step of inputting the temperature data of the point to be inserted into the vector machine model comprises: Obtaining minimum temperature data and maximum temperature data corresponding to the melting curve; Dividing the temperature interval between the minimum temperature data and the maximum temperature data into a preset number of sub-temperature intervals; Based on the interval endpoints of the sub-temperature interval, temperature data of the point to be inserted is obtained, and the temperature data of the point to be inserted is input into the vector machine model.

10. The method according to claim 1, characterized in that Determining standard reference data of the candidate protein based on the protein concentration difference significance and the protein concentration data of the candidate protein, comprising: determining a dissolution curve of the candidate protein based on the protein concentration data of the candidate protein; Determining the average goodness of fit of the dissolution curves of the candidate protein based on the dissolution curves corresponding to each drug concentration of the candidate protein; Based on the average goodness of fit and the significance of the protein concentration difference, the standard reference data of the candidate protein is determined.

11. The method according to claim 10, characterized in that The step of determining the average goodness of fit of the dissolution curves of the candidate protein based on the dissolution curves corresponding to each drug concentration of the candidate protein comprises: Determining actual data, predicted data, and average data of the actual data of the protein concentration of the candidate protein in the candidate protein set based on the dissolution curves of the candidate protein at the drug concentration; determining a goodness of fit of the candidate protein at the drug concentration based on the actual data of the protein concentration, the predicted data, and an average of the actual data; The goodness of fit at each drug concentration is averaged to obtain the average goodness of fit of the dissolution curve of the candidate protein.

12. The method according to claim 1, characterized in that The step of determining the protein concentration difference significance corresponding to the candidate protein set based on the temperature data and protein concentration data of each candidate protein in the candidate protein set includes: determining the area under the curve of the melting curve corresponding to each candidate protein based on the temperature data and protein concentration data of each candidate protein in the candidate protein set; The areas under the curves of the melting curves are fused to obtain the fused area under the curves; Based on the area under the fusion curve and the distribution function, the significance of the difference in protein concentration of the candidate proteins is determined.

13. A device for determining a drug target, characterized in that: The device comprises: An acquisition module is used to obtain a protein heat map corresponding to the candidate protein; wherein the protein concentration in the protein heat map changes with the change of drug concentration and temperature; An extraction module, configured to extract features of the protein heat map based on a feature extraction model to obtain heat map features corresponding to the candidate protein; wherein the feature extraction model is obtained by unsupervised training based on protein heat map samples; a clustering module, configured to perform clustering processing on the candidate proteins based on the heat map features corresponding to the candidate proteins to obtain a plurality of candidate protein sets; A determination module is configured to determine the protein concentration difference significance corresponding to the candidate protein set based on the temperature data and protein concentration data of each candidate protein in the candidate protein set, determine the standard reference data of the candidate protein based on the protein concentration difference significance and the protein concentration data of the candidate protein, and determine the drug target based on the standard reference data of each candidate protein in the multiple candidate protein sets.

14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

Citation Information

Cited By

  • Photocatalytic efficiency measuring method and device based on laser speckle and deep learning

    CN121275751A