Method and device for determining drug target action relationship based on artificial intelligence

By acquiring and extracting drug molecular image data and protein sequence data, combining knowledge maps and two-task prediction models, the accuracy of confirming the relationship between drug target action in the prior art is solved, and the efficiency of drug use is improved.

CN114360639BActive Publication Date: 2025-05-02PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210028223.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-11
Publication Date
2025-05-02
Estimated Expiration
2042-01-11

AI Technical Summary

Technical Problem

The prior art cannot accurately determine the relationship between drug targets, resulting in low drug use efficiency.

Method used

By obtaining drug molecular image data and protein sequence data, molecular structure characterization information and protein target characterization information are extracted, and feature fusion coefficients are obtained based on the knowledge graph. Then, the fused features are predicted using the trained two-task prediction model to obtain the drug target action relationship.

Benefits of technology

It improves the accuracy of confirming the relationship between drug target action, increases the diversity of drug target action relationship, meets the need for confirming drug targets in multiple diseases, and thus improves drug use efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360639B_ABST
    Figure CN114360639B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for determining drug target action relationship based on artificial intelligence, which relates to the field of intelligent medical processing technology, and the main purpose is to solve the problem that the existing drug target action relationship cannot be accurately confirmed. It includes: obtaining drug molecule image data and protein sequence data of the target drug; extracting molecular structure representation information from the drug molecule image data, and extracting protein target representation information from the protein sequence data; obtaining feature fusion coefficients matching the molecular structure representation information and the protein target representation information from the knowledge graph, and performing feature fusion on the molecular structure representation information and the protein target representation information based on the feature fusion coefficient; performing prediction processing on the molecular structure representation information and the protein target representation information after feature fusion based on the trained dual-task prediction model, and obtaining the prediction results as the drug target action relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent medical treatment technology, and in particular to a method and device for determining a drug target action relationship based on artificial intelligence. Background Art

[0002] In recent years, the application field of intelligent medical technology has gradually developed from clinical treatment to drug research and development. More and more artificial intelligence technologies are involved in the analysis of the applicability of drugs to different diseases, so as to accurately find drug targets. In particular, the molecular structure of drugs is studied, so as to determine the relationship between suitable drug targets and different diseases based on drug characteristics, and then use them as the basis for treatment.

[0003] At present, the determination of the action relationship of existing drug targets for different diseases is usually determined directly based on the chemical properties of drug molecules. However, the action relationship of drug targets confirmed based on chemical properties for diseases is relatively simple and cannot be effectively applied to the confirmation of drug targets for multiple diseases, which makes the confirmation accuracy of the action relationship of drug targets poor, resulting in low drug use efficiency. Therefore, a drug target action relationship determination method based on artificial intelligence is urgently needed to solve the above problems. Summary of the invention

[0004] In view of this, the present invention provides a method and device for determining the drug-target action relationship based on artificial intelligence, the main purpose of which is to solve the problem that the existing drug-target action relationship cannot be accurately confirmed.

[0005] According to one aspect of the present invention, a method for determining a drug-target interaction relationship based on artificial intelligence is provided, comprising:

[0006] Obtain drug molecule image data and protein sequence data of target drugs;

[0007] Extracting molecular structure representation information from the drug molecule image data, and extracting protein target representation information from the protein sequence data;

[0008] Acquire a feature fusion coefficient matching the molecular structure representation information and the protein target representation information from the knowledge graph, and perform feature fusion on the molecular structure representation information and the protein target representation information based on the feature fusion coefficient;

[0009] Based on the trained dual-task prediction model, the molecular structure characterization information and the protein target characterization information after feature fusion are predicted and processed, and the obtained prediction results are used as the drug-target action relationship.

[0010] Furthermore, before extracting the molecular structure characterization information from the drug molecule image data, the method further includes:

[0011] Construct an isomorphic network model of unlabeled compound graphs;

[0012] Using the adjacency matrix and attribute information in the drug molecule graph training data, as well as the connection edges, as input parameters of the unlabeled compound graph isomorphic network model to perform model training, and obtain a trained molecular feature prediction model;

[0013] The extracting of molecular structure characterization information from the drug molecule image data comprises:

[0014] The drug molecule image data is predicted and processed based on the trained molecular feature prediction model to obtain molecular structure characterization information.

[0015] Furthermore, before extracting protein target characterization information from the protein sequence data, the method further includes:

[0016] Constructing an unlabeled protein sequence language network model;

[0017] Using the protein sequence training data for word embedding as input parameters of the unlabeled protein sequence language network model to perform model training, thereby obtaining a trained protein sequence target prediction model;

[0018] The extracting protein target characterization information from the protein sequence data comprises:

[0019] The protein sequence data is predicted based on the trained protein sequence target prediction model to obtain protein target characterization information.

[0020] Furthermore, before obtaining the feature fusion coefficient matching the molecular structure representation information and the protein target representation information from the knowledge graph, the method further includes:

[0021] Constructing a knowledge graph based on a magnetic resonance diffusion tensor imaging dataset, wherein the knowledge graph contains at least two nodes corresponding to different molecular structure representation information and different protein target representation information, respectively, wherein the two nodes are linked through an action relationship;

[0022] The step of acquiring the feature fusion coefficient matching the molecular structure representation information and the protein target representation information from the knowledge graph includes:

[0023] The interaction relationship corresponding to the molecular structure characterization information and the protein target characterization information is searched from the knowledge graph, and the interaction relationship is converted into a feature fusion coefficient, wherein the feature fusion coefficient includes an interaction relationship coefficient and an affinity coefficient between the molecular structure and the protein target.

[0024] Furthermore, before the molecular structure characterization information and the protein target characterization information after feature fusion are predicted based on the trained dual-task prediction model and the obtained prediction results are used as the drug-target action relationship, the method further includes:

[0025] A two-layer feedforward neural network model is constructed, and the two-layer feedforward neural network model is trained based on feature fusion training sample data to obtain a dual-task prediction model that has completed training, wherein the input parameters of the dual-task prediction model are fused molecular structure characterization information and protein target characterization information, and the dual-task prediction model is used to perform dual output processing including classification prediction tasks and regression prediction tasks to obtain a drug-target action relationship including interaction relationship prediction results and affinity prediction results.

[0026] Furthermore, after the molecular structure characterization information and the protein target characterization information after feature fusion are predicted based on the trained dual-task prediction model and the prediction results are used as the drug-target action relationship, the method further includes:

[0027] Retrieving a preset drug target associated structure image database, wherein the preset drug target associated structure image database stores drug molecule associated structure image data matching different interaction relationship coefficients and different affinity coefficients;

[0028] Drug molecule associated structure image data matching the drug target action relationship is searched from the preset drug target associated structure image database and outputted.

[0029] Furthermore, the method further comprises:

[0030] If drug molecule associated structure image data matching the drug target action relationship is not found in the preset drug target associated structure image database, the drug target action relationship including the interaction relationship prediction result and the affinity prediction result is output to indicate the matching of artificial drug target action relationship.

[0031] According to another aspect of the present invention, there is provided a device for determining a drug-target interaction relationship based on artificial intelligence, comprising:

[0032] An acquisition module, used to acquire drug molecule image data and protein sequence data of a target drug;

[0033] An extraction module, used to extract molecular structure representation information from the drug molecule image data, and to extract protein target representation information from the protein sequence data;

[0034] A determination module, used to obtain a feature fusion coefficient matching the molecular structure representation information and the protein target representation information from the knowledge graph, and perform feature fusion on the molecular structure representation information and the protein target representation information based on the feature fusion coefficient;

[0035] The processing module is used to perform prediction processing on the molecular structure characterization information and the protein target characterization information after feature fusion based on the trained dual-task prediction model, and the obtained prediction results are used as the drug-target action relationship.

[0036] Furthermore, the device further comprises: a first construction module, a first training module,

[0037] The first construction module is used to construct an unlabeled compound graph isomorphic network model;

[0038] The first training module is used to perform model training using the adjacency matrix and attribute information and connection edges in the drug molecule graph training data as input parameters of the unlabeled compound graph isomorphic network model to obtain a trained molecular feature prediction model;

[0039] The processing unit is used to perform prediction processing on the drug molecule image data based on the trained molecular feature prediction model to obtain molecular structure representation information.

[0040] Furthermore, the device further comprises: a second construction module, a second training module,

[0041] The second construction module is used to construct an unlabeled protein sequence language network model;

[0042] The second training module is used to perform model training using word embedding of protein sequence training data as input parameters of the unlabeled protein sequence language network model to obtain a trained protein sequence target prediction model;

[0043] The processing unit is used to perform prediction processing on the protein sequence data based on the trained protein sequence target prediction model to obtain protein target characterization information.

[0044] Furthermore, the device also includes:

[0045] A third construction module is used to construct a knowledge graph based on the magnetic resonance diffusion tensor imaging dataset, wherein the knowledge graph contains at least two nodes corresponding to different molecular structure representation information and different protein target representation information, wherein the two nodes are linked through an action relationship;

[0046] The determination module is specifically used to search the interaction relationship corresponding to the molecular structure representation information and the protein target representation information from the knowledge graph, and convert the interaction relationship into a feature fusion coefficient, wherein the feature fusion coefficient includes the interaction relationship coefficient and the affinity coefficient between the molecular structure and the protein target.

[0047] Furthermore, the device also includes:

[0048] The third training module is used to construct a two-layer feedforward neural network model, and train the two-layer feedforward neural network model based on feature fusion training sample data to obtain a dual-task prediction model that has completed training, wherein the input parameters of the dual-task prediction model are the fused molecular structure characterization information and the protein target characterization information, and the dual-task prediction model is used to perform dual output processing including classification prediction tasks and regression prediction tasks to obtain a drug-target action relationship including interaction relationship prediction results and affinity prediction results.

[0049] Furthermore, the device also includes:

[0050] A retrieval module, used to retrieve a preset drug target associated structure image database, wherein the preset drug target associated structure image database stores drug molecule associated structure image data matching different interaction relationship coefficients and different affinity coefficients;

[0051] The output module is used to search for drug molecule associated structure image data matching the drug target action relationship from the preset drug target associated structure image database and output it.

[0052] Furthermore, the output module is also used to output the drug target action relationship including the interaction relationship prediction results and the affinity prediction results if the drug molecule association structure image data matching the drug target action relationship is not found in the preset drug target association structure image database, so as to indicate the matching of the artificial drug target action relationship.

[0053] According to another aspect of the present invention, a storage medium is provided, wherein at least one executable instruction is stored in the storage medium, and the executable instruction enables a processor to perform operations corresponding to the above-mentioned artificial intelligence-based drug-target action relationship determination method.

[0054] According to another aspect of the present invention, there is provided a computer device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus;

[0055] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute operations corresponding to the above-mentioned artificial intelligence-based drug-target action relationship determination method.

[0056] By means of the above technical solution, the technical solution provided by the embodiment of the present invention has at least the following advantages:

[0057] The present invention provides a method and device for determining a drug target action relationship based on artificial intelligence. Compared with the prior art, the embodiment of the present invention obtains drug molecule image data and protein sequence data of a target drug; extracts molecular structure representation information from the drug molecule image data, and extracts protein target representation information from the protein sequence data; obtains a feature fusion coefficient matching the molecular structure representation information and the protein target representation information from a knowledge graph, and performs feature fusion on the molecular structure representation information and the protein target representation information based on the feature fusion coefficient; performs prediction processing on the molecular structure representation information and the protein target representation information after feature fusion based on a trained dual-task prediction model, and uses the obtained prediction result as the drug target action relationship, thereby increasing the diversity of confirmation of the drug target-disease action relationship, meeting the requirements of drug targets for multiple diseases, thereby improving the confirmation accuracy of the drug target action relationship, and greatly improving the efficiency of drug use.

[0058] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0060] Figure 1 A flow chart of a method for determining a drug-target interaction relationship based on artificial intelligence provided by an embodiment of the present invention is shown;

[0061] Figure 2 A schematic diagram of the structure of a dual-task prediction model provided by an embodiment of the present invention is shown;

[0062] Figure 3 A flow chart of another method for determining drug-target interaction relationship based on artificial intelligence provided by an embodiment of the present invention is shown;

[0063] Figure 4 A flow chart of another method for determining drug-target interaction relationship based on artificial intelligence provided by an embodiment of the present invention is shown;

[0064] Figure 5 A schematic diagram of the structure of a protein sequence target prediction model provided by an embodiment of the present invention is shown;

[0065] Figure 6 A block diagram of a drug-target action relationship determination device based on artificial intelligence provided by an embodiment of the present invention is shown;

[0066] Figure 7 A schematic structural diagram of a computer device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0067] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0068] The embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0069] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0070] Based on this, in one embodiment, Figure 1As shown, a method for determining drug-target interaction relationships based on artificial intelligence is provided, and the method is applied to a computer device such as a server as an example for explanation, wherein the server can be an independent server, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms, such as intelligent medical systems, digital medical platforms, etc. The above method includes the following steps:

[0071] 101. Obtain drug molecule image data and protein sequence data of the target drug.

[0072] In an embodiment of the present invention, the execution subject may be an intelligent management system with data processing functions, such as an intelligent medical system, a data medical platform, etc. Exemplarily, the current execution subject is an intelligent medical system, the target drug is a related drug suitable for matching the drug characteristics with the protein for drug target action relationship, and correspondingly, the drug molecular structure image data of the target drug is a molecule of the target drug represented by a graph structure, wherein the image content in the drug molecular structure image data is the atomic-chemical bond structure of the target drug molecule, and the characteristic content of the molecular structure such as the spatial characteristics, atomic number, charge number, etc. in the form of nodes and edges can be abstracted from the image content, and the protein sequence data is used to represent the data representing the protein composed of 20 different letters (amino acids) arranged and combined, and the length of the protein sequence is generally hundreds or thousands, and the protein sequence data is the sorting content of the letters corresponding to all amino acids, for example, the letters of glycine-g, alanine-b, valine-j, etc. represent the sorting content as b1-j2-g3-b4..., so as to perform feature extraction based on the drug molecular image data and protein sequence data.

[0073] It should be noted that the drug molecular structure image data in the embodiment of the present invention is obtained by loading the drug molecular structure image data of the target drug generated by the intelligent medical system as the current execution subject based on the computer software for making molecular structure diagrams. At this time, the operator can obtain the drug molecular structure image data matching the target drug based on the drug database already stored in the current intelligent medical system, or can make it through the molecular structure making application and obtain it in the specified file format in the intelligent medical system, which is not specifically limited in the embodiment of the present invention. At the same time, the protein sequence data can be pre-entered by the operator, or directly loaded based on the existing protein sequence data, which is not specifically limited in the embodiment of the present invention.

[0074] 102. Extract molecular structure representation information from the drug molecule image data, and extract protein target representation information from the protein sequence data.

[0075] In order to improve the matching accuracy based on the feature fusion coefficient between drug molecular image data and protein sequence data, features are extracted for drug molecular image data and protein sequence data respectively, that is, molecular structure representation information is extracted from drug molecular image data, and protein target representation information is extracted from protein sequence. Among them, molecular structure representation information is used to describe the content of the molecular structure of the main or specific features in drug molecular image data, and protein target representation information is used to describe the content of the protein target of the main or specific features in protein sequence data. Specifically, the protein target is the binding site of the drug expected to be with the protein sequence, so as to determine whether it has an action relationship with the target drug.

[0076] 103. Obtain a feature fusion coefficient that matches the molecular structure representation information and the protein target representation information from the knowledge graph, and perform feature fusion on the molecular structure representation information and the protein target representation information based on the feature fusion coefficient.

[0077] In an embodiment of the present invention, in order to improve the accuracy of determining the action relationship between a drug and a protein target, before processing based on a dual-task prediction model, the feature fusion coefficients corresponding to the molecular structure characterization information and the protein target characterization information after feature extraction are determined based on the knowledge graph to perform feature fusion. Specifically, the knowledge graph is constructed based on a magnetic resonance diffusion tensor imaging data set DTI, and the knowledge graph contains at least two nodes corresponding to different molecular structure characterization information and different protein target characterization information, and the action relationship between each two nodes is represented by a link, so that the feature fusion coefficient is calculated based on the action relationship between the links, and feature fusion is performed according to this feature fusion coefficient. At this time, the length of the link is used to characterize the size of the action relationship, such as the longer the link, the smaller the corresponding action relationship. At the same time, by pre-configuring a threshold for the length of the link, if the link length is a "several" multiple of the threshold, the feature fusion coefficient is a "several" tenth, thereby calculating the feature fusion coefficient. After the feature fusion coefficient is determined, in order to process the molecular structure characterization information and the protein target characterization information by the dual-task prediction model, feature fusion is performed, that is, when the molecular structure characterization information and the protein target characterization information are converted into feature vectors, the feature vectors of the two are normalized to a numerical interval according to the feature fusion coefficient, so as to obtain the input parameters that can be used as the dual-task prediction model for prediction processing. For example, the molecular structure characterization information and the protein target characterization information are converted into feature vectors respectively, that is, the text or image data is converted into a vector matrix. Then, during the conversion, the feature fusion coefficient is used as the conversion coefficient, and each vector matrix is ​​multiplied to obtain the feature vector matrix of the molecular structure characterization information and the protein target characterization information within a numerical range, so as to use this feature vector matrix as the input parameter of the dual-task prediction model, and the embodiment of the present invention does not make specific limitations.

[0078] 104. Based on the trained dual-task prediction model, the molecular structure characterization information and the protein target characterization information after feature fusion are predicted and processed, and the obtained prediction results are used as the drug-target action relationship.

[0079] In the embodiment of the present invention, the dual-task prediction model is a hybrid neural network model with two input parameters and one output result, such as Figure 2As shown, the molecular structure characterization information Drug1 containing the molecular structure and the protein target characterization information Target1 containing the protein sequence are converted into feature vectors respectively, and then used as two input parameters for model input, and prediction processing is performed based on the trained dual-task prediction model, so as to obtain the prediction result as the drug-target interaction relationship. At this time, the drug-target interaction relationship is represented by the interaction relationship value between the drug molecular feature and the protein target feature. For example, the drug molecular feature of the target drug a includes the atom s-chemical bond 2, and the drug-target interaction relationship with the protein target b4-j6 (alanine-b, valine-j, corresponding to the fourth ranking of alanine and the sixth ranking of valine) is 0.4, which means that the drug-target interaction relationship is poor at this time, or it can be determined by presetting the drug-target interaction relationship threshold whether there is a strong drug-target interaction relationship, which is not specifically limited in the embodiment of the present invention.

[0080] In one embodiment of the present invention, in order to further define and illustrate, as Figure 3 As shown, before extracting the molecular structure characterization information from the drug molecule image data in step 102, the method further includes:

[0081] 201. Constructing an isomorphic network model of unlabeled compound graphs;

[0082] 202. Performing model training using the adjacency matrix and attribute information and connection edges in the drug molecule graph training data as input parameters of the unlabeled compound graph isomorphic network model to obtain a trained molecular feature prediction model;

[0083] Correspondingly, extracting molecular structure characterization information from the drug molecule image data specifically includes:

[0084] 203. Perform prediction processing on the drug molecule image data based on the trained molecular feature prediction model to obtain molecular structure characterization information.

[0085] In an embodiment of the present invention, a graph isomorphism network model is trained on a large amount of unlabeled drug molecule graph training data to obtain a general graph isomorphism network model for migration, so as to support training sample data of indefinite data. An embodiment of the present invention constructs an unlabeled compound graph isomorphism network model, that is, an unlabeled compound GIN model is constructed through a graph isomorphism network GIN model (GraphIsomorphism Network). Among them, the input parameter of the unlabeled compound GIN model is the structural content of image data with graph nodes or edge attributes, that is, the adjacency matrix A of the image data and the corresponding attribute information X. At the same time, the chemical molecular graph training data of the drug is used as the sample data for model training. GIN is based on the adjacency matrix of the molecular image data and the attribute information of each graph node (such as atoms), as well as the information of the connecting edges (such as chemical bonds) between them. In each iteration, each graph node updates its own information by aggregating the features of neighbor nodes and its own features in the previous layer, and usually also performs nonlinear transformation on the aggregated information. By stacking multiple layers of networks, each graph node can obtain the neighbor node information within the corresponding number of hops. For drug molecule image data, the latent vector of a single graph node cannot represent the chemical molecule well. In order to represent the overall information of the molecule from the topological structure of the image data, the information vector representation of the entire image data is finally obtained through pooling. That is, a latent vector rich in structural information is used to represent the overall information representation of the image data, thereby completing the model training of the unlabeled compound graph isomorphic network model to obtain a molecular feature prediction model. After the molecular feature prediction model is obtained, it is processed based on the drug molecule image data that needs to be processed to obtain the molecular structure representation information of feature extraction.

[0086] Among them, the data expression form of the molecular feature prediction model is:

[0087]

[0088] in, is the representation content of each graph node, h g is the representation content of the entire molecular graph, K is the number of model iteration layers, V is the graph node, G is the number of graph nodes or the number of connecting edges, and ε is the graph coefficient.

[0089] In one embodiment of the present invention, in order to further define and illustrate, as Figure 4 As shown, before extracting the molecular structure characterization information from the drug molecule image data in step 102, the method further includes:

[0090] 301. Constructing an unannotated protein sequence language network model;

[0091] 302. Using the protein sequence training data for word embedding as input parameters of the unlabeled protein sequence language network model to perform model training, thereby obtaining a trained protein sequence target prediction model;

[0092] Correspondingly, extracting molecular structure characterization information from the drug molecule image data specifically includes:

[0093] 303. Perform prediction processing on the protein sequence data based on the trained protein sequence target prediction model to obtain protein target characterization information.

[0094] In an embodiment of the present invention, a language network model is trained on a large amount of unlabeled protein sequence training data to obtain a general language network model for migration, so as to support training sample data of indefinite data. An unlabeled protein sequence language network model is constructed in an embodiment of the present invention, that is, an unlabeled protein sequence language network model is constructed through a language representation model BERT (Bidirectional Encoder Representation from Transformer). Among them, the unlabeled protein sequence language network model BERT is trained on a large amount of unlabeled protein sequence training data. Since proteins are composed of 20 different letters (such as amino acids) arranged and combined, and the length of the protein sequence is hundreds or thousands, each protein sequence data is word embedded and used as the input parameter of the unlabeled protein sequence language network model Bert for training, such as Figure 5 As shown, a protein sequence target prediction model is finally obtained, and after obtaining the protein sequence data, the protein sequence data is predicted and processed based on the protein sequence target prediction model to obtain protein target characterization information.

[0095] In one embodiment of the present invention, for further limitation and explanation, before step 103 obtains feature fusion coefficients matching the molecular structure representation information and the protein target representation information from the knowledge graph, the method further includes: constructing a knowledge graph based on a magnetic resonance diffusion tensor imaging dataset.

[0096] In an embodiment of the present invention, in order to determine the feature fusion coefficient based on the knowledge graph, a knowledge graph is pre-constructed based on a magnetic resonance diffusion tensor imaging dataset. The magnetic resonance diffusion tensor imaging dataset is a dataset obtained by performing diffusion tensor imaging (DTI) on protein cells based on nuclear magnetic resonance imaging (MRI). By scanning and imaging molecules, atoms, etc., image content containing the movement direction between molecules and atoms is obtained, thereby serving as a kind of knowledge information for linking various nodes of different molecular structure representation information and protein target representation information. Specifically, the knowledge graph contains at least two nodes corresponding to different molecular structure representation information and different protein target representation information, respectively, wherein the two nodes are linked by an action relationship, so that the action relationship and the corresponding feature fusion coefficient are determined by the link.

[0097] Correspondingly, the step of obtaining a feature fusion coefficient that matches the molecular structure representation information and the protein target representation information from the knowledge graph includes: searching the knowledge graph for the interaction relationship corresponding to the molecular structure representation information and the protein target representation information, and converting the interaction relationship into a feature fusion coefficient.

[0098] In the embodiment of the present invention, since the drug molecular structure image data is represented by G = (A, X), where A and X represent the adjacency matrix and the feature matrix respectively, the number of graph nodes of the image data is n, and the feature dimension of the graph node is d, if the molecular feature prediction model GIN learns an f-dimensional output for each graph node: H1 = Pool (GIN (A, X)) ∈ R f×1 A protein sequence data is S. If the protein sequence target prediction model Bert learns an f-dimensional output for each sequence: H2 = Bert(S)∈R f×1 On this basis, each molecular structure characterization information and protein target characterization information in the knowledge graph can be represented by H1 and H2 respectively, so as to find the matching interaction relationship. At this time, the interaction relationship is the link corresponding to each node. After the link is determined, the interaction relationship weight is converted based on the distance of this link. The feature fusion coefficient includes the interaction relationship coefficient and affinity coefficient between the molecular structure and the protein target. For example, by pre-configuring the interaction relationship threshold or affinity threshold for the length of the link, if the link length is a "several" multiple of the threshold, the interaction relationship coefficient or affinity coefficient in the feature fusion coefficient is converted to a "several" tenths, so as to serve as the interaction relationship coefficient and affinity coefficient. The embodiment of the present invention does not make specific limitations.

[0099] In one embodiment of the present invention, in order to further limit and illustrate, the step is to predict and process the molecular structure characterization information and the protein target characterization information after feature fusion based on the trained dual-task prediction model, and before the obtained prediction result is used as the drug-target action relationship, the method also includes: constructing a two-layer feedforward neural network model, and training the two-layer feedforward neural network model based on feature fusion training sample data to obtain a dual-task prediction model that has completed the training.

[0100] In order to achieve the prediction of dual objectives through the prediction model of dual input tasks, thereby improving the interaction relationship between the dual objectives, such as Figure 2 As shown, a two-layer feedforward neural network model is constructed to train the dual-task prediction model by taking the molecular structure characterization information fused in the feature fusion training sample data and the protein target characterization information as the input parameters of the dual-task prediction model. Wherein, the dual-task prediction model is used to perform dual output processing including classification prediction tasks and regression prediction tasks to obtain a drug-target action relationship including interaction relationship prediction results and affinity prediction results. At this time, the interaction relationship prediction result is the interaction relationship coefficient between the molecular structure and the protein target, and the affinity prediction result is the relationship between the molecular structure and the protein target when there is no antagonism, that is, it is expressed as the affinity coefficient between the molecular structure and the protein target, wherein the antagonism is that the molecular structure of the drug has a therapeutic or mitigating effect on the protein target, so that the protein target can be treated or antagonized by this drug, and then, the affinity is manifested as the mutual promotion or beneficial effect between the molecular structure of the drug and the protein target. In the embodiment of the present invention, the interaction relationship prediction result and the affinity prediction result are both represented by numerical values, and the embodiment of the present invention is not specifically limited. At this time, the role of the feature fusion coefficient containing the interaction relationship coefficient and the affinity coefficient that matches the molecular structure characterization information and the protein target characterization information determined based on the knowledge graph is used for feature fusion, and the predicted interaction relationship prediction results and affinity prediction results are obtained based on artificial intelligence processing. Therefore, the feature fusion coefficient is different from the predicted interaction relationship prediction results and affinity prediction results, and the embodiment of the present invention does not make specific limitations.

[0101] In one embodiment of the present invention, in order to further limit and illustrate, step 104 performs prediction processing on the molecular structure characterization information and the protein target characterization information after feature fusion based on the trained dual-task prediction model, and the obtained prediction result is used as the drug target action relationship. The method also includes: calling a preset drug target associated structure image database; searching for drug molecule associated structure image data that matches the drug target action relationship from the preset drug target associated structure image database, and outputting it.

[0102] In order to meet the needs of operators to obtain other drugs that have interaction relationships or affinity associations with the target drug, after the drug-target action relationship is processed, matching is performed based on the preset drug-target association structure image database in the intelligent medical system. Wherein, the preset drug-target association structure image database stores drug molecule association structure image data matched with different interaction relationship coefficients and different affinity coefficients. At this time, the interaction relationship coefficient and affinity coefficient in the drug-target action relationship as the prediction result are compared and matched with each interaction relationship coefficient and affinity coefficient in the preset drug-target association structure image database, thereby obtaining matched drug molecule association structure image data. At this time, the matched drug molecule association structure image data can be output as other drugs with an associated effect with the target drug pushed to the operating user, so that the operator can perform other drug operations based on this drug molecule association structure image data, which is not specifically limited in the embodiment of the present invention.

[0103] In one embodiment of the present invention, for further limitation and explanation, it also includes: if the drug molecule association structure image data matching the drug target action relationship is not found in the preset drug target association structure image database, the drug target action relationship including the interaction relationship prediction result and the affinity prediction result is output to indicate the matching of the artificial drug target action relationship.

[0104] In order to improve the effectiveness of determining the effects of drug targets and flexibly determine the drug target action relationship for the drug molecule associated structure image data, when the drug molecule associated structure image data matching the drug target action relationship is not found in the preset drug target associated structure image database of the intelligent medical system, it means that there are no other drug molecules associated with the drug target action relationship system to push in the intelligent medical system. Therefore, the drug target action relationship including the interaction relationship prediction results and the affinity prediction results is directly output to instruct the operator to match the manual drug target action relationship.

[0105] The embodiment of the present invention provides a method for determining a drug-target action relationship based on artificial intelligence. Compared with the prior art, the embodiment of the present invention obtains drug molecule image data and protein sequence data of a target drug; extracts molecular structure characterization information from the drug molecule image data, and extracts protein target characterization information from the protein sequence data; obtains a feature fusion coefficient matching the molecular structure characterization information and the protein target characterization information from a knowledge graph, and performs feature fusion on the molecular structure characterization information and the protein target characterization information based on the feature fusion coefficient; performs prediction processing on the molecular structure characterization information and the protein target characterization information after feature fusion based on a trained dual-task prediction model, and obtains a prediction result as the drug-target action relationship, thereby increasing the diversity of confirmation of the drug target-disease action relationship, meeting the requirements of drug targets for multiple diseases, thereby improving the accuracy of confirmation of the drug-target action relationship, and greatly improving the efficiency of drug use.

[0106] Furthermore, as a response to the above Figure 1 The embodiment of the present invention provides a drug-target action relationship determination device based on artificial intelligence, such as Figure 6 As shown, the device comprises:

[0107] An acquisition module 41 is used to acquire drug molecule image data and protein sequence data of a target drug;

[0108] An extraction module 42, used to extract molecular structure representation information from the drug molecule image data, and to extract protein target representation information from the protein sequence data;

[0109] A determination module 43 is used to obtain a feature fusion coefficient matching the molecular structure representation information and the protein target representation information from the knowledge graph, and perform feature fusion on the molecular structure representation information and the protein target representation information based on the feature fusion coefficient;

[0110] The processing module 44 is used to perform prediction processing on the molecular structure characterization information and the protein target characterization information after feature fusion based on the trained dual-task prediction model, and the obtained prediction results are used as the drug-target action relationship.

[0111] Furthermore, the device also includes: a first construction module, a first training module,

[0112] The first construction module is used to construct an unlabeled compound graph isomorphic network model;

[0113] The first training module is used to perform model training using the adjacency matrix and attribute information and connection edges in the drug molecule graph training data as input parameters of the unlabeled compound graph isomorphic network model to obtain a trained molecular feature prediction model;

[0114] The processing unit is used to perform prediction processing on the drug molecule image data based on the trained molecular feature prediction model to obtain molecular structure representation information.

[0115] Furthermore, the device further comprises: a second construction module, a second training module,

[0116] The second construction module is used to construct an unlabeled protein sequence language network model;

[0117] The second training module is used to perform model training using word embedding of protein sequence training data as input parameters of the unlabeled protein sequence language network model to obtain a trained protein sequence target prediction model;

[0118] The processing unit is used to perform prediction processing on the protein sequence data based on the trained protein sequence target prediction model to obtain protein target characterization information.

[0119] Furthermore, the device also includes:

[0120] A third construction module is used to construct a knowledge graph based on the magnetic resonance diffusion tensor imaging dataset, wherein the knowledge graph contains at least two nodes corresponding to different molecular structure representation information and different protein target representation information, wherein the two nodes are linked through an action relationship;

[0121] The determination module is specifically used to search the interaction relationship corresponding to the molecular structure representation information and the protein target representation information from the knowledge graph, and convert the interaction relationship into a feature fusion coefficient, wherein the feature fusion coefficient includes the interaction relationship coefficient and the affinity coefficient between the molecular structure and the protein target.

[0122] Furthermore, the device also includes:

[0123] The third training module is used to construct a two-layer feedforward neural network model, and train the two-layer feedforward neural network model based on feature fusion training sample data to obtain a dual-task prediction model that has completed training, wherein the input parameters of the dual-task prediction model are the fused molecular structure characterization information and the protein target characterization information, and the dual-task prediction model is used to perform dual output processing including classification prediction tasks and regression prediction tasks to obtain a drug-target action relationship including interaction relationship prediction results and affinity prediction results.

[0124] Furthermore, the device also includes:

[0125] A retrieval module, used to retrieve a preset drug target associated structure image database, wherein the preset drug target associated structure image database stores drug molecule associated structure image data matching different interaction relationship coefficients and different affinity coefficients;

[0126] The output module is used to search for drug molecule associated structure image data matching the drug target action relationship from the preset drug target associated structure image database and output it.

[0127] Furthermore, the output module is also used to output the drug target action relationship including the interaction relationship prediction results and the affinity prediction results if the drug molecule association structure image data matching the drug target action relationship is not found in the preset drug target association structure image database, so as to indicate the matching of the artificial drug target action relationship.

[0128] The embodiment of the present invention provides a drug-target action relationship determination device based on artificial intelligence. Compared with the prior art, the embodiment of the present invention obtains drug molecule image data and protein sequence data of the target drug; extracts molecular structure representation information from the drug molecule image data, and extracts protein target representation information from the protein sequence data; obtains a feature fusion coefficient matching the molecular structure representation information and the protein target representation information from a knowledge graph, and performs feature fusion on the molecular structure representation information and the protein target representation information based on the feature fusion coefficient; performs prediction processing on the molecular structure representation information and the protein target representation information after feature fusion based on a trained dual-task prediction model, and obtains the prediction result as the drug-target action relationship, which increases the diversity of the confirmation of the drug target-disease action relationship, meets the requirements of multi-disease drug targets, thereby improving the confirmation accuracy of the drug-target action relationship and greatly improving the drug use efficiency.

[0129] According to one embodiment of the present invention, a storage medium is provided, wherein the storage medium stores at least one executable instruction, and the computer executable instruction can execute the artificial intelligence-based drug-target action relationship determination method in any of the above method embodiments.

[0130] Figure 7 A schematic diagram of the structure of a computer device provided according to an embodiment of the present invention is shown. The specific embodiment of the present invention does not limit the specific implementation of the computer device.

[0131] like Figure 7As shown, the computer device may include: a processor (processor) 502 , a communication interface (Communications Interface) 504 , a memory (memory) 506 , and a communication bus 508 .

[0132] The processor 502 , the communication interface 504 , and the memory 506 communicate with each other via a communication bus 508 .

[0133] The communication interface 504 is used to communicate with other devices such as clients or other servers.

[0134] The processor 502 is used to execute the program 510, and specifically can execute the relevant steps in the above-mentioned embodiment of the drug-target action relationship determination method based on artificial intelligence.

[0135] Specifically, the program 510 may include program codes, which include computer operation instructions.

[0136] The processor 502 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the computer device may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0137] The memory 506 is used to store the program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0138] The program 510 may be specifically configured to enable the processor 502 to perform the following operations:

[0139] Obtain drug molecule image data and protein sequence data of target drugs;

[0140] Extracting molecular structure representation information from the drug molecule image data, and extracting protein target representation information from the protein sequence data;

[0141] Acquire a feature fusion coefficient matching the molecular structure representation information and the protein target representation information from the knowledge graph, and perform feature fusion on the molecular structure representation information and the protein target representation information based on the feature fusion coefficient;

[0142] Based on the trained dual-task prediction model, the molecular structure characterization information and the protein target characterization information after feature fusion are predicted and processed, and the obtained prediction results are used as the drug-target action relationship.

[0143] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0144] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for determining drug-target interaction relationships based on artificial intelligence, characterized in that: include: Obtain drug molecule image data and protein sequence data of target drugs; Extracting molecular structure representation information from the drug molecule image data, and extracting protein target representation information from the protein sequence data; Acquire a feature fusion coefficient matching the molecular structure representation information and the protein target representation information from the knowledge graph, and perform feature fusion on the molecular structure representation information and the protein target representation information based on the feature fusion coefficient, wherein the feature fusion is to normalize the feature vectors of the molecular structure representation information and the protein target representation information to a numerical range according to the feature fusion coefficient when performing feature vector conversion on the molecular structure representation information and the protein target representation information; Based on the trained dual-task prediction model, the molecular structure characterization information and the protein target characterization information after feature fusion are predicted and processed, and the obtained prediction results are used as the drug-target action relationship; Wherein, before acquiring the feature fusion coefficient matching the molecular structure representation information and the protein target representation information from the knowledge graph, the method further includes: Constructing a knowledge graph based on a magnetic resonance diffusion tensor imaging dataset, wherein the knowledge graph contains at least two nodes corresponding to different molecular structure representation information and different protein target representation information, respectively, wherein the two nodes are linked through an action relationship; The step of acquiring the feature fusion coefficient matching the molecular structure representation information and the protein target representation information from the knowledge graph includes: The interaction relationship corresponding to the molecular structure characterization information and the protein target characterization information is searched from the knowledge graph, and the interaction relationship is converted into a feature fusion coefficient, wherein the feature fusion coefficient includes an interaction relationship coefficient and an affinity coefficient between the molecular structure and the protein target.

2. The method according to claim 1, characterized in that Before extracting the molecular structure characterization information from the drug molecule image data, the method further includes: Construct an isomorphic network model of unlabeled compound graphs; The adjacency matrix and attribute information in the drug molecule graph training data, as well as the connection edges, are used as input parameters of the unlabeled compound graph isomorphic network model to perform model training to obtain a trained molecular feature prediction model; The extracting of molecular structure characterization information from the drug molecule image data comprises: The drug molecule image data is predicted and processed based on the trained molecular feature prediction model to obtain molecular structure characterization information.

3. The method according to claim 1, characterized in that Before extracting protein target characterization information from the protein sequence data, the method further includes: Constructing an unlabeled protein sequence language network model; Using the protein sequence training data for word embedding as input parameters of the unlabeled protein sequence language network model to perform model training, thereby obtaining a trained protein sequence target prediction model; The extracting protein target characterization information from the protein sequence data comprises: The protein sequence data is predicted based on the trained protein sequence target prediction model to obtain protein target characterization information.

4. The method according to claim 1, characterized in that: Before performing prediction processing on the molecular structure characterization information and the protein target characterization information after feature fusion based on the trained dual-task prediction model and obtaining the prediction results as the drug-target action relationship, the method further comprises: A two-layer feedforward neural network model is constructed, and the two-layer feedforward neural network model is trained based on feature fusion training sample data to obtain a dual-task prediction model that has completed training, wherein the input parameters of the dual-task prediction model are fused molecular structure characterization information and protein target characterization information, and the dual-task prediction model is used to perform dual output processing including classification prediction tasks and regression prediction tasks to obtain a drug-target action relationship including interaction relationship prediction results and affinity prediction results.

5. The method according to claim 4, characterized in that After the molecular structure characterization information and the protein target characterization information after feature fusion are predicted based on the trained dual-task prediction model and the prediction results are used as the drug-target action relationship, the method further includes: Retrieving a preset drug target associated structure image database, wherein the preset drug target associated structure image database stores drug molecule associated structure image data matching different interaction relationship coefficients and different affinity coefficients; Drug molecule associated structure image data matching the drug target action relationship is searched from the preset drug target associated structure image database and outputted.

6. The method according to claim 5, characterized in that The method further comprises: If drug molecule associated structure image data matching the drug target action relationship is not found in the preset drug target associated structure image database, the drug target action relationship including the interaction relationship prediction result and the affinity prediction result is output to indicate the matching of artificial drug target action relationship.

7. A drug-target action relationship determination device based on artificial intelligence, characterized in that: include: An acquisition module, used to acquire drug molecule image data and protein sequence data of a target drug; An extraction module, used to extract molecular structure representation information from the drug molecule image data, and to extract protein target representation information from the protein sequence data; A determination module, used to obtain a feature fusion coefficient matching the molecular structure representation information and the protein target representation information from a knowledge graph, and perform feature fusion on the molecular structure representation information and the protein target representation information based on the feature fusion coefficient, wherein the feature fusion is to normalize the feature vectors of the molecular structure representation information and the protein target representation information to a numerical range according to the feature fusion coefficient when performing feature vector conversion on the molecular structure representation information and the protein target representation information; A processing module, used for predicting and processing the molecular structure characterization information and the protein target characterization information after feature fusion based on the trained dual-task prediction model, and obtaining the prediction result as the drug-target action relationship; The device also includes: A third construction module is used to construct a knowledge graph based on the magnetic resonance diffusion tensor imaging dataset, wherein the knowledge graph contains at least two nodes corresponding to different molecular structure representation information and different protein target representation information, wherein the two nodes are linked through an action relationship; The determination module is specifically used to search the interaction relationship corresponding to the molecular structure representation information and the protein target representation information from the knowledge graph, and convert the interaction relationship into a feature fusion coefficient, wherein the feature fusion coefficient includes the interaction relationship coefficient and the affinity coefficient between the molecular structure and the protein target.

8. A storage medium storing at least one executable instruction, wherein the executable instruction enables a processor to execute operations corresponding to the method for determining drug-target interaction relationships based on artificial intelligence as described in any one of claims 1 to 6.

9. A computer device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the method for determining drug-target action relationships based on artificial intelligence as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Drug and target interaction prediction method and device, equipment and storage medium

    CN113160894A

  • Drug-target interaction prediction method and device, equipment and storage medium

    CN113409897A