Drug relocation interpretation text generation method and device, equipment and medium

By constructing a unified medical graph and optimization model, combining breadth-first search and language model, an interpretive text associated with drugs and diseases is generated, which solves the problem of uninterpretation of prediction results in the prior art and improves the reliability of decision-making.

CN120108554APending Publication Date: 2025-06-06UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510071216.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing drug relocation technologies are difficult to provide interpretability of predictive results, making it difficult for scientific researchers to intuitively understand the prediction logic and mechanism, affecting the reliability of decisions.

Method used

By obtaining biomedical data, building a unified medical graph, performing global and local graph processing, building optimization models, calculating predictive association scores between drugs and diseases, and generating interpreted texts in combination with breadth-first search algorithms and language models.

Benefits of technology

The interpretive text description of drug-disease association is realized, which improves the reliability of predictive data decisions and provides an explanation of the biological mechanism of drug relocation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108554A_ABST
    Figure CN120108554A_ABST
Patent Text Reader

Abstract

The invention provides a drug relocation interpretation text generation method and device, equipment and a medium, and relates to the technical field of drug relocation, and the method comprises the following steps: extracting structural features from biomedical data, and carrying out graph construction through the structural features and a relationship set to obtain a medical unified graph; performing global and local graph processing on the medical unified graph, and according to global features and local features, optimizing a network structure through mutual information maximization and calculating an association score between the medicine and the disease through an inner product to obtain a medicine-disease relationship prediction result; performing path search processing according to a breadth-first search algorithm and the drug-disease relationship prediction result, and calculating a search path entity and relationship embedded aggregation matching score to obtain an interpretation path prediction result; and constructing a language model, and inputting the interpretation path prediction result into the language model for text conversion to generate a relocation interpretation text. The problems that prediction logic and mechanism are difficult to understand visually, and the reliability of a decision made by prediction data is affected are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug repositioning, and in particular to a method, device, equipment and medium for generating drug repositioning explanation text. Background Art

[0002] In the field of drug repositioning technology, machine learning algorithms are generally relied on to conduct in-depth analysis and data mining of drug molecular structures, biomarker targets for specific diseases, and related genomic data to accurately locate drug action sites. The aim is to provide a theoretical basis for new drug research and development by analyzing the interaction patterns between drugs and molecules in the body. However, although current drug repositioning technology has made progress in data processing efficiency and prediction accuracy, there is a general lack of interpretability of prediction results, which makes it difficult for researchers to intuitively understand the prediction logic and mechanism, affecting the reliability of decisions made based on prediction data.

[0003] There is an urgent need for a method, device, equipment and medium for generating drug repositioning explanation text to solve the problem that it is difficult to intuitively understand the prediction logic and mechanism, which affects the reliability of decisions made based on prediction data. Summary of the invention

[0004] The purpose of the present invention is to provide a method, device, equipment and medium for generating drug repositioning explanation text to improve the above problems. In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0005] In a first aspect, the present application provides a method for generating a drug repositioning explanation text, comprising:

[0006] Acquire biomedical data, wherein the biomedical data includes a biomedical entity set and a relationship set, wherein the biomedical entity set includes drugs and diseases;

[0007] Extracting molecular structure features from the drugs in the biomedical data, and constructing a graph using the molecular structure features and the relationship set to obtain a unified medical graph;

[0008] Performing global and local graph processing on the medical unified graph respectively to obtain global features and local features;

[0009] According to the global features and the local features, the network structure is optimized by maximizing mutual information to construct an optimization model, and the prediction association score between the drug and the disease is calculated for the optimization model by a preset inner product operation method to obtain a drug-disease relationship prediction result;

[0010] Performing path search processing according to the breadth-first search algorithm and the drug-disease relationship prediction results, and obtaining explanation path prediction results by calculating the aggregate matching scores of entity and relationship embeddings on each search path;

[0011] A language model is constructed according to biomedical data, and the interpretation path prediction result is input into the language model for text conversion to generate a relocated interpretation text, wherein the relocated interpretation text is an explanatory text description of the association between the drug and the disease.

[0012] In a second aspect, the present application also provides a device for generating a drug repositioning explanation text, comprising:

[0013] An acquisition module, used for acquiring biomedical data, wherein the biomedical data includes a biomedical entity set and a relationship set, and the biomedical entity set includes drugs and diseases;

[0014] A first construction module is used to extract molecular structure features of the drugs in the biomedical data, and to construct a graph through the molecular structure features and the relationship set to obtain a unified medical graph;

[0015] The second construction module is used to perform global and local graph processing on the medical unified graph to obtain global features and local features;

[0016] A third construction module is used to optimize the network structure by maximizing mutual information according to the global features and the local features, to construct an optimization model, and to calculate the predicted association score between the drug and the disease for the optimization model by a preset inner product operation method to obtain a drug-disease relationship prediction result;

[0017] A fourth building block is used to perform path search processing according to a breadth-first search algorithm and the drug-disease relationship prediction result, and obtain an explanation path prediction result by calculating the aggregate matching score of the entity and relationship embedding on each search path;

[0018] The fifth construction module is used to construct a language model based on biomedical data, and input the explanation path prediction result into the language model for text conversion to generate a relocated explanation text, wherein the relocated explanation text is an explanatory text description of the association between drugs and diseases.

[0019] In a third aspect, the present application also provides a device for generating a drug repositioning explanation text, comprising:

[0020] Memory for storing computer programs;

[0021] A processor is used to implement the steps of the drug repositioning explanation text generation method when executing the computer program.

[0022] In a fourth aspect, the present application further provides a medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above-mentioned method for generating an explanation text based on drug repositioning are implemented.

[0023] The beneficial effects of the present invention are:

[0024] The present invention uses a unified medical graph, an optimization model and a language model to construct a unified medical graph by combining molecular structure features with relationships, which is used for information integration and expansion of more biomedical knowledge systems; the optimization model realizes high-level, structured information extraction of entities in the unified medical graph, thereby providing key information for drug repositioning prediction and its explanatory analysis; the association score between drugs and diseases is used for association prediction, and the path is searched through the association score between drugs and diseases to obtain an explanation path, which can be used to explain the biological mechanism of the association between drugs and diseases, and the language model is used to generate text for the explanation path to obtain a repositioning explanation text. This solves the problem that it is difficult to intuitively understand the prediction logic and mechanism, which affects the reliability of the decision made by the prediction data.

[0025] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or be understood by implementing the embodiments of the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0027] Figure 1 It is a schematic diagram of the flow of the method for generating the drug repositioning explanation text described in an embodiment of the present invention;

[0028] Figure 2 Schematic diagram of the structure of a device for generating a drug relocation explanation text according to an embodiment of the present invention.

[0029] Marks in the figure: 800, drug relocation explanation text generation device; 801, processor; 802, memory; 803, multimedia component; 804, I / O interface; 805, communication component. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0031] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0032] Embodiment 1:

[0033] This embodiment provides a method for generating a drug repositioning explanation text.

[0034] See also Figure 1 , the figure shows that the method includes steps S1 to S6, including:

[0035] S1: Acquire biomedical data, where the biomedical data includes a biomedical entity set and a relationship set, where the biomedical entity set includes drugs and diseases;

[0036] S2: extracting molecular structure features from the drugs in the biomedical data, and constructing a graph based on the molecular structure features and the relationship set to obtain a unified medical graph;

[0037] In order to clarify the specific method of obtaining the unified medical map, step S2 includes S21 to S24, which are:

[0038] S21: converting the drugs in the biomedical data to obtain a matrix, and extracting features from the matrix to obtain molecular structure features;

[0039] In order to clarify the specific method of obtaining the molecular structure characteristics, step S21 includes S211 to S213, which are:

[0040] S211: digitally assigning drug molecules in the biomedical data, and iteratively calculating molecular structures of the drug molecules to generate molecular fingerprints;

[0041] In this step, a number is assigned to each substructure of the drug molecule in the biomedical data using the ECFP algorithm, the molecular structure of the digitized drug molecule is iteratively calculated, a molecular fingerprint list related to the atomic arrangement order is generated, and a molecular fingerprint is obtained.

[0042] S212: performing polar coordinate sequence conversion on the molecular fingerprint, and generating a medical entity matrix by performing trigonometric function on the polar coordinate sequence, wherein the medical entity matrix includes the relationship between each substructure;

[0043] In this step, the molecular fingerprint is converted into a token sequence X and scaled to ensure that the scaled sequence Each element in It is scaled to the interval [-1,1], and the polar coordinate system is used to represent the scaled sequence. The GAF algorithm is used to calculate the time relationship at different time intervals and generate a medical entity matrix related to the sequence length through trigonometric functions;

[0044] The molecular fingerprint is converted into a polar coordinate system:

[0045]

[0046] In the above formula (1), θ represents the angle in the polar coordinate system. Represents an element after the sequence X is scaled, where X represents the token sequence. represents the scaling sequence, r represents the radius in the polar coordinate system, t i represents the timestamp, and N represents the constant factor;

[0047] Wherein, N is used to represent the length of the marker sequence X, so as to adjust the span of the polar coordinate system;

[0048] The medical entity matrix is:

[0049]

[0050] In the above formula (2), represents a two-dimensional symmetric matrix, express The transposed matrix of represents the scaling sequence, X represents the token sequence, I represents the unit row vector, and represents the square root of a diagonal matrix;

[0051] The two-dimensional symmetric matrix represents a medical entity matrix.

[0052] S213: Inputting the plurality of substructures into a convolutional neural network for training, and obtaining molecular structure features by learning and extracting features of the spatial relationship between the plurality of substructures.

[0053] S22: integrating data according to the molecular structure features and the relationship set to obtain a medical database;

[0054] In this step, the medical database includes drugs, diseases, genes, proteins and phenotypic entities and the relationships between them.

[0055] S23: Connecting the drugs and diseases in the medical database to obtain a drug-disease dichotomy model;

[0056] S24: constructing a graph based on the drug and disease nodes in the drug-disease dichotomy model and the disease information network in the medical database, integrating information by analyzing the attribute mapping relationship between drug nodes, disease nodes and the disease information network, and obtaining a unified medical graph.

[0057] In this step, the drug nodes and disease nodes in the drug-disease dichotomy model are respectively mapped to the corresponding entities in the disease information network, and then their node sets and edge sets are merged to obtain a unified medical graph; the attribute mapping relationship is the type of entity and the semantic type of relationship, and information integration is performed through the drug-disease dichotomy model and the disease information network to expand the network to include more biomedical knowledge.

[0058] S3: performing global and local graph processing on the medical unified graph respectively to obtain global features and local features;

[0059] In order to clarify the specific method of obtaining global features and local features, step S3 includes S31 to S33, which are:

[0060] S31: performing a truncated random walk on each entity in the medical unified graph to generate a random walk path;

[0061] S32: performing entity global embedding learning on the walk sequence in the random walk path according to a preset word skipping model to obtain global features;

[0062] In this step, the skip-gram model is a Skip-Gram model, Φ(e h ) The Skip-Gram model is used to learn the entity's global state. The Skip-Gram objective function is:

[0063]

[0064] In the above formula (3), represents the minimum model parameter, Φ represents the model parameter, -log represents the negative value of the logarithmic function, q r ({e h-w ,…,e h+w}e h |Φ(e h )) indicates that given Φ(e h ), under the set of other entities that co-occur with entity h in window w, e h-w ,…,e h+w and e h The probability of simultaneous occurrence, e h-w ,…,e h+w represents the set of other entities that co-occur with entity h in window w, w represents the window size of the random walk path, e h represents the embedding vector of entity h, Φ(e h ) represents the global embedding of entity h;

[0065] The set of other entities that co-occur with entity h in the window w is h-w ,…,e h+w and e h The probability expression of simultaneous occurrence is:

[0066]

[0067] In the above formula (4), q r ({e h-w ,…,e h+w}e h |Φ(e h )) indicates that given Φ(e h ), under the set of other entities that co-occur with entity h in window w, e h-w ,…,e h+w and e h The probability of simultaneous occurrence, e h-w ,…,e h+w represents the set of other entities that co-occur with entity h in window w, w represents the window size of the random walk path, e h represents the embedding vector of entity h, φ(e h ) represents the global embedding of entity h, Φ represents the model parameters, m represents the index variable, q r (e m |Φ(e h )) represents the global embedding representation Φ(e) for a given entity h h ), we observe that the embedding e of entity m m The conditional probability of

[0068] Wherein, the entity h represents the first entity in the biomedical entity set, and the Φ(e h ) The global view of entities is learned through the Skip-Gram model, which is used to maximize the conditional probability of entities observed in the random walk path.

[0069] Global embeddingglo(e h ) is calculated by minimizing the Skip-Gram objective function and applying the gradient descent algorithm. The global embedding glo(e) of each entity h h ) is obtained through the conversion function;

[0070] The conversion function is:

[0071] glo(e h )=Φ(e h ) (5)

[0072] In the above formula (5), glo(e h ) represents the global embedding of entity h, Φ(e h ) represents the global embedding of entity h;

[0073] The DeepWalk algorithm captures the global neighborhood information of each entity in the medical unified graph and encodes it to obtain a global embedding.

[0074] S33: Perform a multi-step graph convolution operation calculation on the entity triplet according to the graph convolutional neural network to obtain a local feature, wherein the entity triplet is a semantic relationship between the first entity and the second entity.

[0075] In order to clarify the specific method of obtaining the local features, step S33 includes S331 to S334, which are:

[0076] S331: performing translation learning on the entity triples according to the graph convolutional neural network to obtain a triplet score function;

[0077] In this step, the triple score function is:

[0078] trans(h,r,t)=exp((e h +e r )·e t ) (6)

[0079] In the above formula (6), trans(h, r, t) represents the score function of the entity triple (h, r, t), exp represents the natural exponential function, and e h The embedding vector representing entity h;

[0080] Among them, the entity h represents the first entity in the biomedical entity set, the entity t represents the second entity in the biomedical entity set, and the relationship r represents the semantic relationship between the first entity and the second entity; the TransE model learns the entity triple (h, r, t) through the translation principle to obtain the embedding vector of entity h.

[0081] S332: assigning weights to the triple score function and the relationship in the entity triple based on the attention mechanism to obtain an attention weight result, wherein the relationship in the entity triple is a semantic relationship defined in the relationship set;

[0082] In this step, the attention weight result is:

[0083]

[0084] In the above formula (7), P r represents the attention weight, tanh((e h +e r )·e t ) represents the dot product of the embedding vector of entity h and relation r with the embedding vector of entity t, tanh represents the activation function, ∑ (h,r′,t′)∈G tanh((e h +e r′ )·e t′ ) represents the sum of the scores of entity h reaching all possible tail entities through all possible relations, ∑ (h,r′,t′) represents the sum of all consecutive triplets (h, r′, t′) in the medical unified graph, G represents the medical unified graph, tanh((e h +e r′ )·e t′ ) represents the dot product of the sum of the embedding of entity h and the embedding of relation r′ and the embedding vector of entity t′;

[0085] The entity h represents the first entity in the biomedical entity set, the entity t represents the second entity in the biomedical entity set, the relationship r represents the semantic relationship between the first entity and the second entity, and the relationship e represents the semantic relationship between the first entity and the second entity. r′ and e t′ Represents any other relationships and associated nodes that may exist in the Medical Unified Graph.

[0086] S333: performing iterative error calculation on the first entity and the second entity in the entity triplet to obtain a local error term;

[0087] In this step, the first entity is an entity in the biomedical entity set, and the second entity is another entity in the biomedical entity set;

[0088] The real triplet score obtained by constructing the first entity and the second entity by presetting the fake triplet score;

[0089] The forged triplet score is made lower than the real triplet score, and an iterative error calculation is performed on the forged triplet score and the real triplet score through a pairwise ranking loss function to obtain a local error term.

[0090] The pairwise ranking loss function is:

[0091] L 1 =∑ (h,r,t,t′)∈T -lnμ(trans(h,r,t)-trans(h,r,t′)) (8)

[0092] In the above formula (8), L 1 denotes the pairwise ranking loss function, T is the set of all valid triplets, -ln denotes the negative natural logarithm function, μ denotes the Sigmoid function, trans(h,r,t) denotes the score function of the true triplet (h,r,t), trans(h,r,t′) denotes the score function of the forged triplet (h,r,t′), (h,r,t′) denotes the forged triplet generated by randomly replacing one entity in the true triplet;

[0093] Among them, the entity h represents the first entity in the biomedical entity set, the entity t represents the second entity in the biomedical entity set, the relationship r represents the semantic relationship between the first entity and the second entity, and the pairwise ranking loss function is used to train the TransE model.

[0094] S334: Perform multi-step graph convolution operations to calculate and extract local features based on the attention weight results and the local error terms to obtain local features.

[0095] In this step, the local features are:

[0096]

[0097] In the above formula (9), represents the weighted sum of the neighbor embedding vectors of entity h at the l-1th iteration, δ′ represents the set of relationship types associated with entity h, and P r represents the attention weight, represents the embedding vector of the neighbor entity t at the l-1 layer, loc(e h ) represents the local embedding of entity h, represents the embedding vector of the l-th layer of entity h, ReLU represents the activation function, W represents the trainable parameters, represents the embedding vector of entity h at layer l-1;

[0098] Among them, the ReLU activation function is used to introduce nonlinearity; the entity h represents the first entity in the biomedical entity set, the entity t represents the second entity in the biomedical entity set, and the relationship r represents the semantic relationship between the first entity and the second entity.

[0099] S4: According to the global features and the local features, the network structure is optimized by maximizing mutual information to construct an optimization model, and the prediction association score between the drug and the disease is calculated for the optimization model by a preset inner product operation method to obtain a drug-disease relationship prediction result;

[0100] In order to clarify the specific method of obtaining the drug-disease relationship prediction result, step S4 includes S41 to S44, which are:

[0101] S41: Optimizing the network structure by maximizing mutual information according to the global features and the local features to obtain an optimized model;

[0102] In this step, the loss function for maximizing the mutual information is:

[0103]

[0104] In the above formula (10), L 2 represents the mutual information maximization loss function, represents the sum of all entities h in the medical unified graph, represents the set of all entity embedding vectors, -log represents the negative value of the logarithmic function, τ represents a hyperparameter to prevent overfitting, s is a function that measures the affinity of any local embedding and global embedding of an entity, loc(e h ) represents the local embedding of entity h, e h represents the embedding vector of entity h, glo(e h ) represents the global embedding of entity h;

[0105] The entity h represents the first entity in the biomedical entity set.

[0106] Furthermore, the local features and the global features are fused, and the mutual information maximization strategy is implemented through the optimization layer to optimize the network structure and obtain an optimized model; high-level and structured information extraction of entities in the medical unified graph is achieved, thereby providing key information for the prediction of drug repositioning and its explanatory analysis.

[0107] S42: updating the entity embedding in the medical unified graph based on the optimization model to obtain embedding vectors of drug entities and disease entities;

[0108] S43: performing prediction calculation on the embedding vectors of the drug entity and the disease entity according to a preset inner product calculation method to obtain the relationship between the drug and the disease;

[0109] In this step, the association score is:

[0110]

[0111] In the above formula (11), represents the association score, Represents the drug embedding vector e u and the disease embedding vector e d The dot product of u represents the drug embedding vector, e d represents the disease embedding vector;

[0112] Among them, the association score Used to determine the relationship between drugs and diseases;

[0113] S44: construct a model based on the drug-disease relationship and the optimization model to obtain a prediction model, and adaptively optimize the prediction model by combining the pairwise ranking loss function, the association prediction loss function and the mutual information maximization loss function to obtain a drug-disease relationship prediction result.

[0114] In this step, the pairwise ranking loss function, the association prediction loss function and the mutual information maximization loss function are constructed to obtain the overall objective function;

[0115] The loss function of the association prediction is:

[0116]

[0117] In the above formula (12), L 3 represents the loss function of association prediction, For all triples (u,d i ,d j ), -ln represents the negative natural logarithm function, μ represents the Sigmoid function, represents drug u and disease d i The predicted value of represents drug u and disease d j The predicted value of .

[0118] The overall objective function is:

[0119]

[0120] In the above formula (13), represents the overall objective function, L 1 represents the pairwise ranking loss function, L 2 represents the mutual information maximization loss function, L 3 represents the loss function of association prediction, represents the regularization term, Θ is the set of model parameters;

[0121] The regularization term is used to prevent the model from overfitting.

[0122] S5: performing path search processing according to the breadth-first search algorithm and the drug-disease relationship prediction result, and obtaining an explanation path prediction result by calculating the aggregate matching scores of the entity and relationship embeddings on each search path;

[0123] In order to clearly explain the specific method of obtaining the path prediction result, step S5 includes S51 to S53, which are specifically:

[0124] S51: performing path search processing on the drug entities and disease entities in the biomedical data according to a breadth-first search algorithm to obtain a drug and disease path set;

[0125] S52: performing aggregate information calculation on the entity triples and the entity and relationship embeddings on each search path to obtain a path matching score;

[0126] In this step, the path matching score is:

[0127]

[0128] In the above formula (14), score(p) represents the matching score of the path, p represents a path from drug to disease, (h, r′, t′) represents the continuous triples on the path, and tanh((e h +e r′ )·e t′ ) means ((e h +e r′ )·e t′ )The dot product of the embedding vector, e h The embedding vector representing entity h, e r′ Represents the embedding vector of relation r′, e t′ Embedding vector representing entity t′;

[0129] The continuous triples on the path represent a series of triples constituting the path, and the triples have a sequential relationship on the path and are connected to each other;

[0130] The entity h represents a first entity in the biomedical entity set, the entity t represents a second entity in the biomedical entity set, and the relationship r represents a semantic relationship between the first entity and the second entity.

[0131] S53: Sort the drug and disease pathway sets based on the pathway matching scores to obtain an explanation pathway prediction result.

[0132] In this step, the drug and disease pathway sets are sorted according to the pathway matching scores, and the K pathways with the highest matching scores are used as explanations for the prediction results. The K pathways are the biological mechanisms that best explain the association between the drug and the disease.

[0133] S6: construct a language model based on biomedical data, input the explanation path prediction result into the language model for text conversion to generate a relocated explanation text, wherein the relocated explanation text is an explanatory text description of the association between the drug and the disease.

[0134] In order to clarify the specific method of obtaining the relocated interpretation text, step S6 includes S61 to S63, which are:

[0135] S61: extracting drug names, disease names and professional terms from the biomedical data, and performing word frequency statistics on the drug names, disease names and professional terms to obtain a medical vocabulary;

[0136] In this step, words that appear frequently in drug names, disease names, and professional terms and are important for explaining the drug's mechanism of action are selected to construct a medical vocabulary, and a unique index is assigned to each word in the medical vocabulary.

[0137] S62: interpreting the medical vocabulary based on the deep learning model, combining the medical vocabulary masking and the multi-head self-attention mechanism in the deep learning model to allow focusing on different positions simultaneously when processing the sequence to construct a language model;

[0138] In this step, the language model uses the cross entropy loss function to measure the difference between the probability distribution of the model output and the true target sequence. The Adam optimization algorithm is used to update the parameters of the language model, and neural network regularization techniques and weight decay are applied to prevent overfitting.

[0139] Preferably, the language model comprises an encoder and a decoder;

[0140] The encoder consists of multiple identical layers, each of which contains a multi-head self-attention mechanism and a feed-forward network. The multi-head self-attention mechanism allows the model to focus on different positions simultaneously when processing a sequence. The feed-forward network further processes the medical vocabulary through linear transformations and nonlinear activation functions to introduce nonlinear characteristics. Each layer is followed by residual connections and layer normalization to stabilize the training process and improve the model's ability to learn deep representations.

[0141] The decoder consists of multiple layers identical to the encoder layers, but masks are added in the self-attention layers to avoid information leakage. In addition, the decoder includes an encoder-decoder attention layer that enables the decoder to focus on the most relevant parts of the encoder's output. Like the encoder, each layer of the decoder ends with a residual connection and layer normalization. At the end of the decoder, a linear layer and soft max layer are added to transform the decoder's output into a probability distribution over a medical vocabulary to generate explanatory text. Beam search is adopted as a decoding strategy to improve the quality and coherence of the generated text.

[0142] S63: Input the interpretation path prediction result into the language model for text conversion to generate a relocated interpretation text.

[0143] In this step, deep biological insights and mechanistic explanations are provided for the drug repositioning mission.

[0144] Embodiment 2:

[0145] This embodiment provides a device for generating a drug relocation explanation text, the device comprising:

[0146] An acquisition module, used for acquiring biomedical data, wherein the biomedical data includes a biomedical entity set and a relationship set, and the biomedical entity set includes drugs and diseases;

[0147] A first construction module is used to extract molecular structure features of the drugs in the biomedical data, and to construct a graph through the molecular structure features and the relationship set to obtain a unified medical graph;

[0148] To clarify the specific methods of obtaining the unified medical diagram, the following are provided:

[0149] A first processing unit is used to convert the drugs in the biomedical data to obtain a matrix, and obtain molecular structure features by extracting features from the matrix;

[0150] A second processing unit is used to integrate data according to the molecular structure characteristics and the relationship set to obtain a medical database;

[0151] A third processing unit is used to connect the drugs and diseases in the medical database to obtain a drug-disease dichotomy model;

[0152] The fourth processing unit is used to construct a graph based on the drug and disease nodes in the drug-disease dichotomy model and the disease information network in the medical database, and to integrate information by analyzing the attribute mapping relationship between the drug nodes, disease nodes and the disease information network to obtain a unified medical graph.

[0153] The second construction module is used to perform global and local graph processing on the medical unified graph to obtain global features and local features;

[0154] A third construction module is used to optimize the network structure by maximizing mutual information according to the global features and the local features, to construct an optimization model, and to calculate the predicted association score between the drug and the disease for the optimization model by a preset inner product operation method to obtain a drug-disease relationship prediction result;

[0155] A fourth building block is used to perform path search processing according to a breadth-first search algorithm and the drug-disease relationship prediction result, and obtain an explanation path prediction result by calculating the aggregate matching score of the entity and relationship embedding on each search path;

[0156] The fifth construction module is used to construct a language model based on biomedical data, and input the explanation path prediction result into the language model for text conversion to generate a relocated explanation text, wherein the relocated explanation text is an explanatory text description of the association between drugs and diseases.

[0157] It should be noted that, regarding the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment of the method, and will not be elaborated here.

[0158] Embodiment 3:

[0159] Corresponding to the above method embodiment, this embodiment further provides a device for generating a drug relocation explanation text. The device for generating a drug relocation explanation text described below and the method for generating a drug relocation explanation text described above can refer to each other.

[0160] Figure 2 8 is a block diagram of a drug relocation explanation text generating device 800 according to an exemplary embodiment. Figure 2 As shown, the drug relocation explanation text generating device 800 may include: a processor 801 and a memory 802. The drug relocation explanation text generating device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.

[0161] The processor 801 is used to control the overall operation of the drug repositioning explanation text generation device 800 to complete all or part of the steps in the above-mentioned drug repositioning explanation text generation method. The memory 802 is used to store various types of data to support the operation of the drug repositioning explanation text generation device 800, and these data may include, for example, instructions for any application or method operating on the drug repositioning explanation text generation device 800, and application-related data, such as contact data, sent and received messages, pictures, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, which is used to receive external audio signals. The received audio signal may be further stored in the memory 802 or sent via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 804 provides an interface between the processor 801 and other interface modules, and the above-mentioned other interface modules can be keyboards, mice, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 805 is used for wired or wireless communication between the drug relocation interpretation text generation device 800 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding communication component 805 can include: Wi-Fi module, Bluetooth module, NFC module.

[0162] In an exemplary embodiment, the drug relocation explanation text generation device 800 can be implemented by one or more application specific integrated circuits (Application Specific Integrated Circuit, referred to as ASIC), digital signal processors (Digital Signal Processor, referred to as DSP), digital signal processing devices (Digital Signal Processing Device, referred to as DSPD), programmable logic devices (Programmable Logic Device, referred to as PLD), field programmable gate arrays (Field Programmable Gate Array, referred to as FPGA), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned drug relocation explanation text generation method.

[0163] Embodiment 4:

[0164] Corresponding to the above method embodiment, this embodiment further provides a medium, and the medium described below and the drug repositioning explanation text generation method described above can correspond to each other.

[0165] A medium stores a computer program, which, when executed by a processor, implements the steps of the method for generating a drug repositioning explanation text in the above method embodiment.

[0166] The medium may specifically be a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or other medium that can store program codes.

[0167] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0168] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A method for generating drug repositioning explanation text, characterized in that: include: Acquire biomedical data, wherein the biomedical data includes a biomedical entity set and a relationship set, wherein the biomedical entity set includes drugs and diseases; Extracting molecular structure features from the drugs in the biomedical data, and constructing a graph using the molecular structure features and the relationship set to obtain a unified medical graph; Performing global and local graph processing on the medical unified graph respectively to obtain global features and local features; According to the global features and the local features, the network structure is optimized by maximizing mutual information to construct an optimization model, and the prediction association score between the drug and the disease is calculated for the optimization model by a preset inner product operation method to obtain a drug-disease relationship prediction result; Performing path search processing according to the breadth-first search algorithm and the drug-disease relationship prediction results, and obtaining explanation path prediction results by calculating the aggregate matching scores of entity and relationship embeddings on each search path; A language model is constructed according to biomedical data, and the interpretation path prediction result is input into the language model for text conversion to generate a relocated interpretation text, wherein the relocated interpretation text is an explanatory text description of the association between the drug and the disease.

2. The method for generating drug repositioning explanation text according to claim 1, characterized in that: Extracting molecular structure features from the drugs in the biomedical data, and constructing a graph based on the molecular structure features and the relationship set to obtain a unified medical graph, including: Converting the drugs in the biomedical data to obtain a matrix, and extracting features from the matrix to obtain molecular structure features; Performing data integration according to the molecular structure characteristics and the relationship set to obtain a medical database; Connecting the drugs and diseases in the medical database to obtain a drug-disease dichotomy model; A graph is constructed based on the drug and disease nodes in the drug-disease dichotomy model and the disease information network in the medical database, and information is integrated by analyzing the attribute mapping relationship between drug nodes, disease nodes and the disease information network to obtain a unified medical graph.

3. The method for generating drug repositioning explanation text according to claim 1, characterized in that: The medical unified graph is processed globally and locally to obtain global features and local features, including: Performing a truncated random walk on each entity in the medical unified graph to generate a random walk path; Performing entity global embedding learning on the walk sequence in the random walk path according to a preset word skipping model to obtain global features; A multi-step graph convolution operation is performed on the entity triplet according to the graph convolutional neural network to obtain a local feature, wherein the entity triplet is a semantic relationship between the first entity and the second entity.

4. The method for generating drug repositioning explanation text according to claim 1, characterized in that: According to the global features and the local features, the network structure is optimized by maximizing mutual information, an optimization model is constructed, and the prediction association score between the drug and the disease is calculated for the optimization model by a preset inner product operation method to obtain a drug-disease relationship prediction result, including: Optimizing the network structure by maximizing mutual information according to the global features and the local features to obtain an optimized model; Based on the optimization model, the entity embedding in the medical unified graph is updated to obtain embedding vectors of drug entities and disease entities; Predicting and calculating the embedding vectors of the drug entity and the disease entity according to a preset inner product operation method to obtain the relationship between the drug and the disease; A model is constructed based on the drug-disease relationship and the optimization model to obtain a prediction model, and the prediction model is adaptively optimized by combining the pairwise ranking loss function, the association prediction loss function and the mutual information maximization loss function to obtain a drug-disease relationship prediction result.

5. The method for generating drug repositioning explanation text according to claim 3, characterized in that: A path search process is performed according to the breadth-first search algorithm and the drug-disease relationship prediction result, and an explanation path prediction result is obtained by calculating the aggregate matching score of the entity and relationship embedding on each search path, including: Performing path search processing on the drug entities and disease entities in the biomedical data according to a breadth-first search algorithm to obtain a drug and disease path set; Aggregate information calculation is performed on the entity triples and the entity and relationship embeddings on each search path to obtain a path matching score; The drug and disease pathway sets are sorted based on the pathway matching scores to obtain an explanation pathway prediction result.

6. The method for generating drug repositioning explanation text according to claim 1, characterized in that: A language model is constructed according to biomedical data, and the interpretation path prediction result is input into the language model for text conversion to generate a repositioned interpretation text, wherein the repositioned interpretation text is an explanatory text description of the association between the drug and the disease, including: Extracting drug names, disease names and professional terms from the biomedical data, and performing word frequency statistics on the drug names, disease names and professional terms to obtain a medical vocabulary; The medical vocabulary is interpreted based on a deep learning model, and the medical vocabulary masking and the multi-head self-attention mechanism in the deep learning model allow different positions to be focused on simultaneously when processing a sequence to obtain a language model; The interpretation path prediction result is input into the language model for text conversion to generate a relocated interpretation text.

7. A device for generating drug repositioning explanation text, characterized in that: include: An acquisition module, used for acquiring biomedical data, wherein the biomedical data includes a biomedical entity set and a relationship set, and the biomedical entity set includes drugs and diseases; A first construction module is used to extract molecular structure features of the drugs in the biomedical data, and to construct a graph through the molecular structure features and the relationship set to obtain a unified medical graph; The second construction module is used to perform global and local graph processing on the medical unified graph to obtain global features and local features; A third construction module is used to optimize the network structure by maximizing mutual information according to the global features and the local features, to construct an optimization model, and to calculate the predicted association score between the drug and the disease for the optimization model by a preset inner product operation method to obtain a drug-disease relationship prediction result; A fourth building block is used to perform path search processing according to a breadth-first search algorithm and the drug-disease relationship prediction result, and obtain an explanation path prediction result by calculating the aggregate matching score of the entity and relationship embedding on each search path; The fifth construction module is used to construct a language model based on biomedical data, and input the explanation path prediction result into the language model for text conversion to generate a relocated explanation text, wherein the relocated explanation text is an explanatory text description of the association between drugs and diseases.

8. The device for generating drug relocation explanation text according to claim 7, characterized in that: The first building block comprises: A first processing unit is used to convert the drugs in the biomedical data to obtain a matrix, and obtain molecular structure features by extracting features from the matrix; A second processing unit is used to integrate data according to the molecular structure characteristics and the relationship set to obtain a medical database; A third processing unit is used to connect the drugs and diseases in the medical database to obtain a drug-disease dichotomy model; The fourth processing unit is used to construct a graph based on the drug and disease nodes in the drug-disease dichotomy model and the disease information network in the medical database, and to integrate information by analyzing the attribute mapping relationship between the drug nodes, disease nodes and the disease information network to obtain a unified medical graph.

9. A device for generating drug repositioning explanation text, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for generating a drug repositioning explanation text as claimed in any one of claims 1 to 6 when executing the computer program.

10. A medium, characterized in that : A computer program is stored on the medium, and when the computer program is executed by the processor, the steps of the method for generating a drug repositioning explanation text as described in any one of claims 1 to 6 are implemented.