A method and system for detecting adverse drug events based on a knowledge-enhanced neural network model

By introducing external medical knowledge into the drug adverse reaction detection model and combining deep learning technology, the existing models have solved the shortcomings in detection accuracy and robustness, and more efficient drug adverse reaction detection is achieved.

CN119442027BActive Publication Date: 2025-06-13GUANGDONG UNIVERSITY OF FOREIGN STUDIES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411500936.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2025-06-13
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

The existing drug adverse reaction detection model mainly relies on the analysis of internal text characteristics, neglecting the introduction and integration of external medical knowledge, resulting in insufficient accuracy and robustness of the detection.

Method used

Using a method based on knowledge-enhanced neural network model, the part-of-speech annotation, dependency and entity span information of biomedical texts are obtained by integrating entity description information and deep learning technology in the external medical knowledge base, and combining pre-trained language models and Transformer networks to capture the global semantic information of the text and the local characteristics of entity keywords, and output the joint detection results of entity relationships.

Benefits of technology

It significantly improves the accuracy and robustness of drug adverse reaction detection, can more effectively capture complex language patterns and medical background knowledge, and enhances knowledge expression ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119442027B_ABST
    Figure CN119442027B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for detecting drug adverse events based on a knowledge-enhanced neural network model, which relates to the technical field of natural language processing. It combines the pre-trained language model BioBERT, convolutional neural network and Transformer network, and enhances the model's understanding ability of biomedical entities and their relationships by obtaining keywords related to biomedical text entities from an external knowledge base. By adopting a joint learning strategy, it realizes the synchronous optimization of entity recognition and relationship extraction, thus significantly improving the accuracy of drug adverse reaction detection and the robustness of the model, and can effectively handle complex semantic relationships.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and more specifically, to a method and system for detecting adverse drug events based on a knowledge-enhanced neural network model. Background Art

[0002] Currently, the detection and analysis of adverse drug reactions (ADRs) are key issues in modern medicine and drug regulation. With the continuous deepening of biomedical research and the rapid growth of medical data, especially text data from electronic health records, drug databases, and academic literature, how to efficiently identify and process key medical information in these texts has become a major challenge. The timely discovery and accurate detection of adverse drug reactions are of great significance for drug safety, patient treatment, and medical decision-making, and can help reduce drug-related medical risks and ensure patient safety.

[0003] In recent years, the development of deep learning technology has brought new ideas for adverse drug reaction detection. Models such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memory networks (LSTMs) can automatically learn complex features in text, capture potential semantic associations, and improve the detection accuracy. In addition, pre-trained language models based on self-attention mechanisms (such as BioBERT) have shown excellent performance in processing biomedical texts, and these models can capture the complex relationships between drugs and adverse reactions through large-scale pre-training.

[0004] Although significant progress has been made in the application of deep learning in adverse drug reaction detection, there are still some problems. Most existing models mainly focus on the internal feature analysis of text, ignoring the introduction and integration of external knowledge, while the expression of adverse drug reactions is usually highly related to specific medical backgrounds and clinical contexts. Therefore, relying solely on text data for detection is difficult to comprehensively improve the accuracy and robustness of detection.

[0005] Therefore, how to effectively combine external knowledge in adverse drug reaction detection to improve the accuracy and robustness of adverse drug reaction detection is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0006] In view of this, the present invention provides a method and system for detecting adverse drug events based on a knowledge-enhanced neural network model, which improves the accuracy and robustness of adverse drug reaction detection by integrating entity description information in an external medical knowledge base and deep learning technology.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A method for detecting adverse drug events based on a knowledge-enhanced neural network model, comprising:

[0009] S1 Obtain the biomedical text to be detected and perform preprocessing to obtain a part-of-speech tagging tensor, a dependency relation tensor, and an entity span tensor;

[0010] S2 Based on an external knowledge base, obtain the description information corresponding to the entities in the biomedical text and select keywords to obtain an entity keyword set;

[0011] S3 Process the biomedical text and the entity keyword set based on a pre-trained language model to obtain a text semantic vector and a keyword embedding vector respectively;

[0012] S4 Based on the part-of-speech tagging tensor, the dependency relation tensor, the entity span tensor, and the text semantic vector, fuse them and input them into a Transformer network to obtain a global text semantic representation;

[0013] S5 Input the keyword embedding vector into a convolutional neural network to obtain keyword local features;

[0014] S6 Based on the entity span tensor, the global text semantic representation, and the keyword local features, output an entity relationship joint detection result.

[0015] Preferably, the preprocessing in S1 specifically includes:

[0016] Perform word segmentation and syntactic analysis on the biomedical text, extract the part-of-speech tags of the words and the dependency relations of the sentences to generate a part-of-speech tag sequence and a dependency relation sequence, and mark the start and end positions of the entities;

[0017] The part-of-speech tag sequence and the dependency relation sequence are correspondingly transformed into a part-of-speech tag matrix and the dependency relation matrix based on word indexing;

[0018] Obtain entity span information based on the start and end positions of the entities;

[0019] Based on the part-of-speech tag matrix, the dependency relation matrix, and the entity span information, correspondingly transform them into the part-of-speech tagging tensor, the dependency relation tensor, and the entity span tensor.

[0020] Preferably, obtaining the entity keyword set in S2 specifically includes:

[0021] Calculate the TF-IDF scores for each word in the description information respectively;

[0022] Sort all the words in descending order based on the corresponding TF-IDF scores to obtain a sorting result;

[0023] Select the top pre - set number of words as entity keywords based on the sorting result to form the entity keyword set.

[0024] Preferably, the TF - IDF score calculation formula is:

[0025]

[0026] where f t,D represents the number of times the word t appears in the description information D, f t',D represents the total number of times all words (including t itself) appear in the description information D, N represents the total number of all relevant documents extracted from the external knowledge base, Z represents the set of documents extracted from the external knowledge base for a specific entity, and |{D'∈Z:t∈D'}| represents the number of documents containing the word t.

[0027] Preferably, obtaining the text semantic vector and the keyword embedding vector in S3 specifically includes:

[0028] Perform word embedding processing on the biomedical text based on the pre - trained language model to obtain the text semantic vector V:

[0029] V = BioBERT(Tokens);

[0030] where BioBERT represents the pre - trained language model, and Tokens represents the vocabulary sequence in the pre - processed biomedical text;

[0031] Perform word embedding processing on the entity keyword set based on the pre - trained language model to obtain the keyword embedding vector K E :

[0032] K E = BioBERT(Keywords);

[0033] where Keywords represents the entity keyword set.

[0034] Preferably, obtaining the text global semantic representation in S4 specifically includes:

[0035] Fuse the part - of - speech tagging tensor T POS , the dependency relationship tensor T DEP , the entity span tensor T Span and the text semantic vector V to obtain the comprehensive semantic feature representation F:

[0036]

[0037] where represents the feature concatenation operation;

[0038] Based on the input of the comprehensive semantic feature representation F into the Transformer network, the global semantic representation H of the text is obtained T :

[0039] H T = Transformer(F);

[0040] Among them, Transformer represents the Transformer network.

[0041] Preferably, the output of the entity relationship joint detection result in S6 specifically includes:

[0042] Based on the global semantic representation H of the text T and the local keyword features are concatenated to obtain joint features;

[0043] Based on the joint features and the entity span tensor T Span perform a max pooling operation to obtain the entity features corresponding to each entity;

[0044] Based on the entity features, an entity category probability distribution is obtained;

[0045] Based on any two of the entity features, a relationship category probability distribution is obtained;

[0046] Based on the entity category probability distribution and the relationship category probability distribution as the entity relationship joint detection result.

[0047] Preferably, the method for obtaining the entity category probability distribution is:

[0048] Based on the entity features, map them to the entity category space through a fully connected layer, and predict the entity category probability distribution through a first activation function;

[0049] The method for obtaining the relationship category probability distribution is:

[0050] Based on any two of the entity features, concatenate them to obtain entity pair relationship features;

[0051] Based on the entity pair relationship features, map them to the entity category space through the fully connected layer, and predict the relationship category between the two entities through a second activation function to obtain the relationship category probability distribution.

[0052] Preferably, collect the biomedical text to be detected from biomedical databases and literature resources, perform text parsing through natural language processing tools, and output the relationship between drug entities and adverse reaction entities as the entity relationship joint detection result.

[0053] A drug adverse event detection system based on a knowledge-enhanced neural network model, comprising: a data processing module, an entity keyword generation module, a pre-trained language model processing module, and a model construction and result output module;

[0054] The data processing module is used to obtain the biomedical text to be detected and perform preprocessing to obtain a part-of-speech tagging tensor, a dependency relationship tensor, and an entity span tensor;

[0055] The entity keyword generation module is used to obtain the description information corresponding to the entities in the biomedical text based on an external knowledge base and select keywords to obtain an entity keyword set;

[0056] The pre-trained language model processing module is used to process the biomedical text and the entity keyword set based on a pre-trained language model to correspondingly obtain a text semantic vector and a keyword embedding vector;

[0057] The model construction and result output module is used to construct a Transformer network and a convolutional neural network, fuse the part-of-speech tagging tensor, the dependency relationship tensor, the entity span tensor, and the text semantic vector and input them into the Transformer network to obtain a text global semantic representation; input the keyword embedding vector into the convolutional neural network to obtain keyword local features; based on the entity span tensor, the text global semantic representation, and the keyword local features, output an entity relationship joint detection result.

[0058] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a drug adverse event detection method and system based on a knowledge-enhanced neural network model. By combining the entity description information in the external medical knowledge base, as well as the part-of-speech tags and dependency relationships in the text, a deep understanding of drug adverse reactions is achieved. Compared with models that do not utilize external knowledge sources, the present invention has stronger knowledge expression ability and higher accuracy; the present invention adopts a multi-layer strategy, combining the pre-trained language model BioBERT, the Transformer network, and the convolutional neural network (CNN). By integrating the global semantic information of the text and the local semantic features of entity keywords, the model can effectively capture complex language patterns and medical background knowledge, greatly improving the accuracy and robustness of drug adverse reaction detection; the present invention can effectively identify the relationship between drug entities and adverse reactions in biomedical texts, providing reliable technical support for drug safety research and medical decision-making. Description of the Drawings

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.

[0060] Figure 1 Flowchart of a method for detecting adverse drug reactions based on a knowledge-enhanced neural network model provided by the present invention.

[0061] Figure 2 Schematic diagram of the overall framework topology of a method for detecting adverse drug reactions based on a knowledge-enhanced neural network model provided by the present invention.

[0062] Figure 3 Schematic diagram of the structure of the pre-trained language model BioBERT provided by the present invention.

[0063] Figure 4 Schematic diagram of the topology of the combination of the Transformer network and the convolutional neural network (CNN) provided by the present invention.

[0064] Figure 5 Schematic diagram of the structure of a system for detecting adverse drug reactions based on a knowledge-enhanced neural network model provided by the present invention. Detailed implementation manners

[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0066] Embodiment 1

[0067] As Figure 1 shown, the embodiment of the present invention discloses a method for detecting adverse drug events based on a knowledge-enhanced neural network model, including:

[0068] S1 Obtain the biomedical text to be detected and perform preprocessing to obtain a part-of-speech tagging tensor, a dependency relationship tensor, and an entity span tensor;

[0069] S2 Based on an external knowledge base, obtain the description information corresponding to the entities in the biomedical text and select keywords to obtain an entity keyword set;

[0070] S3 processes the biomedical text and the set of entity keywords based on the pre-trained language model, and correspondingly obtains the text semantic vector and the keyword embedding vector;

[0071] S4 fuses the part-of-speech tagging tensor, the dependency relation tensor, the entity span tensor and the text semantic vector and inputs them into the Transformer network to obtain the global text semantic representation;

[0072] S5 inputs the keyword embedding vector into the convolutional neural network to obtain the keyword local features;

[0073] S6 outputs the joint detection result of entity relations based on the entity span tensor, the global text semantic representation and the keyword local features.

[0074] Embodiment 2

[0075] As Figure 2 shown, the embodiment of the present invention discloses a method for detecting adverse drug events based on a knowledge-enhanced neural network model, including:

[0076] S1 obtains the biomedical text to be detected and performs preprocessing to obtain the part-of-speech tagging tensor, the dependency relation tensor and the entity span tensor:

[0077] Preferably, the preprocessing in S1 specifically includes:

[0078] Performs word segmentation and syntactic analysis on the biomedical text, extracts the part-of-speech tags of the vocabulary and the dependency relations of the sentences to generate a part-of-speech tag sequence and a dependency relation sequence, and marks the start and end positions of the entities;

[0079] The part-of-speech tag sequence and the dependency relation sequence are correspondingly transformed into a part-of-speech tagging matrix and a dependency relation matrix based on the word index of the position number (i.e., the word number) of each word in the sentence or text;

[0080] Obtains the entity span information based on the start and end positions of the entities;

[0081] Based on the part-of-speech tagging matrix, the dependency relation matrix and the entity span information, they are correspondingly transformed into the part-of-speech tagging tensor, the dependency relation tensor and the entity span tensor.

[0082] Preferably, the word index is the position number (i.e., the word number) of each word in the sentence or text, and this index is used to maintain the order and position information of each word in subsequent processing, so that each word corresponds to its part-of-speech tag, dependency relation and other features.

[0083] Preferably, in this embodiment, the Stanford CoreNLP tool is used to perform syntactic parsing on biomedical texts, extract the part-of-speech tags of vocabulary and the dependency relationship structure of sentences. The Stanford CoreNLP tool generates a part-of-speech tag sequence POS = {p 1 , p 2 , …, p n} and a dependency relationship sequence DEP = {d 1 , d 2 , …, d n} for the vocabulary in each text.

[0084] Preferably, the part-of-speech tag sequence is mapped to a part-of-speech tag matrix based on the word index of the position number (i.e., the word number) of each word in the sentence or text in the text: pos_matrix ∈ R n×k , where n represents the number of words and k represents the dimension of the tags; the dependency relationship sequence is transformed into a dependency relationship matrix dep_matrix ∈ R n×m , where m represents the number of types of dependency relationships.

[0085] Preferably, according to the start and end positions of the labeled entity, the span information Span of the entity is obtained, which describes the start and end positions of the entity in the biomedical text, and is expressed as Span = {start E , end E}, where start E and end E respectively represent the start position and the end position of the entity E in the biomedical text.

[0086] Preferably, the part-of-speech tagging tensor is expressed as: T POS = Tensor(pos_matrix); the dependency relationship tensor is expressed as: T DEP = Tensor(dep_matrix); the entity span tensor is expressed as: T Span = Tensor(Span).

[0087] S2 Obtain the description information corresponding to the entity in the biomedical text based on the external knowledge base and select keywords to obtain the entity keyword set:

[0088] Preferably, the external knowledge base includes: DrugBank, PubMED, and STRING.

[0089] Preferably, for each entity E labeled in the biomedical text, the detailed description information D corresponding to the entity is obtained by querying the external knowledge base E ; the description information D E includes the background knowledge of the entity E, such as the mechanism of action and side effects of the drug, and is expressed as: DE = Query_KnowledgeBase(E).

[0090] Preferably, the entity keyword set obtained in S2 specifically includes:

[0091] Calculate the TF-IDF score for each word in the description information respectively;

[0092] Sort all the words in descending order based on the corresponding TF-IDF scores to obtain a sorting result;

[0093] Select the top preset number of words as entity keywords based on the sorting result to form an entity keyword set.

[0094] Preferably, the TF-IDF score calculation formula is:

[0095]

[0096] where f t,D represents the number of times the word t appears in the description information D, f t',D represents the total number of times all words (including t itself) appear in the description information D, N represents the total number of all relevant documents extracted from the external knowledge base, Z represents the document set extracted from the external knowledge base for a specific entity, and |{D'∈Z:t∈D'}| represents the number of documents containing the word t.

[0097] Preferably, score the words in the description information based on the TF-IDF algorithm, sort all the words in descending order based on the TF-IDF scores, and select the top q words as the most representative entity keywords based on the sorting result to obtain the entity keyword set Keywords: Keywords = {k 1 .k 2 ,…,k q}, k q represents the qth keyword extracted from the sorting.

[0098] S3 processes the biomedical text and the entity keyword set based on a pre-trained language model, and correspondingly obtains a text semantic vector and a keyword embedding vector:

[0099] Preferably, the text semantic vector and the keyword embedding vector obtained in S3 specifically include:

[0100] The structure of the pre-trained language model BioBERT is as Figure 3 shown. Perform word embedding processing on the biomedical text based on the pre-trained language model to obtain the text semantic vector V:

[0101] V = BioBERT(Tokens);

[0102] Among them, BioBERT represents a pre-trained language model, and Tokens represents the sequence of words in the pre-processed biomedical text;

[0103] Perform word embedding processing on the entity keyword set based on the pre-trained language model to obtain the keyword embedding vector K E :

[0104] K E = BioBERT(Keywords);

[0105] Among them, Keywords represents the entity keyword set.

[0106] Preferably, the text semantic vector V ∈ R n×h , Tokens = {t 1 , t 2 , …, t n}, where n represents the number of words, h represents the hidden layer representation dimension of the BioBERT model, and t n represents the nth word in the pre-processed biomedical text; the keyword embedding vector K E ∈ R q×h , and q represents the number of keywords.

[0107] S4 Input the fused part-of-speech tagging tensor, dependency relation tensor, entity span tensor, and text semantic vector into the Transformer network to obtain the global text semantic representation:

[0108] Preferably, obtaining the global text semantic representation in S4 specifically includes:

[0109] Fuse based on the part-of-speech tagging tensor T POS , the dependency relation tensor T DEP , the entity span tensor T Span and the text semantic vector V to obtain the comprehensive semantic feature representation F:

[0110]

[0111] Among them, represents the feature concatenation operation;

[0112] Input the comprehensive semantic feature representation F into the Transformer network to obtain the global text semantic representation H T :

[0113] H T = Transformer(F);

[0114] Among them, Transformer represents the Transformer network.

[0115] Preferably, the comprehensive semantic feature representation F is input into the Transformer network, and the multi-head self-attention mechanism is used to capture the long-range dependencies of the text while preserving the lexical order information. Position annotation is performed through position encoding to generate the global semantic representation H of the text T , H T ∈R n×h .

[0116] S5 inputs the keyword embedding vector into a convolutional neural network to obtain the local keyword features:

[0117] Preferably, based on the keyword embedding vector K E is input into a convolutional neural network for local feature extraction. The local context information between keywords is captured through one-dimensional convolutional operations, and then the most significant features are extracted through max-pooling operations to generate the local keyword features K CNN :

[0118] K CNN = CNN(K E );

[0119] Among them, CNN represents the convolutional neural network.

[0120] S6 outputs the joint entity relationship detection result based on the entity span tensor, the global semantic representation of the text, and the local keyword features.

[0121] Preferably, the output of the joint entity relationship detection result in S6 specifically includes:

[0122] Based on the global semantic representation H of the text T and the local keyword features K CNN are concatenated to obtain the joint features;

[0123] Based on the joint features and the entity span tensor T Span perform max-pooling operations to obtain the entity features corresponding to each entity;

[0124] Based on the entity features, obtain the entity category probability distribution;

[0125] Based on any two entity features, obtain the relationship category probability distribution;

[0126] Based on the entity category probability distribution and the relationship category probability distribution as the joint entity relationship detection result.

[0127] Preferably, the joint features represent the concatenation operation, H fused ∈R(n+q)×h .

[0128] Preferably, the combined feature H fused After being processed by both the Transformer network and the CNN, the fused feature representation can contain global and local semantic information, which plays an important role in accurately predicting the relationship between drug entities and adverse reaction entities.

[0129] Preferably, based on the combined feature H fused and the entity span tensor T Span perform a max pooling operation to obtain the entity feature corresponding to each entity:

[0130]

[0131] where EntityFeatures i represents the entity feature corresponding to the i-th entity, MaxPooling represents the max pooling operation, which is used to extract the most significant features from the span range of the i-th entity, represents the entity span tensor of the i-th entity;

[0132] Preferably, the method for obtaining the entity class probability distribution is:

[0133] Map the entity features to the entity class space through a fully connected layer and predict the entity class probability distribution through the first activation function.

[0134] Preferably, in this embodiment, the first activation function uses the Softmax function. The extracted entity features EntityFeatures i are processed through a fully connected layer and mapped to the entity class space to generate the logits (log-odds) representation entity_logits of the entity class i , which refers to the unnormalized scores output by the neural network. Use the Softmax function to normalize the logits to generate the entity class probability distribution for each entity as:

[0135] entity_logits i = Softmax(FC(EntityFeatures i ))

[0136] where entity_logits i represents the entity class probability distribution of the i-th entity, Softmax represents the Softmax function, which is used to convert the logits into a probability distribution and ensure that the sum is 1, so as to intuitively represent the probability of each class, and FC represents the fully connected layer.

[0137] Preferably, the method for obtaining the probability distribution of the relationship category is as follows:

[0138] Concatenate any two entity features to obtain entity pair relationship features;

[0139] Based on the entity pair relationship features, map them to the entity category space through a fully connected layer, and predict the relationship category between the two entities through a second activation function to obtain the probability distribution of the relationship category.

[0140] Preferably, concatenating any two entity features to obtain entity pair relationship features:

[0141]

[0142] where EntityFeatures i and EntityFeatures j respectively represent the entity features of the i-th entity and the entity features of the j-th entity.

[0143] Preferably, in this embodiment, the second activation function uses the Sigmoid function. Map the entity pair relationship feature RelationType ij to the relationship category space through a fully connected layer, and use the Sigmoid function to predict the relationship category between the two entities to obtain the probability distribution of the relationship category:

[0144] relation_logits ij = Sigmoid(FC(RelationFeatures ij ));

[0145] where relation_logits ij represents the probability distribution of the relationship category, which is used to represent the probability distribution of the relationship between the i-th entity and the j-th entity. Sigmoid represents the Sigmoid function, and the Sigmoid function restricts the output between 0 and 1, which is used to represent the probability of the existence of the relationship.

[0146] Preferably, take the entity category probability distribution of each entity and the relationship category probability distribution between entity pairs as the entity relationship joint detection result, and the entity relationship joint detection result is expressed as:

[0147] Output = (entity_logits, relation_logits)

[0148] Among them, entity_logits represents the class prediction logits (log-odds) of all entities, that is, the entity class probability distribution of all entities, and relation_logits represents the relation prediction logits (log-odds) of all entity pairs, that is, the relation class probability distribution of all entity pairs. Output contains the class predictions of all entities and the relation prediction results of all entity pairs.

[0149] The output result Output covers the classification probabilities of entities and the classification probabilities of relations between entities, which are the direct products of entity recognition and relation classification tasks. With the help of this joint prediction layer, the model can provide accurate classifications for entities and relations in biomedical texts on the premise of comprehensively considering various feature information, thus facilitating the accurate extraction of entities and relations in the study of drug adverse reactions.

[0150] Preferably, biomedical texts to be detected are collected from biomedical databases and literature resources, and text parsing is performed through natural language processing tools to output the relationship between drug entities and adverse reaction entities as the joint detection result of entity relations.

[0151] Preferably, steps S1 - S6 jointly construct a drug adverse reaction detection model. By inputting the text to be detected, the joint detection result of entity relations is finally output.

[0152] Obtain biomedical texts to be detected from external biomedical databases (such as PubMed, DrugBank, STRING, etc.) and clinical reports; the biomedical texts are used to detect the relationship between drug entities and their adverse reaction-related entities.

[0153] Input the biomedical text into the trained drug adverse reaction detection model for processing. Through the above-mentioned steps S1 - S6 for processing, the process from entity recognition to relation extraction is completed; the model can identify drug and adverse reaction entities in the text and capture the semantic and logical relationships between these entities.

[0154] Output the prediction results of the drug adverse reaction detection model. The model makes a joint prediction on the drug entities and their adverse reaction entities identified in the biomedical text and outputs the relationship between entities, including the association information between drugs and adverse reactions.

[0155] Preferably, the adverse drug reaction detection model includes: a Transformer network and a convolutional neural network CNN, which are used to capture the long-distance dependencies in the text and the local semantic features of entity keywords. By combining the entity keyword features through a custom attention mechanism, the model can more accurately perform the joint prediction of drugs and adverse reactions. The Softmax function is used to predict the category of the entity, and the Sigmoid function is used to predict the relationship between entities, and the results of adverse drug event detection are output.

[0156] ; It can accurately identify drug entities and entities related to their adverse reactions, and output the relationship prediction between these entities as the joint detection result of entity relationships in the adverse drug reaction detection task.

[0157] Preferably, in terms of model construction, the present invention uses a multi-layer neural network model, which includes a pre-trained language model, a Transformer network, and a convolutional neural network CNN. First, the model embeds the text through BioBERT and integrates the biomedical text and external entity keyword embeddings. Then, the Transformer network is used to model the global semantic information of the text, while the convolutional neural network extracts the local features of entity keywords. Finally, the model performs relationship prediction through the span information and joint feature representation of entities, generating the relationship prediction result between drug entities and adverse reaction entities. In order to improve the accuracy and robustness of the model, an optimization algorithm is used to optimize the model, and finally an adverse drug reaction detection model is obtained.

[0158] Preferably, it also includes the testing and evaluation of the model: the test results of the model are comprehensively analyzed through evaluation metrics such as accuracy, recall, and F1 value. By adjusting the hyperparameters of the pre-trained language model BioBERT, the Transformer network, and the convolutional neural network (CNN), the model performance is optimized and improved, further enhancing the accuracy and robustness of adverse drug reaction detection.

[0159] Preferably, the AdamW optimization algorithm is used to optimize the model performance, and the specific rules are as follows:

[0160]

[0161] where m t and v t respectively represent the estimated value of the first moment and the estimated value of the second moment, and β 1 and β 2 respectively represent the decay factors, usually set to 0.9 and 0.999 correspondingly. denotes the gradient of the loss function L with respect to the model parameters θ, η denotes the learning rate, and ε denotes a small constant to prevent division by zero, usually set to 1×10 -8 ; θ t represents the model parameters at time t.

[0162] After each training phase is completed, the weights of the model are adjusted according to the gradient of the loss function to optimize the performance of the model.

[0163] Preferably, when evaluating the performance of the model in this embodiment, accuracy, recall rate, and F-value are used as the measurement criteria:

[0164] Accuracy is expressed as:

[0165]

[0166] Recall rate is expressed as:

[0167]

[0168] The F1 score corresponding to the F-value is:

[0169]

[0170] Among them, TP, FP, and FN represent the number of true positives, false positives, and false negatives respectively.

[0171] In this embodiment, part-of-speech tags, dependency relationships, etc. are extracted from the biomedical text to be detected through the Stanford CoreNLP tool, converted into indices, and then these information are transformed into a part-of-speech tagging matrix and a dependency relationship matrix. Subsequently, in the word embedding layer, the vocabulary in the text is embedded into vector representations, and at the same time, the span information of entities is extracted and transformed into corresponding vectors. These vectors are input into the encoding layer, and the encoding layer combines the Transformer network and the convolutional neural network (CNN) to process the embedded vectors of the text and entities.

[0172] The encoding layer captures the global context information in the text through the Transformer network and extracts the local semantic features of entity keywords through the convolutional neural network. The model integrates the word embedding representation of BioBERT and the keyword features extracted from the knowledge base, and combines the part-of-speech tagging matrix and the dependency relationship matrix to generate a comprehensive feature representation. Through the multi-head self-attention mechanism of the Transformer, the long-distance dependency relationships in the text are modeled; at the same time, the convolutional neural network extracts the local features of entity keywords, and forms a joint feature representation by splicing the global semantic representation of the text and the local features of the keywords.

[0173] The model combines the span information of entities for joint prediction, uses the Softmax function to predict entity categories, uses the Sigmoid function to predict the relationships between entities, and outputs the relationship prediction results between drug and adverse reaction entities. The model is trained and optimized through optimization algorithms (such as AdamW) to further improve the accuracy and robustness of the model, and finally achieve efficient prediction of the drug adverse reaction detection task.

[0174] Example 3

[0175] Based on the above embodiments, in a specific embodiment, in order to verify the effectiveness of the model of the present invention, multiple groups of comparative experiments were carried out:

[0176] First, on the publicly available Adverse Drug Event (ADE) dataset, the performance of different network models in the drug adverse reaction detection task was compared in the experiment.

[0177] The different comparative network models selected in this embodiment include: the BiLSTM-SDP model based on bidirectional long short-term memory network (BiLSTM) and shortest dependency path (SDP), the multi-head attention mechanism neural network model, the neural network model based on entity span prediction, the pre-trained language model SpanBERT based on spans, the ERSTG model based on graph neural network, the recurrent neural network GRU and knowledge-span encoding, and the Transformer model using the biomedical pre-trained language model BioBERT.

[0178] The present invention adopts a neural network model based on the pre-trained language model BioBERT, combines the entity keyword information in the external knowledge base, captures the global context information of the text through the Transformer network, and combines the convolutional neural network (CNN) to extract the local features of the entity keywords. The topological schematic diagram of the combination of the Transformer network and the convolutional neural network (CNN) in the present invention is as Figure 4 shown.

[0179] In terms of model construction, the present invention uses a multi-layer neural network model, which includes a pre-trained language model, a Transformer network, and a convolutional neural network CNN. First, the model embeds the text through BioBERT, integrating biomedical text and external entity keyword embeddings. Then, the Transformer network is used to model the global semantic information of the text, while the convolutional neural network extracts the local features of the entity keywords. Finally, the model makes relationship predictions through the span information of the entities and the joint feature representation, generating the relationship prediction results between drug entities and adverse reaction entities. To improve the accuracy and robustness of the model, an optimization algorithm is used to tune the model, and finally a drug adverse reaction detection model is obtained.

[0180] The experimental dataset is the publicly available ADE dataset, which mainly includes the annotations of drug entities and adverse reaction entities and their relationships. This dataset is used to train and evaluate the drug adverse reaction detection model.

[0181] To ensure a balance between model training efficiency and hardware resource utilization, the training batch size (train_batch_size) in the experiments of the present invention is set to 8; to ensure higher accuracy during the evaluation process, the evaluation batch size (eval_batch_size) is set to 1; referring to the common parameters in BERT-related research and verifying their effectiveness in this task through experiments, the learning rate (learning_rate) is set to 4×10 -5 ; to help the model make a smooth transition in the early stage of training, the learning rate warmup ratio (warmup_ratio) is set to 10%.

[0182] In terms of model structure design, the present invention sets the multi-head attention mechanism in the Transformer network, selecting 4 attention heads (heads) so that the model can process information from different semantic spaces in parallel; the model hidden layer dimension (hidden_size) is set to 768, and the feedforward network dimension (feedforward_size) is set to 2048 to ensure that the model has sufficient expressive power; during model training, to prevent overfitting, the dropout technique is used, and the dropout ratio of the attention layer (attention_dropout) and the dropout ratio of the BERT model (bert_dropout) in the model are both set to 0.1 to ensure improving the model's robustness while maintaining strong learning ability; to optimize the accuracy of relationship prediction, the present invention sets the relationship filter threshold (relation_filter_threshold) to 0.4, which is obtained from the results of cross-validation and can achieve a good balance between accuracy and recall.

[0183] The experiment was conducted on an Nvidia RTX 1080Ti GPU with 12GB of video memory, ensuring efficient training and evaluation of large-scale datasets and deep models. The above-mentioned multiple comparison models were constructed and tested on the publicly available Adverse Drug Event (ADE) dataset. The results of the comparative experiment are shown in Table 1:

[0184] Table 1 Results of the comparative experiment

[0185]

[0186]

[0187] As can be seen from the experimental data in Table 1, the knowledge-enhanced neural network model proposed in the present invention achieved F1 scores of 93.64% and 88.07% in the named entity recognition and relation extraction tasks respectively, significantly exceeding other baseline models. The experimental results verified the superior performance of the model of the present invention in the adverse drug reaction detection task. Especially after integrating the keywords and semantic information of the external knowledge base, the model can more accurately capture entities and their relationships in complex biomedical texts, providing stronger recognition and reasoning capabilities for adverse drug reaction detection.

[0188] The experimental results of the influence of different knowledge sources on the model performance are shown in Table 2:

[0189] Table 2 Experimental results of different knowledge sources

[0190]

[0191] As can be seen from the results in Table 2, different knowledge sources have a significant impact on model performance: when using STRING (a protein interaction database), the model achieves the highest F1 score in the named entity recognition task, reaching 93.71%, and the F1 score in the relation extraction task is 88.07%. This indicates that the protein interaction information provided by the STRING database is particularly useful for capturing relationships between entities, significantly enhancing the overall performance of the model; in contrast, DrugBank (a drug database) also performs relatively well, with an F1 score of 93.37% for entity recognition and 87.38% for relation extraction; while PubMED (a biomedical literature database) performs slightly worse in entity recognition and relation extraction tasks due to its extensive non-specialized information, achieving F1 scores of 92.93% and 86.57% respectively; when using PubCHEM (a chemical substance and bioactivity database), the model's performance is similar to that of DrugBank, with an F1 score of 93.40% for entity recognition and 87.69% for the relation extraction task; Wikipedia, as a non-professional knowledge base, although its performance in the entity recognition task is close to that of other knowledge sources, its score in the relation extraction task is slightly lower, with an F1 score of 87.05%.

[0192] The experimental results show that selecting knowledge sources highly relevant to the task can significantly improve the performance of the model. In particular, the specialized biomedical information provided by STRING (a protein interaction database) significantly enhances the model's performance in the adverse drug reaction detection task. This further validates the importance of the selection of external knowledge sources for biomedical text processing tasks, especially in the relation extraction task, where the information quality of the knowledge source directly affects the prediction accuracy of the model.

[0193] The experimental results of the impact of different word representation methods on model performance are shown in Table 3:

[0194] Table 3 Experimental Results of Different Word Representation Methods

[0195]

[0196] As can be seen from the experimental results in Table 3, different word representation methods have a significant impact on the model performance: in the named entity recognition (NER) task, BioBERT (a pre-trained model in the biomedical field) achieved the highest F1 score of 93.71%, and in the relation extraction (RE) task, the F1 score was 88.07%. This result indicates that BioBERT, as a pre-trained model optimized specifically for the biomedical field, can better capture the technical terms and complex semantics in biomedical texts, especially performing outstandingly in handling adverse drug reaction detection tasks; in contrast, SciBERT (a pre-trained model in the scientific literature field) also showed excellent performance in the entity recognition and relation extraction tasks, achieving F1 scores of 92.15% and 85.03% respectively, indicating its strong generalization ability in processing scientific literature corpora; ClinicalBERT (a pre-trained model in the clinical field) performed well in the relation extraction task with an F1 score of 86.41% due to its focus on clinical corpora, while the F1 score in the entity recognition task was 92.85%, showing its ability to capture entities and relationships in clinical texts; SpanBERT (a span-based pre-trained model) emphasizes the capture of entity span information, with F1 scores of 91.50% (NER) and 84.51% (RE), demonstrating its practicality in entity recognition and relation extraction tasks; in contrast, BERT (a general pre-trained model), due to the lack of adaptation to the biomedical field, although it can provide good baseline performance, performed slightly weaker in processing biomedical texts, with an F1 score of 89.91% in the entity recognition task and an F1 score of 82.02% in the relation extraction task.

[0197] The experimental results verify the key role of domain-specific pre-trained models (such as BioBERT) in improving the performance of biomedical text processing tasks.

[0198] The experimental results of the impact of different keyword selection strategies on the model performance are shown in Table 4:

[0199] Table 4 Experimental results of different keyword selection strategies

[0200]

[0201] As can be seen from the experimental results in Table 4, different keyword selection strategies have a significant impact on the model performance in the entity recognition and relation extraction tasks:

[0202] The F1 score of TF-IDF (Term Frequency-Inverse Document Frequency) in the entity recognition task is 93.71%, and the F1 score in the relation extraction task is 88.07%, showing the best performance. This indicates that TF-IDF can effectively extract keywords related to entities and provide the most representative external knowledge information for the model, significantly improving the model's performance in complex biomedical texts.

[0203] LDA (Latent Dirichlet Allocation) achieved F1 scores of 93.62% (entity recognition) and 88.04% (relation extraction) by identifying latent topics in the text, demonstrating its ability to capture domain-specific semantic features in keyword extraction and thus providing useful background knowledge for the model.

[0204] TextRank, as a graph-structure-based keyword ranking algorithm, has an F1 score of 93.44% in the entity recognition task and an F1 score of 87.49% in the relation extraction task. Although TextRank can capture keywords in the text through the graph structure, it is slightly less professional in the biomedical field, so its performance is slightly lower than that of TF-IDF and LDA.

[0205] Rake (Rapid Automatic Keyword Extraction), as an algorithm for extracting keywords by analyzing word frequency and co-occurrence information, has an F1 score of 87.69% in the relation extraction task, slightly higher than TextRank. This indicates that Rake has certain advantages in quickly extracting high-frequency keywords, especially without relying on complex topic models.

[0206] The experimental results show that different keyword selection strategies significantly affect the model's performance in the drug adverse reaction detection task.

[0207] Example 4

[0208] As Figure 5 shown, the present invention proposes a drug adverse event detection system based on a knowledge-enhanced neural network model, including: a data processing module, an entity keyword generation module, a pre-trained language model processing module, and a model construction and result output module;

[0209] The data processing module is used to obtain the biomedical text to be detected and perform preprocessing to obtain a part-of-speech tagging tensor, a dependency relation tensor, and an entity span tensor;

[0210] The entity keyword generation module is used to obtain the description information corresponding to the entities in the biomedical text based on an external knowledge base and select keywords to obtain an entity keyword set;

[0211] A pre-trained language model processing module for processing biomedical texts and entity keyword sets based on a pre-trained language model to obtain text semantic vectors and keyword embedding vectors respectively;

[0212] A model construction and result output module for constructing a Transformer network and a convolutional neural network, fusing the part-of-speech tagging tensor, dependency relation tensor, entity span tensor and text semantic vectors and inputting them into the Transformer network to obtain the global semantic representation of the text; inputting the keyword embedding vectors into the convolutional neural network to obtain the local features of the keywords; and outputting the joint detection result of entity relations based on the entity span tensor, the global semantic representation of the text and the local features of the keywords.

[0213] Preferably, it further includes a detection module;

[0214] The detection module is used to input the biomedical text to be detected into a trained drug adverse event detection model. The model combines the text and keyword information to detect drug entities and adverse reaction entities and outputs the result of drug adverse event detection.

[0215] Embodiment 5

[0216] Based on the same inventive concept, the present invention also provides a computer device, including a processor, a communication interface, a memory and a communication bus. Among them, the processor, the communication interface and the memory complete mutual communication through the communication bus;

[0217] The memory is used to store a computer program;

[0218] When the processor is used to execute the program stored in the memory, it can implement a method for detecting drug adverse events based on a knowledge-enhanced neural network model as described in Embodiment 1 or 2.

[0219] The electronic device may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can call the logical instructions in the memory to execute a method for detecting drug adverse events based on a knowledge-enhanced neural network model as described in Embodiment 1 or 2.

[0220] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0221] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.

[0222] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting adverse drug events based on a knowledge-enhanced neural network model, characterized in that: include: S1 obtains the biomedical text to be tested and preprocesses it to obtain a part-of-speech tagging tensor, a dependency tensor, and an entity span tensor; S2 obtains description information corresponding to the entity in the biomedical text based on an external knowledge base and selects keywords to obtain an entity keyword set; S3 processes the biomedical text and the entity keyword set based on a pre-trained language model to obtain a text semantic vector and a keyword embedding vector accordingly; S4 is based on the fusion of the part-of-speech tagging tensor, the dependency tensor, the entity span tensor and the text semantic vector, and inputs them into the Transformer network to obtain a global semantic representation of the text; S5 inputs the keyword embedding vector into a convolutional neural network to obtain a local feature of the keyword; S6 outputs an entity relationship joint detection result based on the entity span tensor, the global semantic representation of the text and the local features of the keywords; S6 outputs the entity relationship joint detection results, including: Based on the global semantic representation H of the text T and the local features of the keyword to obtain a joint feature; Based on the joint feature and the entity span tensor T Span Perform the maximum pooling operation to obtain the entity features corresponding to each entity; Obtaining entity category probability distribution based on the entity features; Obtaining a relationship category probability distribution based on any two of the entity features; Based on the entity category probability distribution and the relationship category probability distribution as the entity relationship joint detection result; The entity category probability distribution acquisition method is: Based on the entity features, the entity features are mapped to the entity category space through a fully connected layer, and the entity category probability distribution is obtained by prediction through a first activation function; The method for obtaining the probability distribution of relationship categories is: Based on any two of the entity features, concatenation is performed to obtain an entity pair relationship feature; Based on the entity pair relationship features, the fully connected layer is used to map them to the entity category space, and the relationship category between the two entities is predicted through a second activation function to obtain the relationship category probability distribution; The biomedical texts to be tested are collected from biomedical databases and literature resources, and the texts are parsed using natural language processing tools. The relationships between drug entities and adverse reaction entities are output as the entity relationship joint detection results.

2. A method for detecting adverse drug events based on a knowledge-enhanced neural network model according to claim 1, characterized in that: The preprocessing in S1 specifically includes: Based on the biomedical text, word segmentation and syntactic analysis are performed to extract the part-of-speech tags of the vocabulary and the dependency relationships of the sentences, generate a part-of-speech tag sequence and a dependency relationship sequence, and mark the start and end positions of the entities; The part-of-speech tag sequence and the dependency relationship sequence are converted into a part-of-speech tag matrix and a dependency relationship matrix based on word index correspondence; Obtain entity span information based on the start and end positions of the entity; Based on the part-of-speech tag matrix, the dependency matrix and the entity span information, they are correspondingly converted into the part-of-speech tag tensor, the dependency tensor and the entity span tensor.

3. The method for detecting adverse drug events based on a knowledge-enhanced neural network model according to claim 1, characterized in that: The entity keyword set obtained in S2 includes: Calculate the TF-IDF score for each word in the description information; Sort all words in descending order based on the corresponding TF-IDF scores to obtain a sorting result; A preset number of words are selected based on the sorting results as entity keywords to form the entity keyword set.

4. The method for detecting adverse drug events based on a knowledge-enhanced neural network model according to claim 3, characterized in that: The TF-IDF score calculation formula is: Among them, f t,D represents the number of occurrences of word t in the description information D, f t',D represents the total number of occurrences of all words in the description information D, including t itself, N represents the total number of all relevant documents extracted from the external knowledge base, Z represents the set of documents extracted from the external knowledge base for a specific entity, and |{D'∈Z:t∈D'}| represents the number of documents containing word t.

5. The method for detecting adverse drug events based on a knowledge-enhanced neural network model according to claim 3, characterized in that: The text semantic vector and keyword embedding vector are obtained in S3, including: The biomedical text is subjected to word embedding processing based on the pre-trained language model to obtain the text semantic vector V: V = BioBERT(Tokens); Among them, BioBERT represents the pre-trained language model, and Tokens represents the word sequence in the pre-processed biomedical text; Based on the pre-trained language model, the entity keyword set is processed by word embedding to obtain the keyword embedding vector K E : K E =BioBERT(Keywords); Among them, Keywords represents the entity keyword set.

6. A method for detecting adverse drug events based on a knowledge-enhanced neural network model according to claim 5, characterized in that: The global semantic representation of the text is obtained in S4, including: Based on the part-of-speech tagging tensor T POS , the dependency tensor T DEP , the entity span tensor T Span Fermented with the text semantic vector V, a comprehensive semantic feature representation F is obtained: in, Represents feature concatenation operation; Based on the comprehensive semantic feature representation F, input into the Transformer network to obtain the global semantic representation H of the text T : H T =Transformer(F); Among them, Transformer represents the Transformer network.

7. A drug adverse event detection system based on a knowledge-enhanced neural network model, using a drug adverse event detection method based on a knowledge-enhanced neural network model as described in any one of claims 1 to 6, characterized in that: include: Data processing module, entity keyword generation module, pre-trained language model processing module and model construction and result output module; The data processing module is used to obtain the biomedical text to be detected and perform preprocessing to obtain a part-of-speech tagging tensor, a dependency tensor and an entity span tensor; The entity keyword generation module is used to obtain description information corresponding to the entity in the biomedical text based on an external knowledge base and select keywords to obtain an entity keyword set; The pre-trained language model processing module is used to process the biomedical text and the entity keyword set based on the pre-trained language model to obtain a text semantic vector and a keyword embedding vector accordingly; The model construction and result output module is used to construct a Transformer network and a convolutional neural network, fuse the part-of-speech tagging tensor, the dependency tensor, the entity span tensor and the text semantic vector, and input them into the Transformer network to obtain a global semantic representation of the text; input the keyword embedding vector into the convolutional neural network to obtain local features of the keywords; and output entity relationship joint detection results based on the entity span tensor, the global semantic representation of the text and the local features of the keywords.

Citation Information

Patent Citations

  • Financial entity relationship extraction system and method in combination with priori knowledge

    CN115687634A

  • Entity relation mining method based on biomedical literature

    US20230007965A1