Adverse drug reaction event identification method and system based on multivariate knowledge mixed retrieval enhancement

By building a method that combines a multi-knowledge base and a large language model, the problem of accurately identifying adverse drug reaction events in clinical medical records has been solved, the recognition efficiency and recall rate have been improved, and a solid foundation has been provided for clinical drug safety and new drug research.

CN120767012AActive Publication Date: 2025-10-10CENT SOUTH UNIV

Patent Information

Application Number
CN202511280583.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-10
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Existing technologies lack accuracy and comprehensiveness when identifying adverse drug reaction events from clinical medical records, and reliance on manual identification is inefficient. Existing machine learning methods lack external knowledge support.

Method used

Construct a multi-knowledge base, including a drug concept knowledge base, a drug adverse reaction knowledge base, a drug adverse reaction event knowledge base, and a drug field text knowledge base. Combined with a large language model, enhance the identification of drug adverse reaction events through multi-knowledge retrieval, and use ontology semantic relationships and vector similarity retrieval to deeply mine relevant knowledge.

Benefits of technology

It has significantly improved the accuracy and recall rate of identifying adverse drug reaction events, ensuring the timely capture of more potential events, and laying the foundation for clinical drug safety and new drug research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120767012A_ABST
    Figure CN120767012A_ABST
Patent Text Reader

Abstract

The invention discloses an adverse drug reaction event recognition method and system based on multivariate knowledge mixed retrieval enhancement, and the method comprises the steps: extracting drug entities from a to-be-recognized clinical disease course record, segmenting the disease course record, and obtaining a drug entity set and a sentence set; for each extracted drug entity, retrieving drug concept knowledge having a hyponymy relationship with the drug entity and drug adverse reaction knowledge having an adverse reaction relationship with the drug entity in a pre-constructed multivariate knowledge base; for each sentence obtained through segmentation, searching suspected adverse drug reaction events meeting a first similarity requirement and drug field text knowledge meeting a second similarity requirement in a pre-constructed multivariate knowledge base; and finally, calling a large language model, taking all the retrieved knowledge as reference knowledge, and identifying the adverse drug reaction event from the to-be-identified clinical disease course record. According to the invention, the accuracy and reliability of adverse drug reaction event identification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of medical information processing, and particularly relates to a drug adverse reaction event identification method and system based on multi-element knowledge hybrid retrieval enhancement. BACKGROUND

[0002] Drug adverse reactions (ADR, Adverse Drug Reaction) are harmful reactions that occur during drug use and are unrelated to the purpose of drug treatment.

[0003] In the process of clinical diagnosis and treatment, timely and accurate identification of drug adverse reaction events (ADE, Adverse Drug Event) from the medical records written by clinicians is crucial for patient safety, optimization of drug treatment plans, clinical new drug trial research, potential drug adverse reaction identification, and drug adverse reaction reporting. Clinical medical records contain rich patient medication information and adverse reaction-related descriptions, but there is currently a lack of effective methods to accurately identify drug adverse reaction events from these unstructured texts. Traditional methods mainly rely on manual identification, which is inefficient and prone to errors. Existing automatic identification methods based on machine learning or deep learning lack rich external knowledge support and rely on large-scale high-quality labeled corpus, which has limitations in identification accuracy and comprehensiveness. SUMMARY

[0004] The application provides a drug adverse reaction event identification method and system based on multi-element knowledge hybrid retrieval enhancement, which improves the accuracy and reliability of drug adverse reaction event identification.

[0005] To achieve the above technical purposes, the application adopts the following technical solutions: A drug adverse reaction event identification method based on multi-element knowledge hybrid retrieval enhancement, comprising: Step 1: Obtain the clinical medical record to be identified, extract the drug entities and segment the medical record to obtain the drug entity set and the sentence set, respectively; Step 2: For each drug entity extracted, search in the pre-constructed multi-element knowledge base: drug concept knowledge with hierarchical relationship, and drug adverse reaction knowledge with adverse reaction relationship; For each sentence obtained by segmentation, search in the pre-constructed multi-element knowledge base: suspected drug adverse reaction events belonging to the N1 most similar and satisfying the first similarity threshold, and drug field text knowledge belonging to the N2 most similar and satisfying the second similarity threshold; Step 3: Call a large language model, use all the knowledge retrieved in step 2 as reference knowledge, and identify drug adverse reaction events from the clinical medical record to be identified.

[0006] Further, the multi-element knowledge base comprises a drug concept knowledge base DRCK, a drug adverse reaction knowledge base ADRK, a drug adverse reaction event knowledge base ADEK and a drug domain text knowledge base DDPK; the drug concept knowledge base DRCK and the drug adverse reaction knowledge base ADRK are respectively used for storing drug concept knowledge and drug adverse reaction knowledge; the drug adverse reaction event knowledge base ADEK and the drug domain text knowledge base DDPK are respectively used for storing suspected drug adverse reaction events and drug domain text knowledge.

[0007] Further, the drug concept knowledge base DRCK stores the hyponym-hypernym concept relationship of each drug entity in the form of triples; The drug adverse reaction knowledge base ADRK stores the relationship between each drug entity and an adverse reaction entity in the form of triples; The drug adverse reaction event knowledge base ADEK is used for storing suspected drug adverse reaction events, and each suspected drug adverse reaction event comprises a sentence description, a drug entity, an adverse reaction entity and a label; The drug domain text knowledge base DDPK stores drug domain text knowledge of each drug entity.

[0008] Further, the hyponym-hypernym concept relationship of each drug entity is constructed in the drug concept knowledge base DRCK using the ontology semantic relationship SubClassOf; Using the depth-first search method, starting from a target drug entity, all hypernym drug concepts of the target drug entity are retrieved upwards and all hyponym drug concepts of the target drug entity are retrieved downwards in the drug concept knowledge base DRCK using the ontology semantic relationship SubClassOf.

[0009] Further, the types of the adverse reaction entity are disease, symptom, sign, test result and examination result, and the relationship between the drug entity and the adverse reaction entity is extracted from medical literature, drug instructions and clinical research reports.

[0010] Further, when retrieving suspected drug adverse reaction events, based on the vector similarity between a target sentence and the sentence description of each suspected drug adverse reaction event in the drug adverse reaction event knowledge base ADEK, N1 suspected drug adverse reaction events with the highest vector similarity and satisfying the similarity not less than a first similarity threshold are selected.

[0011] Furthermore, when retrieving pharmaceutical field text knowledge, based on the vectorized similarity between the target sentence and the pharmaceutical field text knowledge of each pharmaceutical entity in the pharmaceutical field text knowledge base DDPK, the N2 pharmaceutical field text knowledge with the highest vector similarity and satisfying a similarity not less than a second similarity threshold are selected.

[0012] Furthermore, in step 3, system instructions, reference knowledge, output requirements, and user input are filled in according to a preset structured prompt word framework template to call the large language model to identify adverse reaction events; The system instructions require the large language model to identify adverse drug reaction event information from clinical course records using reference knowledge as context; The output requirements are used to specify the output data format of the identified adverse reaction events; The user input is the clinical course record to be identified.

[0013] A drug adverse reaction event recognition system based on multi-knowledge hybrid retrieval enhancement, including: A preprocessing module is used to extract drug entities from the clinical course records to be identified and segment the course records to obtain a drug entity set and a sentence set respectively; The multi-knowledge base module includes the drug concept knowledge base DRCK, the adverse drug reaction knowledge base ADRK, the adverse drug reaction event knowledge base ADEK, and the drug domain text knowledge base DDPK, which are used to store drug concept knowledge, adverse drug reaction knowledge, suspected adverse drug reaction events, and drug domain text knowledge respectively; The knowledge retrieval module is used to: (1) for each extracted drug entity, search in the multivariate knowledge base: drug concept knowledge with a hierarchical relationship with it, and drug adverse reaction knowledge with an adverse reaction relationship with it; (2) for each segmented sentence, search in the pre-constructed multivariate knowledge base: suspected drug adverse reaction events that have the largest N1 similarities with it and meet a first similarity threshold, and drug field text knowledge that has the largest N2 similarities with it and meet a second similarity threshold; The large language model module is used to: use all retrieved knowledge as reference knowledge to identify adverse drug reaction events from the clinical course records to be identified.

[0014] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor implements any of the above-mentioned methods for identifying adverse drug reaction events based on multi-knowledge hybrid retrieval enhancement.

[0015] Compared with the prior art, the present invention has the following beneficial effects: First, the present invention innovatively constructs a multi-dimensional knowledge base, which deeply integrates the hierarchical relationship between drug concepts, triple knowledge related to adverse drug reactions, real clinical adverse drug reaction event cases and professional text knowledge in the field, building a solid and rich knowledge base for the accurate identification of adverse drug reaction events.

[0016] Secondly, the present invention adopts a multi-knowledge-driven retrieval-enhanced generation paradigm, which can deeply mine and fully utilize the relevant knowledge obtained from the retrieval, fully activate the potential of the large language model, and significantly enhance its recognition efficiency of adverse drug reaction events in clinical course records. In this way, the present invention not only improves the accuracy of adverse drug reaction event identification, but also significantly improves the recall rate, ensuring that more potential adverse drug reaction events can be captured in a timely and accurate manner.

[0017] Thirdly, the present invention can lay a solid and reliable foundation for the research of downstream tasks such as ensuring clinical drug safety, promoting clinical research and trials of new drugs, identifying and predicting potential adverse drug reactions, and intelligent reporting of adverse drug reactions, and has important clinical application value and scientific research significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 4 is a flow chart of the method for identifying adverse drug reaction events according to an embodiment of the present invention.

[0019] Figure 2 This is an example of a drug concept knowledge base described in an embodiment of the present invention.

[0020] Figure 3 This is the structured prompt word framework template and example described in the embodiment of the present invention, wherein Figure 3 (a) and Figure 3 (b) Corresponding structured prompt word frame templates and instances respectively.

[0021] Figure 4 This is the definition of the ADE JSON Schema specification described in the embodiment of the present invention.

[0022] Figure 5 This is the recognition result of adverse drug reaction events based on the ADE JSON Shema specification in an embodiment of the present invention.

[0023] Figure 6 This is an example of identifying adverse drug reaction events from clinical course records.

[0024] Figure 7 This is the second example of identifying adverse drug reaction events from clinical course records.

[0025] Figure 8This is the third example of identifying adverse drug reaction events from clinical course records.

[0026] Figure 9 This is Example 4 of identifying adverse drug reaction events from clinical course records. DETAILED DESCRIPTION

[0027] This embodiment is based on the technical solution of the present invention, provides a detailed implementation method and specific operation process, and further explains the technical solution of the present invention.

[0028] This embodiment provides a method for identifying adverse drug reaction events based on multi-knowledge hybrid retrieval enhancement, referring to Figure 1 As shown, the following steps are included: Step 1: Obtain the clinical course records to be identified, extract drug entities and segment the course records to obtain a drug entity set and a sentence set respectively.

[0029] This embodiment specifically adopts BiLSTM-CRF deep learning network to train and build drug entity recognition model based on the annotated drug entity corpus, so as to The drug entity set D is identified in the equation, and the identification method is defined as follows: ; in, It is an entity recognition model built based on BiLSTM-CRF. It is a method for realizing drug entity recognition. D is obtained from clinical course records. The set of drug entities identified in the

[15] dataset. The BiLSTM-CRF deep learning network is a prior art, and reference can be made to the document "Bidirectional LSTM-CRF Models for Sequence Tagging". This invention will not elaborate on this in detail.

[0030] Hypothetical clinical course records to be identified For example, "After receiving intravenous levofloxacin, the patient developed itchy skin and rash on his hands. The patient was suspected of being allergic to levofloxacin and the drug was discontinued. The patient was advised to purchase cefuroxime for anti-infection treatment and have his urine checked again after anti-infection treatment." The identified entity set D is ["levofloxacin", "cefoxitin"].

[0031] In order to effectively retrieve relevant knowledge from the knowledge base, this embodiment uses punctuation segmentation and sliding window mode to analyze clinical course records. Perform sentence segmentation to obtain the segmented sentence set S. The segmentation method is defined as follows: ; Here, chunk is the segmentation method for clinical records; win_size represents the window size, which defaults to 3; step represents the sliding window step size, which defaults to 2; and flag represents the symbol to be segmented, which defaults to ",?;.". This means that medical_record is segmented according to the symbol in flag. The resulting segmented sentence set S is obtained. This segmentation method ensures overlap between the segmented sentences, avoids information loss, and provides support for subsequent knowledge base retrieval.

[0032] above For example, the sentence set S can be obtained = ["After intravenous infusion of levofloxacin, the patient developed itchy skin and rash on his hands, which may be considered as levofloxacin allergy", "Since levofloxacin allergy is considered, it should be discontinued. Pay attention to self-purchased cefuroxime for anti-infection treatment", "Pay attention to self-purchased cefuroxime for anti-infection treatment. After anti-infection treatment, check urine routine"].

[0033] Step 2: Knowledge retrieval.

[0034] On the one hand, for each extracted drug entity, the following drug concept knowledge with a hierarchical relationship and drug adverse reaction knowledge with an adverse reaction relationship are retrieved in the pre-constructed multivariate knowledge base; on the other hand, for each segmented sentence, the following suspected drug adverse reaction events with the greatest similarity to the N1 and meeting the first similarity threshold are retrieved in the pre-constructed multivariate knowledge base, as well as drug field text knowledge with the greatest similarity to the N2 and meeting the second similarity threshold are retrieved.

[0035] First, the pre-built multi-knowledge base in this embodiment includes a drug concept knowledge base DRCK, an adverse drug reaction knowledge base ADRK, an adverse drug reaction event knowledge base ADEK, and a drug domain text knowledge base DDPK.

[0036] (1) Construct the Drug Concept Knowledge Base (DRCK).

[0037] The ontology semantic relationship SubClassOf is a key attribute in the ontology that represents the hierarchical relationship between classes and has transitiveness. Therefore, this embodiment uses the ontology semantic relationship SubClassOf to construct the hierarchical and subordinate concept relationships of each drug entity in the drug concept knowledge base DRCK.

[0038] Each piece of drug concept knowledge c (c∈DRCK) in DRCK stores the relationship between each drug entity and adverse reaction entity in the form of a triple. The pattern is expressed as: ; in, Represents a drug entity.

[0039] For example: <aspirin, subClassOf, nonsteroidal anti-inflammatory drugs> means "aspirin is a nonsteroidal anti-inflammatory drug", <nonsteroidal anti-inflammatory drugs, subClassOf, anti-inflammatory drugs> means "nonsteroidal anti-inflammatory drugs are anti-inflammatory drugs", and based on the transitivity of subClassOf, it can be inferred that "aspirin is also an anti-inflammatory drug".

[0040] The subClassOf relationship is used to clarify the hierarchical classification of different dimensions between drug entities, such as Figure 2 As shown, this supports reasoning and querying. Furthermore, the drug concept knowledge in this embodiment is stored in a graph database (such as Neo4j), with each drug entity as a node and subClassOf as an edge connecting two drug entity nodes with a subordinate relationship.

[0041] (2) Construct an adverse drug reaction knowledge base (ADRK).

[0042] Each piece of adverse drug reaction knowledge r (r∈ADRK) in ADRK is also described by triples, which can be expressed as: ; Among them, drug is the drug entity, has_adr indicates the existence of an adverse reaction relationship, and reaction represents the adverse reaction entity, which may be entities such as disease, symptoms, signs, test results, and examination results.

[0043] When constructing this knowledge graph, the relationships between drugs and adverse reaction entities are extracted from a large number of data sources such as medical literature, drug instructions, and clinical research reports. For example, the following knowledge can be extracted from drug instructions: <Aspirin, has_adr, gastric mucosal injury and bleeding> <Penicillin, has_adr, anaphylactic shock> Primaquine, has_adr, hemolytic anemia

[0044] The adverse drug reaction knowledge graph is stored in a graph database (such as Neo4j), with each drug entity and adverse reaction entity as a node, and has_adr as an edge connecting the corresponding drug entity and adverse reaction entity nodes.

[0045] (3) Construct an adverse drug event knowledge base (ADEK).

[0046] When constructing the ADEK, this embodiment annotates and extracts suspected adverse drug reaction event information from a large number of clinical records, including positive and negative examples. Positive examples are records that are clearly adverse drug reaction events, such as "The patient developed a rash after using amoxicillin"; negative examples are records that appear to be adverse drug reactions but are not, such as "The patient developed cold symptoms while using the drug, but the symptoms were unrelated to the drug used." Each adverse drug reaction event case includes: the attribute sentence, which represents the sentence description of the suspected adverse drug reaction event; the attribute label, which represents a positive or negative example; the attribute drugs, which represents the drug entity; and the attribute reactions, which represents the adverse reaction entity. An example of ADEK is shown in Table 1.

[0047] ; For any adverse drug reaction event knowledge base, represented as a (a∈ADEK), a.sentence represents the description sentence of the suspected adverse reaction event in the case knowledge; a.drugs represents the drug entity that causes the adverse reaction in the case knowledge; a.reactions represents the adverse reaction entity caused by the drug; a.label indicates whether the current case knowledge is a positive example or a negative example.

[0048] (4) Construct a drug domain text knowledge base (DDPK, Drug Domain Paragraph Knowledge).

[0049] The drug domain text knowledge base DDPK is mainly derived from text paragraphs describing drugs in drug instructions, literature, and textbooks. For any piece of drug domain text knowledge, it is represented as p (p∈DDPK). Example knowledge of p is as follows: "Ibuprofen is a widely used non-steroidal anti-inflammatory drug (NSAID). Its main function is to relieve mild to moderate pain, such as headaches, joint pain, toothaches, and dysmenorrhea. It can also effectively reduce fever symptoms caused by colds or flu. It exerts antipyretic, analgesic and anti-inflammatory effects by inhibiting the synthesis of prostaglandins in the body. Ibuprofen usually appears in the form of oral tablets, capsules or liquids, which is convenient for patients of different ages to use. Although ibuprofen is effective in relieving pain and reducing fever, long-term or excessive use may cause gastrointestinal discomfort, such as stomach pain, nausea, and may even cause serious side effects such as gastric ulcers or bleeding. Therefore, when using ibuprofen, be sure to follow the doctor's instructions or the recommended dosage on the drug instructions, and avoid increasing or decreasing the dosage on your own to ensure safe and effective use."

[0050] In this step 2, for each extracted drug entity and each segmented sentence, relevant knowledge is retrieved from the pre-built multivariate knowledge base.

[0051] (1) Retrieve drug concept knowledge in DRCK.

[0052] Hierarchical search is performed using the subClassOf relationship to obtain relevant drug concept hierarchical information. During the search process, the subClassOf relationship chain of the target drug entity is traversed to determine the position of the target drug entity in the concept hierarchy and its related superordinate and subordinate concepts.

[0053] Specifically, the depth-first search (DFS) algorithm is used to traverse the relationship chain in DRCK, starting from the target drug entity, tracing upward to the top-level concept to obtain all superordinate concepts; and exploring downward to explore all subordinate concepts.

[0054] Assume that the target drug entity to be retrieved is d (d∈D), and retrieve all the superordinate concept entity sets of drug concept entity d through the depth-first search algorithm. and the set of all subordinate concept entities , expressed as: ; ; by Figure 2 Taking the drug concept knowledge as an example, assuming that the target drug entity to be retrieved is d = "small molecule drug", then Ud = {drug, organic compound, chemical substance}, Ld = {non-font anti-inflammatory drug, aspirin, ibuprofen, antibiotic, penicillin}.

[0055] For the extracted drug entity set D, the triple set C of drug concept knowledge is retrieved from the drug concept knowledge DRCK, which is expressed as: ; Among them, the drug concept knowledge c(c∈C) satisfies or .

[0056] (2) Retrieve adverse drug reaction knowledge from ADRK.

[0057] Based on the extracted drug entity set D, the adverse reaction entities related to each target drug entity are retrieved in ADRK. The corresponding "reaction" entity can be obtained by matching the "drug" entity. The query statement of the graph database is used for retrieval. The query formula can be expressed as follows: Let the target drug be d (d∈D), and the retrieved adverse reaction entity set be E d .

[0058] ; For example, searching for the adverse reaction entity corresponding to "aspirin" yields E = ["gastrointestinal bleeding", "allergic reaction", etc.].

[0059] For the drug entity set D, the method of retrieving the adverse drug reaction knowledge set R in triple form from the adverse drug reaction knowledge ADRK is expressed as: ;

[0060] Among them, the adverse drug reaction knowledge r (r∈R) satisfies .

[0061] (3) Search for adverse drug reaction events in ADEK.

[0062] Using a vector retrieval method, we calculated the vector similarity between the segmented sentence s (s∈S) in the clinical course record to be identified and the sentence description (i.e., a.sentence) in each suspected adverse drug reaction event a (a∈ADEK) in the ADEK. First, using a text vector model, we converted the sentence s in the clinical course record into an n-dimensional vector A (n=768). Simultaneously, we converted the sentence description a.sentence of the suspected adverse drug reaction event in the ADEK into a vector B of the same n-dimensionality. We then used cosine similarity to calculate the cosine distance between the two vectors using the following formula: ; in, This is an indicator that measures the degree of directional similarity between two vectors, and its value range is [-1, 1]. Values ​​closer to 1 indicate more similar directions; values ​​closer to -1 indicate more opposite directions; and values ​​closer to 0 indicate nearly orthogonal (independent) directions. In accordance with the requirements of this invention, directional similarity is required. Values ​​closer to 1 indicate more similar, and therefore more valuable, ADE case knowledge retrieved.

[0063] In this embodiment, the first similarity threshold is set to simScore1, and the most similar N1 adverse drug reaction events are retrieved.

[0064] The sentence s and the statement description a.sentence of the suspected adverse drug reaction event are collectively referred to as text t. The text t is converted into a vector of fixed dimension w (w=768) Expressed as: ; According to the sentence set S obtained by segmenting the clinical course records, the topN1 is retrieved from ADEK, and the similarity is greater than or equal to the first similarity threshold simScore1 to obtain the drug adverse reaction event set A: ; For any retrieved adverse drug reaction event a (a∈A), there exists s(s∈S) that satisfies ConsineSimilarity(Va,Vs)>=simScore1, and the size of set A is less than or equal to topN1.

[0065] (4) Retrieve drug domain text knowledge in DDPK.

[0066] The same vector retrieval method is used to calculate the vector similarity between the segmented sentence s (s∈S) of the clinical course record to be identified and the drug domain knowledge text p (p∈DDPK). The calculation formula is the same as that introduced in the retrieval of adverse drug reaction events in ADEK. The text knowledge p that meets the second similarity threshold simScore2 and topN2 (the default is 5) is used as the retrieval result.

[0067] Based on the sentence set S obtained by segmenting the clinical course records, the topN2 is retrieved from DDPK, and the similarity is greater than or equal to the second similarity threshold simScore2 to obtain the drug domain knowledge set P: ; For any retrieved pharmaceutical field text knowledge (p∈P), there exists s (s∈S) that satisfies ConsineSimilarity(Vp,Vs)>=simScore2, and the size of the set P is less than or equal to topN2.

[0068] In step 3, the large language model is called, and all the knowledge retrieved in step 2 is used as reference knowledge to identify adverse drug reaction events from the clinical course records to be identified.

[0069] The structured prompt word frame template preset in this embodiment is as follows: Figure 3 As shown in (a), it is mainly divided into the following four parts: (1) System instructions: Clarify the system instruction requirements, that is, use reference knowledge as context to identify adverse drug reaction event information from clinical course records.

[0070] (2) Reference knowledge: Provide a reference knowledge template and retrieve drug concept knowledge C, adverse drug reaction knowledge R, ADE case knowledge A, and drug domain knowledge P from the multivariate knowledge base (DRCK, ADRK, ADEK, DDPK) based on the input medical_record. The retrieved knowledge needs to be filled into the corresponding parameters.

[0071] (3) Output requirements: The data format of the output generated by the instructions must be clearly specified. JSON serialization must be used and must comply with the ADE JSON Schema specification.

[0072] (4) User input: The user input part requires the input of the clinical course record to be analyzed, and the parameter medical_record is replaced by the content of the clinical course record to be analyzed.

[0073] Fill in the system instructions, reference knowledge, output requirements and user input according to the preset structured prompt word framework template to obtain a structured prompt word framework instance, such as Figure 3 (b) shown.

[0074] Based on the clinical course record to be identified, medical_record, and the retrieved drug concept knowledge set C, drug adverse reaction knowledge set R, drug adverse reaction event set A, and drug domain text knowledge set P, a structured prompt word framework is constructed as follows: ; After completing the construction of the structured prompt word framework instance, the large language model LLM (such as GPT-4o, Qwen, DeepseekV3) can be called to generate it, thereby identifying the adverse drug reaction event information results ADEs.

[0075] .

[0076] In order to accurately express and standardize the ADE recognition output, this embodiment uses JSON Schema to formulate the output specification of the structured prompt word framework. Through analysis, the adverse drug reaction events identified from a given clinical course record have the following characteristics: (1) May describe 0 or more adverse drug reaction events.

[0077] (2) Each adverse drug reaction event will have a drug entity that causes the adverse reaction, which may be 0 or more. When it is 0, it means that the drug entity is not clearly specified in the medical record.

[0078] (3) Each adverse drug reaction event will describe the adverse reaction entity that occurs after taking the drug. The adverse reaction entity may be one or more.

[0079] (4) There is a clear causal relationship between the drug entity and the adverse reaction entity in each adverse drug reaction event.

[0080] In view of the characteristics of the above-mentioned adverse drug reaction events, the present invention adopts JSON serialization output mode and formulates the output specification ADE JSON Schema, such as Figure 4 As shown, the description is as follows: (a) Root structure: It is an array containing 0 or more ADEs. If there is no adverse drug reaction event in the medical record to be analyzed, the recognition result is an empty array []. (b) Each ADE structure contains three fields: 1) sentence: a string, a sentence describing an ADE in the medical record, such as "the patient developed a rash after taking aspirin"; 2) drugs: string array, a list of drug entities identified in ADE, such as [“aspirin”]; 3) reactions: String array, a list of adverse reaction entities identified in the ADE. This list of adverse reaction entities is caused by drugs and has a causal relationship between them. For example, ["rash"] means that rash is an adverse reaction caused by taking the drug aspirin.

[0081] By formulating the above-mentioned ADE JSON Schema specification, the core information elements of adverse drug reaction events can be systematically and standardizedly represented, while providing clear and verifiable data specification constraints for the target identification tasks of downstream methods.

[0082] Based on the clinical record medical_record example, the adverse drug reaction events are identified according to the ADE JSONSchema output as follows Figure 5 shown.

[0083] Based on the method presented in this paper, 6024 clinical records with ADEs annotated were used as a reference set. The recognition results of the method presented in this paper were compared with the standard results. A 10-fold cross-validation analysis method was used to evaluate the effectiveness of the method presented in this paper using precision, recall, and the comprehensive evaluation metric F1. The experimental evaluation results are shown in Table 2: ; As shown in Table 2, the method presented in this paper was evaluated using three large language models: Enire 3.5-8k, Deepseek V3, and Qwen-Turbo. The Deepseek V3-based model achieved a recognition accuracy of 98.38%, a recall of 92.97%, and an overall F1 score of 0.956. These experimental results demonstrate the effectiveness of the method presented in this paper.

[0084] Next, combined with a specific clinical case analysis, the beneficial effects of the present invention are exemplified as follows: (1) The method of the present invention can effectively identify adverse drug reaction events and can also infer specific drug entities based on clinical context semantic relationships, such as Figure 6 shown.

[0085] (2) The method of the present invention can effectively identify adverse drug reaction events. When there are multiple drug entities, it can effectively analyze the causal relationship between the drug entity and the adverse reaction entity and exclude drug entities that do not cause adverse reactions, such as Figure 7 shown.

[0086] (3) The method of the present invention can effectively identify adverse drug reaction events, understand abnormal test index values, and convert them into qualitative test results as adverse reaction entities, such as Figure 8 shown.

[0087] (4) The method of the present invention can effectively identify adverse drug reaction events and can identify multiple adverse drug reaction events described in the medical records, such as Figure 9 shown.

[0088] The above embodiments are preferred embodiments of the present invention. Ordinary technicians in this field can also make various changes or improvements on this basis. Without departing from the overall concept of the present invention, these changes or improvements should fall within the scope of protection required by the present invention.

Claims

1. A method for identifying adverse drug reaction events based on multi-knowledge hybrid retrieval enhancement, characterized in that: include: Step 1: Obtain the clinical course records to be identified, extract drug entities from them, and segment the course records to obtain a drug entity set and a sentence set respectively; Step 2: For each extracted drug entity, search the pre-built multivariate knowledge base for: drug concept knowledge with a hierarchical relationship with it, and drug adverse reaction knowledge with an adverse reaction relationship with it; For each segmented sentence, search the pre-built multivariate knowledge base for: suspected adverse drug reaction events that have the largest N1 similarities and meet the first similarity threshold, and pharmaceutical field text knowledge that has the largest N2 similarities and meets the second similarity threshold; In step 3, the large language model is called, and all the knowledge retrieved in step 2 is used as reference knowledge to identify adverse drug reaction events from the clinical course records to be identified.

2. The method for identifying adverse drug reaction events according to claim 1, wherein: The multi-knowledge base includes a drug concept knowledge base DRCK, a drug adverse reaction knowledge base ADRK, a drug adverse reaction event knowledge base ADEK and a drug domain text knowledge base DDPK; the drug concept knowledge base DRCK and the drug adverse reaction knowledge base ADRK are used to store drug concept knowledge and drug adverse reaction knowledge, respectively; the drug adverse reaction event knowledge base ADEK and the drug domain text knowledge base DDPK are used to store suspected drug adverse reaction events and drug domain text knowledge, respectively.

3. The method for identifying adverse drug reaction events according to claim 2, wherein: The drug concept knowledge base DRCK stores the hierarchical and hyponymous concept relationships of each drug entity in the form of triples; The adverse drug reaction knowledge base ADRK stores the relationship between each drug entity and adverse reaction entity in the form of triples; The adverse drug reaction event knowledge base ADEK is used to store suspected adverse drug reaction events, each suspected adverse drug reaction event includes a statement description, a drug entity, an adverse reaction entity and a label; The drug domain text knowledge base DDPK stores drug domain text knowledge of each drug entity.

4. The method for identifying adverse drug reaction events according to claim 3, wherein: Using the ontology semantic relationship SubClassOf, construct the hierarchical and hyponymous concept relationships of each drug entity in the drug concept knowledge base DRCK; A depth-first search method was adopted, starting from the target drug entity, and using the ontology semantic relationship SubClassOf in the drug concept knowledge base DRCK, all superordinate drug concepts of the target drug entity were retrieved upward, and all subordinate drug concepts of the target drug entity were retrieved downward.

5. The method for identifying adverse drug reaction events according to claim 3, wherein: The types of adverse reaction entities are: diseases, symptoms, signs, test results, and examination results. The relationships between drug entities and adverse reaction entities are extracted from medical literature, drug instructions, and clinical research reports.

6. The method for identifying adverse drug reaction events according to claim 3, wherein: When retrieving suspected adverse drug reaction events, based on the vectorized similarity between the target sentence and the statement description of each suspected adverse drug reaction event in the adverse drug reaction event knowledge base ADEK, the N1 suspected adverse drug reaction events with the highest vector similarity and satisfying the similarity not less than the first similarity threshold are selected.

7. The method for identifying adverse drug reaction events according to claim 3, wherein: When retrieving pharmaceutical domain text knowledge, based on the vectorized similarity between the target sentence and the pharmaceutical domain text knowledge of each pharmaceutical entity in the pharmaceutical domain text knowledge base DDPK, the N2 pharmaceutical domain text knowledge with the highest vector similarity and a similarity not less than a second similarity threshold are selected.

8. The method for identifying adverse drug reaction events according to claim 1, wherein: Step 3: Fill in system instructions, reference knowledge, output requirements, and user input according to the preset structured prompt word framework template to call the large language model to identify adverse reaction events; The system instructions require the large language model to identify adverse drug reaction event information from clinical course records using reference knowledge as context; The output requirements are used to specify the output data format of the identified adverse reaction events; The user input is the clinical course record to be identified.

9. A drug adverse reaction event recognition system based on multi-knowledge hybrid retrieval enhancement, characterized by: include: A preprocessing module is used to extract drug entities from the clinical course records to be identified and segment the course records to obtain a drug entity set and a sentence set respectively; The multi-knowledge base module includes the drug concept knowledge base DRCK, the adverse drug reaction knowledge base ADRK, the adverse drug reaction event knowledge base ADEK, and the drug domain text knowledge base DDPK, which are used to store drug concept knowledge, adverse drug reaction knowledge, suspected adverse drug reaction events, and drug domain text knowledge respectively; The knowledge retrieval module is used to: (1) for each extracted drug entity, search in the multivariate knowledge base: drug concept knowledge with a hierarchical relationship with it, and drug adverse reaction knowledge with an adverse reaction relationship with it; (2) for each segmented sentence, search in the pre-constructed multivariate knowledge base: suspected drug adverse reaction events that have the largest N1 similarities with it and meet a first similarity threshold, and drug field text knowledge that has the largest N2 similarities with it and meet a second similarity threshold; The large language model module is used to: use all retrieved knowledge as reference knowledge to identify adverse drug reaction events from the clinical course records to be identified.

Citation Information

Patent Citations

  • Intelligent question answering method for adverse drug reaction by fusing multi-channel text features

    CN108984699A

  • Untoward drug reaction monitoring and early warning method

    CN118280603A

  • Medical Prediction Method and System Based on Semantic Graph Network

    US20220277858A1

Cited By

  • Assessment and enhancement method and device for LLMs concept mutual exclusion recognition capability

    CN121543698A

  • Drug adverse event grading prediction method and device based on local data

    CN121790017A

  • Untoward effect attribution analysis method and system based on multi-modal constraint decoding

    CN121938659A

  • Adverse reaction attribution analysis method and system based on multi-modal constraint decoding

    CN121938659B