Reasonable drug use system and method based on natural language processing and drug knowledge graph

By establishing a rational drug use system based on natural language processing and drug knowledge graphs, the challenges of rational drug use for pharmacists and patients have been addressed. This system automates prescription review and drug information retrieval, thereby improving drug safety and efficiency.

CN115293161BActive Publication Date: 2026-07-24GUANGZHOU ZHONGKANG INFORMATION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU ZHONGKANG INFORMATION CO LTD
Filing Date
2022-08-19
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In the current technology, pharmacists in Internet hospitals and DTP pharmacies lack clinical knowledge and experience, which makes it difficult to prescribe drugs rationally. Patients lack drug knowledge, online consultations with doctors are not timely, and drug search relies on limited search engines, which poses risks of duplicate medication and prescription discrepancies.

Method used

A rational drug use system based on natural language processing and drug knowledge graph is adopted, including an automatic prescription review module, a drug recommendation module, an information backtracking module, and a drug information query module. The drug knowledge graph is used to review the rationality of prescriptions, recommend drugs, and query drug information, and the drug knowledge graph is updated through a natural language processing model.

Benefits of technology

It has automated the review of drug rationality and the recommendation of medication, improved the efficiency of pharmacists, ensured the rationality of prescriptions, provided convenient access to drug knowledge, reduced the risk of duplicate medication and prescription discrepancies, and improved patient medication safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115293161B_ABST
    Figure CN115293161B_ABST
Patent Text Reader

Abstract

The application discloses a rational drug use system and method based on natural language processing and a drug knowledge graph, and constructs a rational drug use system based on natural language processing and a drug knowledge graph. The system can automatically perform data management on drug instruction book data and form a drug knowledge graph, and realizes the functions of the rational drug use system based on the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of natural language processing and knowledge graph technology, specifically to a rational drug use system and method based on natural language processing and drug knowledge graph. Background Technology

[0002] In the era of big data, data is the cornerstone of innovative business models and cutting-edge technological development. In recent years, with the popularization of the big data concept, many enterprises have increasingly attached importance to internal data governance, breaking down the barriers between data from different sub-modules and further integrating them to maximize the utilization of enterprise data resources.

[0003] In the pharmaceutical field, artificial intelligence (AI) technology can quickly and effectively process unstructured instruction manual data to form a pharmaceutical knowledge graph. Based on this knowledge graph, many different applications can be derived, such as the rational drug use system functions of internet hospitals, pharmaceutical knowledge service platforms for DTP pharmacists, and intelligent pharmaceutical services for patients. In internet hospital scenarios, doctors have the ability to prescribe medications online. On the one hand, it is necessary to avoid risks such as duplicate medication, the influence of drug vendors, and prescription drugs not matching their indications; on the other hand, it is necessary to verify the compliance of prescriptions to prevent misprescription. The rational drug use alert function plays an important role here, reminding doctors and pharmacists of risk information when prescribing medications. In DTP pharmacy scenarios, pharmacists need to improve their service level to patients, but due to a lack of clinical knowledge and experience, they need to spend more time keeping up with pharmaceutical knowledge. However, the high volume of business in DTP pharmacies leads to heavy workloads for pharmacists, forcing many to use their personal time to learn about pharmaceuticals. However, due to limited background knowledge, keeping up with pharmaceutical knowledge is a significant challenge for them. In this scenario, pharmaceutical knowledge service platforms become crucial. Based on the latest pharmaceutical knowledge graphs and user-friendly search interfaces, pharmacists can easily access pharmaceutical information, significantly saving their time. In the patient scenario, due to a lack of background knowledge, patients are often unfamiliar with information such as drug administration methods and contraindications. While users have limited access to knowledge through search engines and online doctor consultations are often time-consuming, intelligent pharmaceutical services based on human-computer dialogue can help users quickly understand drug information. Furthermore, deploying pharmaceutical services on platforms like WeChat mini-programs is also an important way to expand user traffic. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention aims to provide a rational drug use system and method based on natural language processing and drug knowledge graphs.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] A rational drug use system based on natural language processing and drug knowledge graph includes an automatic prescription review module, a drug recommendation module, a prompt information backtracking module, a drug information query module, and a data update and maintenance module;

[0007] The automatic prescription review module is used in medical settings to automatically review the rationality of prescriptions after a doctor issues a prescription, based on the input of the patient's basic information, medical information, and information about the specific drugs in the prescription, using a drug knowledge graph. The review includes assessing whether the drugs in the prescription are consistent with the symptoms and disease diagnosis in the medical information; whether there are interactions between the drugs in the prescription and between the drugs in the prescription and the patient's current medications; whether the patient's population type is within the contraindications group for the drugs in the prescription; whether the patient has a history of allergies to the drugs in the prescription; whether there are duplicate medications in the prescription; and whether the dosage and administration of the drugs in the prescription are consistent with the dosage and administration of the corresponding drugs in the drug knowledge graph. The patient's basic information includes age, gender, population type, allergy history, and recent medication history. The medical information includes symptoms and disease diagnosis results. The information about the specific drugs in the prescription includes the generic name, manufacturer, specifications, approval number, and dosage and administration of each drug.

[0008] Recommended medication module: Used in diagnosis and treatment scenarios or medication consultation scenarios, it generates a list of recommended medications based on the input patient's basic information and disease and symptom information using a drug knowledge graph;

[0009] Information tracing module: This module provides a query interface for the sources of prescription rationality review results and recommended drug lists. These sources include the original drug instructions and an interaction database. For the original instructions, the results of prescription rationality review and recommended drug lists come from the indications, interactions, contraindications, and dosage of the drug instructions. The interaction database explicitly indicates the interactions between components. The recommended drug module or automatic prescription review module searches the interaction database for interactions by looking up the drug components in the drug knowledge graph. Therefore, tracing the source of interaction results involves querying the detailed interaction information in that database.

[0010] Drug Information Query Module: This module provides a query interface for querying drug information, including the original instructions, drug knowledge graph, adverse reactions, indications, contraindications, and dosage.

[0011] Data update and maintenance module: This module provides an upload interface and interface for uploading drug instructions. It uses a natural language processing model to process the text data of the uploaded drug instructions and update the drug knowledge graph accordingly.

[0012] Furthermore, the recommended medication module uses a drug knowledge graph to filter suitable medications based on the input patient's disease and symptom information. Then, it filters the filtered medications based on the patient's population type, allergy history, and recent medication use, removing medications that are unsuitable for the patient's corresponding population type, may cause allergies, or overlap with or interact with the patient's currently taken medications. The remaining medications after filtering are sorted according to the sorting rules set by the user to generate a recommended medication list.

[0013] Furthermore, when querying the original instructions, drug knowledge graph, or usage and dosage, specific information is returned based on the drug's generic name, manufacturer, specifications, and approval number. Querying adverse reactions, indications, and contraindications supports both forward and reverse queries. Forward queries return information based on the drug's generic name, manufacturer, specifications, and approval number, while reverse queries return a specific list of drugs based on a particular adverse reaction, indication, or contraindication.

[0014] Furthermore, the data update and maintenance module supports uploading drug instruction manuals in various data formats, including PDF, IMG, and text. For text data, the data update and maintenance module directly uses a natural language processing model to process the uploaded instruction manual text data and update the drug knowledge graph accordingly. For PDF and IMG type data, the data update and maintenance module uses image processing technology to extract the content and convert it into text, and then uses a natural language processing model to process the extracted text data.

[0015] Furthermore, the data update and maintenance module provides data quality control functions, allowing users to compare the results of automated text annotation and graph generation with the original text of the instruction manual to ensure data quality.

[0016] Furthermore, the data update and maintenance module provides a drug catalog integration function. Users can upload drug data, including fields such as generic name, manufacturer, specifications, and approval number, through this function. The data update and maintenance module uses a natural language processing model to automatically match the uploaded drug data to the existing drug catalog. For drugs that are not in the original drug catalog, the data update and maintenance module automatically structures the uploaded drug instructions to form a drug knowledge graph. Finally, the drug data integration is completed and the data is put into use.

[0017] The present invention also provides a method for constructing the above system, the specific process of which is as follows:

[0018] I. Designing a Drug Knowledge Graph:

[0019] Based on the business requirements of the rational drug use system, a drug knowledge graph structure containing corresponding fields was designed, including ingredients, indications, adverse reactions, contraindications, interactions, and usage and dosage.

[0020] II. Design a labeling system:

[0021] The labeling system is defined based on the design of a drug knowledge graph;

[0022] III. Data Labeling:

[0023] After the labeling system design is completed, annotation tools are used to annotate the sample text data of the drug instructions. The data annotation work is divided into tasks according to the field design of the drug knowledge graph structure. Each different field includes three annotation tasks: named entity annotation, entity relationship annotation, and annotation term alignment. Named entity annotation is to define the scope and select the entity type of specific words based on the entity tags, including continuous entity annotation and non-continuous entity annotation. Entity relationship annotation specifies the directed relationship type between annotated named entities. Entity alignment annotation determines which standardized term the entity is aligned to based on the named entity and entity type tags.

[0024] IV. Natural Language Processing Model Construction:

[0025] The natural language processing model automatically parses text data obtained from user-uploaded drug instructions, extracts entity and entity relationship information, and further aligns it to standard terminology to construct a drug knowledge graph; the construction process of the natural language processing model is as follows:

[0026] 4.1 Constructing an information extraction module

[0027] The information extraction module includes models for drug text classification, discourse structure analysis, BERT encoding tasks, entity extraction, and joint relation extraction.

[0028] The drug text classification task is used to classify the input text data of drug instructions into traditional Chinese medicine and chemical drugs.

[0029] The BERT encoding task is used to encode using massive unsupervised corpora or pre-trained BERT adapted to the medical field;

[0030] The text structure analysis task is used to divide the input drug instruction manual text into text medical blocks semantics. The drug instruction manual text involves multiple fields such as indications, contraindications, usage and dosage. The text structure analysis task is used to divide different text segments into fields and identify the medical semantic role to which a certain text segment belongs.

[0031] The medical entity recognition task is used to extract medical-related entities mentioned in the input drug instruction manual text;

[0032] The medical relation extraction task is used to identify and determine the specific relationships between entity pairs in an input drug instruction manual text.

[0033] The model integrating the above tasks includes a BERT encoder module, a head entity labeling module, a tail entity labeling module, a document extraction module, a document type classification module, and a loss function calculation module. The entire process is optimized through a shared BERT encoding layer and multi-task joint learning. The data processing procedure of the model is as follows:

[0034] Suppose we have a text sequence X = (x1, x2, x3, ..., x...). n ), x t The t-th character or word represents the position t, and n represents the length of the text sequence.

[0035] S4.1.1 The BERT encoder module adds a predefined start character [CLS] to the beginning of the text sequence X, and then performs BERT encoding:

[0036] H=BERT([CLS]+X).............(1);

[0037] Where H represents the hidden state of the text sequence X after being encoded by the function BERT(·), H∈R n×d , where n represents the number of characters or words in the text sequence X, and d represents the dimension of the encoded vector;

[0038] S4.1.2 The header entity labeling module calculates the boundary probabilities of header entity labels as follows:

[0039]

[0040]

[0041] in, Let represent the probability of the i-th character or word indicating the start of the header entity. W represents the probability that the i-th character or word indicates the end of the header entity. s ∈R d×1 W e ∈R d×1 ,b s ∈R,b e ∈R represents the parameters to be learned, R represents the set of real numbers, and σ(·) represents the activation function;

[0042] Therefore, the likelihood function of the head entity can be calculated:

[0043]

[0044] Where s represents the head entity, and I(·) represents the indicator function. A marker indicating the start or end position of the i-th character or word, with a value of {0, 1};

[0045] S4.1.3 The calculation relationship and tail entity boundary of the tail entity marking module are as follows:

[0046]

[0047]

[0048] in, This represents the probability that the i-th character or word indicates the start of the last entity. v represents the probability that the i-th character or word indicates the end of the entity. k ∈R d Let d represent the k-th head entity vector, and h represent the encoding dimension. i ∈H represents the implicit vector of a character or word; if the entity consists of multiple characters or words, an average value is calculated first. Let represent the parameters that the model needs to learn, m represent the number of relation types, R represent the set of real numbers, and σ(·) represent the activation function.

[0049] Therefore, the likelihood functions of the relations and entities can be calculated:

[0050]

[0051] Where o represents the tail entity, and I(·) represents the indicator function. A marker indicating the start or end position of the i-th character or word, with a value of {0, 1};

[0052] S4.1.3 Assume the output label sequence is y = (y1, y2, y3, ..., y4). n The document extraction module calculates the total score of the document analysis sequence as follows:

[0053]

[0054] Where A∈R n×n Let W be the transition matrix. crf ∈R d×n ,b crf ∈R n Indicates the parameters to be learned;

[0055] The probability of a text sequence corresponding to a target sequence is calculated as follows:

[0056]

[0057] Among them, Y x Represents the set of possible target sequences;

[0058] In the prediction phase, the Viterbi algorithm is used to find the optimal sequence:

[0059]

[0060] S4.1.4, The instruction manual type classification module calculates the probability of text category:

[0061] p c =W c ×h cls +b c ...(11);

[0062] Among them, W c ∈R d ,b c ∈R represents the parameters to be learned, h cls The implicit vector representing [CLS];

[0063] S4.1.5 The loss function calculation module calculates the loss function as follows:

[0064] L = -(l s+o +l crf +l c )......................(12);

[0065] in,

[0066]

[0067]

[0068]

[0069] Where M represents the total number of samples, n represents the text length, and λ and γ are regularization parameters;

[0070] 4.2 Constructing a terminology standardization module

[0071] The terminology standardization module is mainly based on a two-stage terminology standardization model. First, candidate terms are recalled from standard terms, and then the candidate terms are finely ranked and calculated. The construction process of the terminology standardization module is as follows:

[0072] S4.2.1 Collect Chinese and English terminology corpora and open terminology databases, organize them and construct a medical terminology database. The fields of the medical terminology database include Unified Code (CUI), English standard words, English synonyms, Chinese standard words, and Chinese synonyms.

[0073] S4.2.2. Based on indexing tools or using custom indexes, create indexes for Chinese and English. When querying Chinese terms, translate the Chinese terms into English and use both English and Chinese as input.

[0074] 4.2.3. Establish a keyword search engine and integrate recall scores through multi-channel recall. recall Assume s1 and s2 are two terms:

[0075] s recall =α1×s bm25 +α2×s Jaccard +α3×s MED +α4×s DICE .......(16)

[0076] Where α1, α2, α3, and α4 represent the weights of each scoring path, and α1 + α2 + α3 + α4 = 1;

[0077] Among them, BM25 score bm25 The calculation is as follows:

[0078]

[0079]

[0080] Where, ω i This represents the i-th word of the query term s1; f i It is a word ω i The frequency of occurrence of term s2, k1 and b are adjustment factors, len(·) is the function for calculating sentence length, avgsl is the average length of all documents in the index; N represents the total number of documents in the index, n(ω i ) contains ω i The number of documents;

[0081] Jaccard coefficients Jaccard The calculation is as follows:

[0082]

[0083] Where A and B represent the word segmentation sets of s1 and s2 respectively, and · represents the number of elements in the set;

[0084] Edit distance (MED) similarity score MED The calculation is as follows:

[0085]

[0086] Where len(·) is the function for calculating sentence length, and d(s1,s2) represents the edit distance between s1 and s2;

[0087] DICE Distance Similarity Ratings DICE The calculation is as follows:

[0088]

[0089] Where A and B represent the word segmentation sets of s1 and s2 respectively, and |·| represents the number of elements in the set;

[0090] S4.2.4 Training the fine-ranking model:

[0091] S4.2.4.1 Constructing Positive Samples: Divide all words with the same concept CUI code in the terminology database terms_db into two groups: the first group is the Chinese term set, and the second group is the English term set. Then, pair each term set to form term training pairs, which constitute the positive sample set set. + ;

[0092] S4.2.4.2 Constructing Negative Samples: Traverse each set element Q in the concept CUI code, calculate the similarity of the preferred words in the medical terminology database terms_db using formula (16), then take the top 100 preferred words and form a term pair set with Q, and finally remove the positive sample set. + The elements obtain the negative sample set set. - ;

[0093] S4.2.4.3 Constructing training samples: Randomly shuffle the set... + , set - From set at a ratio of 1:10 - Take negative samples and merge them into the set. + We obtain a training sample set, which, in addition to term pairs, includes the label {0, 1}, where 0 represents a negative sample and 1 represents a positive sample.

[0094] S4.2.4.4. Using the BERT model for scoring: The term pairs in the set constitute the input sequence "[cls]s1[seq]s2", and y selects {0, 1} as the sequence classification task to form examples. The loss function is the cross-entropy loss function.

[0095] S4.2.4.5. Divide the samples (examples) into training and test sets for training; evaluate the model using the F1 score, select the model with the highest F1 score from the test set and save it, forming a dual-model std_model (Chinese / English).zh and std_model en ;

[0096] S4.2.5 Two-stage terminology standard prediction:

[0097] S4.2.5.1 Input the standardized Chinese term "Q" and translate it into "Q". EN ;

[0098] S4.2.5.2 Based on formula (16), Q is retrieved from the Chinese preferred words and Chinese synonyms of the medical terminology database terms_db, the Top 30 are selected, term pairs are constructed, and input into the std_model of the second stage. zh In the model, scores are obtained, and the top 5 candidate sets C are selected. zh-top5 Similarly, Q EN The top 30 terms were retrieved from the medical terminology database terms_db, along with their English synonyms. These terms were then used to construct term pairs and input into the std_model for the second stage. en In the model, scores are obtained, and the top 5 candidate sets C are selected. en-top5 ;

[0099] S4.2.5.3, Integrate the final results; set the weighting λ for Chinese and English scoring. zh ,λ en ,

[0100] For C respectively zh-top5 With C en-top5 Calculate the score, then average the scores of CUI for the same concept to form the final standardized set C, and select the top 1 from C as the best standardized result;

[0101] V. Constructing a Drug Knowledge Graph

[0102] After being parsed by a natural language processing model, the drug instruction manual can produce a basic graph, which mainly consists of defined entity labels, entities, relation labels, and relations. The transformation from the basic graph to the final target graph also requires a data processing process, using RDFS / OWL technology to describe the graph schema and using SWRL language to reason about the basic graph to finally form the target graph.

[0103] VI. Building a Continuous Data Update Module

[0104] We use a microservice architecture to encapsulate deep learning natural language processing application services, and build the MLOPS process based on Docker, Kubernetes, and GitLab technologies. We use this technology to automate service deployment and continuously process data. When a new medical institution or pharmacy is needed, it only needs to upload the drug instruction manual data. The data continuous update module will automatically use the natural language processing model to process the data, build a drug knowledge graph, and deploy drug-related applications.

[0105] Furthermore, in the drug knowledge graph, the structure includes blank nodes to expand for different scenarios.

[0106] Furthermore, in the drug knowledge graph, the ingredient field is used to record information including the main ingredients and excipients of the drug; the indication field is used to record the diseases and symptoms to which the drug is applicable; the adverse reaction field is used to record information including the name, type, and frequency level of adverse reactions of the drug; the contraindication field is used to record contraindications to the drug, contraindications to use corresponding to symptoms and diseases, and contraindications to use in specific populations; the interaction field is used to record specific interaction information between drug components and between drug categories; and the dosage and administration field is used to record information on the purpose of drug use, route of administration, time of administration, population, frequency of administration, number of administrations, dosage type, dose value per administration, and dosage unit.

[0107] Furthermore, in the above method, the label system design is an iterative design process. When the performance of the natural language processing model does not meet the requirements, the label system needs to be dynamically modified.

[0108] The beneficial effects of this invention are as follows: This invention constructs a rational drug use system based on natural language processing and drug knowledge graph. This system can automatically perform data governance on drug instruction data and form a drug knowledge graph, and realize the functions of the rational drug use system based on the knowledge graph. Attached Figure Description

[0109] Figure 1 This is a flowchart illustrating the construction process of the natural language processing model in Embodiment 2 of the present invention. Detailed Implementation

[0110] The present invention will be further described below. It should be noted that this embodiment is based on the present technical solution and provides detailed implementation methods and specific operation processes, but the protection scope of the present invention is not limited to this embodiment.

[0111] Example 1

[0112] This embodiment provides a rational drug use system based on natural language processing and drug knowledge graph, including an automatic prescription review module, a drug recommendation module, a prompt information backtracking module, a drug information query module, and a data update and maintenance module;

[0113] The automatic prescription review module is used in medical settings to automatically review the rationality of prescriptions after a doctor issues a prescription, based on the input of the patient's basic information, medical information, and information about the specific drugs in the prescription, using a drug knowledge graph. The review includes assessing whether the drugs in the prescription are consistent with the symptoms and disease diagnosis in the medical information; whether there are interactions between the drugs in the prescription and between the drugs in the prescription and the patient's current medications; whether the patient's population type is a contraindication group for the drugs in the prescription; whether the patient has a history of allergies to the drugs in the prescription; whether there are any duplicate medications in the prescription; and whether the dosage and administration of the drugs in the prescription are consistent with the dosage and administration of the corresponding drugs in the drug knowledge graph. The patient's basic information includes age, gender, population type (pregnant woman, child, elderly, etc.), allergy history, and recent medication history; the medical information includes symptoms and disease diagnosis; and the information about the specific drugs in the prescription includes the generic name, manufacturer, specifications, approval number, and dosage and administration for each drug.

[0114] The recommended medication module generates a list of recommended medications based on the input patient's basic information and disease and symptom information in diagnosis and treatment or medication consultation scenarios. Specifically, the module uses a drug knowledge graph to filter suitable medications based on the input patient's disease and symptom information. Then, it filters the filtered medications based on the patient's population type, allergy history, and recent medication use, removing medications that are unsuitable for the patient's corresponding population type, may cause allergies, or overlap with or interact with medications the patient is currently taking. The remaining medications are then sorted according to user-defined sorting rules to generate the recommended medication list. Filtering based on population type and allergy history prevents recommending medications to individuals explicitly listed as unsuitable in the drug's instructions. Filtering based on recent medication use prevents recommending medications that interact with the patient's current medications and also avoids duplicate medication use.

[0115] The information backtracking module provides the source of prescription rationality review results and recommended drug lists. The sources include the original drug instructions and an interaction database. In both clinical and recommended drug scenarios, there is a need to view the source of information regarding prescription rationality review results or recommendations to ensure the rationality and safety of medication use. Information backtracking has two directions: the original drug instructions and the interaction database. For the original instructions, the results of prescription rationality review and recommended drug lists come from fields such as indications, interactions, contraindications, and dosage in the drug instructions. The interaction database explicitly indicates the interactions between components. The recommended drug module or the automatic prescription review module searches the interaction database for interactions by looking up the drug's components in the drug knowledge graph. Therefore, tracing the source of interaction results involves querying the detailed interaction information in that database.

[0116] The drug information query module provides a query interface for drug information, including the original instructions, drug knowledge graph, adverse reactions, indications, contraindications, and dosage. It's important to note that drug information queries are needed in both clinical and pharmacy settings to assist healthcare professionals and pharmacy staff in retrieving relevant drug information based on a patient's specific condition. The original instructions, drug knowledge graph, and dosage information return specific information based on the drug's generic name, manufacturer, specifications, and approval number. Adverse reactions, indications, and contraindications support both forward and reverse queries. Forward queries return information based on the drug's generic name, manufacturer, specifications, and approval number, while reverse queries return a list of specific drugs based on a particular adverse reaction, indication, or contraindication.

[0117] The data update and maintenance module provides an upload interface for drug instruction manuals and supports various data formats, including PDF, IMG, and text. For text data, the module uses a natural language processing model to process the uploaded manual text and update the drug knowledge graph accordingly. It also provides data quality control, allowing users to compare the automatically labeled text and generated knowledge graph with the original instruction manual to ensure data quality. For PDF and IMG data, the module uses image processing technology to extract the content and convert it into text, then processes the extracted text data using a natural language processing model.

[0118] In addition, to facilitate system deployment across different institutions, the data update and maintenance module provides a drug catalog integration function. Users can upload drug data, including fields such as generic name, manufacturer, specifications, and approval number, through this function. The data update and maintenance module uses a natural language processing model to automatically match the uploaded drug data to existing drug catalogs using an algorithm. For drugs not already in the catalog, the module automatically structures the uploaded drug instructions to form a drug knowledge graph; ultimately, the drug data integration is completed, and the system is ready for use.

[0119] Example 2

[0120] This embodiment provides a method for constructing the system described in Embodiment 1, the specific process of which is as follows:

[0121] The system described in Example 1 is a rational drug use system based on natural language processing and knowledge graph technology. This example will explain the construction method of the system described in Example 1 from several aspects, including drug knowledge graph design, tag system design, data annotation, natural language processing model construction, drug knowledge graph construction, and continuous data update process construction.

[0122] I. Designing a Drug Knowledge Graph

[0123] The system described in Example 1 is based on knowledge graph technology to provide various rational drug use information prompts. The knowledge graph design is well-suited to the application scenarios. Considering the business requirements of the rational drug use system, a drug knowledge graph structure containing corresponding fields is designed, including ingredients, indications, adverse reactions, contraindications, interactions, and dosage. Since drug usage scenarios vary under different conditions, including different patients and different times, the drug knowledge graph structure includes blank nodes to expand for different scenarios. The ingredient field records information including the main ingredients and excipients of the drug; the indication field records the diseases and symptoms to which the drug is applicable; the adverse reaction field records information including the name, type, and frequency level of adverse reactions; the contraindication field records contraindications related to allergies, symptoms, and diseases, as well as contraindications for specific populations; the interaction field records specific interaction information between drug ingredients and between drug categories; and the dosage field records information on the drug's purpose of use, route of administration, time of administration, population, frequency of administration, number of administrations, dosage type, dose value per administration, and dosage unit.

[0124] II. Design a labeling system:

[0125] This embodiment defines the tagging system based on the design of a drug knowledge graph, aiming to facilitate the conversion of annotation results into the drug knowledge graph. On the other hand, the difficulty of natural language processing must also be considered when designing tags. The tagging system design is also an iterative process; when the performance of the natural language processing model does not meet the requirements, the tagging system needs to be dynamically modified. Through this iterative process, a standard tagging system is ultimately formed.

[0126] III. Data Labeling:

[0127] After the labeling system design is completed, annotation tools are used to annotate the sample text data of the drug instructions. Data annotation includes two stages: annotation and review. Administrators create tasks and assign them to specific annotators. After completing the annotation, the annotators submit the annotated data to the administrator for review. Data that fails the review is returned to the annotators for inspection and modification, while data that passes the review can be used for subsequent natural language processing model training. Data annotation tasks are divided according to the fields designed in the drug knowledge graph structure. Each different field includes three annotation tasks: named entity annotation, entity relationship annotation, and annotation terminology alignment. Named entity annotation uses annotation tools to define the scope and select entity types for specific words based on entity tags, including continuous and non-continuous entity annotation. Entity relationship annotation specifies the directed relationship type between annotated named entities. Entity alignment annotation determines which standardized terminology an entity should be aligned to based on the named entity and entity type tags. To more comprehensively align entities to standard terms, this embodiment uses three alignment methods: equivalent, superordinate, and subordinate.

[0128] Furthermore, during the annotation process, a self-learning approach is used to dynamically evaluate the annotators' results and provide them with the annotated data with the highest model perplexity, allowing them to focus on strengthening their annotation efforts. Data annotation and model training are continuously alternated to achieve better model performance. Additional data annotation tasks are assigned to data where the model's performance is poor.

[0129] IV. Natural Language Processing Model Construction

[0130] This embodiment uses a natural language processing model to automatically parse text data obtained from user-uploaded drug instruction manuals, extracting entity and entity relationship information, and further aligning it to standard terminology to construct a drug knowledge graph. Therefore, the text data processing flow for drug instruction manuals is mainly divided into two modules: an information extraction module and a medical terminology alignment module. For example... Figure 1 As shown, the construction process of the natural language processing model is as follows:

[0131] 4.1 Constructing an information extraction module

[0132] This embodiment is based on the model in paper [1], and improves it into a model that includes tasks such as drug text classification, chapter structure analysis, entity extraction and relation joint extraction, as an information extraction module.

[0133] The drug text classification task is used to classify the input text data of drug instructions into traditional Chinese medicine and chemical drugs.

[0134] The BERT encoding task utilizes massive unsupervised corpora or pre-trained BERT models adapted to the medical domain for encoding. The BERT model structure originates from the BERT (Bidirectional Encoder Representations from Transformers) model proposed by Google AI in 2018. [2] Compared to the earliest algorithms for calculating language models, and the subsequent word vector technology Word2vec... [3] It is more semantically expressive and informative.

[0135] The discourse structure analysis task is used to segment the input drug instruction manual text into medical semantic blocks. The drug instruction manual text involves multiple fields such as indications, contraindications, and dosage. The discourse structure analysis task is used to segment different text segments into these fields and identify the medical semantic roles to which a particular text segment belongs.

[0136] The medical entity recognition task is used to extract medical-related entities mentioned in the input drug instruction text, such as diseases, symptoms, and drugs.

[0137] The medical relation extraction task is used to identify and determine the specific relationships between entity pairs in an input drug instruction manual. For example, "excessive stomach acid" (symptom) leads to "stomach pain" (symptom), which is a relationship where excessive stomach acid causes the symptom of stomach pain.

[0138] The model integrating the above tasks comprises a BERT encoder module, a head entity labeling module, a tail entity labeling module, a document extraction module, a document type classification module, and a loss function calculation module. The entire process utilizes a shared BERT encoder layer and back-optimizes the model through joint learning of multiple tasks. The data processing procedure of the model is as follows:

[0139] Suppose we have a text sequence X = (x1, x2, x3, ..., x...). n ), x t This represents the character or word at position t, and n represents the length of the text sequence.

[0140] S4.1.1 The BERT encoder module adds a predefined start character [CLS] to the beginning of the text sequence X, and then performs BERT encoding:

[0141] H=BERT([CLS]+X)............(1)

[0142] Where H represents the hidden state of the text sequence X after being encoded by the function BERT(·), H∈R n×d , where n represents the number of characters or words in the text sequence X, and d represents the dimension of the encoded vector;

[0143] S4.1.2 The header entity labeling module calculates the boundary probabilities of header entity labels as follows:

[0144]

[0145]

[0146] in, Let represent the probability of the i-th character or word indicating the start of the header entity. W represents the probability that the i-th character or word indicates the end of the header entity. s ∈R d×1 W e ∈R d×1 ,b s ∈R,b e ∈R represents the parameters to be learned, R represents the set of real numbers, and σ(·) represents the activation function.

[0147] Therefore, the likelihood function of the head entity can be calculated:

[0148]

[0149] Where s represents the head entity, and I(·) represents the indicator function. A marker indicating the start or end position of the i-th character or word, with a value of {0, 1}.

[0150] S4.1.3 The calculation relationship and tail entity boundary of the tail entity marking module are as follows:

[0151]

[0152]

[0153] in, This represents the probability that the i-th character or word indicates the start of the last entity. v represents the probability that the i-th character or word indicates the end of the entity. k ∈R dLet d represent the k-th head entity vector, and h represent the encoding dimension. i ∈H represents the implicit vector of a character or word; if the entity consists of multiple characters or words, an average value is calculated first. Let represent the parameters that the model needs to learn, m represent the number of relation types, R represent the set of real numbers, and σ(·) represent the activation function.

[0154] Therefore, the likelihood functions of the relations and entities can be calculated:

[0155]

[0156] Where o represents the tail entity, and I(·) represents the indicator function. A marker indicating the start or end position of the i-th character or word, with a value of {0, 1}.

[0157] S4.1.3 Assume the output label sequence is y = (y1, y2, y3, ..., y4). n The document extraction module calculates the total score of the document analysis sequence as follows:

[0158]

[0159] Where A∈R n×n Let W be the transition matrix. crf ∈R d×n ,b crf ∈R n This indicates the parameters to be learned.

[0160] The probability of a text sequence corresponding to a target sequence is calculated as follows:

[0161]

[0162] Among them, Y x This represents the set of possible target sequences.

[0163] In the prediction phase, the Viterbi algorithm is used to find the optimal sequence:

[0164]

[0165] S4.1.4, The instruction manual type classification module calculates the probability of text category:

[0166] p c =W c ×h cls +b c ............(11)

[0167] Among them, W c ∈R d ,bc ∈R represents the parameters to be learned, h cls The latent vector represents [CLS]. S4.1.5, the loss function calculation module calculates the loss function as follows:

[0168] L = -(l s+o +l crf +l c )......................(12)

[0169] in,

[0170]

[0171]

[0172]

[0173] Where M represents the total number of samples, n represents the text length, and λ and γ are regularization parameters.

[0174] 4.2 Constructing a terminology standardization module

[0175] Reference [4] states that the terminology standardization module is mainly based on a two-stage terminology standard model. First, candidate terms are recalled from standard terms, and then the candidate terms are finely ranked and calculated. This embodiment is based on this model framework and combines the Chinese medical terminology database to expand and comprehensively recall terms on an order of magnitude. The construction process of the terminology standardization module is as follows:

[0176] S4.2.1 Collect Chinese and English terminology corpora and open terminology databases, and construct a medical terminology database terms_db after sorting. The fields of the medical terminology database include Unicode CUI, English standard words, English synonyms, Chinese standard words, Chinese synonyms, etc.

[0177] S4.2.2. Based on indexing tools (such as ES) or using custom indexes, create indexes for Chinese and English. When querying Chinese terms, translate the Chinese terms into English and use both English and Chinese as input.

[0178] 4.2.3. Establish a keyword search engine and integrate recall scores through multi-channel recall. reca ll Assume s1 and s2 are two terms:

[0179] s recall =α1×s bm25 +α2×s Jaccard +α3×s MED +α4×s DICE .......(16)

[0180] Where α1, α2, α3, and α4 represent the weights of each scoring path, and α1 + α2 + α3 + α4 = 1;

[0181] BM25 ratings bm25 The calculation is as follows:

[0182]

[0183]

[0184] Where, ω i This represents the i-th word of the query term s1; f i It is a word ω i The frequency of occurrence of term s2, k1 and b are adjustment factors, len(·) is the function for calculating sentence length, avgsl is the average length of all documents in the index; N represents the total number of documents in the index, n(ω i ) contains ω i The number of documents.

[0185] Jaccard coefficients Jaccard The calculation is as follows:

[0186]

[0187] Where A and B represent the word segmentation sets of s1 and s2 respectively, and · represents the number of elements in the set.

[0188] Edit distance (MED) similarity score MED The calculation is as follows:

[0189]

[0190] Here, len(·) is the function for calculating sentence length, and d(s1,s2) represents the edit distance between s1 and s2.

[0191] DICE Distance Similarity Ratings DICE The calculation is as follows:

[0192]

[0193] Where A and B represent the word segmentation sets of s1 and s2 respectively, and · represents the number of elements in the set.

[0194] S4.2.4 Training the fine-ranking model:

[0195] S4.2.4.1 Constructing Positive Samples. Divide all words with the same concept CUI code in the terminology database (terms_db) into two groups: the first group is the Chinese terminology set, and the second group is the English terminology set. Then, pair each terminology set to form term training pairs, thus constructing the positive sample set `set`. + ;

[0196] S4.2.4.2 Constructing negative samples. Traverse each set element Q in the concept CUI code, calculate the similarity of the preferred words in the medical terminology database terms_db using formula (16), then take the top 100 preferred words and form a term pair set with Q, and finally remove the positive sample set. + The elements obtain the negative sample set set. - ;

[0197] S4.2.4.3 Constructing Training Samples. Randomly shuffle the set... + , set - From set at a ratio of 1:10 - Take negative samples and merge them into the set. + We obtain a training sample set, which, in addition to term pairs, includes the label {0, 1}, where 0 represents a negative sample and 1 represents a positive sample.

[0198] S4.2.4.4 The BERT model is used for scoring. The terms in the set constitute the input sequence “[cls]s1[seq]s2”, and y selects {0, 1} as the sequence classification task, forming examples. The loss function is the cross-entropy loss function.

[0199] S4.2.4.5. Divide the samples (examples) into training and test sets for training; evaluate the model using the F1 score, select the model with the highest F1 score from the test set and save it, forming a dual-model std_model (Chinese / English). zh and std_model en .

[0200] S4.2.5 Two-stage terminology standard prediction:

[0201] S4.2.5.1 Input the standardized Chinese term "Q" and translate it into "Q". EN ;

[0202] S4.2.5.2 Based on formula (16), Q is retrieved from the Chinese preferred words and Chinese synonyms of the medical terminology database terms_db, the Top 30 are selected, term pairs are constructed, and input into the std_model of the second stage. zh In the model, scores are obtained, and the top 5 candidate sets C are selected. zh-top5 Similarly, QEN The top 30 terms were retrieved from the medical terminology database terms_db, along with their English synonyms. These terms were then used to construct term pairs and input into the std_model for the second stage. en In the model, scores are obtained, and the top 5 candidate sets C are selected. en-top5 ;

[0203] S4.2.5.3, Integrate the final results. Set the weighting λ for the scores in Chinese and English. zh ,λ en ,

[0204] For C respectively zh-top5 With C en-top5 Calculate the score, then average the scores of CUI for the same concept to form the final standardized set C, and select the top 1 from C as the best standardized result.

[0205] V. Constructing a Drug Knowledge Graph

[0206] After being parsed by a natural language processing model, the drug instruction manual yields a basic graph, primarily composed of defined entity tags, entities, relation tags, and relations. However, this graph is not the final target graph. Transforming the basic graph into the final target graph requires further data processing. This embodiment uses RDFS / OWL / SWRL technologies for this transformation. Considering that knowledge graphs built based on RDFS / OWL technology have multiple data graph reasoning languages, RDFS / OWL technology is used to describe the graph schema, and SWRL is used to reason about the basic graph to ultimately form the target graph.

[0207] VI. Building a Continuous Data Update Module

[0208] This embodiment uses a microservice architecture to encapsulate deep learning natural language processing application services, and builds an MLOPS (Multi-Localization Optimization) process based on Docker, Kubernetes, and GitLab technologies. This technology automates service deployment and continuous data processing. When a new medical institution or pharmacy is presented, only the drug instruction manual data needs to be uploaded. The continuous data update module will automatically use a natural language processing model to process the data, build a drug knowledge graph, and deploy drug-related applications.

[0209] References:

[0210] [1] Wei, Zhepei, et al. "A novel cascade binary tagging framework for relational triple extraction." arXiv preprint arXiv: 1909.03227 (2019).

[0211] [2] Devlin, Jacob, et al. "Bert: Pre-training of deep bidirectional transformers for language understanding." arXiv preprint arXiv: 1810.04805 (2018).

[0212] [3]Mikolov, Tomas, et al. "Efficient estimation of word representations in vector space." arXiv preprint arXiv: 1301.3781 (2013).

[0213] [4] Sun et al. "Standardization of clinical terminology based on BERT." Journal of Chinese Information Processing 35.4(2021):8.

[0214] For those skilled in the art, various corresponding changes and modifications can be made based on the above technical solutions and concepts, and all such changes and modifications should be included within the protection scope of the claims of this invention.

Claims

1. A method for constructing a rational drug use system based on natural language processing and a drug knowledge graph, characterized in that, The rational drug use system based on natural language processing and drug knowledge graph includes an automatic prescription review module, a drug recommendation module, a prompt information backtracking module, a drug information query module, and a data update and maintenance module. The automatic prescription review module is used in medical settings to automatically review the rationality of prescriptions after a doctor issues a prescription, based on the input of the patient's basic information, medical information, and information about the specific drugs in the prescription, using a drug knowledge graph. The review includes assessing whether the drugs in the prescription are consistent with the symptoms and disease diagnosis in the medical information; whether there are interactions between the drugs in the prescription and between the drugs in the prescription and the patient's current medications; whether the patient's population type is within the contraindications group for the drugs in the prescription; whether the patient has a history of allergies to the drugs in the prescription; whether there are duplicate medications in the prescription; and whether the dosage and administration of the drugs in the prescription are consistent with the dosage and administration of the corresponding drugs in the drug knowledge graph. The patient's basic information includes age, gender, population type, allergy history, and recent medication history. The medical information includes symptoms and disease diagnosis results. The information about the specific drugs in the prescription includes the generic name, manufacturer, specifications, approval number, and dosage and administration of each drug. Recommended medication module: Used in diagnosis and treatment scenarios or medication consultation scenarios, it generates a list of recommended medications based on the input patient's basic information and disease and symptom information using a drug knowledge graph; Information traceability module: This module provides a query interface for the sources of prescription rationality review results and recommended drug lists. The sources include the original drug instructions and interaction databases. For the original instructions, the results of prescription rationality review and recommended drug list come from the indications, interactions, contraindications, and dosage of the drug instructions; the interaction database clearly points out the interactions between components, and the recommended drug module or prescription automatic review module searches the interaction database for the interaction of drug components in the drug knowledge graph. Therefore, when tracing the interaction results, it is to query the detailed interaction information in the database. Drug Information Query Module: This module provides a query interface for querying drug information, including the original instructions, drug knowledge graph, adverse reactions, indications, contraindications, and dosage. Data update and maintenance module: This module is used to upload drug instructions and provides an upload interface. It uses a natural language processing model to process the text data of the uploaded drug instructions and update the drug knowledge graph accordingly. The specific process of the method is as follows: S1. Design a drug knowledge graph: Based on the business requirements of the rational drug use system, a drug knowledge graph structure containing corresponding fields was designed, including ingredients, indications, adverse reactions, contraindications, interactions, and usage and dosage. S2. Design a labeling system: The labeling system is defined based on the design of a drug knowledge graph; S3, Data Labeling: After the labeling system design is completed, annotation tools are used to annotate the sample text data of the drug instructions. The data annotation work is divided into tasks according to the field design of the drug knowledge graph structure. Each different field includes three annotation tasks: named entity annotation, entity relationship annotation, and annotation term alignment. Named entity annotation is to define the scope and select the entity type of specific words based on the entity tags, including continuous entity annotation and non-continuous entity annotation. Entity relationship annotation specifies the directed relationship type between annotated named entities. Entity alignment annotation determines which standardized term the entity is aligned to based on the named entity and entity type tags. S4. Natural Language Processing Model Construction: The natural language processing model automatically parses text data obtained from user-uploaded drug instructions, extracts entity and entity relationship information, and further aligns it to standard terminology to construct a drug knowledge graph; the construction process of the natural language processing model is as follows: S4.1, Constructing the Information Extraction Module The information extraction module includes models for drug text classification, discourse structure analysis, BERT encoding tasks, entity extraction, and joint relation extraction. The drug text classification task is used to classify the input text data of drug instructions into traditional Chinese medicine and chemical drugs. The BERT encoding task is used to encode using massive unsupervised corpora or pre-trained BERT adapted to the medical field; The text structure analysis task is used to divide the input drug instruction manual text into text medical blocks semantics. The drug instruction manual text involves roles such as indications, contraindications, and usage and dosage. The text structure analysis task is used to divide different text segments into fields and identify the medical semantic roles to which a certain text segment belongs. The medical entity recognition task is used to extract medical-related entities mentioned in the input drug instruction manual text; The medical relation extraction task is used to identify and determine the specific relationships between entity pairs in an input drug instruction manual text. The model integrating the above tasks includes a BERT encoder module, a head entity labeling module, a tail entity labeling module, a document extraction module, a document type classification module, and a loss function calculation module. The entire process is optimized through a shared BERT encoding layer and multi-task joint learning. The data processing procedure of the model is as follows: Suppose we have a text sequence as follows: , Indicates the first The character or word at each position, Indicates the length of the text sequence; S4.1.1 The BERT encoder module adds a set start character at the beginning of the text sequence X. Then perform BERT encoding: ; in, Represents a text sequence After function Encoded hidden state , Represents a text sequence The number of characters or words, This represents the dimension of the encoded vector. S4.1.2 The header entity labeling module calculates the boundary probabilities of header entity labels as follows: ; in, For the first Each character or word represents the probability of starting a header entity. For the first Each character or word represents the probability of the head entity ending. This represents the parameters to be learned. Represents the set of real numbers. Indicates the activation function; Therefore, the likelihood function of the head entity can be calculated: ; in, Indicates the head entity. Indicates an indicator function, A marker indicating the start or end position of the i-th character or word, with a value of {0, 1}; S4.1.3 The calculation relationship and tail entity boundary of the tail entity marking module are as follows: ; in, Indicates the first Each character or word represents the probability of starting the last entity. Indicates the first Each character or word represents the probability of the tail entity ending. Indicates the first Individual entity vectors Indicates the dimension of the encoding. The implicit vector representing character or word i; if the entity consists of multiple characters or words, perform an average calculation first. These represent the parameters that the model needs to learn. Indicates the number of relation types. Represents the set of real numbers. Indicates the activation function; Thus, the likelihood functions of the relations and entities can be calculated: ; in, Indicates the tail entity. Indicates an indicator function, A marker indicating the start or end position of the i-th character or word, with a value of {0, 1}; S4.1.4, Assume the output label sequence is The document extraction module calculates the total score of the document analysis sequence as follows: ; in, Let be the transition matrix. Indicates the parameters to be learned; The probability of a text sequence corresponding to a target sequence is calculated as follows: ; in, Represents the set of target sequences; In the prediction phase, the Viterbi algorithm is used to find the optimal sequence: ; S4.1.5, The instruction manual type classification module calculates the probability of text category: ; in, Indicates the parameters to be learned. express The implicit vector; S4.1.6 The loss function calculation module calculates the loss function as follows: ; in, ; ; ; Where M represents the total number of samples, and n represents the text length. For regularization parameters; S4.2, Constructing a Terminology Standardization Module The terminology standardization module is mainly based on a two-stage terminology standardization model. First, candidate terms are recalled from standard terms, and then the candidate terms are finely ranked and calculated. The construction process of the terminology standardization module is as follows: S4.2.1 Collect Chinese and English terminology corpora and open terminology databases, organize them and construct a medical terminology database. The fields of the medical terminology database include Unified Code (CUI), English standard words, English synonyms, Chinese standard words, and Chinese synonyms. S4.2.

2. Based on indexing tools or using custom indexes, create indexes for Chinese and English. When querying Chinese terms, translate the Chinese terms into English and use both English and Chinese as input. S4.2.3 Establish a keyword search engine and integrate recall scoring through multi-channel recall. Assuming and For two terms: in, These represent the weights of each scoring method, and ; Among them, BM25 score The calculation is as follows: in, Representing query terms The Each word segmentation; It is a word In terms Frequency of occurrence in and As a regulating factor, To calculate the sentence length function, The average length of all documents in the index; This represents the total number of documents in the index. For includes The number of documents; Jaccard coefficient The calculation is as follows: in, They represent and The word segmentation set, Indicates the number of elements in the set; Edit distance (MED) similarity score The calculation is as follows: in, To calculate the sentence length function, express and Edit distance; DICE Distance Similarity Rating The calculation is as follows: in, They represent and The word segmentation set, Indicates the number of elements in the set; S4.2.4 Training the fine-ranking model: S4.2.4.1 Constructing Positive Samples: Divide all words with the same concept CUI code in the terminology database terms_db into two groups: the first group is the Chinese terminology set, and the second group is the English terminology set. Then, pair each terminology set to form term training pairs, which constitute the positive sample set. ; S4.2.4.2 Constructing negative samples: Traverse each set element Q in the concept CUI code, calculate the similarity of the preferred words in the medical terminology database terms_db using formula (16), then take the top 100 preferred words and form a term pair set with Q, and finally remove the positive samples. The elements yield a negative sample set. ; S4.2.4.3 Constructing training samples: Randomly shuffle each sample. , At a ratio of 1:10 from Take negative samples and merge them into , obtain training samples , In addition to the terminology pairs, the {0, 1} label has been added, where 0 represents a negative sample and 1 represents a positive sample; S4.2.4.

4. Using the BERT model for scoring: For the input sequence "[cls]s1[seq]s2", the terminology in the example forms the sequence classification task, y selects {0, 1}, and the loss function is the cross-entropy loss function. S4.2.4.

5. Divide the samples into training and test sets for training; evaluate the model using the F1 score, select the model with the highest F1 score from the test set and save it, forming a dual-model system (Chinese-English). and ; S4.2.5 Two-stage terminology standard prediction: S4.2.5.1 Input requires standardized Chinese terminology. Translated into ; S4.2.5.2 Based on formula (16), The top 30 terms were retrieved from the medical terminology database terms_db, using the preferred Chinese terms and their synonyms. These terms were then used to construct term pairs and input into the second phase of the database. In the model, scores are obtained, and the top 5 candidates are selected. Similarly, The top 30 terms were retrieved from the medical terminology database terms_db, along with their English synonyms, and then used to construct term pairs for the second phase of research. In the model, scores are obtained, and the top 5 candidates are selected. ; S4.2.5.3, Integrate the final results; set weighted scoring for Chinese and English. , respectively and Calculate the scores, and then average the scores of CUI for the same concept to form the final standardized set. The top 1 from C is taken as the best standardized result; S5. Constructing a Drug Knowledge Graph After the drug instruction manual is parsed by a natural language processing model, a basic graph is obtained, which mainly consists of defined entity labels, entities, relation labels, and relations. The transformation from the basic graph to the final target graph also requires a data processing process. The graph schema is described using RDFS / OWL technology, and the SWRL language is used to reason about the basic graph to finally form the target graph. S6. Build a continuous data update module We use a microservice architecture to encapsulate deep learning natural language processing application services, and build the MLOPS process based on Docker, Kubernetes, and GitLab technologies. We use this technology to automate service deployment and continuously process data. When a new medical institution or pharmacy is needed, it only needs to upload the drug instruction manual data. The data continuous update module will automatically use the natural language processing model to process the data, build a drug knowledge graph, and deploy drug-related applications.

2. The method according to claim 1, characterized in that, In the drug knowledge graph, the structure includes blank nodes to expand different scenarios.

3. The method according to claim 1, characterized in that, In the drug knowledge graph, the ingredient field is used to record information including the main ingredients and excipients of the drug; the indication field is used to record the diseases and symptoms to which the drug is applicable; and the adverse reaction field is used to record information including the name, type, and frequency level of the adverse reactions of the drug. The contraindications field is used to record contraindications to the drug, contraindications to use for symptoms and diseases, and contraindications to use for specific populations; the interaction field is used to record specific interaction information between drug components and between drug categories; the dosage and administration field is used to record information on the purpose of drug use, route of administration, time of administration, population, frequency of administration, number of administrations, dosage type, dose value per administration, and dosage unit.

4. The method according to claim 1, characterized in that, Tag system design is an iterative process. When the performance of the natural language processing model does not meet the requirements, the tag system needs to be dynamically modified.

5. The method according to claim 1, characterized in that, The recommended medication module uses a drug knowledge graph to filter suitable medications based on the input patient's disease and symptom information. Then, it filters the filtered medications based on the patient's population type, allergy history, and recent medication use, removing medications that are unsuitable for the patient's corresponding population type, may cause allergies, or overlap with or interact with the patient's currently taken medications. The remaining medications are sorted according to the sorting rules set by the user to generate a recommended medication list.

6. The method according to claim 1, characterized in that, When querying the original instructions, drug knowledge graph, or usage and dosage, specific information is returned based on the drug's generic name, manufacturer, specifications, and approval number. Querying adverse reactions, indications, and contraindications supports both forward and reverse queries. Forward queries return information based on the drug's generic name, manufacturer, specifications, and approval number, while reverse queries return a specific list of drugs based on a particular adverse reaction, indication, or contraindication.

7. The method according to claim 1, characterized in that, The data update and maintenance module supports uploading drug instruction manuals in various data formats, including PDF, IMG, and text. For text data, the module directly uses a natural language processing model to process the uploaded instruction manual text data and update the drug knowledge graph accordingly. For PDF and IMG data, the module uses image processing technology to extract the content and convert it into text, and then uses a natural language processing model to process the extracted text data.

8. The method according to claim 1 or 7, characterized in that, The data update and maintenance module provides data quality control functions, allowing users to compare the results of automated text annotation and graph generation with the original text of the instruction manual to ensure data quality.

9. The method according to claim 1, characterized in that, The data update and maintenance module provides a drug catalog integration function. Users can upload drug data, including fields such as generic name, manufacturer, specifications, and approval number, through this function. The data update and maintenance module uses a natural language processing model to automatically match the uploaded drug data to the existing drug catalog. For drugs that are not in the original drug catalog, the data update and maintenance module automatically structures the uploaded drug instructions to form a drug knowledge graph. Finally, the drug data integration is completed and the data is put into use.