A drug adverse reaction monitoring and early warning method
By constructing a multi-dimensional patient attribute labeling system and drug adverse reaction knowledge map, combined with machine learning algorithms and real-time data monitoring, the problem of inaccurate individual risk assessment of drug adverse reactions in the existing technology is solved, early identification and early warning of drug adverse reactions is achieved, and drug safety and treatment effect are improved.
Patent Information
- Application Number
- CN202410522406.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-04-28
AI Technical Summary
The prior art is difficult to accurately identify and early warning of individual risks of adverse drug reactions, and lacks quantitative indicators and standards, which mainly rely on the empirical judgment of physicians.
By constructing a multi-dimensional patient attribute labeling system and drug adverse reaction knowledge map, combining machine learning algorithms and real-time data monitoring, personalized assessment and early warning of the risk of adverse reactions in patients. Specific steps include: classifying and annotating the patient's attributes and building a patient's attribute tag system; using decision trees or random forest algorithms to train the correlation model between patient's attributes and adverse reaction risks; building a knowledge graph of adverse drug reactions to semantic association and fusion; obtaining the attribute information of individual patients, semantically matches with nodes in the knowledge graph to judge their adverse reaction risk levels; generating personalized drug use suggestions and monitoring plans; dynamically collecting real-time data such as physiological indicators, and by comparing with the knowledge graph, early identification and early warning of drug adverse reactions.
It has achieved accurate identification and evaluation of the risks of adverse drug reactions in different patient groups, prevented and reduced the occurrence of adverse reactions in advance, improved the safety and treatment effect of drugs, and achieved personalized treatment.
Smart Images

Figure CN118280603B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a drug adverse reaction monitoring and early warning method. Background Art
[0002] The incidence and manifestation patterns of adverse drug reactions vary significantly among different patient groups, but the inherent laws of such differences are difficult to reveal through conventional statistical methods. Patients' genetic background, medical history, and concomitant medications are closely related to the risk of adverse reactions, but it is currently impossible to accurately quantify the extent to which these factors affect individual risks. In addition, due to the diverse phenotypes and complex mechanisms of adverse drug reactions, which involve multiple biological processes such as drug metabolism, immune response, and organ toxicity, a single biomarker is difficult to fully reflect the susceptibility differences of patients. For patients who have already experienced adverse reactions, it is necessary to analyze the severity of the reaction, the involvement of organ systems, suspected drugs, etc., to infer the possible causes of the occurrence. In clinical practice, physicians need to comprehensively assess the risk of adverse reactions based on the patient's demographic characteristics, medical history, medication status, and other information, but there is a lack of quantitative indicators and standards, and they mainly rely on empirical judgment. Summary of the invention
[0003] The present invention provides a method for monitoring and early warning of adverse drug reactions, which mainly includes:
[0004] Classify and label patient groups according to different regions, ages, and gender attributes, build a multi-dimensional patient attribute label system, and through combined analysis of patient attribute labels, explore the differences in the incidence and manifestation patterns of adverse drug reactions in different patient groups;
[0005] Combined with the mined difference rules, the decision tree or random forest algorithm is used to train the association model between patient attributes and adverse reaction risks. By analyzing the feature importance of the association model, the strength of the correlation between the attribute characteristics of the patient's genetic background, previous medical history, and combined medication and the adverse reaction risk is determined;
[0006] Based on the attribute features that are strongly associated with adverse reaction risks, a knowledge graph of adverse drug reactions is constructed to semantically associate and fuse multi-source heterogeneous data including patient attributes, drug information, adverse reaction phenotypes, and biological mechanisms. Through multi-hop reasoning and path analysis of the knowledge graph of adverse drug reactions, the association between different factors is revealed.
[0007] Obtain the attribute information of individual patients, including region, age, gender, genetic background, past medical history, and concomitant medications. By semantically matching the corresponding nodes in the knowledge graph, determine the similarity between the patient and known adverse reaction cases and estimate the adverse reaction risk level.
[0008] Generate personalized medication recommendations and monitoring plans based on the attributes of high-risk patients and the corresponding association information in the knowledge graph, including adjusting drug dosages, avoiding specific drug combinations, and strengthening monitoring of specific adverse reactions to reduce the probability and severity of adverse reactions;
[0009] During the medication process, the patient's physiological indicators, symptom reports, and real-time data of test results are dynamically collected, and compared with the adverse reaction phenotypes in the knowledge graph in real time to achieve early identification and early warning of adverse drug reactions. If suspected adverse reaction signals are found, the doctor is notified to intervene and deal with them;
[0010] For newly identified adverse drug reaction cases, extract information including clinical manifestations, occurrence time, and severity, and determine whether the newly identified adverse drug reaction cases are new adverse reaction phenotypes or special population effects by comparing and analyzing with existing cases in the knowledge graph. Feedback the judgment results to the knowledge graph to update the adverse drug reaction knowledge graph.
[0011] Based on the updated knowledge graph of adverse drug reactions, the group distribution patterns and individual difference patterns of adverse drug reactions are explored, and the common characteristics of high-risk groups are identified through cluster analysis methods. Prevention and intervention strategies are then formulated accordingly to achieve management and personalized treatment of adverse drug reactions.
[0012] Combined with the biological mechanism information in the adverse drug reaction knowledge graph, analyze the adverse drug reaction cases that have occurred.
[0013] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:
[0014] The present invention discloses a method for monitoring and early warning of adverse drug reactions. Accurately identify and evaluate the risk of adverse drug reactions in different patient groups, prevent and reduce the occurrence of these adverse reactions in advance, classify and label patient groups according to different regions, ages and gender attributes, and build a multi-dimensional patient attribute label system for analyzing the incidence and performance pattern differences of adverse reactions. Use decision trees or random forest machine learning algorithms to train the association model between patient attributes and adverse reaction risks, and determine the key factors affecting adverse reaction risks through feature importance analysis. Construct a knowledge graph of adverse drug reactions, semantically associate and fuse multi-source heterogeneous data, and reveal the association and influence mechanism of different factors through reasoning and path analysis of knowledge graphs. Obtain the attribute information of individual patients, estimate their adverse reaction risk level by matching with nodes in the knowledge graph, and generate personalized medication recommendations and monitoring plans. During the patient's medication process, dynamically collect real-time data such as physiological indicators, and compare them with the knowledge graph to achieve early identification and early warning of adverse drug reactions. Analyze newly identified adverse reaction cases to determine whether they are new reaction phenotypes or special population effects, and update the knowledge graph. Based on the updated knowledge graph, the characteristics of high-risk groups are identified through cluster analysis methods, and corresponding prevention and intervention strategies are formulated. The present invention realizes personalized assessment and early warning of patient adverse reaction risks by constructing a multi-dimensional patient attribute label system and adverse drug reaction knowledge graph, combined with machine learning algorithms and real-time data monitoring. In addition, the present invention can dynamically update the knowledge graph to reflect the latest clinical information and adverse drug reaction cases, thereby improving drug safety and treatment effects, and realizing precision medicine and personalized treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 The present invention is a flowchart of a method for monitoring and early warning of adverse drug reactions.
[0016] Figure 2 The figure is a schematic diagram of a method for monitoring and early warning of adverse drug reactions of the present invention.
[0017] Figure 3 This is another schematic diagram of a method for monitoring and early warning of adverse drug reactions of the present invention. DETAILED DESCRIPTION
[0018] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0019] like Figure 1-3 In this embodiment, a method for monitoring and early warning of adverse drug reactions may specifically include:
[0020] Step S101, classify and label patient groups according to different regions, ages and gender attributes, build a multi-dimensional patient attribute label system, and through combined analysis of patient attribute labels, explore the differences in the incidence and manifestation patterns of adverse drug reactions in different patient groups.
[0021] According to the attributes of the patient group, including region, age, and gender, the patient group is classified and labeled to obtain a multi-dimensional patient attribute label system; the drug usage data of the patient group is obtained, and the drug usage data includes drug type, dosage, frequency, and course of treatment, and the drug usage data is associated with the patient attribute label system to construct a patient-drug attribute matrix; through combined analysis of the patient-drug attribute matrix, the difference patterns of the incidence and manifestation patterns of adverse drug reactions in different patient groups are explored.
[0022] Specifically, according to different attributes of patients, such as region, age, gender, etc., the patient groups are classified and labeled to build a multi-dimensional patient attribute label system. Obtain the patient's drug use data, including drug type, dosage, frequency, course of treatment, etc., associate the drug attributes with the patient attributes to form a patient-drug attribute data set. Preprocess and feature engineering are performed on the patient-drug attribute data set, including data cleaning, missing value processing, attribute encoding, feature selection, etc., to prepare for subsequent data mining analysis. By performing combined analysis on the patient-drug attribute data set, the different laws of the incidence and manifestation patterns of adverse drug reactions in different patient groups are mined. Association rule mining algorithms such as Apriori and FP-growth and sequential pattern mining algorithms such as GSP, SPADE, and PrefixSpan are used to discover the association rules and time series patterns between different patient attribute combinations and the incidence and manifestation patterns of adverse drug reactions. The specific implementation process of the association rule mining algorithm includes frequent item set generation and association rule generation, and the specific implementation process of the sequential pattern mining algorithm includes sequential pattern generation. When constructing a patient attribute label system, a multi-dimensional attribute combination can be used. For example, patients can be divided into East China, North China, South China, Southwest China, Northwest China and other regions according to region, divided into infants, children, adolescents, young adults, middle-aged and elderly groups according to age, and divided into males and females according to gender, forming a 3×5×2 three-dimensional attribute label matrix. When obtaining patient drug use data, multi-source heterogeneous data such as patient prescription information, medical insurance records, and drug purchase records can be collected, and they can be mapped into a unified patient-drug attribute table using data fusion technology. For attributes with many missing values, interpolation, regression or matrix decomposition methods can be used to fill them; for categorical attributes, one-hot encoding or ordinal encoding can be used for digitization; for numerical attributes, normalization or standardization methods can be used for dimensionless processing; for high-dimensional attributes, principal component analysis or factor analysis dimensionality reduction methods can be used for feature selection. When mining association rules, the Apriori algorithm can be used, with the minimum support threshold set to 0.05 and the minimum confidence threshold set to 0.8, to obtain the association rule that the probability of dizziness adverse reaction is 0.12 when the patient is over 60 years old, has a history of hypertension, and takes antihypertensive drugs. When mining sequence patterns, the PrefixSpan algorithm can be used, with the time window set to 30 days and the minimum support threshold set to 0.1, to obtain the medication sequence pattern that the patient uses antipyretic and analgesic drug B within 3 days after using antibiotic drug A, and then uses antidiarrheal drug C.
[0023] Step S102, combining the mined difference rules, using a decision tree or random forest algorithm to train an association model between patient attributes and adverse reaction risks, and determining the strength of correlation between the attribute characteristics including the patient's genetic background, past medical history, and combined medication and the adverse reaction risk by analyzing the feature importance of the association model.
[0024] Based on the attribute data of the patient group, the attribute data includes genetic background, past medical history, and combined medication information, a patient attribute data set is constructed; the patient attribute data set is preprocessed to obtain a preprocessed patient attribute data set; features related to adverse reactions are extracted from the patient attribute data set to form an adverse reaction risk feature set; a machine learning algorithm is used to train the association model between the patient attributes and the adverse reaction risk with the adverse reaction risk feature set as input and the adverse reaction occurrence as output; based on the association model, the influence weight of each patient attribute feature on the adverse reaction risk is calculated, and the weights are sorted to determine the correlation strength between the patient's genetic background, past medical history, combined medication factors and the adverse reaction risk; according to the correlation strength sorting results, several patient attribute features with the strongest correlation are selected as key indicators for adverse reaction risk prediction, and a risk prediction model is constructed; the patient attribute data set and adverse reaction-related data are regularly updated, and the association model and risk prediction model are retrained.
[0025] Specifically, the patient's attribute data, including genetic background, past medical history, combined medication and other information, are collected to construct a structured patient attribute data set. The patient attribute data set is preprocessed, including data cleaning, missing value filling, data encoding and other operations to ensure data quality and availability. Features related to adverse reactions, such as age, gender, allergy history, liver and kidney function indicators, gene polymorphism, etc., are extracted from the patient attribute data set to form an adverse reaction risk feature set. Machine learning algorithms, such as decision trees and random forests, are used to train the association model between patient attributes and adverse reaction risks with the adverse reaction risk feature set as input and the adverse reaction occurrence as output. When training the association model, it is necessary to reasonably select the hyperparameters of the algorithm, such as the maximum depth of the decision tree, the minimum number of samples of leaf nodes, the node splitting standard, etc. Grid search, random search, Bayesian optimization and other methods can be used for parameter tuning. At the same time, appropriate model evaluation indicators, such as accuracy, AUC value, F1-score, etc., should be selected to judge the model performance. The feature importance analysis function of the association model is used to calculate the influence weight of each patient attribute feature on the adverse reaction risk. For different types of machine learning models, corresponding feature importance calculation methods can be used, such as feature importance calculation based on tree models, regression coefficients based on linear models, and sensitivity analysis based on neural networks. According to the weight order, the correlation strength between factors such as patient genetic background, past medical history, and combined medication and adverse reaction risk is determined. According to the correlation strength ordering results, the top N patient attribute features with the strongest correlation are selected as key indicators for adverse reaction risk prediction to build a risk prediction model. When selecting a risk prediction model, models such as Logistic regression, support vector machine, and neural network commonly used in the medical field can be considered. Since there is often a sample imbalance problem in medical data, it is necessary to use oversampling, undersampling, and cost-sensitive learning methods to process it. After the risk prediction model is trained, the risk prediction model should be evaluated using indicators such as sensitivity, specificity, precision, and recall to ensure its effectiveness in practical applications. Patient attribute data sets and adverse reaction-related data should be updated regularly, association models and risk prediction models should be retrained, model performance should be continuously optimized, and the accuracy and real-time performance of adverse reaction risk prediction should be improved. By continuously improving the patient attribute-adverse reaction association model and risk prediction model, medical institutions can more accurately and efficiently identify high-risk patients, optimize clinical decision-making, reduce the incidence of adverse reactions, and improve patient medication safety. When collecting patient attribute data, data can be obtained from multiple sources such as the hospital's electronic medical record system and health management system, and data integration technology can be used to clean, convert and integrate data from different sources and formats to form a unified patient attribute data set.For example, the patient's gene polymorphism information is extracted from the gene test report, such as the CYP2C9*3 allele frequency of 0.15; the patient's allergy history information is extracted from the previous medical records, such as the penicillin allergy history positive rate of 0.2; the patient's combined medication information is extracted from the doctor's order record, such as the proportion of simultaneous use of non-steroidal anti-inflammatory drugs and warfarin is 0.05. When extracting features, statistical methods such as chi-square test and mutual information can be used to calculate the correlation between each attribute feature and the occurrence of adverse reactions, and select features with a P value less than 0.01 or a mutual information greater than 0.05 as risk feature sets. Features with a P value less than 0.01 or a mutual information greater than 0.05 are considered to be strongly associated with the risk of adverse reactions. When training the association model, 20,000 patient data are used as the training set and 5,000 as the test set, and a 5-fold cross validation method is used for training and evaluation. For the random forest model, the optimal hyperparameters obtained by grid search are the number of trees as 100, the maximum tree depth as 8, the minimum number of samples required for node splitting as 50, and the node splitting quality evaluation standard as the Gini index. The AUC value of the trained random forest model on the test set was 0.85, and the F1 score was 0.78. In the feature importance analysis, the Mean Decrease Impurity method was used to obtain the importance scores of age, gene polymorphism, liver function abnormality, and number of combined medications, which were 0.25, 0.18, 0.15, and 0.12, respectively. Based on this, the top 10 important features were selected to construct the Logistic regression prediction model. Due to the low incidence of adverse reactions, which is a class imbalance problem, the SMOTE algorithm was used to oversample the minority class samples with adverse reactions to increase their proportion to 40% to improve the prediction sensitivity of the model. The sensitivity of the obtained Logistic regression model on the test set was 0.75, the specificity was 0.90, the accuracy was 0.82, and the AUC value was 0.88. The model was updated every six months, and the model was retrained using the newly collected 10,000 patient data, and the performance indicators of the model were continuously tracked to dynamically adjust the model parameters and feature selection strategies.
[0026] Step S103, based on the attribute features that are strongly associated with the adverse reaction risk, a knowledge graph of adverse drug reactions is constructed, and semantic association and fusion of multi-source heterogeneous data including patient attributes, drug information, adverse reaction phenotypes, and biological mechanisms are performed. Through multi-hop reasoning and path analysis of the knowledge graph of adverse drug reactions, the association between different factors is revealed.
[0027] Multi-source heterogeneous data related to adverse drug reactions are parsed and extracted, wherein the multi-source heterogeneous data include unstructured and semi-structured data; based on the structured data, semantic modeling and ontology construction are performed by defining entity types and relationship types between entities to obtain an ontology architecture of an adverse drug reaction knowledge graph; words and phrases belonging to defined entity types and relationship types are identified from the multi-source heterogeneous data, and semantic mapping and association between the multi-source heterogeneous data and the knowledge graph are performed; structural similarities between nodes are captured on the ontology architecture of the adverse drug reaction knowledge graph to discover implicit associations between patient attributes, drug information, adverse reaction phenotypes, and biological mechanism factors.
[0028] Specifically, based on multi-source heterogeneous data related to adverse drug reactions, including patient electronic medical records, drug instructions, drug vigilance reports, biomedical literature, etc., these unstructured and semi-structured data are parsed and extracted to obtain structured data such as patient attributes, drug information, adverse reaction phenotypes, and biological mechanisms. Semantic modeling and ontology construction are performed on the extracted structured data to define entity types such as patients, drugs, adverse reactions, and biological mechanisms, as well as the relationship types between entities, such as "patient-use-drug", "drug-induced-adverse reaction", "adverse reaction-involved-biological mechanism", etc., to form the ontology architecture of the knowledge graph of adverse drug reactions. Using named entity recognition methods such as rule-based, dictionary-based or machine learning and relationship extraction methods based on pattern matching, dependency analysis or deep learning, words and phrases belonging to defined entity types and relationship types are identified from multi-source data, and semantic mapping and association between multi-source heterogeneous data and knowledge graphs are achieved by matching and associating with defined entity types and relationship types in the knowledge graph. Through knowledge graph representation learning methods such as TransE, TransR, and ComplEx, entities and relationships in the knowledge graph are embedded into a low-dimensional vector space, so that semantically similar entities and relationships are close in the vector space, and semantically different entities and relationships are far away in the vector space, so that the semantic association strength between entities can be measured and predicted through vector operations. Multi-hop reasoning algorithms based on random walks such as DeepWalk, node2vec, MetaPath2Vec, and HIN2Vec are applied on the knowledge graph to generate node sequences by random walks on the graph, capturing the structural similarity between nodes, thereby discovering implicit associations between factors such as patient attributes, drug information, adverse reaction phenotypes, and biological mechanisms. At the same time, rule-based or path-based reasoning methods such as Horn clause-based logical reasoning and Prolog-based rule reasoning can also be used to infer unknown entities or relationships using explicit rules or path patterns in the knowledge graph. Adverse drug reaction prediction can be performed based on graph neural networks, and end-to-end graph neural networks are used to treat adverse drug reaction prediction as a link prediction task. By embedding entities and relationships in the knowledge graph as additional features of drug and patient attributes, and inputting them into the graph neural network together with other structured features, we directly apply models such as graph convolutional networks (GCN) and graph attention networks (GAT) on the knowledge graph to predict whether there is a relationship of "producing adverse reactions" between patient nodes and drug nodes. During model training, we use negative sampling methods to construct positive and negative samples, use optimization targets such as cross entropy loss functions, and use drug-adverse reaction association matrices for semi-supervised learning, thereby achieving risk prediction of adverse drug reactions.Based on the visualization exploration and analysis tool of the knowledge graph of adverse drug reactions, users can perform entity search, relationship query, path analysis and other operations through a graphical interface, revealing the complex association chain between factors such as patient attributes, drug information, adverse reaction phenotypes, and biological mechanisms, identifying high-risk groups and high-risk drugs for adverse reactions, and assisting clinical decision-making and drug vigilance. Through the construction, association analysis, risk prediction and visualization application of the knowledge graph of adverse drug reactions, we can understand the influencing factors and mechanisms of adverse drug reactions more comprehensively and systematically, and provide intelligent decision-making support for personalized medication and precision medicine. When collecting adverse drug reaction data, we can extract the medication and diagnosis information of 50,000 patients from the electronic medical record systems of 10 hospitals, and use regular expressions and dictionary matching methods to identify 2,000 drugs and 500 adverse reactions from unstructured medical orders and diagnosis records; then use the conditional random field (CRF) algorithm to perform named entity recognition on drug and adverse reaction entities, and the F1 value reaches 0.85. When constructing the ontology, you can refer to the adverse drug reaction ontologies such as CIDP and SIDER, and use medical concept dictionaries such as MeSH and SNOMEDCT to semantically annotate entities and relationships, and build a knowledge graph containing 50,000 entities and 100,000 relationships. When embedding the knowledge graph, you can use the TransE algorithm to learn the 100-dimensional vector embedding representation of 50,000 entities and 100,000 relationships, and use the Markov Chain Monte Carlo (MCMC) method for negative sampling, set the negative sampling rate to 5, and train for 1,000 epochs to obtain low-dimensional semantic representations of entities and relationships. When multi-hop reasoning, you can use the MetaPath2Vec algorithm to generate a random walk sequence based on the meta-path of "patient-drug-adverse reaction-biological mechanism", and use the Skip-gram model for node embedding representation learning to obtain the correlation measurement of different drugs and adverse reactions at the biological mechanism level. When predicting adverse drug reactions, a graph convolutional network model with 5 graph convolution layers and 2 fully connected layers can be constructed, with drug and patient attribute features as input and whether adverse reactions occur as output. The 5-fold cross-validation method is used to train and test the model on 50,000 patient data, and the average AUC value and average accuracy of the model are 0.92 and 0.85, respectively. Through the knowledge graph visualization tool based on the Neo4j graph database, multi-dimensional exploration and analysis of the relationship between entities such as drugs, adverse reactions, patient attributes, and biological mechanisms can be achieved.
[0029] Step S104, obtaining the attribute information of individual patients, including region, age, gender, genetic background, past medical history, and concomitant medications. By performing semantic matching with the corresponding nodes in the knowledge graph, the similarity between the patient and known adverse reaction cases is determined, and the adverse reaction risk level is estimated.
[0030] The patient attribute information is structured and semantically annotated to form a patient attribute feature set centered on the patient; the patient attribute feature set is mapped to the adverse drug reaction knowledge graph, and semantic matching is performed with the patient nodes in the adverse drug reaction knowledge graph, and the similarity between the patient and the existing patient nodes in the adverse drug reaction knowledge graph is determined by feature similarity calculation to obtain a patient group similar to the patient; by counting the association strength of the patient group similar to the patient with different drugs and different adverse reactions in the adverse reaction knowledge graph, a triple association network with patient-drug-adverse reaction as the main body is formed, and the patient group similar to the patient is obtained by using Apriori association rule mining algorithm is used to discover frequent association patterns in the patient group; based on the mined frequent association patterns, a drug adverse reaction risk assessment model is constructed in combination with the severity of the adverse drug reaction and the patient's own risk factors, and a risk prediction model suitable for the patient group is trained; using the trained risk prediction model, the patient's attribute characteristics are predicted to obtain the risk probability distribution of different adverse reactions to different drugs in the patient, and the drugs are sorted and graded according to the risk probability to form a personalized drug adverse reaction risk report; the patient's personalized risk report is fed back to the doctor, and decision support is assisted based on the knowledge graph.
[0031] Specifically, detailed attribute information of patients, including region, age, gender, genetic background, past medical history, combined medication, etc., can be obtained from multi-source heterogeneous data such as hospital information systems, electronic medical records, and consultation records, and this information can be structured and semantically annotated to form a patient-centered attribute feature set. The patient attribute feature set is mapped to the knowledge graph of adverse drug reactions, and semantic matching is performed with the patient nodes in the graph. The similarity between the patient and the existing patient nodes in the knowledge graph is determined by the feature similarity calculation method, so as to obtain a patient group similar to the patient. Common feature similarity calculation methods include Euclidean distance, Manhattan distance, jaccard coefficient, cosine similarity, TF-IDF, Word2Vec, etc. Different similarity measurement methods can be used for different types of features. In addition, the metric learning method can also be considered to minimize the distance of similar samples and maximize the distance of heterogeneous samples by optimizing the similarity function. For the target patients and their similar patient groups, drug-ADR association patterns are mined in a larger scope. Frequent drug combination-adverse reaction association patterns are discovered through association rule mining algorithms such as Apriori, FP-Growth, Eclat, H-mine, etc., and more abundant association pattern types are mined by combining methods such as temporal association rules and fuzzy association rules. When applying association rules, the rules can be weighted in combination with patient similarity. The more similar the group is to the target patient, the higher the weight of the association pattern. According to the mined association patterns, combined with the severity of adverse drug reactions and the patient's own risk factors, a personalized adverse drug reaction risk assessment model is constructed. Common risk assessment models include logistic regression (LR), support vector machine (SVM), decision tree (DT), etc. The performance of the model is improved through feature engineering and parameter tuning. In view of the sample imbalance problem of adverse drug reaction data, data balancing techniques such as oversampling, undersampling, SMOTE, as well as cost-sensitive learning, FocalLoss and other methods can be used to improve the model's prediction ability for minority samples. In addition, the time series information of adverse drug reactions can be added to build a time series classification or survival analysis model to capture the dynamic evolution characteristics of adverse reactions. Using the trained risk assessment model, the attribute characteristics of the target patient are predicted to obtain the risk probability distribution of the patient's adverse reactions to different drugs, and the drugs are sorted and graded according to the risk probability to form a personalized adverse drug reaction risk report. The patient's personalized risk report is fed back to the doctor, and auxiliary decision support is provided based on the knowledge graph of adverse drug reactions. Using rule-based reasoning (such as horn clauses) and embedding-based reasoning (such as TransE) and other methods, combined with association rule mining and representation learning, multi-step reasoning is performed from drug indications to candidate drugs and then to drug risks.For a given indication, a group of candidate drugs is first matched from the knowledge graph, and then inappropriate drug combinations are filtered out using rules such as drug interactions and incompatibility contraindications. The risk level of the drug combination is estimated using information such as the correlation strength and probability of adverse drug reactions, and personalized drug recommendations and risk warnings are given in combination with factors such as the patient's own contraindications and medication history. The patient's medication feedback information is continuously tracked, and the occurrence of adverse reactions is collected, and it is used as new sample data to be included in the dynamic training of the personalized risk assessment model. Using active learning, the samples that have the greatest improvement on the model are selected for annotation, and the model performance is continuously optimized; using incremental learning, the parameters of the personalized risk model are dynamically adjusted to adapt to changes in the patient's own status. At the same time, considering the dynamic evolution of the knowledge graph, entities and relationships such as drugs, indications, and adverse reactions are regularly updated to keep the knowledge graph synchronized with the actual situation, thereby forming a closed-loop system of dynamic incremental learning and continuously improving the intelligent level of adverse drug reaction risk prediction and warning. For a target patient, the patient's medical records, diagnosis information, prescription information, etc. in the past year are extracted from the hospital HIS system. After data cleaning and structural processing, a patient feature vector containing 120 attribute fields is formed. The patient feature vector is semantically encoded using a pre-trained medical domain word vector model, such as ClinicalBERT, to obtain a 768-dimensional patient semantic vector representation. In the adverse drug reaction knowledge graph, which contains 10,000 patient nodes and 50,000 adverse drug reaction associated edges, the Faiss vector index library is used to perform K nearest neighbors, K=100, on the semantic vectors of the patient nodes to quickly retrieve and obtain the 100 patients most similar to the target patient as the similar patient group. On the 10,000 adverse drug reaction associated data of the similar patient group, the FP-Growth algorithm is used to mine frequent itemsets and association rules, setting the minimum support to 0.01 and the minimum confidence to 0.5, and mining 50 association rules of drug combinations leading to adverse reactions. According to the confidence and lift of association rules, the rules are sorted and screened, and the rules are weighted according to the patient similarity, and the top-10 association rules with the highest weights are obtained as the risk patterns of focus. A personalized risk assessment model is constructed using the logistic regression algorithm, with the attribute characteristics of the target patient and similar patients as input, and the occurrence of adverse drug reactions as 0 / 1, 0 means no adverse drug reactions, 1 means adverse drug reactions as output. The imbalance problem is handled by L1 regularization and SMOTE oversampling. After 10-fold cross validation, a risk prediction model with an average AUC of 0.85 is obtained. The model is used to predict the medication risk of the target patient, and the probability of adverse reactions to different drugs, such as Stevens-Johnson syndrome, is obtained. High-risk drugs are marked according to the probability threshold, such as 0.3.Combined with the rules in the knowledge graph of drug indications, ingredients, contraindications, adverse reactions, etc., high-risk drugs are inferred and interpreted. For example, cephalosporins belong to penicillin antibiotics and are contraindicated due to the patient's history of penicillin allergy. The use of non-steroidal anti-inflammatory drugs may induce gastric bleeding, which requires risk warnings and regular reviews. A personalized medication risk warning report is formed and pushed to clinicians for reference.
[0032] Step S105, based on the attribute characteristics of patients with high risk levels and the corresponding associated information in the knowledge graph, generates personalized medication recommendations and monitoring plans, including adjusting drug dosages, avoiding specific drug combinations, and strengthening monitoring of specific adverse reactions to reduce the probability and severity of adverse reactions.
[0033] Through the adverse drug reaction knowledge graph, the known adverse reaction phenotypes associated with the drug are obtained; according to the physiological indicator data collected at each time point in the patient's dynamic data stream, the abnormal characteristics of the physiological indicators are extracted, and according to the physiological indicators involved in the known adverse reaction phenotypes in the adverse drug reaction knowledge graph, it is judged whether the patient has abnormalities in the corresponding physiological indicators; for the self-reported symptom data collected in the patient's dynamic data stream, according to the clinical manifestations involved in the known adverse reaction phenotypes in the adverse drug reaction knowledge graph, it is judged whether the patient has corresponding symptoms; for the test result data collected in the patient's dynamic data stream, the test result is matched with the clinical manifestation in the adverse drug reaction knowledge graph to judge whether the patient has corresponding test abnormalities; according to the abnormal physiological indicators, abnormal symptom manifestations and test abnormalities of the patient within a certain time window, it is matched with the adverse reaction phenotype in the adverse drug reaction knowledge graph to obtain the mapping results of different adverse drug reactions occurring in the patient.
[0034] Specifically, the patient's risk level is determined based on the patient's attribute characteristics and the output results of the adverse drug reaction risk prediction model. According to the precision-recall curve of the model on the validation set, the risk probability threshold corresponding to the optimal balance point, such as 0.6, can be selected to trigger the generation process of personalized medication recommendations and monitoring plans for high-risk patients. From the knowledge graph of adverse drug reactions, the associated information such as drug contraindications, high-risk populations, dangerous drug combinations, common adverse reactions, etc. related to the patient's attribute characteristics is retrieved to form a personalized risk factor list for the patient, which may include drug allergy history, specific gene mutations, liver and kidney function abnormalities, etc., and the associated information is quickly retrieved and matched through methods such as knowledge graph embedding representation. For the drugs currently used by the patient, the drug-drug interaction rules in the knowledge graph, such as QT interval prolongation risk and drug-adverse reaction risk scoring rules, such as hepatotoxicity risk, are used to determine whether there are high-risk drugs or high-risk drug combinations, formally expressed as first-order logic rules, and suggestions for alternative drugs or dosage adjustments are given. According to the patient's personalized risk factors, the corresponding medication principles are matched from the knowledge graph, such as "avoid drugs containing X ingredients", "the dosage of Y drugs does not exceed Zmg / day" and monitoring program templates, such as "monthly liver function tests" and "see a doctor in time when symptom A appears", and combined with the patient's specific situation to form personalized medication recommendations and monitoring program texts. Using the knowledge graph-based medical order text generation technology, the semantic vectors of entities and relationships are first obtained through knowledge graph embedding, and then natural language text is generated through sequence-to-sequence models, such as Seq2Seq+Attention or pre-trained language models, such as GPT-2 and T5. The diversity and accuracy of the generation are controlled by methods such as BeamSearch and Top-kSampling. The structured information such as patient attributes, personalized risk factors, medication recommendations and monitoring programs is used as input to generate fluent and easy-to-understand personalized medical order texts. The generation model can also be optimized based on the doctor's feedback through reinforcement learning methods. Send the generated personalized medical advice text to clinicians for review and modification, design a friendly medical advice editing interface, allow doctors to modify and evaluate the generated medical advice, and feed it back to the algorithm model for parameter fine-tuning. The algorithm can also use active learning to autonomously select the medical advice samples that need the most feedback from doctors to improve the efficiency of optimization. According to the feedback from doctors, dynamically adjust the parameters and architecture of the generated model to continuously improve the quality and accuracy of personalized medical advice. Regularly collect feedback information such as the treatment effect and adverse reaction of high-risk patients, such as updating each patient once a week or month, or updating immediately when a new adverse reaction event occurs in the patient. Through time series analysis, change point detection and other methods, automatically identify significant changes in the patient's status and start the update process. At the same time, the dynamic evolution of the knowledge graph itself should also be considered, such as the emergence of new drugs and new side effects, and the entities and relationships of the graph should be updated regularly.Dynamically update the patient's risk level and personalized medication plan to form a closed-loop intelligent medication optimization process, and continuously reduce the probability and severity of adverse drug reactions. For a 65-year-old male patient with a history of hypertension and type 2 diabetes, the adverse drug reaction risk prediction model shows that his risk of serious adverse reactions is 0.75, which exceeds the preset threshold of 0.6 and is judged as a high-risk patient. Retrieve information related to attributes such as "65 years old", "male", "hypertension", and "type 2 diabetes" in the knowledge graph, and find that such patients are prohibited from taking non-steroidal anti-inflammatory drugs, sulfonylurea hypoglycemic drugs, etc., and are prone to adverse reactions such as gastrointestinal bleeding and hypoglycemia, forming a personalized risk factor list for patients. According to the patient's current list of drugs, the drug interaction rule "diuretics + glucocorticoids → hypokalemia risk increased by 30%" in the knowledge graph, as well as the adverse reaction rule "acetaminophen daily dose > 4g → liver damage risk increased by 2 times", found that the patient has potential high-risk drug combinations and overdose medication, and gave suggestions such as "reduce acetaminophen to 3g / day and monitor blood potassium regularly". The medication principle for diabetic patients "giving priority to metformin and avoiding meglitinides" and the monitoring plan for cardiovascular disease patients "following up blood pressure and heart rate changes monthly, and performing dynamic electrocardiogram examinations when necessary" were matched from the knowledge graph. Combined with the patient's condition, the Seq2Seq+Attention model was used to generate personalized medication recommendations: "Metformin 1.5g / d is recommended, liver and kidney function is tested regularly; anti-inflammatory and analgesic drugs such as ibuprofen are suspended; follow up monthly and monitor changes in blood pressure and heart rate." After receiving the generated medical advice, the doctor modified "regular liver and kidney function testing" to "liver function testing every 2 weeks and kidney function testing every 4 weeks" based on the patient's latest examination results and changes in the condition, and evaluated the overall rationality of the medical advice as 4 / 5 points. According to the doctor's feedback, the Attention weight of the Seq2Seq model was adjusted through the reinforcement learning algorithm to improve the coverage of key information. One month later, the patient reported "occasionally mild dizziness, but no other discomfort" during a follow-up visit. Based on the patient's complaints, physical signs and test results, and taking into account individual differences and daily fluctuations in blood sugar data, a dynamic threshold can be set, such as the lowest blood sugar value in the individual's history plus a certain safety range; when the real-time monitored blood sugar value is lower than the dynamic threshold, it indicates that the risk of hypoglycemia is significantly increased; when it is judged that the risk of hypoglycemia is significantly increased, the personalized medication plan is automatically updated to "recommend reducing metformin to 1.0g / d and adding repaglinide 25mg / d; increasing the frequency of self-blood sugar monitoring; and replenishing sugar in time when hypoglycemia occurs." Compared with the initial plan, the patient's risk of severe hypoglycemia was reduced by 65%.
[0035] Step S106, during the patient's medication process, dynamically collect the patient's physiological indicators, symptom reports, and real-time data of test results, and compare them with the adverse reaction phenotypes in the knowledge graph in real time to achieve early identification and early warning of adverse drug reactions. If suspected adverse reaction signals are found, notify the doctor to intervene and deal with them.
[0036] Acquire adverse drug reaction cases verified by clinicians, extract structured adverse drug reaction triplets based on patient attributes, drug use information, and adverse reaction manifestations, and update the adverse drug reaction knowledge graph; obtain the empirical knowledge summarized by clinicians in the process of diagnosing and treating adverse drug reactions, including that certain drugs are more likely to cause adverse reactions in specific populations, and enrich the semantic representation of the drugs, adverse reaction phenotypes, population characteristic entities and relationships by learning low-dimensional vector representations of empirical knowledge; retrain the adverse drug reaction risk prediction model based on the updated adverse drug reaction knowledge graph to enable it to capture the interaction patterns between drugs, populations, and adverse reactions; embed the optimized adverse drug reaction knowledge graph into the clinical decision support system. When doctors formulate medication plans for patients, they retrieve historical adverse reaction cases that are similar to patient attributes and related to the drugs to be used, indicate potential risks, and give alternative medication recommendations to assist doctors in making personalized treatment decisions.
[0037] Specifically, according to the entity types such as drugs, indications, adverse reactions, clinical manifestations, and relationship types such as drugs-adverse reactions and adverse reactions-clinical manifestations in the knowledge graph of adverse drug reactions, based on technologies such as ontology mapping and natural language processing, structured adverse reaction knowledge is extracted from heterogeneous data sources such as drug instructions, adverse reaction case reports, and medical literature to form a high-quality and high-coverage knowledge base of adverse drug reactions. For each patient taking medication, with the help of wearable devices, mobile applications and other means, the physiological indicators after medication, such as heart rate, blood pressure, body temperature, etc., self-reported symptoms such as rash, vomiting, dizziness, etc., and regular test results such as blood routine, urine routine, liver and kidney function, etc., are collected in real time, and pre-processing such as data cleaning, noise removal, missing value interpolation, and semantic standardization are performed to construct a patient dynamic data stream indexed by time. The patient's drug use information, such as drug name, dosage, frequency, course of treatment, etc., is associated with its dynamic data stream to form a high-dimensional tensor form of "patient-drug-time-feature" time series data as input for subsequent adverse reaction identification and early warning. Based on the knowledge graph of adverse drug reactions, the patient's dynamic data stream is extracted and semantically mapped in real time. First, the word vector model in the medical field is trained based on the corpus, such as Word2Vec, Glove, BERT, etc., to obtain the distributed semantic representation of words; then, the word vector is mapped to the standardized medical concept space using the semantic types and concept relationships of medical ontologies such as UMLS and SNOMEDCT, so as to achieve the semantic matching of the patient's dynamic features to the adverse reaction phenotype in the knowledge graph; in the case of ambiguity in semantic matching, context information, part-of-speech tagging, etc. can also be used for disambiguation. Disambiguation is to determine the exact meaning of a polysemous word in a specific context through a specific method, so as to automatically associate the abnormal physiological indicators, symptom manifestations, abnormal test examinations and other information collected by the patient at different time points with the adverse reaction phenotype in the knowledge graph. Comprehensive factors such as drugs, dosages, and stages of the disease are used to judge the risk probability of a patient having a certain adverse drug reaction phenotype within a given time window, such as 1 day, 3 days, and 7 days. You can choose statistical process control methods such as Shewhart control charts and EWMA charts to project the patient's multidimensional time series data into a low-dimensional subspace for modeling and anomaly detection; you can also use distance-based methods such as DTW dynamic time warping distance and t-SNE random neighborhood embedding, and density-based methods such as LOF local anomaly factor; you can also try to use deep learning models such as LSTM, GRU, and Attention to auto-encode and reconstruct the time series and calculate the reconstructed residual as the anomaly score. When the risk probability exceeds the preset threshold, a warning signal for adverse drug reactions is generated and pushed to clinicians for identification.Based on the patient's dynamic data and warning signals, combined with the patient's symptoms, signs, auxiliary examinations and other clinical clues, clinicians use knowledge graphs to assist in diagnosing whether the patient has actually had an adverse drug reaction, identify possible allergic drugs and severity, and formulate corresponding treatment plans such as discontinuation, reduction, and substitution. Feedback the adverse drug reaction cases diagnosed and handled by doctors to the knowledge graph, and optimize it through the model layer and representation layer: the model layer uses knowledge fusion technologies such as rules, clustering, and embedding to continuously mine and supplement new adverse drug reaction entities, relations, and attributes from the doctor's feedback data to expand the coverage of the knowledge graph; the representation layer uses representation learning models based on translation, such as TransE and TransR, semantic matching, such as RESCAL and DistMult, and graph neural networks, such as GCN and GAT, to learn low-dimensional dense vector representations of entities and relations from structures such as adverse drug reaction bipartite graphs and multi-relation heterogeneous graphs to improve the performance of semantic matching and reasoning. Finally, a dynamically growing and continuously optimized knowledge graph of adverse drug reactions is formed and applied to clinical decision support, providing doctors with personalized medication optimization suggestions and improving the accuracy and timeliness of early warning. From heterogeneous data sources such as the World Wide Web, scientific literature databases, and the FDA adverse event reporting system, the instructions of 2,000 commonly used drugs, 1,000 adverse drug reaction research papers, and 5,000 adverse reaction case reports are crawled, and the BERT-BiLSTM-CRF named entity recognition model is used to extract the entities such as drugs, indications, adverse reactions, and clinical manifestations, with an F1 value of 91.2%; then the BiLSTM-Attention relationship extraction model is used to identify the association between entities, with an F1 value of 87.5%. For a given patient, physiological signals such as heart rate, blood pressure, and body surface temperature are collected once every 5 minutes through a wearable smart bracelet, and denoised by a genetic algorithm-optimized Kalman filter, reducing the noise level by 75%. A questionnaire on drug adverse reaction-related symptoms is pushed out regularly every day, and self-reports from patients are collected and analyzed for sentiment tendency using the Bo-W model, with an accuracy of 82.3%. The blood routine, urine routine, liver and kidney function test results of patients are selected every week, and abbreviations and synonyms are standardized and mapped to construct a high-dimensional sparse patient-drug-time-feature tensor with a span of 90 days and 1.2 million records. Based on Word2Vec word vectors and RadLex ontology, the patient's symptom description, such as "palpitations", is mapped to standard adverse reaction concepts, such as "tachycardia", with an ambiguity resolution accuracy of 93.1%. The stagewise-LSTM model is used to predict the risk of drug adverse reactions for the patient's multidimensional time series data, and the ROC-AUC value within a 5-day time window reaches 0.927, exceeding the traditional Shewhart control chart method by 23.5%.For a patient who took haloperidol to treat schizophrenia, he had symptoms such as drowsiness, muscle rigidity, and tremor for three consecutive days. The phenotypic similarity with the extrapyramidal reaction caused by haloperidol in the knowledge graph was 0.852, triggering an early warning signal. The doctor asked about the medical history and checked the patient's blood drug concentration test values. Excluding other causes such as infection, the severity of the adverse reaction was determined to be moderate. The haloperidol dose was reduced by 20% and diphenhydramine hydrochloride was added for symptomatic treatment. After one week, the patient's symptoms improved. The case was automatically labeled as a drug-adverse reaction triple (haloperidol, cause, extrapyramidal reaction). The embedding representation of the relevant entities in the knowledge graph was updated through the attention mechanism learning, and the overall accuracy of the adverse reaction prediction model was improved by 1.7%.
[0038] Step S107: For newly identified adverse drug reaction cases, extract information including clinical manifestations, occurrence time, and severity. By comparing and analyzing with existing cases in the knowledge graph, determine whether the newly identified adverse drug reaction cases are new adverse reaction phenotypes or special population effects, and feed the judgment results back to the knowledge graph to update the adverse drug reaction knowledge graph.
[0039] Based on each identified adverse drug reaction case, the patient's demographic characteristics, drug name and usage and dosage, adverse reaction onset and end time, clinical manifestation description, and laboratory test index elements are extracted from the case to obtain a structured representation in the form of triples; based on the extracted adverse reaction onset and end time, the duration of the adverse reaction is calculated; based on the extracted adverse reaction severity, measures taken, and prognosis, the severity of the case is graded to obtain a structured feature representation of a single case; the structured representation of the case is mapped to the existing adverse drug reaction ontology in the knowledge graph, and its correlation with the existing adverse drug reaction ontology is calculated. Semantic similarity of reaction types; judging whether the case is a new adverse reaction phenotype based on the comparison result of the semantic similarity with a preset threshold; for cases determined to be known adverse reaction phenotypes, comparing with the demographic characteristics, drug usage and dosage, concomitant medication, and underlying disease attributes of existing cases under the phenotype, judging whether the case is a special population effect case based on the difference between the case and the existing cases in a certain attribute dimension; based on the judgment result, adding the structured representation of the new phenotype case or the special population effect case to the knowledge graph in the form of a new node or a new relationship, and updating the topological structure and semantic space distribution of the graph.
[0040] Specifically, for each case of adverse drug reaction identified, named entity recognition and relation extraction technologies based on deep learning, such as sequence annotation models such as BiLSTM-CRF and BERT-CRF and end-to-end relation extraction models based on attention mechanisms, are used to extract key elements such as patient demographic characteristics, drug names and usage and dosage, adverse reaction start and end time, clinical manifestation description, and laboratory test indicators from structured forms of cases, such as adverse event report forms, and unstructured texts, such as medical records and patient complaints, and are structured in the form of triples (entity 1, relationship, entity 2). The named entity recognition and relation extraction models used need to be trained on large-scale medical text corpora, which come from medical textbooks, guidelines, literature, etc., and clinical experts should annotate entities and relationships. According to the start and end time of the adverse reaction, its duration is calculated; according to the clinical manifestations, laboratory tests, measures taken, and prognosis of the adverse reaction, the severity of the case is graded, such as level 1-5, with reference to the commonly used grading standards such as CTCAE and WHO, to obtain a comprehensive and multi-dimensional structured feature representation of a single case. The Wu-Palmer semantic similarity algorithm based on WordNet and the cosine similarity semantic similarity algorithm based on Word2Vec are used to match the structured representation of the case with the existing adverse drug reaction ontology in the knowledge graph, and calculate its semantic similarity with the existing adverse reaction type. If the similarity is lower than the preset threshold, such as 0.5, it is determined to be a new adverse reaction phenotype, otherwise it is classified as a new case of a known adverse reaction phenotype. For cases classified as known adverse reaction phenotypes, the patient's demographic characteristics, including age, gender, race, physiological and pathological characteristics, including liver and kidney function, comorbidities, medication history, genetic characteristics, including drug metabolizing enzyme genotype, target polymorphism, and other dimensions are further compared with existing cases under this phenotype. Distribution deviation test or interaction test statistical methods are used to determine whether new cases are significantly different from known cases in different feature dimensions. For example, if the distribution of population characteristics deviates by more than 3 standard deviations, it is determined whether it is a special population effect case. For cases determined to be new phenotypes or special population effects, refer to ontology evolution methods such as IMHO and MELT, add their structured representations to the knowledge graph in the form of new nodes (new phenotypes) or new relationships (special population effects), and update the topological structure of the knowledge graph by considering constraints such as consistency and redundancy of the ontology. Based on the TransE or TransR knowledge graph representation learning model, retrain the low-dimensional dense vector representation of all entities and relationships to optimize their distribution position in the semantic space. When the knowledge graph is expanded and updated, new adverse drug reaction cases are used as incremental data to retrain or fine-tune the existing multi-classification prediction model based on random forests. The input features include a combination of patient demographics, physiological and pathological characteristics, genetic characteristics and drug characteristics, and the output is the type of adverse reaction.Through grid search and other methods, the hyperparameters of random forests, such as the number of trees and maximum depth, are optimized, and the impact of the addition of new samples on model performance and resource consumption is weighed to improve its predictive ability for new phenotypes and new populations. A multi-dimensional feedback evaluation mechanism can be used to regularly count the number of newly discovered adverse reaction phenotypes, the types of drugs covered, the number of cases involved, the clinical verification rate and other objective indicators, such as the number of new adverse reaction phenotypes, the types of drugs covered, the number of cases involved, and the clinical verification rate, to evaluate the effect of knowledge graph expansion and update. At the same time, through questionnaires, interviews and other methods, subjective feedback from doctors, patients, pharmaceutical companies and other parties on newly discovered adverse reaction types is collected, and the effectiveness, timeliness and ease of use of knowledge graph expansion are evaluated from the perspectives of clinical application value and drug vigilance influence. According to objective indicators and subjective feedback, the ontology evolution rules, judgment thresholds, etc. are dynamically adjusted and optimized. The newly discovered adverse reaction types are promptly fed back to drug regulatory agencies, pharmaceutical companies, clinicians, etc. through emails, website announcements, etc., to promote the update of drug instructions, safety re-evaluation and rational use of drugs, and form a socialized collaborative mechanism for drug adverse reaction monitoring. For a case of rheumatoid arthritis treated with aspirin, the BiLSTM-CRF model of the bidirectional long short-term memory network (BiLSTM) combined with the conditional random field (CRF) trained on 5,000 medical records was used to extract 60 key entities from the electronic medical record text, including patient age 65 years old, gender female, disease rheumatoid arthritis, drug aspirin, 0.1, dose 0.1, route of administration oral, frequency of administration once a day, adverse reaction gastric ulcer, and occurrence time range 2022-01-01 to 2022-01-14, with an F1 of 91%; The score of 1 is the harmonic mean of precision and recall, which comprehensively reflects the precision of the model in identifying key entities. Precision is used to reduce false positives and recall, and recall is used to reduce missed negatives. The F1 score of 91% means that the model performs well in identifying and extracting key entities, and can accurately find most of the real entities in a large amount of text, and can effectively avoid incorrectly labeling irrelevant text as entities. It was determined that the adverse reaction lasted for 14 days, and the fifth edition of the CTCAE5.0 General Adverse Event Evaluation Criteria reached level 3. The Wu-Palmer semantic similarity calculation was performed on this case and 500 cases under the adverse reaction phenotype of "aspirin-gastric ulcer" in the knowledge graph, with an average score of 0.62, which is higher than the threshold of 0.5, so it is classified as a known adverse reaction. Further comparison found that the age of the patient in this case was 20 years older than the known case concentration trend of 45±8 years old, exceeding 3 times the standard deviation, suggesting that elderly female patients with rheumatoid arthritis may be a high-risk group for aspirin-induced gastric ulcer. Based on this, a special population effect association from "elderly women-rheumatoid arthritis" to "aspirin-gastric ulcer" is added to the knowledge graph, and the TransE model is used to update the embedding representation of the relevant entities and feed it back to the aspirin adverse reaction prediction model.The model was built based on 500 decision trees, integrating 20 features such as age, gender, comorbidities, and drugs. After the new case was included, the AUC was improved from 0.78 to 0.81. The emergency plan for severe adverse reactions was initiated for this case, a drug alert report was submitted to the drug regulatory authorities, and the pharmaceutical company was notified to evaluate the revision of the instructions. This was completed within 1 month, effectively shortening the original 2-3 month response process and improving patient safety.
[0041] Step S108, based on the updated adverse drug reaction knowledge graph, the group distribution rules and individual difference patterns of adverse drug reactions are mined, the common characteristics of high-risk groups are identified through cluster analysis methods, and corresponding prevention and intervention strategies are formulated to achieve the management and personalized treatment of adverse drug reactions.
[0042] From the updated knowledge graph of adverse drug reactions, extract the patient's demographic characteristics, clinical characteristics, molecular biological characteristics, and drug characteristic attribute information to form a patient-drug feature matrix; perform dimensionality reduction processing on the patient-drug feature matrix to obtain the mapping result of the matrix in two-dimensional space or three-dimensional space, and preliminarily judge whether there are group differences and individual outliers based on the distribution of different patient groups in the feature space; perform cluster analysis on the patient-drug feature matrix, divide the patient group into several subgroups, and obtain the characteristic distribution and common characteristics of the members within each subgroup; based on the patient subgroups obtained by clustering, explore the association patterns of different adverse reactions caused by different drugs in different subgroups; based on the common characteristics of patient subgroups, judge whether the patient subgroup is a high-risk group for a certain drug adverse reaction; for the identified high-risk group, formulate a monitoring plan and intervention strategy for drug use, incorporate it into the clinical decision support system, and provide prompts for doctors to carry out personalized prescription design and pharmaceutical management.
[0043] Specifically, from the updated knowledge graph of adverse drug reactions, multidimensional attribute information such as patient demographic characteristics, clinical characteristics, molecular biological characteristics, and drug characteristics are extracted. With reference to medical informatics standards such as OMOPCDM and FHIR, these attributes are semantically mapped and standardized to form a structured feature matrix. Nonlinear dimensionality reduction algorithms such as t-SNE and autoencoder are used to map the high-dimensional patient-drug feature matrix into a two-dimensional or three-dimensional space, and specific visualization techniques such as parallel coordinates and radar charts are combined to display the distribution of different patient groups in the feature space from multiple angles. According to the visualization results, appropriate clustering algorithms such as k-medoids and hierarchical clustering are selected, and the optimal number of clusters is determined by the silhouette coefficient indicator. The patient group is divided into several subgroups, and the characteristic distribution and common characteristics of the members within each subgroup are obtained. For each patient subgroup, the types of drugs, dosages, frequencies, and combination medications used by its members, as well as the types, organ distribution, and severity of adverse drug reactions, were collected. Data mining methods such as statistical tests, association rules, and frequent item sets were used to measure the strength and significance of drug-adverse reaction associations, and to mine the association patterns of different adverse reactions caused by different drugs in different subgroups. The association patterns were interpreted and verified in combination with clinical pharmacology knowledge. For the clustered patient subgroups, based on their common characteristics, genome-wide association analysis (GWAS) and expression quantitative trait loci (eQTL) analysis were used to conduct in-depth analysis from the perspectives of molecular biology such as drug metabolizing enzyme genotypes, target polymorphisms, and drug transporters. With the help of bioinformatics tools such as gene enrichment analysis, such as GO, KEGG, protein interaction network analysis (PPI), and cistrome analysis, common pathways and biomarkers of high-risk populations were explored, possible pharmacogenomic mechanisms were inferred, and verified using knowledge bases and literature mining platforms such as PharmGKB and GLAD4U. Develop targeted medication monitoring plans and intervention strategies for identified high-risk populations. Optimize the dosage and frequency of medication based on the patient's pharmacokinetic / pharmacodynamic (PK / PD) model; use the physiologically based pharmacokinetic (PBPK) model to predict potential drug interaction risks; and conduct companion diagnostic marker testing and interpretation for specific targeted therapies. Integrate individualized medication regimens into clinical decision support systems to prompt doctors to conduct personalized prescription design and pharmaceutical management, and dynamically monitor treatment efficacy and safety.For cases with severe adverse reactions, based on the common characteristics of the subgroups to which they belong, such as patients with specific disease backgrounds, genetic characteristics, and concomitant medications, causal inference tools such as Bradford-Hill Criteria or Naranjo Scale are used to combine epidemiological evidence with biological mechanisms, trace possible causes, improve the credibility of inferences, and use evidence-based medicine to conduct targeted treatment and prognosis assessments for individual cases. New evidence obtained during the diagnosis and treatment of individual cases is promptly fed back to the knowledge graph and clustering model to achieve a precise management closed loop from group to individual, from data to knowledge, and from association to causality. From the knowledge graph containing 1,000 commonly used drugs, 50 million prescription records, 5 million adverse reaction reports, and 10 million electronic medical records of 500,000 patients are extracted, and OMOPCDM is used to standardize information such as drug names, indications, and dosages to form a structured matrix containing 200 dimensional features. The t-SNE algorithm was used to reduce the feature matrix to a 3D space. The 3D scatter plot showed that there were three obvious patient clusters, corresponding to children, adults, and the elderly. The k-medoids clustering algorithm was further used to determine the optimal k value of 10 in combination with the silhouette coefficient, and the patients were divided into 10 subgroups. The aspirin use rate of each subgroup was 30% in the elderly group and the incidence of gastrointestinal bleeding was 1% in the elderly group. The Fisher exact test method was used to calculate the odds ratio and p value, and it was found that the correlation between the two was the highest in the elderly group (odds ratio = 3.2, p < 0.01). The GWAS analysis of this subgroup found that the mutation frequency of a SNP site rs1057910 located on the CYP2C9 gene was significantly higher than that of other subgroups, 40% vs 10%, p < 0.001, suggesting that this gene mutation may be closely related to gastrointestinal bleeding caused by aspirin. For elderly patients with this gene mutation, it is recommended to reduce the dose of aspirin from 100 mg / d to 75 mg / d and perform regular endoscopic examinations. A 65-year-old patient with a CYP2C9 gene mutation who developed severe gastrointestinal bleeding after taking regular doses of aspirin was retrospectively analyzed for evidence such as medication history, genetic test results, and endoscopic results. The Naranjo scale was scored as "very likely" 7 points, supporting the case analysis results. At the same time, the association strength in the knowledge graph was updated accordingly to form a more accurate drug safety model.
[0044] Step S109, combining the biological mechanism information in the adverse drug reaction knowledge graph to analyze the adverse drug reaction cases that have occurred.
[0045] Based on the patient's adverse drug reaction report information, analyze the patient's attributes according to the patient's medical history, genetic background, lifestyle, and history of drug use. Check all the drugs used by the patient to determine whether there are known drug interactions. Use the biological mechanism information in the drug knowledge graph to analyze how the patient's genetic background affects the metabolic pathways of specific drugs or the sensitivity of drug targets. Determine the changes in biological markers of adverse reactions through laboratory tests and imaging test results, and analyze the association between changes in biological markers and the molecular mechanism of drug action. Compare the incidence of adverse drug reactions in specific groups to analyze whether the patient belongs to a high-risk group. Analyze management and intervention records, evaluate the impact of symptomatic treatment on adverse reactions, and determine whether the effects of intervention measures meet expectations.
[0046] For example, a case report may detail a 35-year-old female patient who developed severe liver damage after taking a new chemotherapy drug. The report may include her basic physical measurements, such as height 165cm, weight 60kg, and other relevant clinical information. In the electronic health record system, it may be discovered that the patient has a genetic G6PD deficiency, which is detailed in past medical records. This genetic factor may affect her ability to metabolize certain drugs. Using the drug interaction database may reveal that the patient is taking cimetidine (an acid-suppressing drug) and chemotherapy drugs at the same time, and cimetidine may slow the metabolism of the chemotherapy drug, resulting in increased blood drug concentrations, and a normal dose may become an overdose. By analyzing the drug knowledge graph, it is found that the patient has a polymorphism in the CYP2C19 gene, which makes her less capable of metabolizing certain drugs than ordinary people. The drug concentration at a standard dose may be 5 times the usual in such patients. According to the patient's liver and kidney function test results, the creatinine level is found to have increased from 0mg / dL to 9mg / dL, indicating impaired kidney function. This may mean that the ability to excrete drugs is reduced, resulting in drug accumulation. Laboratory tests showed that the patient's liver enzyme ALT level increased from the normal value of 40U / L to 200U / L, which is one of the biomarkers of drug liver injury. By analyzing the clinical database, it was found that the incidence of adverse liver reactions of this chemotherapy drug in Asian women was 2%, while the incidence in other groups was less than 5%. The case report may record that the patient received N-acetylcysteine (NAC) treatment after liver injury, and liver function improved within one week after treatment, and ALT levels dropped to 100U / L.
[0047] The above is only a preferred embodiment of the present invention. It should be pointed out that ordinary technicians in this technical field can make several improvements and supplements without departing from the principle of the present invention. These improvements and supplements should also be regarded as the scope of protection of the present invention.
Claims
1. A method for monitoring and early warning of adverse drug reactions, characterized in that: The method comprises: Patient groups are classified and labeled according to different geographical, age and gender attributes, and a multi-dimensional patient attribute labeling system is constructed to obtain patients' drug use data, including drug types, dosages, frequencies and courses of treatment. Drug attributes are associated with patient attributes to form a patient-drug attribute data set, and the patient-drug attribute data set is preprocessed and feature engineered, including data cleaning, missing value processing, attribute coding and feature selection. By combining and analyzing the preprocessed and feature-engineered patient-drug attribute data set, association rule mining algorithms and sequence pattern mining algorithms are used to discover the association rules and time series patterns between different combinations of patient attributes and the incidence and manifestation patterns of adverse drug reactions, and to mine the differences in the incidence and manifestation patterns of adverse drug reactions in different patient groups. Based on the attribute data of the patient group, including genetic background, past medical history and combined medication information, a patient attribute data set is constructed, and the patient attribute data set is preprocessed to obtain the preprocessed patient attribute data set. Combined with the difference rules mined, features related to adverse reactions are extracted from the preprocessed patient attribute data set to form an adverse reaction risk feature set; with the adverse reaction risk feature set as input and the adverse reaction occurrence as output, a decision tree or random forest algorithm is used to train the association model between patient attributes and adverse reaction risk. By analyzing the feature importance of the association model, the influence weight of each patient attribute feature on the adverse reaction risk is calculated, and the attributes including the patient's genetic background, previous medical history, and combined medication are sorted according to the weight size to determine the correlation strength with the adverse reaction risk; according to the correlation strength sorting results, the top N patient attribute features with the strongest correlation are selected as the key indicators for adverse reaction risk prediction, and a risk prediction model is constructed; Based on the attribute features that are strongly associated with the risk of adverse reactions, a knowledge graph of adverse drug reactions is constructed. Multi-source heterogeneous data including patient attributes, drug information, adverse reaction phenotypes, and biological mechanisms are semantically associated and fused. Named entity recognition methods and relationship extraction methods based on pattern matching, dependency analysis, or deep learning are used to identify words and phrases belonging to defined entity types and relationship types from multi-source data. By matching and associating with the defined entity types and relationship types in the knowledge graph, semantic mapping and association between multi-source heterogeneous data and the knowledge graph are achieved. The association between different factors is revealed through multi-hop reasoning and path analysis of the knowledge graph of adverse drug reactions. The knowledge graph is used to represent Learning methods embed entities and relationships in knowledge graphs into low-dimensional vector spaces, so that semantically similar entities and relationships are close in the vector space, and semantically different entities and relationships are far away in the vector space, so that the semantic association strength between entities can be measured and predicted through vector operations; multi-hop reasoning algorithms based on random walks are applied to knowledge graphs, and node sequences are generated by random walks on the graph to capture the structural similarity between nodes, thereby discovering implicit associations between patient attributes, drug information, adverse reaction phenotypes, and biological mechanism factors; rule-based or path-based reasoning methods are used to infer unknown entities or relationships using explicit rules or path patterns in knowledge graphs; Obtain the attribute information of individual patients, including region, age, gender, genetic background, past medical history, and concomitant medications. By semantically matching the corresponding nodes in the knowledge graph, determine the similarity between the patient and known adverse reaction cases and estimate the adverse reaction risk level. Generate personalized medication recommendations and monitoring plans based on the attributes of high-risk patients and the corresponding association information in the knowledge graph, including adjusting drug dosages, avoiding specific drug combinations, and strengthening monitoring of specific adverse reactions to reduce the probability and severity of adverse reactions; During the medication process, the patient's physiological indicators, symptom reports, and real-time data of test results are dynamically collected, and compared with the adverse reaction phenotypes in the knowledge graph in real time to achieve early identification and early warning of adverse drug reactions. If suspected adverse reaction signals are found, the doctor is notified to intervene and deal with them; For newly identified adverse drug reaction cases, extract information including clinical manifestations, occurrence time, and severity, and determine whether the newly identified adverse drug reaction cases are new adverse reaction phenotypes or special population effects by comparing and analyzing with existing cases in the knowledge graph. Feedback the judgment results to the knowledge graph to update the adverse drug reaction knowledge graph. Based on the updated knowledge graph of adverse drug reactions, the group distribution patterns and individual difference patterns of adverse drug reactions are explored, and the common characteristics of high-risk groups are identified through cluster analysis methods. Prevention and intervention strategies are then formulated accordingly to achieve management and personalized treatment of adverse drug reactions. Combined with the biological mechanism information in the adverse drug reaction knowledge graph, analyze the adverse drug reaction cases that have occurred.
2. The method according to claim 1, wherein: The method of obtaining the attribute information of individual patients, including region, age, gender, genetic background, past medical history, and concomitant medication, determines the similarity between the patient and known adverse reaction cases by semantically matching with corresponding nodes in the knowledge graph, and estimates the adverse reaction risk level, including: Structural representation and semantic annotation of patient attribute information to form a patient-centered patient attribute feature set; Mapping the patient attribute feature set to the adverse drug reaction knowledge graph, performing semantic matching with the patient node in the adverse drug reaction knowledge graph, and determining the similarity between the patient and the existing patient nodes in the adverse drug reaction knowledge graph by feature similarity calculation to obtain a patient group similar to the patient; By counting the association strengths of patient groups similar to the patient with different drugs and different adverse reactions in the adverse reaction knowledge graph, a triple association network with patient-drug-adverse reaction as the main body is formed, and the Apriori association rule mining algorithm is used to discover frequent association patterns in the patient group; Based on the frequently mined association patterns, the adverse drug reaction risk assessment model is constructed in combination with the severity of adverse drug reactions and the patient's own risk factors, and a risk prediction model suitable for the patient group is trained; The trained risk prediction model is used to predict the patient's attribute characteristics, obtain the risk probability distribution of different adverse reactions to different drugs, and sort and grade the drugs according to the risk probability to form a personalized adverse drug reaction risk report; The patient's personalized risk report is fed back to the doctor, and decision support is assisted based on the knowledge graph.
3. The method according to claim 1, wherein: The method generates personalized medication recommendations and monitoring plans based on the attribute characteristics of patients with high risk levels and the corresponding associated information in the knowledge graph, including adjusting drug dosages, avoiding specific drug combinations, and strengthening monitoring of specific adverse reactions to reduce the probability and severity of adverse reactions, including: Obtaining known adverse reaction phenotypes associated with the drug through the drug adverse reaction knowledge graph; Extracting abnormal features of physiological indicators based on the physiological indicator data collected at each time point in the patient's dynamic data stream, and judging whether the patient has abnormalities in corresponding physiological indicators based on the physiological indicators involved in the known adverse reaction phenotypes in the adverse drug reaction knowledge graph; For the self-reported symptom data collected in the patient's dynamic data stream, judging whether the patient has corresponding symptoms according to the clinical manifestations involved in the known adverse reaction phenotypes in the adverse drug reaction knowledge graph; For the test result data collected in the patient's dynamic data stream, the test result is matched with the clinical manifestations in the adverse drug reaction knowledge graph to determine whether the patient has a corresponding test abnormality; According to the abnormal physiological indicators, abnormal symptom manifestations and abnormal test examinations of the patient within a certain time window, they are matched with the adverse reaction phenotypes in the adverse drug reaction knowledge graph to obtain the mapping results of different adverse drug reactions occurring in the patient.
4. The method according to claim 1, wherein: During the course of medication, the patient's physiological indicators, symptom reports, and real-time data of test results are dynamically collected, and compared with the adverse reaction phenotypes in the knowledge graph in real time to achieve early identification and early warning of adverse drug reactions. If a suspected adverse reaction signal is found, the doctor is notified to intervene and deal with it, including: Obtain adverse drug reaction cases verified by clinicians, extract structured adverse drug reaction triples based on patient attributes, drug use information, and adverse reaction manifestations, and update the adverse drug reaction knowledge graph; Acquire the empirical knowledge summarized by clinicians in the process of diagnosing and treating adverse drug reactions, including that certain drugs are more likely to cause adverse reactions in specific populations, and enrich the semantic representation of the drugs, adverse reaction phenotypes, population characteristic entities and relationships by learning low-dimensional vector representations of empirical knowledge; Based on the updated knowledge graph of adverse drug reactions, the adverse drug reaction risk prediction model is retrained to enable it to capture the interaction patterns between drugs, populations, and adverse reactions; The optimized adverse drug reaction knowledge graph is embedded in the clinical decision support system. When doctors formulate medication plans for patients, historical adverse reaction cases similar to the patient's attributes and related to the drugs to be used are retrieved, potential risks are indicated, and alternative medication recommendations are given to assist doctors in making personalized treatment decisions.
5. The method according to claim 1, wherein: For the newly identified adverse drug reaction cases, extract the clinical manifestations, occurrence time, and severity information, and judge whether the newly identified adverse drug reaction cases are new adverse reaction phenotypes or special population effects by comparing and analyzing with the existing cases in the knowledge graph, and feed the judgment results back to the knowledge graph to update the adverse drug reaction knowledge graph, including: Based on each identified adverse drug reaction case, the patient's demographic characteristics, drug name and usage and dosage, adverse reaction onset and end time, clinical manifestation description, and laboratory test index elements were extracted from the case to obtain a structured representation in the form of triples; According to the extracted start and end time of the adverse reaction, the duration of the adverse reaction is calculated; According to the extracted adverse reaction severity, measures taken and outcomes, the severity of the cases was graded to obtain a structured feature representation of a single case; Map the structured representation of the case with the existing adverse drug reaction ontology in the knowledge graph, and calculate its semantic similarity with the existing adverse reaction types; Judging whether the case is a new adverse reaction phenotype according to a comparison result between the semantic similarity and a preset threshold; For cases determined to have known adverse reaction phenotypes, the demographic characteristics, drug usage and dosage, concomitant medications, and underlying disease attributes of existing cases under this phenotype are compared, and whether the case is a special population effect case is determined based on the difference between the case and the existing cases in a certain attribute dimension; According to the judgment results, the structured representation of new phenotypic cases or special population effect cases is added to the knowledge graph in the form of new nodes or new relationships, and the topological structure and semantic space distribution of the graph are updated.
6. The method according to claim 1, wherein: Based on the updated knowledge graph of adverse drug reactions, the group distribution patterns and individual difference patterns of adverse drug reactions are mined, the common characteristics of high-risk groups are identified through cluster analysis methods, and corresponding prevention and intervention strategies are formulated to achieve the management and personalized treatment of adverse drug reactions, including: From the updated adverse drug reaction knowledge graph, extract patient demographic characteristics, clinical characteristics, molecular biological characteristics and drug characteristic attribute information to form a patient-drug characteristic matrix; Perform dimensionality reduction on the patient-drug feature matrix to obtain the mapping result of the matrix in two-dimensional or three-dimensional space. According to the distribution of different patient groups in the feature space, preliminarily judge whether there are group differences and individual outliers. Performing cluster analysis on the patient-drug feature matrix, dividing the patient population into several subgroups, and obtaining the feature distribution and common characteristics of the members within each subgroup; Based on the patient subgroups obtained by clustering, the association patterns of different adverse reactions caused by different drugs in different subgroups are explored; Based on the common characteristics of patient subgroups, determine whether the patient subgroup is at high risk of adverse drug reactions; for the identified high-risk population, formulate drug monitoring plans and intervention strategies, incorporate them into the clinical decision support system, and provide doctors with tips for personalized prescription design and pharmaceutical management.
7. The method according to claim 1, wherein: The combination of biological mechanism information in the adverse drug reaction knowledge graph and analysis of adverse drug reaction cases that have occurred includes: Based on the patient's adverse drug reaction report information, analyze the patient's attributes according to the patient's medical history, genetic background, lifestyle habits and drug use history; Review all medications the patient is taking to determine if there are any known drug interactions; Using the biological mechanism information in the drug knowledge graph, we can analyze how the patient's genetic background affects the metabolic pathways of specific drugs or the sensitivity of drug targets; Determine changes in biomarkers of adverse reactions through laboratory tests and imaging test results, and analyze the association between changes in biomarkers and the molecular mechanisms of drug action; Compare the incidence of adverse drug reactions in specific groups to analyze whether the patient belongs to the high-risk group; Analyze management and intervention records, evaluate the impact of symptomatic treatment on adverse reactions, and determine whether the effectiveness of intervention measures meets expectations.
Citation Information
Patent Citations
Drug adverse reaction data mining method based on decision-making tree
CN111883219A
Method for detecting drug sex difference adverse reaction signal
CN112185547A