A method for evaluating real-world adverse drug use in a hospital based on a knowledge base
By constructing a bad drug evaluation model based on the knowledge base, using patient indicators and large text data to extract explicit and implicit features for standardization, the problem of incomplete processing of data for bad drug monitoring in the prior art is solved, and the accurate evaluation and prediction of bad drug use is achieved, especially in the case of combined drug use, the accuracy and prediction ability of monitoring are improved.
Patent Information
- Application Number
- CN202011125971.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-10-20
AI Technical Summary
In the monitoring of adverse drug use in the prior art, the complexity and diversity of real event data lead to insufficient comprehensive data processing and lack of practicality. It is difficult for traditional methods to accurately monitor adverse events caused by combination medication, and it is also low in sensitivity to low-probability incidents.
A adverse drug evaluation model is constructed based on the knowledge base. By obtaining the patient's index data and large text data, the explicit and implicit features are extracted, the machine learning model is trained after standardization processing, and the adverse reaction characteristics are obtained using the drug knowledge base, combined with time series for splicing, and standardized features are formed for evaluation.
Accurate monitoring of adverse medications is achieved, especially in the case of combined medication, which can identify adverse events caused by specific medications, reduce drug screening errors, ensure patient health, and make a certain degree of prediction of the first medication use of new patients.
Smart Images

Figure CN112366002B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer application technologies, and more particularly, to a method for evaluating real-world adverse drug reactions in a hospital based on a knowledge base. Background Art
[0002] According to statistics, about 5% of inpatients are admitted to the hospital due to improper medication, and among the non-accidental deaths worldwide, about 1 / 7 are caused by irrational drug use. Rational drug use not only requires attention to the management of prescription and doctor's order systems, but also depends on the understanding of drug knowledge, the awareness of the situation, and the coping methods, etc.
[0003] Currently, in the process of monitoring adverse drug reactions, due to the complexity and diversity of real-world event data, as well as the inconsistency of the adverse reaction characteristics of each drug, the data processing for inferring adverse drug reactions is not comprehensive enough, resulting in inaccurate monitoring results of real-world data and lack of practicality. Moreover, in the process of monitoring adverse drug reactions, adverse events that deviate from the guidelines are often sporadic and low-probability events, and traditional statistical algorithms have very low sensitivity to such events and are difficult to incorporate into the standards of adverse drug reactions. In addition, adverse events caused by combined drug use are more serious than the above-mentioned single adverse drug use cases, but are also more difficult to monitor. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention proposes a method for constructing an evaluation model and predicting real-world adverse drug reactions in a hospital based on a knowledge base.
[0005] In a first aspect, the present invention provides a method for constructing an evaluation model for real-world adverse drug reactions in a hospital based on a knowledge base, the method comprising the following steps:
[0006] Obtain indicator data, large text data, and adverse drug reaction identification data of multiple patients, wherein the indicator data includes biochemical indicators, the large text data includes case texts, and the large text data includes at least part of the indicator data;
[0007] Extract the indicator data from the large text data as explicit features;
[0008] Based on the drug knowledge base, obtain implicit features indicating adverse drug reaction characteristics according to the large text data;
[0009] Concatenate the indicator data, the explicit features, and the implicit features according to a time series to obtain standardized features;
[0010] Use the standardized features as samples and the adverse drug reaction identification data as labels to train a preset machine learning model to obtain an evaluation model for real-world adverse drug reactions in a hospital based on a knowledge base.
[0011] Further, the adverse drug use identification data includes event data indicating whether the patient has an adverse reaction; based on the drug knowledge base, the implicit features indicating the characteristics of adverse drug reactions obtained from the large text data include:
[0012] Obtaining multiple feature words indicating terms related to adverse reactions based on the drug knowledge base;
[0013] Determining the TF-IDF information of each of the feature words in the large text data;
[0014] Determining the difference feature of the feature words according to the TF-IDF information;
[0015] Using the difference feature as a sample and the event data as a label to train a preset neural network to obtain a pre-trained model;
[0016] Taking the output of the neurons in the calibrated layer of the pre-trained model as the implicit feature.
[0017] Further, the determining the difference feature of the feature words according to the TF-IDF information includes:
[0018] Determining the mean value of the feature words in the known adverse event samples according to the TF-IDF information;
[0019] Determining the word frequency of the feature words in the sample to be predicted;
[0020] Taking the standard deviation of the word frequency and the mean value as the difference feature.
[0021] Further, the splicing the index data, the explicit feature, and the implicit feature according to the time series to obtain the standardized feature includes:
[0022] Performing normalization processing on the index data, the explicit feature, and the implicit feature;
[0023] Splicing the data after normalization processing according to the time series.
[0024] Further, the splicing the data after normalization processing according to the time series includes:
[0025] Splicing according to the following formula:
[0026]
[0027] where Y” represents the standardized feature, Y’ represents the implicit feature, X represents the index data, X’ represents the explicit feature, t i represents the time starting point, and Δt represents the average time increment. Represents a splicing function.
[0028] Furthermore, the adverse drug identification data includes probability information of at least one drug causing adverse reactions.
[0029] In a second aspect, the present invention provides an apparatus for constructing an evaluation model for real-world adverse drug use in a hospital based on a knowledge base. The apparatus includes a memory and a processor; the memory is used to store a computer program; the processor is used to, when executing the computer program, implement the method for constructing an evaluation model for real-world adverse drug use in a hospital based on a knowledge base as described above.
[0030] In a third aspect, the present invention provides a method for evaluating real-world adverse drug use in a hospital based on a knowledge base. The method includes the following steps:
[0031] Input the index data and large text data of the calibrated patient into the adverse drug use evaluation model constructed by the method for constructing an evaluation model for real-world adverse drug use in a hospital based on a knowledge base as described above. Among them, the index data includes biochemical indexes, and the large text data includes case texts;
[0032] Map the output of the adverse drug use evaluation model to at least one drug to determine whether the drug causes adverse reactions.
[0033] In a fourth aspect, the present invention provides an apparatus for evaluating real-world adverse drug use in a hospital based on a knowledge base. The apparatus includes a memory and a processor; the memory is used to store a computer program; the processor is used to, when executing the computer program, implement the method for evaluating real-world adverse drug use in a hospital based on a knowledge base as described above.
[0034] In a fifth aspect, the present invention provides a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the method for constructing an evaluation model for real-world adverse drug use in a hospital based on a knowledge base as described above, or implements the method for evaluating real-world adverse drug use in a hospital based on a knowledge base as described above.
[0035] The beneficial effects of the evaluation model construction, evaluation method, device and storage medium for real-world adverse drug use in hospitals based on a knowledge base are as follows: By obtaining the index data and large text data of patients and performing corresponding processing, the finally obtained model can be used for monitoring real-world data adverse drug use. Especially when adverse reactions occur while using multiple drugs simultaneously, it can effectively identify which drug caused the adverse event, thereby reducing the drug screening for patients in real-world data and determining whether it is an adverse event caused by multiple combined drugs, making the screening results more accurate. When, for example, a patient's medication needs to change or multiple drugs need to be used, based on the constructed model, a relatively accurate prediction can be made on whether there will be adverse drug reactions, so as to make corresponding preparations in advance and ensure the health of the patient. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 It is a flowchart showing the method for constructing an evaluation model for real-world adverse drug use in hospitals based on a knowledge base according to an embodiment of the present invention;
[0038] Figure 2 It is a flowchart showing the evaluation method for real-world adverse drug use in hospitals based on a knowledge base according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The following describes the principles and features of the present invention in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0040] As Figure 1 shown, a method for constructing an evaluation model for real-world adverse drug use in hospitals based on a knowledge base according to an embodiment of the present invention includes the following steps:
[0041] S11. Obtain the index data, large text data, and adverse drug use identification data of multiple patients. Among them, the index data includes biochemical indexes, the large text data includes case texts, and the large text data includes at least part of the index data.
[0042] Specifically, data can be obtained from the hospital database, including but not limited to basic patient information (e.g., gender, age, etc.), test index data (e.g., blood biochemical data, blood pressure, blood lipids, etc.), case (e.g., including symptoms, medication, prescriptions, etc.), adverse reaction conclusions and levels. Since each patient has its corresponding information, data as shown in Table 1 below can be obtained.
[0043] Table 1
[0044]
[0045] Among them, ID represents the patient number, and ellipsis represents the corresponding text information that is not listed.
[0046] Indicator data and big text data are two categories in data collection. For example, age, some biochemical indicators such as the amount of hemoglobin in blood routine, etc. The diagnosis time, hemoglobin and white blood cell count in Table 1 can be used directly and are mostly indicator data in diagnostic tests, while the subsequent large text data such as admission diagnosis and ward rounds records are considered big text data.
[0047] All data are original data. Among them, indicator data can be temporarily unprocessed, but large text data needs to be structured and extracted. For example, firstly, large text information such as cases is structured, and effective features are extracted using entity recognition plus rules, including but not limited to diagnosis, positive modification, negative modification, location, symptoms, drugs (prescriptions), diseases, course of disease, examination values, examination names, treatment methods, populations, etc., and adverse reaction conclusions, adverse reaction levels, and final suspected adverse drug use are obtained.
[0048] S12, extracting the indicator data from the large text data as an explicit feature.
[0049] Specifically, big text data usually consists of two parts. One part is explicit features. For example, in a case, "The patient came to our hospital for a routine blood test, hemoglobin N1..." appears. Although this situation belongs to big text data, the indicator data is implicit in it. After structured extraction, it is regarded as the indicator data equivalent to the corresponding data in Table 1. This process is explicit feature extraction, and the result after extraction is also in the same form as the indicator data.
[0050] S13, based on the drug knowledge base, obtaining implicit features indicating adverse drug reaction characteristics according to the big text data.
[0051] Specifically, another part of the big text data is implicit features, which can be used to indicate adverse drug reaction characteristics.
[0052] S14, concatenating the indicator data, the explicit features and the implicit features according to a time series to obtain standardized features.
[0053] Specifically, since the preprocessed data standards may not be unified and the data has temporal characteristics, before being used as model input, normalization and concatenation processing are first performed to obtain more accurate standardized features.
[0054] S15. Use the standardized features as samples and the adverse drug use identification data as labels to train a preset machine learning model to obtain an adverse drug use evaluation model.
[0055] In this embodiment, by obtaining the patient's indicator data and large text data and performing corresponding processing, the finally obtained model can be used for monitoring adverse drug use in real-world data. Especially when adverse reactions occur while using multiple drugs simultaneously, it can effectively identify which drug caused the adverse event, thereby reducing the drug screening for patients in real-world data and determining whether it is an adverse event caused by multiple combined drugs, making the screening results more accurate. For example, when a patient's medication needs to change or when multiple drugs need to be used, based on the constructed model, a relatively accurate prediction can be made on whether a drug will cause an adverse reaction, so as to make corresponding preparations in advance and ensure the patient's health.
[0056] In addition, since the input data comes from real-world data in the hospital and is usually relatively rich, it ensures the robustness of the model. Even if some indicators are missing, it is allowed. Therefore, it is also possible to make a certain degree of prediction on whether adverse reactions will occur in the initial medication of new patients.
[0057] Preferably, the adverse drug use identification data includes event data indicating whether the patient has an adverse reaction; the obtaining of implicit features indicating adverse drug use reaction characteristics based on the drug knowledge base according to the large text data includes:
[0058] Obtain multiple feature words indicating adverse reaction-related terms based on the drug knowledge base.
[0059] Determine the TF-IDF (Term Frequency–Inverse Document Frequency) information of each of the feature words in the large text data.
[0060] Determine the difference features of the feature words according to the TF-IDF information.
[0061] Use the difference features as samples and the event data as labels to train a preset neural network to obtain a pre-trained model.
[0062] Use the neuron outputs of the calibrated number of layers of the pre-trained model as the implicit features.
[0063] Specifically, we first need to use the text as a quantitative indicator according to the dimension of adverse reaction-related terms. We extract some adverse reaction-related terms from the drug knowledge base (including but not limited to rescue drugs, diagnostic feature descriptions, and landmark terms for suspected adverse events, etc.), and calculate the frequency of single text feature words and TF-IDF information as the original features for subsequent calculation of distribution.
[0064] More specifically, the word frequency of the characteristic word is determined according to a first formula, and the first formula is:
[0065]
[0066] Among them, TF i,j represents the frequency of feature word i in text j, n i,j represents the number of times feature word i appears in text j, Σ k n k,j Represents all words Σ in text j k Number of occurrences.
[0067] The reverse text probability of the feature word is determined according to a second formula, where the second formula is:
[0068]
[0069] Among them, IDF i,j represents the inverse text probability of feature word i in text j, |D| represents the total number of cases, |{d∈D:t∈d}| represents the total number of valid cases containing the calibrated feature word, feature word t exists in a single case d, and the case d belongs to the complete set D. For example, there are 3 cases containing "A" among 5 stable cases, and this number is 3.
[0070] The final TF-IDF information is:
[0071] TFIDF i,j =TF i,j ×IDF i,j .
[0072] Preferably, determining the difference feature of the feature word according to the TF-IDF information includes:
[0073] The mean value of the feature words in the known adverse event samples is determined according to the TF-IDF information.
[0074] Determine the word frequency of the feature word in the sample to be predicted.
[0075] The standard deviation of the word frequency and the mean is used as the difference feature.
[0076] More specifically, the big data text calculates the standard deviation according to the distribution of the above-mentioned bad trigger words in the positive example samples, and then calculates the difference feature p i , that is, the word frequency and TF-IDF information of a certain feature word i in the case to be predicted and the mean value of a certain feature word i in the known bad event samples of the standard deviation. The calculation method is as follows:
[0077]
[0078] where x i can be understood as that in the text to be predicted, the mean value while is the mean value of x in the training set i .
[0079] After that, the difference feature p i is used as the output, and a neural network is introduced. The final label of whether it is a bad event is used as the output y i .
[0080] The neural network Y = w i P + b i , Y′ = σ i (Y) parameters are used as the pre-trained model. And the neuron output of a certain layer from the bottom (selected according to the depth of the neural network, generally the second layer from the bottom) is used as the implicit feature Y′ of the large text data of the subsequent model. In actual operation, the 16-dimensional feature output by the 16 neurons of the second layer from the bottom has the best effect.
[0081] In the TF-IDF part, the input is the full sample, which can be expressed as shown in Table 2.
[0082] Table 2
[0083] ID Headache Poor sleep Nausea Acid regurgitation …… 0001 1.7 0 2.32 0.98 …… 0002 0.3 0.045 0 0 ……
[0084] Among them, Table 2 does not list all the feature items in detail, so ellipsis is used to represent.
[0085] Then, the model extracts further, that is, taking the content in the above table as the input and outputting N-dimensional features, in the following form:
[0086] 0001: [-0.54, 3.23, 2.34, …]
[0087] 0002: …
[0088] This is the implicit feature of each large text data.
[0089] Preferably, the splicing of the index data, the explicit feature, and the implicit feature according to the time series to obtain the standardized feature includes:
[0090] Normalize the said indicator data, the explicit features, and the implicit features.
[0091] Concatenate the normalized data according to the time series.
[0092] Specifically, due to the differences in the measurement methods, existing forms, and types of different indicators, for example, some are texts, some are numbers, some values are very large, and some values are very small, subsequent calculations will be overly affected by an indicator with a large single measurement unit. Therefore, normalization is performed first to make the contribution weights of each indicator equal at the initial stage of the model as much as possible.
[0093] The above-mentioned indicator data, explicit features, and implicit features can be divided into two parts according to quantifiable information and non-quantifiable information. For non-quantifiable information, divide it into presence or absence and standardize it to 1 or 0. For quantifiable information, standardize it all to the 0-1 interval.
[0094] The method of standardization is to perform a standardized linear transformation on a column of data, so that
[0095] where is the standardized value, x i is the original value, and approximate values of k and w are obtained, satisfying:
[0096]
[0097] 3σ = 0.5,
[0098] so that the finally standardized data conforms to the Gaussian distribution.
[0099] Preferably, the adverse drug use identification data includes probability information of at least one drug causing adverse reactions.
[0100] For the treatment of suspected adverse drug use categories, each category takes values from 0 to 1 to fit the probability distribution. For the treatment of suspected adverse drug use, for a single identified adverse drug use in a certain case, it is considered that the value of this category is 1, and the values of other categories are 0. If there are multiple combined adverse drug uses or it is suspected in a certain case, then the value of each adverse drug use is 0.8. Here, it is not appropriate to select 0.5 because the positive examples of adverse drug use are much less than the negative examples, so the weight value is between 0.7 and 0.9. Finally, the value of each category is used as the gold standard of the final prediction model.
[0101] Preferably, the concatenating the normalized data according to the time series includes:
[0102] Concatenate according to the following formula:
[0103]
[0104] Among them, Y” represents the standardized feature, Y’ represents the latent feature, X represents the index data, X’ represents the explicit feature, and t i represents the starting point of time, and Δt represents the average time increment. represents the splicing function.
[0105] Specifically, it is necessary to splice the above-mentioned text latent feature Y’ (16-dimensional in this case) with the index data and explicit feature that have been normalized, and perform horizontal dimension expansion in combination with time information. For the latent feature, the time point is used as the time feature, and for the indicators in other time periods, the starting point of time plus the average time increment is used as the time feature. The form is as shown in the above formula, where mainly a time weight is given to the above-mentioned normalized indicators. The time point data uses the time point as the time weight, such as 3:00, and the time period data uses the starting time point plus the average duration as the time weight.
[0106] Finally, the feature sequence is input into the RNN network in chronological order to obtain the final layer, as shown in the following example:
[0107]
[0108] Among them,
[0109] In the prediction model part, a five-classification layer is connected after the final layer, and each node performs cross-entropy to obtain the likelihood probability as the possible probability of this class.
[0110] During the final training, the FGM adversarial perturbation x adv = x + r adv can be added to improve the robustness of the model. The perturbation interval is:
[0111]
[0112] Among them, x is the original sample, r is the random perturbation, and g is the gradient of the loss function L.
[0113] Finally, a large amount of known gold standard data is prepared, and the standardized feature x i and the gold standard label y i are obtained through the above data processing and input into the model to learn the relevant parameters, and finally a trained model is obtained as the model for subsequent prediction.
[0114] In another embodiment of the present invention, an evaluation model construction device for real-world adverse drug use in a hospital based on a knowledge base includes a memory and a processor; the memory is used to store a computer program; the processor is used to, when executing the computer program, implement the above-mentioned method for constructing an evaluation model for real-world adverse drug use in a hospital based on a knowledge base.
[0115] As Figure 2 shown, a method for evaluating real-world adverse drug reactions in a hospital based on a knowledge base according to an embodiment of the present invention includes the following steps:
[0116] S21, input the index data and large text data of the calibrated patient into the adverse drug reaction evaluation model constructed by the method for constructing an adverse drug reaction evaluation model in a hospital based on a knowledge base as described above, wherein the index data includes biochemical indexes, and the large text data includes case texts.
[0117] S22, map the output of the adverse drug reaction evaluation model to at least one drug to determine whether the drug causes adverse reactions.
[0118] Specifically, the relevant data of the calibrated patient can be processed step by step, such as in steps S12, S13, etc. above, and input into the trained model to obtain the probability of multi-classification. Then, the probability of each class is mapped to a certain type of drug as the determination value.
[0119] If the probability of the class with the highest probability exceeds 0.8 and no other classes exceed it, it is considered that this type of drug is a single definite adverse drug reaction; if the probability of the class with the highest probability exceeds 0.5 but does not exceed 0.8 and no other classes exceed it, it is considered that this type of drug is a single suspected adverse drug reaction; if the probability of multiple classes of drugs exceeds 0.8, it is considered that these drugs are multiple combined adverse drug reactions; if the probability of multiple classes of drugs exceeds 0.5 but does not exceed 0.8, it is considered that these drugs are suspected multiple combined adverse drug reactions.
[0120] Taking a multi-classification output as an example, it can be expressed in the form shown in Table 3.
[0121] Table 3
[0122] ID Aspirin Digoxin Xueshuantong …… 0001 0.213 0.852 0.034 ……
[0123] Among them, the probability of digoxin exceeds 0.8, and it can be determined that it is a single definite adverse drug reaction. In addition, some drugs are not shown, so they are represented by ellipsis.
[0124] In this embodiment, it is applicable to situations such as when a patient uses multiple drugs and adverse reactions occur, and it is necessary to investigate the drugs, etc. By obtaining the patient's indicator data and large text data and performing corresponding processing, the finally obtained model can be used for the monitoring of adverse drug use in real-world data. Especially when adverse reactions occur while using multiple drugs simultaneously, it can effectively identify which drug use has caused the adverse event, thereby reducing the drug investigation for patients in real-world data and determining whether it is an adverse event caused by multiple combined drug uses, making the investigation results more accurate. When, for example, the patient's drug use needs to change or when multiple drugs need to be used, based on the constructed model, a relatively accurate prediction can be made on whether there will be adverse reactions caused by drugs, so as to make corresponding preparations in advance and ensure the patient's health. In addition, since the input data comes from the in-hospital real-world data, which is usually relatively rich, it ensures the robustness of the model. Even if some indicators are missing, it is allowed. Therefore, to a certain extent, it is also possible to predict whether adverse reactions will occur in the initial drug use of new patients.
[0125] In another embodiment of the present invention, an evaluation device for in-hospital real-world adverse drug use based on a knowledge base includes a memory and a processor; the memory is used to store a computer program; the processor is used to, when executing the computer program, implement the evaluation method for in-hospital real-world adverse drug use based on the knowledge base as described above.
[0126] In another embodiment of the present invention, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the evaluation model construction method for in-hospital real-world adverse drug use based on the knowledge base as described above, or implements the evaluation method for in-hospital real-world adverse drug use based on the knowledge base as described above.
[0127] Readers should understand that in the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0128] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for constructing an evaluation model for real-world adverse drug use in a hospital based on a knowledge base, characterized in that, Comprising: Obtaining indicator data, large text data, and adverse drug use identification data of multiple patients, wherein the indicator data includes biochemical indicators, the large text data includes case texts, and the large text data includes at least part of the indicator data; Extracting the indicator data therein from the large text data as explicit features; Based on a drug knowledge base, obtaining implicit features indicating adverse drug reaction characteristics according to the large text data; Splicing the indicator data, the explicit features, and the implicit features according to a time series to obtain standardized features; Using the standardized features as samples and the adverse drug use identification data as labels to train a preset machine learning model to obtain an adverse drug use evaluation model; The adverse drug use identification data includes event data indicating whether an adverse reaction has occurred in the patient; the obtaining implicit features indicating adverse drug reaction characteristics based on the drug knowledge base according to the large text data includes: Obtaining multiple feature words indicating adverse reaction related terms based on the drug knowledge base; Determining the TF-IDF information of each of the feature words in the large text data; Determining the difference features of the feature words according to the TF-IDF information; Using the difference features as samples and the event data as labels to train a preset neural network to obtain a pre-trained model; Taking the neuron output of the calibrated layer number of the pre-trained model as the implicit feature; The determining the difference features of the feature words according to the TF-IDF information includes: Determining the mean value of the feature words in known adverse event samples according to the TF-IDF information; Determining the word frequency of the feature words in the sample to be predicted; Taking the standard deviation of the word frequency and the mean value as the difference feature.
2. The method for constructing an evaluation model for real-world adverse drug use in a hospital based on a knowledge base according to claim 1, wherein, The splicing the indicator data, the explicit features, and the implicit features according to a time series to obtain standardized features includes: Performing normalization processing on the indicator data, the explicit features, and the implicit features; Splicing the data after normalization processing according to a time series.
3. The method for constructing an evaluation model for in-hospital real-world adverse drug use based on a knowledge base according to claim 2, wherein The splicing the data after normalization processing according to a time series includes: Performing splicing according to the following formula: ; Among them, Y’’ represents the standardized feature, Y’ represents the implicit feature, X represents the index data, X’ represents the explicit feature, t i represents the time starting point, Δ t represents the average time increment, φ represents the splicing function.
4. The method for constructing an evaluation model for real-world adverse drug use in a hospital based on a knowledge base according to claim 1, wherein The adverse drug use identification data includes probability information of at least one drug causing an adverse reaction.
5. An apparatus for constructing an evaluation model for real-world adverse drug use in a hospital based on a knowledge base, characterized in that, Including a memory and a processor; the memory is used for storing a computer program; the processor is used for, when executing the computer program, implementing the method for constructing an evaluation model for in-hospital real-world adverse drug use based on a knowledge base as described in any one of claims 1 to 4.
6. A method for evaluating real-world adverse drug events in a hospital based on a knowledge base, characterized in that, Comprising: Inputting the indicator data and large text data of a calibrated patient into an adverse drug use evaluation model constructed by the method for constructing an evaluation model for in-hospital real-world adverse drug use based on a knowledge base as described in any one of claims 1 to 4, wherein the indicator data includes biochemical indicators and the large text data includes case texts; Mapping the output of the adverse drug use evaluation model to at least one drug to determine whether the drug causes an adverse reaction.
7. An evaluation device for real-world adverse drug use in a hospital based on a knowledge base, characterized in that, It includes a memory and a processor; the memory is used for storing a computer program; the processor is used for implementing the method for evaluating real-world adverse drug use in a hospital based on a knowledge base as described in claim 6 when executing the computer program.
8. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the method for constructing an evaluation model for real-world adverse drug use in a hospital based on a knowledge base as described in any one of claims 1 to 4, or implements the method for evaluating real-world adverse drug use in a hospital based on a knowledge base as described in claim 6.
Citation Information
Patent Citations
Method for establishing blood transfusion adverse reaction database, storage system and active early warning system
CN111063448A
Tele-analytics based treatment recommendations
US20140278475A1
Method and device for building patent knowledge base, computer equipment and storage medium
CN108763445A
Electronic medical record data processing method and device, electronic equipment and readable medium
CN111383726A