Method and System for Temporal-based Data Fusion and Prediction of Drug-induced Liver Injury in Children

Through the time-series-based data fusion and prediction method, the MUIRFTI model is used to process unstructured information of drug-induced liver injury in children, which solves the problem of missing timing information and improves the prediction accuracy of DILI.

CN119541890BActive Publication Date: 2025-06-17BEIJING CHILDRENS HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510104574.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-17
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The missing timing information of events related to drug-induced liver injury (DILI) in children in the prior art affects the accuracy of identification of DILI and leads to insufficient accuracy of DILI prediction.

Method used

The fusion and prediction method of children's drug-induced liver injury data based on time is adopted. By collecting and cleaning historical medical data, a clinical database is constructed, and the unstructured information is marked and structured using the MUIRFTI model to fuse the timing information to train the machine learning prediction model.

Benefits of technology

By fusion of timing information, the amount of information and logic of the research data set is improved, a more effective machine learning prediction model is trained, and the prediction accuracy of DILI is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541890B_ABST
    Figure CN119541890B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for data fusion and prediction of drug-induced liver injury in children based on time series, which relates to the technical field of information processing of drug-induced liver injury in children. The method mainly includes: collecting the historical medical data of inpatients as a longitudinal data set, arranging it horizontally in chronological order, and constructing a clinical database; based on the related event phenotypes of drug-induced liver injury in children, performing entity relationship annotation on the unstructured information in the research data set; constructing and training a MUIRFTI model for structured mapping reconstruction of unstructured text based on time series information; training a machine learning prediction model to predict whether drug-induced liver injury occurs; inputting the medical data of the measured child into the machine learning prediction model to predict whether the child will suffer from drug-induced liver injury. This solution can greatly improve the information content and logic of the research data set, so as to train a better-performing machine learning prediction model for predicting drug-induced liver injury in children.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing of drug-induced liver injury in children, and in particular to a method and system for data fusion and prediction of drug-induced liver injury in children based on time series. Background Art

[0002] The pathogenesis of drug-induced liver injury (DILI) is complex, and most of the time it is the result of multiple mechanisms acting sequentially or together, including direct drug hepatotoxicity and idiosyncratic hepatotoxicity. Direct drug hepatotoxicity refers to the direct damage to the liver caused by drugs and / or their metabolites ingested into the human body, which can further cause other immune and inflammatory responses. Idiosyncratic hepatotoxicity refers to increased susceptibility to drug-induced liver injury in individuals due to factors such as abnormal individual drug metabolism, drug-mediated immune damage, or individual genetic differences.

[0003] Children are still in the developmental stage and are highly susceptible to DILI. Monitoring DILI in children has important clinical and public health significance. In recent years, in order to overcome the limitations of passive monitoring, the use of electronic medical record (EMR) data for active DILI monitoring has become a research hotspot. However, the lack of temporal information of DILI-related events (including symptoms, signs, etc.) will seriously affect the accuracy of DILI identification. Therefore, studying the temporal processing strategy of DILI-related events is a scientific problem that needs to be solved urgently to improve the accuracy of DILI prediction. Summary of the invention

[0004] The object of the present invention is to provide a time series-based data fusion and prediction method and system for drug-induced liver injury in children to solve at least one of the above-mentioned technical problems existing in the prior art.

[0005] In the first aspect, in order to solve the above technical problems, the present invention provides a method for fusion and prediction of drug-induced liver injury in children based on time series, comprising:

[0006] Step 1. Collect historical medical data of hospitalized children as a longitudinal data set, arrange them horizontally in chronological order, and construct a clinical database; perform data cleaning and standardization on the clinical database; calculate the sample size of the prediction model for drug-induced liver injury in children, and extract the research data set from the clinical database according to the inclusion and exclusion criteria; based on the relevant event phenotypes of drug-induced liver injury in children, annotate the entity relationships of the unstructured information in the research data set, and process the missing values ​​and outliers of the structured information in the research data set; in the research data set, use whether drug-induced liver injury occurs as a label, and divide it into a training set and a test set according to a preset ratio.

[0007] In a feasible implementation manner, the historical medical data includes electronic medical record (EMR) data of hospitalized children.

[0008] In a feasible implementation, the data cleaning includes data desensitization, screening out errors and duplicate records, etc.

[0009] In a feasible implementation, the standardization process includes making the diagnostic names and drug names consistent with the term specifications such as the International Classification of Diseases and Related Health Problems, 10th Revision (ICD-10) and the Anatomical Therapeutic Chemical Classification (ATC) codes of drugs.

[0010] In a feasible implementation, the calculation method of the sample size includes:

[0011] Based on literature retrieval and expert consultation, select the specific manifestations of liver injury and the number of related clinical events to be included;

[0012] Calculate the sample size based on the four-step method for calculating the sample size of prediction models published in BMJ literature.

[0013] In a feasible implementation, the inclusion and exclusion criteria include:

[0014] Include children within a preset number of years and with the first laboratory biochemical examination indicators at admission: alkaline phosphatase (ALP), alanine aminotransferase (ALT), aspartate aminotransferase (AST), total bilirubin (TB), and gamma-glutamyl transferase (GGT) within the normal value range; the age range of the children is from 1 month to 18 years old;

[0015] Exclude children with underlying diseases affecting liver-related functions at admission, and children with abnormal liver function before using the target drug.

[0016] In a feasible implementation, the research data set includes a standard positive event group and a control group:

[0017] The standard positive event group is to select the historical medical data of children with DILI from the longitudinal data set through the ADR causality assessment method and the RUCAM scale;

[0018] The control group is randomly selected from the longitudinal data set according to a preset control ratio, and includes the historical medical data of children without DILI who have used the target drug (the drug that induces DILI).

[0019] In a feasible implementation, the relevant event phenotypes are determined through multiple rounds of Delphi expert consultation based on the literature in this field, such as "Common Clinical Medical Terms 2023 Edition", etc., and combined with the writing characteristics of the electronic medical records of inpatients, and include symptom indicators, laboratory examination indicators, demographic indicators, target drug use indicators, and disease indicators;

[0020] The symptom indicators include jaundice, nausea, vomiting, rash, skin pruritus, abdominal discomfort, upper abdominal discomfort, right upper quadrant pain, fever, anorexia, decreased appetite, fatigue, weakness, mild jaundice, moderate jaundice, severe jaundice, and incomplete subsidence of jaundice;

[0021] The laboratory test indicators include alkaline phosphatase, alanine aminotransferase, aspartate aminotransferase, total bilirubin, and gamma-glutamyl transferase;

[0022] The demographic indicators include age, height, weight, and gender;

[0023] The target drug use indicators include dose, dosage form, administration route, whether to change the drug, and concomitant medications;

[0024] The disease indicators include disease discharge diagnosis and complications.

[0025] In a feasible implementation manner, the unstructured information includes text data such as the front page information of the medical record, admission record, daily course, handover record, ward round record, and first postoperative course in the electronic medical record of inpatients.

[0026] In a feasible implementation manner, the specific method for entity relationship annotation includes:

[0027] Segment and tokenize the unstructured information to obtain a corpus set; for the corpus set, perform non-standard entity naming and corresponding relationship extraction; the entity naming includes: symptoms (jaundice, nausea, vomiting, rash, skin pruritus, abdominal discomfort, upper abdominal discomfort, right upper quadrant pain, fever, anorexia, decreased appetite, fatigue, weakness, mild jaundice, moderate jaundice, severe jaundice, and incomplete subsidence of jaundice) and the severity levels of the corresponding symptoms; the drug names and attributes (dose, frequency, course of treatment, etc.) before and after the occurrence of these symptoms;

[0028] Then, through the pre-trained MedBERT model, perform standard entity naming and corresponding relationship extraction on the corpus set.

[0029] In a feasible implementation manner, the processing of missing values and outliers includes:

[0030] First, uniformly label the missing values, and then perform statistical analysis, filling, or elimination;

[0031] For outliers, first identify them according to the variable characteristics and coding table, and handle the outliers that cannot be identified as missing values.

[0032] Step 2: Construct and train a medical unstructured information based temporal reconstruction and fusion model, abbreviated as the MUIRFTI model, for structurally mapping and reconstructing unstructured text based on temporal information; input the unstructured text of the research dataset into the trained MUIRFTI model, perform structural reconstruction according to the rules of each module in the MUIRFTI model, and after obtaining the reconstructed structured text, fuse it with the original structured text in the research dataset to obtain a modeling dataset based on temporal fusion.

[0033] In a feasible implementation, the MUIRFTI model includes a text correction module (TC), a text structuring module (TS), a medical facts and events extraction module (MFEE), and a time information extraction module (TIE) arranged in sequence;

[0034] The text correction module includes a correction strategy library for correcting spelling mistakes in the text according to regular strategies through natural language processing (NLP) tools;

[0035] The text structuring module includes a grammar rule library for adjusting grammar mistakes in the text and segmenting sentence granularity according to grammar rules through natural language processing (NLP) tools; the grammar rule library includes rules for describing symptoms of children's cases and terms related to drug-induced liver injury events; in this way, the writing characteristics of symptom descriptions of children's cases can be taken into account to determine noun terms related to drug-induced liver injury events.

[0036] The medical facts and events extraction module includes a key value model (KV), a time event description model (TED), a medical research knowledge base, and a labeled corpus for extracting medical facts and events in the text according to the mapping of time sequence through natural language processing (NLP) tools to obtain a structured database of clinical events with DILI as the target outcome;

[0037] The time information extraction module includes a time expression rule library for processing time expressions in the structured database of clinical events through natural language processing (NLP) tools to construct a time development model of the child's condition (TDM-PC).

[0038] In a feasible implementation, the specific formula of the key value model includes:

[0039] ;

[0040] where represents the key value extraction result of the th statement; represents the key parameter name of the th statement; represents the The key parameter values of a statement; Indicates the number of keywords in the statement; Indicates the set of key value extraction results; Indicates the item to which the key value belongs; Indicates the number of statements.

[0041] In a feasible implementation manner, the specific formula of the time event model includes:

[0042] ;

[0043] Among them, Indicates the th statement; Indicates the key time in the th statement; Indicates the key action in the th statement; Indicates the key effect in the th statement; Indicates probability.

[0044] In a feasible implementation manner, the time information extraction module identifies the time expressions in the clinical event structured database through natural language processing (NLP) tools.

[0045] In a feasible implementation manner, the time development model of the child's condition is used to identify the key information of the child's condition in the clinical event structured database, and based on the time expressions, sort them in chronological order to reproduce the development process of the disease and clinical events; the key information of the child's condition includes: health status (initial symptoms and diagnosis) information, hospitalization diagnosis information, treatment process information, treatment medications and laboratory test information, clinical follow-up information, and reexamination information.

[0046] In a feasible implementation manner, the method for evaluating the training effect of the MUIRFTI model is: by calculating the change in information entropy between the unstructured text before the conversion of the MUIRFTI model and the structured text after the conversion, quantifying the improvement of text uncertainty, so as to judge the conversion effect of the MUIRFTI model, specifically including:

[0047] Step a1, calculate the total text information entropy of the unstructured text based on the training set , and the specific formula is:

[0048] ;

[0049] Among them, Indicates the total number of phrases in the total text; Represents the preset occurrence probability of a single phrase in the text; the total text consists of sentences, the sentences consist of paragraphs, the paragraphs consist of phrases, and the phrases consist of words;

[0050] Step a2: Calculate the information entropy at the output stage of the text structuring module. The information entropy of the total text at this stage is equal to the sum of the information entropies of each sentence. The specific formula is:

[0051] ;

[0052] Where, Represents the number of sentences in the text; Represents the th paragraph of the th sentence; Represents the th sentence of the th word's preset occurrence probability; in this stage, Is equal to ;

[0053] Calculate the change in information entropy between this stage and the information entropy in step a1 to evaluate the training effect of the text structuring module;

[0054] Step a3: Calculate the information entropy at the output stage of the medical facts and events extraction module. Based on the formula in step a2, it also includes calculating the information entropy of phrases in the text , and the specific formula is:

[0055] ;

[0056] Where, Represents the adjacent word occurrence frequency, and the specific calculation formula is:

[0057] ;

[0058] Where, Represents the length of the word;

[0059] Calculate the change in information entropy between this stage and the information entropy in step a1 to evaluate the training effect of the text structuring module and the medical facts and events extraction module;

[0060] Step a4: Calculate the information entropy at the output stage of the time information extraction module , and the specific formula is:

[0061] ;

[0062] Where, Represents the information entropy of the phrase at the th time node; Indicates the preset occurrence probability of the phrase in the text;

[0063] Calculate the change in information entropy at this stage and the information entropy in step a1 to evaluate the training effect of the MUIRFTI model.

[0064] In a feasible implementation, the method for evaluating the training effect of the MUIRFTI model further includes:

[0065] Step a5: Based on the test set, calculate the information entropy of each stage according to the methods in steps a1 - a4, and compare them with the information entropy in steps a1 - a4 respectively; by comparing the change in information quotient before and after the time series reconstruction process, supplement and correct the strategy library, grammar rule library, medical research knowledge library, and labeled corpus according to the actual corpus of inpatient children's electronic medical record texts and the actual corpus of children's symptom descriptions, so as to improve the MUIRFTI model in combination with the writing characteristics of inpatient children's electronic medical record data.

[0066] Step 3: Train a machine learning prediction model based on different strategy combinations through the training set in the modeling dataset to predict whether drug-induced liver injury occurs; test the machine learning prediction model through the test set in the modeling dataset; evaluate the performance of the machine learning prediction model under different strategy combinations and parameters through evaluation indicators, and optimize to obtain the final machine learning prediction model; the independent variables in the feature set of the machine learning prediction model include the occurrence times of clinical events; the strategy combination includes three time series processing strategies for clinical events, which are used to count the occurrence times of each clinical event, specifically including: independent event set strategy, bundled event set strategy, and weighted event set strategy.

[0067] In a feasible implementation, the machine learning prediction model is a classification model.

[0068] In a feasible implementation, the independent event set strategy means that when calculating the number of times of each clinical event, regardless of the time series relationship with the occurrence of drug-induced liver injury, directly calculate the total occurrence times of the clinical event within a preset time period before the onset of drug-induced liver injury.

[0069] In a feasible implementation, the bundled event set strategy means that when calculating the number of times of each clinical event, consider the time series relationship with the occurrence of drug-induced liver injury, and calculate the occurrence times of the clinical event within each day before the onset of drug-induced liver injury respectively.

[0070] In a feasible implementation manner, the weighted event set strategy means that when calculating the number of occurrences of each clinical event, the temporal relationship with the occurrence of drug-induced liver injury is considered, and it is assumed that the clinical events that occur closer to the onset of drug-induced liver injury contribute more to the occurrence of drug-induced liver injury, and the assigned weight is also greater. That is, the weight assigned to the day before the onset of drug-induced liver injury is 1, and the weight assigned to the dth day before the onset of drug-induced liver injury is 1 / d. Then, the weighted sum is used to calculate the total number of occurrences of this clinical event. The specific formula is:

[0071] ;

[0072] wherein, represents the preset statistical number of days before the onset of drug-induced liver injury; represents the number of occurrences of this clinical event on the dth day before.

[0073] In a feasible implementation manner, the clinical events include prescription drugs, symptoms, and laboratory test results.

[0074] In a feasible implementation manner, in the training set, the number of days before the onset of drug-induced liver injury is divided into several time windows for model training. The reason for such setting is that the closer to the onset of drug-induced liver injury, the more clinical events occur and the higher their importance, and the more intensive attention is required.

[0075] In a feasible implementation manner, the prediction results of the machine learning prediction model are verified by the 10-fold cross-validation method, and the parameters of the machine learning prediction model are optimized.

[0076] In a feasible implementation manner, the evaluation indicators include the area under the precision-recall curve (PR-AUC), precision, recall, and F1-score.

[0077] Step 4: Input the medical data of the tested child into the machine learning prediction model to predict whether the child will suffer from drug-induced liver injury.

[0078] In a second aspect, based on the same inventive concept, the present application further provides a temporal-based data fusion and prediction system for drug-induced liver injury in children, including a data receiving module, a data processing module, and a result generating module;

[0079] The data receiving module is used to receive the medical data of the tested child;

[0080] The data processing module, based on the medical data of the tested child, predicts whether the child will suffer from drug-induced liver injury through the above-mentioned temporal-based data fusion and prediction method for drug-induced liver injury in children;

[0081] The result generating module is configured to send out the prediction result.

[0082] Adopting the above technical solution, the present invention has the following beneficial effects:

[0083] A method and system for data fusion and prediction of children's drug-induced liver injury based on time series provided by the present invention can perform structured processing on unstructured data related to children's drug-induced liver injury, reconstruct it according to time series, and then fuse it with the original structured data, thereby greatly improving the information volume and logic of the research data set, and can train a better machine learning prediction model for predicting children's drug-induced liver injury; this solution also optimizes and compares different time series processing strategies for clinical events, and clarifies that compared with the strategy of simply constructing a model based on structured demographic information, medication information, and laboratory indicators, fully considering the rich time series information contained in the EMR can effectively improve the prediction efficiency and effect of the machine learning prediction model; at the same time, the processing strategy considering the occurrence weight of DILI-related events can further improve the effect of the DILI prediction model compared with simply based on the time series of event occurrence. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0085] Figure 1 It is a flowchart of a method for data fusion and prediction of children's drug-induced liver injury based on time series provided by an embodiment of the present invention;

[0086] Figure 2 It is an architecture diagram of the MUIRFTI model provided by an embodiment of the present invention;

[0087] Figure 3 It is an explanatory diagram of the principle of the disease time development model of children provided by an embodiment of the present invention;

[0088] Figure 4 It is a flowchart of a method for evaluating the training effect of the MUIRFTI model provided by an embodiment of the present invention;

[0089] Figure 5 It is an illustrative diagram of three time series processing strategies for clinical events provided by an embodiment of the present invention;

[0090] Figure 6 It is an illustrative diagram of weight assignment provided by an embodiment of the present invention;

[0091] Figure 7 This is a diagram of a system for data fusion and prediction of children's drug-induced liver injury based on time series provided by an embodiment of the present invention. Detailed implementation manners

[0092] Next, the technical solutions of the present invention will be described clearly and completely with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0093] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0094] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0095] The following further explains and illustrates the present invention in combination with specific implementation manners.

[0096] It should also be noted that the following specific embodiments or specific implementation manners are a series of optimized setting manners listed by the present invention to further explain the specific invention content, and these setting manners can be combined with each other or used in association with each other.

[0097] Embodiment 1:

[0098] As Figure 1 shown, a method for data fusion and prediction of children's drug-induced liver injury based on time series provided in this embodiment includes:

[0099] Step 1: Collect the historical medical data of inpatients as a longitudinal dataset, arrange it horizontally in chronological order, and construct a clinical database; perform data cleaning and standardization on the clinical database; calculate the sample size of the pediatric drug-induced liver injury prediction model, and extract the research dataset from the clinical database according to the inclusion and exclusion criteria; based on the relevant event phenotypes of pediatric drug-induced liver injury, perform entity relationship annotation on the unstructured information in the research dataset, and process missing values and outliers for the structured information in the research dataset; in the research dataset, use whether drug-induced liver injury occurs as a label, and divide it into a training set and a test set according to a preset division ratio.

[0100] Further, the historical medical data includes the electronic medical record (EMR) data of inpatients.

[0101] Further, the data cleaning includes data desensitization, screening out incorrect and duplicate records, etc.

[0102] Further, the standardization processing includes keeping the diagnostic names and drug names consistent with the term specifications such as the International Classification of Diseases, Tenth Revision (ICD-10) and the Anatomical Therapeutic Chemical Classification (ATC) codes of drugs.

[0103] Further, the calculation method of the sample size includes:

[0104] According to literature retrieval and expert consultation, select the specific manifestations of liver injury and the number of related clinical events to be included (23 related event phenotypes can be referred to);

[0105] Based on the four-step method for calculating the sample size of prediction models published in BMJ literature, calculate the sample size.

[0106] Further, the inclusion and exclusion criteria include:

[0107] Include children within a preset number of years and with the first laboratory biochemical examination indicators at admission: alkaline phosphatase (ALP), alanine aminotransferase (ALT), aspartate aminotransferase (AST), total bilirubin (TB), and gamma-glutamyl transferase (GGT) within the normal range;

[0108] Exclude children with underlying diseases affecting liver-related functions at admission, such as children in the hematology-oncology department and the oncology department, and children with abnormal liver function before using the target drug.

[0109] Further, the research dataset includes a standard positive event group and a control group:

[0110] The standard positive event group selects the historical medical data of children with DILI from the longitudinal dataset through the ADR causality assessment method (such as the Naranjo assessment method) and the RUCAM scale;

[0111] The control group is randomly selected from the longitudinal dataset according to a preset control ratio (such as 1:2), and includes the historical medical data of children without DILI who have used the target drug (the drug that induces DILI);

[0112] For example, from January 2009 to February 2022, there were 826 children with DILI in a certain hospital, serving as the standard positive event group; based on the aforementioned sample size calculation method, the calculated sample size was 1783 cases, so a total of about 2478 children were included to ensure subsequent model building.

[0113] Furthermore, the relevant event phenotypes are determined through two rounds of Delphi expert consultation based on the literature "Common Clinical Medicine Terms 2023 Edition" in this field and combined with the writing characteristics of the electronic medical records of inpatients, including symptom indicators, laboratory examination indicators, demographic indicators, target drug use indicators, and disease indicators;

[0114] The symptom indicators include jaundice, nausea, vomiting, rash, skin itching, abdominal discomfort, upper abdominal discomfort, right upper quadrant pain, fever, loss of appetite, decreased appetite, fatigue, weakness, mild jaundice, moderate jaundice, severe jaundice, and incomplete resolution of jaundice;

[0115] The laboratory examination indicators include alkaline phosphatase, alanine aminotransferase, aspartate aminotransferase, total bilirubin, and gamma-glutamyl transferase;

[0116] The demographic indicators include age, height, weight, and gender;

[0117] The target drug use indicators include dosage, dosage form, administration route, whether to change the drug, and concomitant medications;

[0118] The disease indicators include the discharge diagnosis of the disease and complications.

[0119] Furthermore, the unstructured information includes text data such as the front page information of the medical record, admission record, daily course, handover record, ward round record, and first postoperative course in the electronic medical record data of inpatients.

[0120] Furthermore, the specific method of entity relationship annotation includes:

[0121] Segment and tokenize the unstructured information to obtain a corpus; for the corpus, perform non-standard entity naming and corresponding relationship extraction; the entity naming includes: symptoms (jaundice, nausea, vomiting, rash, skin itching, abdominal discomfort, upper abdominal discomfort, right upper abdominal pain, fever, loss of appetite, decreased appetite, fatigue, weakness, mild jaundice, moderate jaundice, severe jaundice, unresolved jaundice) and the severity levels of the corresponding symptoms; the names and attributes of medications (dosage, frequency, course of treatment, etc.) before and after the occurrence of these symptoms;

[0122] Then, use the pre-trained conventional MedBERT model to perform standard entity naming and corresponding relationship extraction on the corpus.

[0123] Furthermore, the missing value and outlier processing includes:

[0124] First, uniformly label the missing values, and then perform statistical analysis, filling, or elimination;

[0125] For outliers, first identify them according to the variable characteristics and coding table, and treat the outliers that cannot be identified as missing values.

[0126] Furthermore, the preset division ratio is 7:3, that is, the research data set is divided into a training set and a test set according to the ratio of 7:3.

[0127] Step 2: Construct and train the MUIRFTI model for structured mapping reconstruction of unstructured text based on temporal information; input the unstructured text of the research data set into the trained MUIRFTI model, perform structured reconstruction according to the rules of each module in the MUIRFTI model, and after obtaining the reconstructed structured text, fuse it with the original structured text in the research data set to obtain a modeling data set based on temporal fusion.

[0128] Furthermore, as Figure 2 shown, the MUIRFTI model includes a text correction module (TC), a text structuring module (TS), a medical facts and events extraction module (MFEE), and a time information extraction module (TIE) arranged in sequence;

[0129] The text correction module includes a correction strategy library for correcting spelling mistakes in the text according to regular strategies through conventional natural language processing (NLP) tools;

[0130] The text structuring module includes a grammar rule library for adjusting grammar mistakes in the text and segmenting sentence granularity according to grammar rules through conventional natural language processing (NLP) tools; the grammar rule library includes rules for describing symptoms in children's cases and terms related to drug-induced liver injury events;

[0131] The medical fact and event extraction module includes a key value model (KV), a time event description model (TED), a medical research knowledge base, and a tagged corpus, and is used to extract medical facts and events in the text through conventional natural language processing (NLP) tools according to the mapping of time sequences, so as to obtain a structured database of clinical events with DILI as the target outcome;

[0132] The time information extraction module includes a time expression rule base, and is used to process the time expressions in the structured database of clinical events through conventional natural language processing (NLP) tools, construct a time development model of the child's condition (TDM-PC), so as to realize the reconstruction of structured data based on time sequences, and longitudinally outline each time scenario corresponding to the change of the child's condition by identifying DILI-related events and time key information in the electronic medical record text of inpatients.

[0133] Furthermore, the specific formula of the key value model includes:

[0134] ;

[0135] where, represents the key value extraction result of the th statement; represents the key parameter name of the th statement; represents the key parameter value of the th statement; represents the number of keywords in the statement; represents the set of key value extraction results; represents the item to which the key value belongs; represents the number of statements;

[0136] For example:

[0137] The statement is "On March 11, 2018, there was a sore throat, and azithromycin was given for anti-infection but ineffective. A blood routine examination was performed in our hospital: WBC: 2.13×10 9 / L, Hb: 102g / L, PLT: 177×10 9 / L";

[0138] After being mapped by the key value model, the following are extracted:

[0139] ;

[0140] ;

[0141] ;

[0142] ;

[0143] 。

[0144] Furthermore, the specific formula of the time event model includes:

[0145] ;

[0146] wherein, represents the th statement; represents the key time in the th statement; represents the key action in the th statement; represents the key effect in the th statement; represents probability;

[0147] For example:

[0148] The statement is "On March 11, 2018, sore throat occurred. Azithromycin was given for anti-infection but ineffective. Blood routine examination was performed in our hospital: WBC: 2.13×10 9 / L, Hb: 102g / L, PLT: 177×10 9 / L";

[0149] After being mapped by the time event model, the following are extracted:

[0150] ;

[0151] ;

[0152] ;

[0153] ;

[0154] wherein, represents the structured information finally formed after time event mapping.

[0155] Furthermore, as Figure 3 shown, the time development model of the child's condition (TDM-PC) is used to identify the key information of the child's condition in the structured database of clinical events, and based on the time expression, sort them in chronological order to reproduce the development process of the disease and clinical events; the key information of the child's condition includes: health status (initial symptoms and diagnosis) information, hospitalization diagnosis information, treatment process information, treatment medications and laboratory test information, clinical follow-up information, and reexamination information.

[0156] Figure 3The five stages of the patient's disease time development model are shown. By continuously recording these five stages, a structural reconstruction of all data (in the electronic medical records of hospitalized children) can be achieved as the child's disease progresses over time, including:

[0157] Phase 1 ) is the health status stage, which requires extracting the health status information of the children at the time of admission from the electronic medical records of the hospitalized children, including baseline symptoms, admission complaints and preliminary diagnosis information; the information at this stage will serve as the initial input of the children's status in the subsequent stages;

[0158] Phase 2 ) is the stage of hospital definite diagnosis, during which the child receives a definite diagnosis from the hospital and records the child's hospitalization diagnosis information;

[0159] Phase 3 ) is the treatment process stage, which records the treatment process information of the child;

[0160] Phase 4 ( ) is the drug treatment stage, which records the child's treatment medication and laboratory test information;

[0161] Phase 5 ) is the stage of discharge and prognosis follow-up for children, which records clinical follow-up information and re-examination information; when children are discharged from the hospital, their health status can be observed based on clinical follow-up information, and the long-term changes in the hospital's treatment effect can be evaluated;

[0162] Through the time development model of the child's condition, the data of the above stages are arranged and sorted in sequence according to the development of the child's condition, which helps to record and analyze the entire course of the hospitalized child from the time series dimension.

[0163] Furthermore, the method for evaluating the training effect of the MUIRFTI model is to quantify the improvement of text uncertainty by calculating the change in information entropy between the unstructured text before the MUIRFTI model conversion and the structured text after the conversion, so as to judge the conversion effect of the MUIRFTI model, such as Figure 4 As shown, specifically including:

[0164] Step a1: Calculate the total text information entropy of unstructured text based on the training set , the specific formula is:

[0165] ;

[0166] in, Indicates the total number of phrases in the total text; Represents the preset occurrence probability of a single phrase in the text; the total text consists of sentences, the sentences consist of sentence segments, the sentence segments consist of phrases, and the phrases consist of words;

[0167] Step a2, calculate the information entropy at the output stage of the text structuring module. The information entropy of the total text at this stage is equal to the sum of the information entropies of each sentence. The specific formula is:

[0168] ;

[0169] Wherein, Represents the number of sentences in the text; Represents the th sentence segment of the th sentence; Represents the preset occurrence probability of the th word in the

[0170] th

[0171] sentence; in this stage, is equal to

[0172] ;

[0173] Wherein, Represents the frequency of adjacent words appearing. The specific calculation formula is:

[0174] ;

[0175] Wherein, Represents the length of the word;

[0176] Calculate the change in information entropy between this stage and the information entropy in step a1 to evaluate the training effect of the text structuring module and the medical fact and event extraction module;

[0177] Step a4, calculate the information entropy at the output stage of the time information extraction module The specific formula is:

[0178] ;

[0179] Wherein, Represents the information entropy of the phrase at the th time node; Indicates the preset occurrence probability of the phrase in the text;

[0180] Calculate the change in information entropy at this stage and the information entropy in step a1 to evaluate the training effect of the MUIRFTI model.

[0181] Furthermore, the method for evaluating the training effect of the MUIRFTI model also includes:

[0182] Step a5: Based on the test set, calculate the information entropy of each stage according to the methods in steps a1 - a4, and compare them with the information entropy in steps a1 - a4 respectively; by comparing the change in information quotient before and after the time series reconstruction process, supplement and correct the strategy library (adjust the unstructured text regularization method), grammar rule library, medical research knowledge library, and labeled corpus according to the actual corpus of inpatient children's electronic medical record texts and the actual corpus of children's symptom descriptions, so as to improve the MUIRFTI model in combination with the writing characteristics of inpatient children's electronic medical record data.

[0183] Furthermore, the MUIRFTI model can be implemented through the statsmodels library of Python 3.11.0.

[0184] Furthermore, the information entropy calculation can be implemented through the Matlab R2021b program.

[0185] Step 3: Through the training set in the modeling dataset, train a machine learning prediction model based on different strategy combinations to predict whether drug-induced liver injury occurs; through the test set in the modeling dataset, test the machine learning prediction model; through evaluation indicators, evaluate the performance of the machine learning prediction model under different strategy combinations and parameters, and optimize to obtain the final machine learning prediction model; the independent variables of the feature set in the machine learning prediction model include the occurrence times of clinical events; the strategy combination includes three clinical event time series processing strategies for counting the occurrence times of each clinical event, specifically including: independent event set strategy, bundled event set strategy, and weighted event set strategy.

[0186] Furthermore, the machine learning prediction model includes classification models such as random forest (RF), support vector machine (SVM), and Naive Bayesian.

[0187] Furthermore, the machine learning prediction model can be implemented through the scikitlearn and SHAP packages of Python 3.10.2.

[0188] Furthermore, as Figure 5As shown, the independent event set strategy means that when calculating the number of occurrences of each clinical event, regardless of the temporal relationship with the occurrence of drug-induced liver injury, directly calculate the total number of occurrences of the clinical event within the preset time period before the onset of drug-induced liver injury. For example, "the symptom (examination) occurred 4 times in the first three days", then the total number of occurrences of this clinical event is recorded as 4;

[0189] The bundled event set strategy means that when calculating the number of occurrences of each clinical event, consider the temporal relationship with the occurrence of drug-induced liver injury, and calculate the number of occurrences of this clinical event within each day before the onset of drug-induced liver injury respectively;

[0190] The weighted event set strategy means that when calculating the number of occurrences of each clinical event, consider the temporal relationship with the occurrence of drug-induced liver injury, and assume that the clinical events that occur closer to the onset of drug-induced liver injury contribute more to the cause of drug-induced liver injury and are assigned greater weights. That is, assign a weight of 1 to the day before the onset of drug-induced liver injury, assign a weight of 1 / d to the dth day before the onset of drug-induced liver injury, and then calculate the total number of occurrences of this clinical event by weighted summation. The specific formula is:

[0191] ;

[0192] where represents the preset statistical number of days before the onset of drug-induced liver injury; represents the number of occurrences of this clinical event on the dth day before.

[0193] Furthermore, the clinical events include prescription drugs, symptoms (examinations), and laboratory examination (measurement) results.

[0194] Furthermore, in the training set, the number of days before the onset of drug-induced liver injury is divided into 10 time windows, namely the 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 14th, 21st, and 30th days, in order to conduct model training; the reason for this setting is that the closer to the onset of drug-induced liver injury, the more clinical events occur and the higher their importance, and the more intensive attention is required.

[0195] Furthermore, as Figure 6 shown, the weights of the weighted event set strategy also include time window weights, individual event weights, and time window and event type weights. The specific calculation methods are as follows:

[0196] The time window weight refers to calculating the weight only considering the distance from the event to the occurrence of the DILI outcome based on the distance from the DILI outcome event;

[0197] The individual event weight refers to calculating the weight only considering the importance of the clinical event (medication, laboratory examination, and symptom);

[0198] The time window and event type weights refer to calculating weights by simultaneously considering the distance of the event from the occurrence of the DILI outcome and the importance degree of different events for DILI prediction.

[0199] Further, the prediction results of the machine learning prediction model are verified by a 10-fold cross-validation method, and the parameters of the machine learning prediction model are optimized.

[0200] Further, the evaluation metrics include the area under the precision-recall curve (PR-AUC), precision, recall, and F1-score.

[0201] Step 4: Input the medical data of the tested child into the machine learning prediction model to predict whether the child will suffer from drug-induced liver injury.

[0202] Embodiment 2:

[0203] As Figure 7 shown, this embodiment provides a time-series-based data fusion and prediction system for drug-induced liver injury in children, including a data reception module, a data processing module, and a result generation module;

[0204] The data reception module is used to receive the medical data of the tested child;

[0205] The data processing module, based on the medical data of the tested child, predicts whether the child will suffer from drug-induced liver injury by the time-series-based data fusion and prediction method for drug-induced liver injury in children as described above;

[0206] The result generation module is used to send out the prediction results.

[0207] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for data fusion and prediction of drug-induced liver injury in children based on time series, characterized in that: include: Step 1: Collect the historical medical data of hospitalized children as a longitudinal data set, arrange them horizontally in chronological order, and build a clinical database; Perform data cleaning and standardization on clinical database; The sample size of the prediction model for drug-induced liver injury in children was calculated, and the research data set was extracted from the clinical database according to the inclusion and exclusion criteria: the research data set included a standard positive event group and a control group: The standard positive event group is selected from the longitudinal data set using the ADR causal evaluation method and the RUCAM scale to obtain historical medical data of children with drug-induced liver injury; The control group is randomly selected from the longitudinal data set according to a preset control ratio, and includes historical medical data of children with non-drug-induced liver injury who have used the target drug; Based on the relevant event phenotypes of drug-induced liver injury in children, the entity relationship of unstructured information is annotated, and missing values ​​and outliers of structured information are processed; Whether drug-induced liver injury occurs is used as a label, and the set is divided into a training set and a test set according to a preset ratio; Step 2: construct and train a MUIRFTI model for reconstructing unstructured text into a structured mapping based on temporal information; input the unstructured text of the research dataset into the trained MUIRFTI model, perform structural reconstruction according to the rules of each module in the MUIRFTI model, obtain the reconstructed structured text, and then fuse it with the original structured text in the research dataset to obtain a modeling dataset based on temporal fusion; The MUIRFTI model includes a text correction module, a text structuring module, a medical fact and event extraction module and a time information extraction module which are arranged in sequence; The text correction module includes a correction strategy library for correcting spelling errors in the text according to regular strategies through natural language processing tools; The text structuring module includes a grammatical rule library, which is used to adjust grammatical errors in the text and segment sentences according to grammatical rules through natural language processing tools; the grammatical rule library includes symptom description rules for children's cases and terms related to drug-induced liver injury events; The medical fact and event extraction module includes a key value model, a time event description model, a medical research knowledge base and a marked corpus, which is used to extract medical facts and events in the text according to the time sequence mapping through a natural language processing tool to obtain a structured database of clinical events with drug-induced liver injury as the target outcome; The time information extraction module includes a time expression rule library, which is used to process the time expression in the clinical event structured database through natural language processing tools to construct a time development model of the child's condition; The patient's condition time development model is used to identify key information of the patient's condition in the clinical event structured database, and sort it in chronological order based on time expression to reproduce the development process of the disease and clinical events; The key information of the child's condition includes: health status information, hospitalization diagnosis information, treatment process information, treatment medication and laboratory test information, clinical follow-up information and re-examination information; The temporal progression model of the patient's condition specifically includes five stages of continuous loop recording: Phase 1 is the health status phase, in which the health status information of the hospitalized children at the time of admission is extracted from the electronic medical record text of the hospitalized children, including baseline symptoms, admission complaints and preliminary diagnosis information; the information in this phase will serve as the initial input of the status of the children in the subsequent phases; Stage 2 is the stage of hospital definite diagnosis, during which the child receives a definite diagnosis from the hospital and the hospitalization diagnosis information of the child is recorded; Stage 3 is the treatment process stage, which records the child's treatment process information; Stage 4 is the drug treatment stage, which records the child's treatment medication and laboratory test information; Stage 5 is the discharge and prognosis follow-up stage for children, which records clinical follow-up information and re-examination information; Step 3: Train the machine learning prediction model through the training set in the modeling data set to predict whether drug-induced liver injury occurs; test the machine learning prediction model through the test set in the modeling data set; evaluate the performance of the machine learning prediction model under different strategy combinations and parameters through evaluation indicators, and optimize to obtain the final machine learning prediction model; The independent variables of the machine learning prediction model include the number of occurrences of clinical events; the strategy combination includes three clinical event temporal processing strategies: independent event set strategy, bundled event set strategy and weighted event set strategy; Step 4: Input the medical data of the tested child into the machine learning prediction model to predict whether the child will suffer from drug-induced liver injury.

2. The method according to claim 1, characterized in that The inclusion and exclusion criteria in step 1 include: Children with the preset years and whose first laboratory biochemical examination indicators at admission were within the normal range: alkaline phosphatase, alanine aminotransferase, aspartate aminotransferase, total bilirubin, and glutamyl transferase; Children with underlying diseases that affect liver-related functions at admission and those with abnormal liver function before using the target drug were excluded.

3. The method according to claim 1, characterized in that The relevant event phenotypes include symptom indicators, laboratory test indicators, demographic indicators, target drug use indicators and disease indicators; The symptom indicators include jaundice, nausea, vomiting, rash, itchy skin, abdominal discomfort, upper abdominal discomfort, right upper abdominal pain, fever, loss of appetite, decreased appetite, fatigue, weakness, mild jaundice, moderate jaundice, severe jaundice, and jaundice that has not completely subsided; The laboratory test indicators include alkaline phosphatase, alanine aminotransferase, aspartate aminotransferase, total bilirubin and glutamyl transferase; The demographic indicators include age, height, weight and gender; The target drug use indicators include dosage, dosage form, route of administration, whether to change medications, and combined medications; The disease indicators include discharge diagnoses and comorbidities.

4. The method according to claim 1, characterized in that: The unstructured information includes the medical record homepage information, admission records, daily medical course, handover records, ward rounds records and the first postoperative medical course in the electronic medical record data of hospitalized children.

5. The method according to claim 1, characterized in that The specific formula of the key value model includes: ; in, Indicates The key value extraction results of the statements; Indicates The key parameter name of the statement; Indicates The key parameter value of each statement; Indicates the number of keywords in a sentence; Represents a key value extraction result set; Indicates the project to which the key value belongs; Indicates the number of statements; The specific formula of the time event description model includes: ; in, Indicates statements; Indicates The key time in a sentence; Indicates The key action in a sentence; Indicates The key effect in each sentence; Represents probability.

6. The method according to claim 1, characterized in that The method for evaluating the training effect of the MUIRFTI model is to quantify the improvement of text uncertainty by calculating the change in information entropy between the unstructured text before the MUIRFTI model conversion and the structured text after the conversion, including: Step a1: Calculate the total text information entropy of unstructured text based on the training set , the specific formula is: ; in, Indicates the total number of phrases in the total text; Indicates the preset occurrence probability of a single phrase in a text; the total text is composed of sentences, the sentences are composed of segments, the segments are composed of phrases, and the phrases are composed of words; Step a2: Calculate the information entropy of the output stage of the text structuring module. The information entropy of the total text at this stage is equal to the sum of the information entropy of each sentence. The specific formula is: ; in, Indicates the number of sentences in the text; Indicates The first sentence of Sentences; Indicates The first sentence of The preset probability of occurrence of words; in this stage, equal ; Calculate the information entropy of this stage and the change in the information entropy in step a1 to evaluate the training effect of the text structuring module; Step a3: Calculate the information entropy of the output stage of the medical facts and events extraction module. Based on the formula in step a2, it also includes calculating the information entropy of phrases in the text. , the specific formula is: ; in, Indicates the frequency of occurrence of adjacent words. The specific calculation formula is: ; in, Indicates the length of the word; Calculate the change in information entropy between this stage and the information entropy in step a1 to evaluate the training effect of the text structuring module and the medical fact and event extraction module; Step a4: Calculate the information entropy of the output stage of the time information extraction module , the specific formula is: ; in, Indicated in The information entropy of the phrase at each time node; Indicates the preset probability of occurrence of the phrase in the text; The information entropy of this stage and the change in the information entropy in step a1 are calculated to evaluate the training effect of the MUIRFTI model.

7. The method according to claim 1, characterized in that The independent event set strategy refers to calculating the total number of occurrences of each clinical event within a preset time period before the onset of drug-induced liver injury when calculating the number of clinical events; The bundled event set strategy refers to calculating the number of occurrences of each clinical event on each day before the onset of drug-induced liver injury when calculating the number of occurrences of the clinical event; The weighted event set strategy refers to calculating the number of each clinical event by assigning a weight of 1 to the day before the onset of drug-induced liver injury, and a weight of 1 / d to the day before the onset of drug-induced liver injury, and then calculating the total number of occurrences of the clinical event by weighted summation. The specific formula is: ; in, Indicates the preset statistical days before the onset of drug-induced liver injury; Indicates the number of occurrences of the clinical event on the previous d days.

8. A time series-based data fusion and prediction system for drug-induced liver injury in children, characterized in that: It includes a data receiving module, a data processing module and a result generating module; The data receiving module is used to receive the medical data of the tested child; The data processing module predicts whether the child will suffer from drug-induced liver injury based on the medical data of the child being tested by using the method described in any one of claims 1 to 7; The result generation module is used to send the prediction results externally.

Citation Information

Patent Citations

  • Clinical decision-making system for acute and critical diseases based on multi-modal pre-training large model

    CN117174298A

  • Children drug-induced liver injury risk identification and prediction method and system

    CN117219275A