NomogramICU Geriatric Disease Risk Scoring Model, Device, and Methodology Integrating Medical Record Text
Patent Information
- Application Number
- CN202211300558.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-24
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-10-24
AI Technical Summary
这些重要有用的信息绝大多数以非结构化的方式呈现,由于缺乏明确的结构,使得这些文本只能由人解释和估计,而不能由任何计算机程序解释和定量评估
[0037](1)针对ICU入院前的临床情况会影响高龄患者的预后和结局,根据EHR 中对患者的病史文本记录内容,抽取主诉、家族史、现病史、入院时用药、既往病史、体检和社会史用于评估患者的入ICU前生理状态/疾病严重程度;
Smart Images

Figure CN115527678B_ABST
Abstract
Description
Technical Field
[0001] This application relates to medical information decision-making technology, and more particularly to a Nomogram ICU elderly risk scoring model, device, and method thereof based on pre-trained deep learning language models and machine learning models that fuse text. Background Technology
[0002] Due to increased life expectancy in many countries, the proportion of elderly patients in intensive care units (ICUs) is expected to increase significantly in the coming decades. Elderly patients (≥80 years old) have become a high-priority group in ICUs in recent years. Previous studies have shown that age at admission and severity of illness only partially explain the survival chances of elderly patients; their pre-ICU admission clinical condition significantly impacts prognosis and outcomes. Healthcare staff often rely on reviewing patients' baseline and pre-admission conditions and conducting interviews with patients and their families to obtain and assess their baseline and pre-admission status. These documents typically contain crucial information about the patient's current condition, symptoms, family history, medical history, procedures performed (e.g., X-rays, laboratory tests), and medications. This vital information is mostly presented in an unstructured manner; due to the lack of a clear structure, these texts can only be interpreted and estimated by humans, not by any computer program for quantitative evaluation. Therefore, extracting and refining the clinical information embedded in narrative reports as prior knowledge for decision-making can help healthcare professionals obtain key information better and faster, enabling a more comprehensive and multi-dimensional assessment of patients' conditions and providing higher quality care. Furthermore, although many studies have analyzed factors associated with increased mortality rates among elderly patients admitted to the ICU, hoping to use these factors to help physicians triage elderly patients and improve care during ICU stays, a more core and urgent need for healthcare professionals is a practical and feasible prognostic assessment tool for elderly patients. This tool should quantitatively assess disease severity and serve as a guide and basis for their decision-making process, treatment planning, and communication with patients and their families. Summary of the Invention
[0003] In view of the above problems, this application aims to propose a Nomogram ICU geriatric disease risk scoring model that integrates medical record text.
[0004] This application presents a Nomogram ICU geriatric disease risk scoring model that integrates medical record text, which includes: a data acquisition module, a data processing module, a BERT-like model calculation module, a multivariate logistic regression model, and a Nomogram output module;
[0005] The data acquisition module is used to obtain the patient's medical record text information before admission to the ICU and the numerical information collected on the first day of admission to the ICU. The medical record text information is unstructured and includes the chief complaint, family history, present illness, medications used upon admission, past medical history, physical examination and social history. The numerical information is structured and includes the Glasgow Coma Scale (GCS) score, whether vasopressors are used, CCI index, whether the patient is in absolute bed rest, whether mechanical ventilation is required, respiratory rate, whether the patient was admitted urgently, shock index, and whether palliative care is selected.
[0006] The data processing module is used to process medical record text information and numerical information; the processing of medical record text information includes: character / letter lowercase conversion, removal of special characters, sentence sliding segmentation, word segmentation, and sentence embedding representation. After processing, the medical record text information is represented as multiple vectors of predetermined length.
[0007] The BERT-like model calculation module calculates the preICU_risk_score, a disease severity assessment before admission to the ICU, based on a selected fine-tuned pre-trained clinical text-like model and multiple pre-defined length encoding vectors of standard medical record text information.
[0008] The multivariate logistic regression model uses preICU_risk_score and GCS score, whether vasopressors are used, CCI index, whether the patient is in absolute bed rest, whether mechanical ventilation is performed, respiratory rate, whether the patient is admitted in an emergency, shock index, and whether palliative care is selected as inputs to calculate the patient's probability of in-hospital mortality and risk level.
[0009] The Nomogram output module outputs a nomogram and disease severity score for the patient based on the input and parameters of the multivariate logistic regression model.
[0010] Preferably, the data acquisition module extracts the patient's chief complaint, family history, present illness history, medications used upon admission, past medical history, physical examination and social history from the electronic health record.
[0011] Preferably, for English medical records, the BERT-type model calculation module is one of Bio-ClinicalBERT, Clinical-Bigbird, Clinical-Longformer, and PubMedBERT models; for Chinese medical records, the BERT-type model calculation module is one of PCL-MedBERT, ChineseBLUE, ChineseEHRBert, and Medbert models.
[0012] Preferably, the risk levels include low risk, medium risk, and high risk;
[0013] Based on the calculated probability of hospitalized death, patients with a probability of 0 to 0.1 are considered low risk, those with a probability of 0.1 to 0.35 are considered medium risk, and those with a probability of 0.35 to 1 are considered high risk.
[0014] Preferably, the disease severity score is calculated according to the following formula:
[0015] Total points=10*preICU_risk_score-1.4116*GCS score+3.5544*vasopressor +1.769*CCI score+1.0344*respiratory rate+2.8915*admission type+18.2479*shock index+4.4857*mechanical ventilation+9.9566*activity status+9.6055*code status;
[0016] Among them, Total points is the severity score of the disease; GCS score is the GCS score; vasopressor indicates whether vasopressors are used (1 if yes, 0 if no); CCI score is the CCI index; respiratory rate is the respiratory rate; admission type indicates whether emergency admission is required (1 if yes, 0 if no); shock index is the shock index; mechanical ventilation indicates whether mechanical ventilation is used (1 if yes, 0 if no); activity status indicates whether absolute bed rest is required (1 if yes, 0 if no); code status indicates whether palliative care is selected (1 if yes, 0 if no).
[0017] Preferably, when segmenting sentences, a sliding window method is used for segmentation.
[0018] This application also aims to propose a Nomogram ICU geriatric disease risk scoring device that integrates medical record text, which is implemented by a computer and configured for the aforementioned Nomogram ICU geriatric disease risk scoring model that integrates medical record text.
[0019] This application also proposes a method for establishing a Nomogram ICU geriatric disease risk scoring model that integrates medical record text, which includes:
[0020] Data acquisition steps: Obtain information for model building from the patient's electronic health record, including the patient's medical record text information before admission to the ICU and the numerical information collected on the first day of admission to the ICU; extract the medical record text information, including chief complaint, family history, present illness, medications used upon admission, past medical history, physical examination and social history; extract the numerical information, including basic information, cognitive function, activity tolerance, vital signs, laboratory tests, treatment interventions, fluid output (i.e., urine output) and commonly used clinical scores;
[0021] Data processing steps: Cleaning, processing, and feature construction are performed on unstructured case text information and structured numerical information, and the dataset is segmented for subsequent model training and evaluation; unstructured data processing includes: lowercase character / letter conversion, removal of special characters, sentence sliding segmentation, word segmentation, and sentence embedding representation; structured data processing includes: outlier removal, data alignment, interpolation, statistical feature construction, and factor variable setting; dataset segmentation includes the preparation of development set, internal validation set, and time-series validation set;
[0022] Model development steps: Obtain the preICU_risk_score using a pre-trained deep learning model; Select important risk factors using a machine learning model; The Nomogram ICU geriatric disease risk scoring model, which integrates medical record text, combines the preICU_risk_score and the selected important risk factors, and trains it using a multivariate logistic regression model to obtain the patient's in-hospital mortality risk probability, risk level, nomogram, and disease severity score.
[0023] Model evaluation steps: Under different validation methods, select clinically relevant performance indicators and compare multiple baseline models for different scenarios / needs; Baseline models include: preICU_risk_score obtained solely from medical record text, important risk factors selected solely from structured data modeling, and clinically commonly used disease severity scoring systems; Validation types / methods include: internal validation and time-series validation; Performance evaluation includes ROC curves, calibration curves, DCA curves, and seven related evaluation indicators, which include: area under the receiver operating system curve, area under the curve enclosed by precision and recall, sensitivity, specificity, F1 score, accuracy and its corresponding 95% confidence interval, and Brier score.
[0024] Preferably, the optimal set of risk variables is obtained based on the entire training set: (1) Univariate logistic analysis was used to obtain the probability, 95% confidence interval and P value of each variable, and clinical variables with P value < 0.05 were selected; (2) The LASSO regression algorithm was used, and 5-fold cross-validation was performed to remove variables when selecting the lambda.1se parameter settings to achieve the most concise variable combination; (3) The important variables were selected again using the forward and backward stepwise algorithms based on the Akaike information criterion; (4) The selected variables were input into the multivariate logistic regression model to obtain the OR value, 95% CI, P value and variable coefficients, and important variables were selected; (5) Variables with small coefficients or high missing rates were excluded.
[0025] The optimal risk variables obtained included GCS score, use of vasopressors, CCI index, absolute bed rest, mechanical ventilation, respiratory rate, emergency admission, shock index, and palliative care. For the continuous variables, correlation analysis was performed to confirm whether they met the requirements for subsequent modeling.
[0026] Preferably, in the data processing step, the data processing module is used to simultaneously process the medical record text information of unstructured data and the numerical record information of structured data. The processed unstructured and structured data are matched and associated through the patient's unique identifier ID, enabling fusion analysis. The unstructured data is embedded and represented as a vector of a predetermined length, representing the information contained in a piece of text. For the structured data, the worst clinical value of each variable on the first day in the ICU is calculated. In addition, frailty index and geriatric nutritional risk index are constructed based on the structured data.
[0027] Preferably, in the model development step, a pre-trained clinical domain BERT model is used to perform downstream tasks such as assessing the severity of the patient's disease and the risk of in-hospital mortality. This aims to mimic a doctor's comprehensive and multi-dimensional perception and assessment of the patient's current condition by inquiring about their medical history. For English medical records, the pre-trained clinical domain BERT model is one of Bio-ClinicalBERT, Clinical-Bigbird, Clinical-Longformer, or PubMedBERT. For Chinese medical records, the pre-trained clinical domain BERT model is one of PCL-MedBERT, ChineseBLUE, ChineseEHRBert, or MedBERT. Through a large amount of medical record text, the pre-trained clinical domain BERT model is fine-tuned to assess the patient's risk of in-hospital mortality. The probability output of the pre-trained clinical domain BERT model serves as the preICU_risk_score, an assessment indicator measuring the patient's condition before admission to the ICU.
[0028] Preferably, in the model development step, a pre-trained BERT model for the clinical domain, used as a deep learning model, and a multivariate logistic regression model, used as a machine learning model, are fused together to achieve the fusion analysis of unstructured and structured data. This allows for the integration and quantification of the patient's chronic / historical condition before admission to the ICU and the acute condition on the day of admission, thereby constructing a scoring model that conforms to clinical behavior habits, is interpretable, and transparent.
[0029] Preferably, the model evaluation step utilizes the model evaluation module to fully evaluate the performance of the obtained risk scoring model;
[0030] The risk scoring model assesses the patient's prognosis based on the patient's condition before admission to the ICU and the condition on the first day of admission to the ICU.
[0031] By combining actual clinical needs, patients using risk scoring models, and existing clinical assessment methods, three usage scenarios were designed to compare the performance differences between the selected scoring system and the model in the corresponding scenarios:
[0032] (1) When a patient is admitted to the ICU, the severity of the patient’s disease and short-term prognosis are assessed based on the patient’s medical record text before admission to the ICU. That is, the preICU_risk_score obtained by the deep learning model is used to assess the patient’s prognosis.
[0033] (2) On the first day of a patient’s admission to the ICU, very little information is available before admission. Based on the patient information collected and measured on the day and the treatment interventions received by the patient, the severity of the patient’s disease and short-term prognosis are assessed. The multivariate logistic regression model of the machine learning model calculates the patient’s in-hospital mortality risk probability and risk level based on the inputs of GCS score, whether vasopressors are used, CCI index, whether the patient is in absolute bed rest, whether mechanical ventilation is performed, respiratory rate, whether the patient is admitted in an emergency, shock index, and whether palliative care is selected. This enables the assessment of the patient’s prognosis.
[0034] (3) Use commonly used clinical disease severity scoring systems such as SOFA and SAPS II scores to assess patient prognosis;
[0035] Based on the comparison of the three types of use cases, and combined with the performance and indicators of clinical concern, the model performance is fully quantitatively evaluated, compared and presented intuitively.
[0036] Technical advantages of the present invention:
[0037] (1) Since the clinical condition before admission to the ICU can affect the prognosis and outcome of elderly patients, the chief complaint, family history, present illness, medication at admission, past medical history, physical examination and social history were extracted from the patient's medical history text record in the EHR to assess the patient's physiological status / disease severity before admission to the ICU.
[0038] (2) Three novel pre-trained BERT models in the clinical domain (Bio-ClinicalBERT, Clinical-Bigbird and Clinical-Longformer) were fine-tuned to perform downstream tasks, namely, to characterize the severity of patients’ disease through information records before patients are admitted to the ICU, and the representation is named preICU risk score.
[0039] (3) The risk assessment model comprehensively incorporated variables related to the characteristics of elderly patients, including frailty, comorbidities, cognitive function (including delirium), nutrition, activity tolerance, and pre-ICU medical history. It also included key information collected on the first day of ICU admission, including vital signs, laboratory tests, urine output, and important treatment interventions received by the patient.
[0040] (4) The risk scoring model / system is presented using a nomogram, which facilitates understanding, absorption, and use by healthcare professionals. It can serve as a reference guide for them in wards or during consultations with patients and their families, and also facilitates the promotion and rapid deployment / embedding of the scoring into the information systems of other hospitals;
[0041] (5) This application is the first to integrate and analyze unstructured medical record text and structured collection and operation record data in the field of disease and health monitoring of elderly patients, realize the integration modeling of deep learning and machine learning methods, and finally build a disease risk scoring system that conforms to the medical and nursing behavior habits and has interpretability.
[0042] (6) This application uses comprehensive evaluation metrics, including ROC curve (measuring discrimination), calibration curve (measuring calibration) and DCA curve (measuring clinical net benefit), to compare and evaluate our model, baseline model and commonly used clinical scores. Our scoring system performs consistently and significantly better than the compared models and scores.
[0043] (7) This application encapsulates the disease risk scoring system into a user-friendly device, which only requires inputting the information required by the 10-class model to obtain the patient's Nomogram output, in-hospital mortality risk probability and current risk level (low, medium, high). Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating the method execution of this application;
[0045] Figure 2 This is a schematic diagram illustrating the application scenarios of a predictive scoring system.
[0046] Figure 3 A diagram illustrating the inclusion and exclusion criteria for patients;
[0047] Figure 4 A schematic diagram comparing the AUROC of clinical knowledge representation models under different hyperparameter settings;
[0048] Figure 5 A diagram illustrating the selection of risk variables using Lasso regression;
[0049] Figure 6 A schematic diagram of the correlation analysis for important risk variables;
[0050] Figure 7 A nomogram for assessing disease severity in elderly ICU patients, combining features of deep learning representations (preICU risk score) and schematic diagrams of clinical variables;
[0051] Figure 8 A schematic diagram of a nomogram illustrating the risk assessment of a patient;
[0052] Figure 9 A schematic diagram showing the comparison of ROC curves between the predictive model and the baseline model and clinical scores during internal and time-series validation.
[0053] Figure 10 A schematic diagram showing the comparison of calibration curves between the predictive model, the baseline model, and the clinical score on the internal validation set and the time-series validation set;
[0054] Figure 11 A schematic diagram showing the comparison of clinical decision curves between the predictive model, the baseline model, and clinical scores on the internal validation set and the time-series validation set;
[0055] Figure 12 This is a schematic diagram of the Nomogram risk scoring device used for early assessment of adverse outcomes in elderly patients in the ICU. Detailed Implementation
[0056] The present application will now be described in detail with reference to the accompanying drawings.
[0057] This invention proposes an approach to assess the severity of illness and in-hospital mortality risk in elderly patients based on their pre-ICU medical records, variables routinely collected on the day of ICU admission, and key treatment interventions received. Through data processing and model development modules, the medical record text is processed, modeled, and represented. Specifically, a pre-trained deep learning language BERT-like model is fine-tuned to complete downstream tasks to assess and quantify the severity of the patient's illness before ICU admission, represented by a preICU risk score. This process is similar to a doctor assessing the severity of a patient's illness by analyzing their pre-ICU admission information, mimicking a doctor's more comprehensive and holistic understanding, assessment, and quantification of the patient's current condition through history taking. The data processing and model development modules also handle structured data processing, modeling, and feature selection to obtain key risk factors. The preICU risks score, after semantic analysis and representation using deep learning, is combined with the key risk factors selected by machine learning to build a multivariate logistic regression model. The prediction model is interpreted and presented in the form of a nomogram, aligning with the behavioral habits of medical staff. Model evaluation covered internal and temporal validation, comparing our scoring system with three types of models / scoring systems: the predICU risk score (built using a deep learning model considering only the patient's pre-ICU condition), a machine learning model (built using only the important risk factors selected by the patient on the first day of ICU admission), and a commonly used clinical disease severity scoring system. The analysis focused on core clinically relevant aspects: discrimination, calibration, and net clinical benefit. Through this process, we ultimately achieved an assessment of disease severity and short-term prognosis (in-hospital outcome) in elderly patients requiring only 10 categories of information input. This provides a reference for physicians to rationally formulate / adjust treatment plans, while also avoiding the need for manual calculation of additional workload for medical staff. The 10 categories of information are: pre-ICU medical record text (all or some of the 7 relevant categories can be entered), GCS score, whether vasopressors were used, CCI index, respiratory rate, whether the admission type was emergency, shock index, whether mechanical ventilation was administered, whether absolute bed rest was required, and code status. Because it only requires 10 categories of information, and 5 of them are "yes / no" judgments, it can be assessed much more conveniently, quickly, and accurately compared to the SOFA score commonly used in clinical practice and the more complex Acute Physiology and Chronic Health Score. It avoids the problem of manually calculating additional workload for medical staff and is easy to deploy and promote in different medical institutions (tertiary / non-tertiary) and on different devices (computers / online calculators / tablets / manual assessment and recording).
[0058] This invention develops and validates a disease severity scoring system for elderly ICU patients based on deep learning and machine learning. In-hospital mortality rate is a key objective we use to assess disease severity. The score is designed by combining the patient's past medical history, description of admission status, and common measurements from the first day in the ICU. Past medical history and description of admission status are typically recorded in the clinical text. Common measurements should include risk variables relevant to elderly patients, such as frailty, cognitive function, and whether palliative care was administered. This study incorporates clinical text, patient baseline information, frailty level, cognitive function, vital signs, laboratory tests, treatments, and urine output information to construct an ICU elderly disease severity scoring system, also known as a "trustworthy AI agent." Figure 2 This presents an overview of the application scenarios for this rating.
[0059] This invention proposes a Nomogram ICU elderly risk scoring system, device, and method for integrating medical record text. Its specific implementation is as follows: Figure 1 As shown, it includes the following steps:
[0060] I. The data acquisition module process in this invention is as follows:
[0061] This study used electronic medical records of elderly patients aged 65 years and older admitted to Beth Israel Deaconess Medical Center (BIDMC) from 2001 to 2019. The data was extracted from MIMIC-III and the updated MIMIC-IV (Medical Information Mart for Intensive Care, MIMIC). Inclusion criteria included: no records of second hospital admissions or second ICU admissions (more than twice); ICU stay of less than 24 hours; and lack of essential measurement records (heart rate, respiratory rate, mean arterial pressure, systolic blood pressure, Glasgow Coma Scale score, body temperature, and oxygen saturation). Patients without a recorded past medical history (i.e., no clinical written records) were further excluded. Our study cohort screening process was as follows: Figure 3 As shown. Further, patients from 2001 to 2016 were selected as our development set for model training and internal validation, while the remaining patients served as a time-series validation set to evaluate the model's performance in future hospital use.
[0062] Considering the characteristics of the elderly patient population and the design objectives of early assessment and ease of clinical use, we collected commonly used textual records before ICU admission and data that were easily measured on the first day in the ICU. Unstructured data such as Figure 1The information shown includes the chief complaint, family history, present illness, medications used upon admission, past medical history, physical examination, and social history. For structured data, specific information includes: basic information (age, body mass index, sex, Charles's comorbidity index, number of days hospitalized before ICU admission, admission type); cognitive function (delirium, Glasgow Coma Scale score); activity tolerance (absolutely bedridden, able to sit, able to stand); vital signs (shock index, respiratory rate, heart rate, inhaled oxygen concentration, mean arterial pressure, systolic blood pressure, oxygen saturation, body temperature); and laboratory tests (albumin, alkaline phosphatase, alanine aminotransferase, anion gap, aspartate aminotransferase, base excess, bicarbonate, bilirubin, blood urea nitrogen, blood urea nitrogen to creatinine ratio, chloride, creatinine, estimated glomerular filtration rate, blood glucose, hematocrit, hemoglobin, international normalized ratio, lactate, lymphocytes, magnesium ions, neutrophils, neutrophil to lymphocyte ratio, partial pressure of carbon dioxide, partial pressure of oxygen, oxygenation index during mechanical ventilation, oxygenation index without mechanical ventilation, platelets, serum potassium, and prothrombin time). Partial thromboplastin time, serum sodium, white blood cell count, treatment interventions (mechanical ventilation, vasopressors, code status), and output (urine volume) were all measured. In addition, clinical scores, including SOFA and SAPSII, were calculated for comparison of predictive performance.
[0063] II. The data processing module process in this invention is as follows:
[0064] Unstructured and structured data were processed using different methods. The extracted text was converted to lowercase letters, and special characters such as "==", "--", and "\n" were removed. Considering the heterogeneity of sentence length, a sliding window approach was used to segment the text to avoid information loss due to excessively long sentences during model training. The window size and sliding stride were then used to investigate and obtain a set of sentences of appropriate length. Each sentence was segmented and further mapped to a corresponding ID based on the dictionary of the selected language model. Finally, it was embedded as a fixed-length vector representing the information contained in a sentence.
[0065] For structured data, values exceeding the physiological boundary values for each variable were removed. On the first day in the ICU, the clinically worst values for each variable were calculated, such as the highest and lowest values of laboratory tests, the lowest and average values of vital signs, and the presence or absence of treatment. Therefore, the data for each patient was consistent. Missing values were imputed in three ways: variables with a missing rate below 30% were calculated using the median of that variable; variables with a missing rate above 30% were imputed with zeros, and the corresponding variables were renamed "flag," such as lactate_flag, to indicate the presence or absence of a measurement; missing FiO2 values were imputed by 21%, and the indicator variable "FiO2_flag" was constructed. Frailty (Fibrillation Index, FI-LAB) is a frailty index assessing the risk of death in older adults, calculated from 21 routine laboratory data, SBP, and DBP. The Geriatric Nutritional Risk Index (GNRI) is a nutritional index assessing the risk of morbidity and death in older patients, calculated from albumin, weight, and height. Categorical variables were further defined as factor variables. After completing the above feature construction, the names used for model training and variable selection are presented according to type and source, as shown in Table 1.
[0066] The processed unstructured and structured data were matched and associated using patient IDs to construct the research dataset. We randomly selected 80% of the patient data from 2001 to 2016 as the training set, and the remaining 20% as the internal validation set. Patients from 2017 to 2019 were used as a separate time-series validation set. We calculated the missing rates of the structured numerical variables in the development set and the time-series validation set before feature construction, as shown in Table 2. Table 3 presents a baseline comparison of patients in the development set and the time-series validation set. This invention analyzed a total of 29,474 elderly critically ill patients. The total study cohort was divided into a development set of 26,473 patients (13.1% mortality rate) and a time-series validation set of 3,001 patients (12.0% mortality rate).
[0067] Table 1. Summary of the names of the included variables
[0068]
[0069]
[0070] Table 2. Missing Feature Information for Development Set and Temporal Validation Set
[0071]
[0072]
[0073]
[0074] Table 3. Comparison of patient baselines in the development set and the temporal validation set.
[0075]
[0076]
[0077]
[0078]
[0079] III. The model development module process in this invention is as follows:
[0080] In this study, we selected three novel pre-trained BERT models (Bidirectional Encoder Representations from Transformers) for our downstream task: characterizing the severity of a patient's disease based on their pre-ICU admission information (mimicking a doctor's more comprehensive and holistic understanding and assessment of the patient's current condition through history taking) – Bio-ClinicalBERT, Clinical-Bigbird, and Clinical-Longformer. Clinical-Bigbird and Clinical-Longformer were designed to reduce the significant memory consumption caused by Full Self-attention through sparse attention mechanisms and enhance the model's ability to model long-term sequence dependencies. We fine-tuned these three pre-trained models based on clinical texts from 85.7% of patients in the training set to assess in-hospital mortality risk. Data from an additional 14.3% of patients was used to evaluate the model's predictive performance under different hyperparameter combinations, while avoiding overfitting. The fine-tuning process was performed on two parallel GPUs. We segmented each patient's text into 120 or 240 words using a sliding step of 10 or 20 words, forming the segmented corpus dataset. Based on fine-tuning methods and some preliminary experiments, we set the epoch and batch size to 2 and 12, respectively, selected the Adam optimizer, and varied the learning rate using the `get_linear_schedule_with_warmup` method. We explored learning rates ranging from 1e-5 to 5e-5 (with intervals of 1e-5). Table 4 details our comparison of the performance of the three aforementioned clinical language models under different hyperparameter settings on internal and temporal validation sets. Figure 4 shows the comparison of the model's AUROC with 95% CI. Considering the model's performance in predicting past and future use, we selected the Clinical-Longformer with a window size of 240 and a learning rate of 1e-5 to obtain the `preICU_risk_score` for each patient.
[0081] Table 4. Detailed comparison of three deep learning clinical BERT models
[0082]
[0083]
[0084]
[0085]
[0086] We chose to obtain the optimal set of risk variables based on the entire training set. We first used univariate logistic regression analysis to obtain the odds (ORs), 95% confidence intervals (CIs), and p-values for each variable, and selected useful clinical variables for assessing disease severity in elderly patients. Then, we used the LASSO regression algorithm with 5-fold cross-validation to further screen key variables. We then used a forward and backward stepwise algorithm based on the Akaike Information Criterion (AIC) to select important variables again. Finally, we input the selected variables into a multivariate LR model to obtain ORs, 95% CIs, p-values, and variable coefficients. To ensure the generalizability and ease of use of our risk score, variables with small coefficients or high missing rates were excluded in subsequent studies. Furthermore, the correlation of continuous variables needed to be checked before modeling the predictive score. In Table 5, univariate logistic regression showed that 10 independent risk factors were not significantly associated with assessing the risk of in-hospital mortality. The process of finding the optimal lambda using LASSO regression is as follows: Figure 5 As shown. When selecting lambda.1se, a LASSO LR model with 5-fold cross-validation was used to remove variables to achieve the most concise variable combination, resulting in 25 variables with non-zero coefficients. Table 6 lists the variables selected using LASSO regression, forward and backward stepwise algorithms, and their corresponding model coefficients. We ultimately adopted nine key important variables (i.e., important risk factors) with absolute coefficients greater than 0.06 and easy to collect, facilitating rapid use by healthcare personnel. These include the lowest GCS score, vasopressor (yes / no), CCI index, average respiratory rate, mechanical ventilation (yes / no), admission type (emergency / non-emergency), hospital code status (yes / no palliative care), optimal activity (absolute bed rest / paralyzed), and shock index. For the continuous variables, the correlation analysis results also meet the requirements for subsequent modeling, such as... Figure 6 As shown.
[0087] Table 5. Univariate analysis of all variables
[0088]
[0089]
[0090]
[0091]
[0092] Table 6. Selection of risk variables using Lasso regression model and stepwise algorithm
[0093]
[0094]
[0095] We used a deep learning-based model to obtain the risk probability based on pre-ICU admission records, named `preICU_risk_score`. If a patient received multiple predictions due to sentence segmentation, the average probability was used as the final result. Combining the `preICU_risk_score` (the probability value was multiplied by 10 for easier calculation and presentation) and nine selected key important variables, we applied a multivariable logistic regression model on the training set to obtain the final predictive model. Based on the results of the developed model, a nomogram was used to visually describe and present the fitted model. Ultimately, we obtained a scoring system that can assess the disease severity of elderly patients early on the first day of ICU care. In Table 7, univariate and multivariate analyses showed that they had a significant impact on hospitalization outcomes in elderly patients (P < 0.001 for all variables except admission type, which was 0.022). Based on the final trained model, we obtained a nomogram, see [see Table 7]. Figure 7 Thresholds for low, medium, and high risk levels can be set based on the patient's risk probability (model output) and the doctor's requirements for accuracy and sensitivity. Figure 8 This example demonstrates the calculation process and final risk level of a patient's disease severity score obtained after inputting data across 10 dimensions. The formula for calculating the disease severity score for an elderly ICU patient is as follows:
[0096] Total points=10*preICU_risk_score-1.4116*GCS score(min)+ 3.5544*vasopressor(use=1,no use=0)+1.769*CCI score+1.0344*respiratory rate(bpm)+2.8915*admission type(urgent=1,no urgent=0)+18.2479*shock index (bpm / mmHg)+4.4857*mechanical ventilation(received=1,not received=0)+9.9566*activitystatus(bed=1,not bed=0)+9.6055*code status(received=1,not received=0)
[0097] Table 7. Univariate and multivariate analyses of selected risk factors for early assessment of mortality
[0098]
[0099] IV. The model evaluation module process in this invention is as follows:
[0100] Our clinical nomogram model was compared with three other models: a deep learning model developed from unstructured data (clinical text) assessing pre-ICU disease severity (preICU risk score); a multivariate logistic regression model developed from structured data using a machine learning model to select nine key important variables; and the SOFA (Sequential Organ Failure Assessment) and SAPSII (Simplified Acute Physiology Score II) scores commonly used in clinical practice. The ROC (Receiver Operating Characteristic) curve was used to evaluate the model's ability to differentiate between survivors and non-survivors. The calibration curve employed 500 resampling iterations to assess the consistency between actual and predicted risk probabilities. The DCA (Decision Curve Analysis) curve evaluated the model's clinical applicability by assessing the net gain obtained when varying probability thresholds. All models were evaluated for performance through internal and time-series validation. Comparison metrics included AUROC (area under the receiver operating curve), AUPRC (area under the curve between precision and recall), sensitivity, specificity, F1 score, and accuracy, along with their corresponding 95% CI (confidence interval) values and Brier score, with 95% of CIs obtained using 500 bootstrap resampling.
[0101] We compared four application scenarios / scoring systems, including our risk score [here named Nomogram score] (clinical notes and variable fusion), structured data scoring (clinical variables), preICU_risk_score (clinical notes), SAPSII, and SOFA (commonly used clinical scores). Figure 9In (a) and (b), ROC curves were used to compare the discriminative power of the four models in internal and time-series validation, with the Nomogram score significantly outperforming the other scores. Table 8 shows a detailed comparison of the metrics using 500 bootstrap iterations. For internal validation, the Nomogram score had an AUROC (95% CI) of 0.84 (0.816–0.861), the structured data score was 0.798 (0.77–0.821), the preICU_risk_score was 0.767 (0.741–0.794), the SAPSII score was 0.764 (0.738–0.792), while the SOFA score was only 0.697 (0.667–0.732). For time-series validation, AUROCs (95% CI) were 0.871 (0.851–0.887), 0.814 (0.791–0.837), 0.806 (0.777–0.831), 0.762 (0.736–0.79), and 0.746 (0.717–0.774), respectively. Figure 10 In (a) and (b), calibration performance is shown using calibration curves. In both internal and time-series validation, the Nomogram score shows an acceptable Brier score (internal: 1.097, time-series: 1.062). Figure 11 In (a) and (b), the Nomogramscore also showed a broader net clinical decision benefit in DCA curves during internal and time-series validation compared to the other four scoring systems.
[0102] Table 8. Performance comparison of the predictive model, baseline model, and clinical scores in internal and time-series validation.
[0103]
[0104]
[0105] Ultimately, we integrated and packaged data acquisition, data processing, model computation, and model output into a user-friendly and easy-to-use device, such as... Figure 12As shown, the system inputs the patient's pre-ICU medical record text (chief complaint, family history, present illness, medications used upon admission, past medical history, physical examination, social history [partial or complete content can be selectively input depending on the actual situation]) and numerical / type information from the first day of ICU admission (GCS score, whether vasopressors were used, CCI index, whether the patient was strictly bedridden, whether mechanical ventilation was performed, respiratory rate, whether the patient was admitted urgently, shock index, whether palliative care was selected). The system then processes, calculates, analyzes, and outputs the patient's Nomogram, in-hospital mortality risk, and risk level (low / medium / high). The risk level is determined by the physician using the system, using a threshold range for different levels, or the system's default range can be selected (low: 0–0.1, medium: 0.1–0.35, high: 0.35–1). Its minimal input, convenient operation, and transparent calculation process make the Nomogram ICU geriatric risk scoring system and device, which integrates medical record text, easy to promote and use. In addition, taking into account the input of medical record text in Chinese format and the personal preferences of different users, we provide scoring systems for English medical record input that are pre-trained using Bio-ClinicalBERT, Clinical-Bigbird, Clinical-Longformer, and PubMedBERT, and scoring systems for Chinese medical record input that are pre-trained using PCL-MedBERT, ChineseBLUE, ChineseEHRBert, and MedBERT.
[0106] Unless otherwise defined, all technical and / or scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this invention relates. The materials, methods, and embodiments mentioned in this application are illustrative only and not restrictive.
[0107] Although the present invention has been described in conjunction with specific embodiments, those skilled in the art can make appropriate substitutions, modifications and changes within the inventive spirit of this application, and such substitutions, modifications and changes still fall within the protection scope of this application.
Claims
1. A Nomogram ICU geriatric disease risk scoring model integrating medical record text, comprising: Data acquisition module, data processing module, BERT-like model calculation module, multivariate logistic regression model, and Nomogram output module; The data acquisition module is used to obtain the patient's medical record text information before admission to the ICU and the numerical information collected on the first day of admission to the ICU. The medical record text information is unstructured and includes the chief complaint, family history, present illness, medications used upon admission, past medical history, physical examination and social history. The numerical information is structured and includes the Glasgow Coma Scale (GCS) score, whether vasopressors are used, CCI index, whether the patient is in absolute bed rest, whether mechanical ventilation is required, respiratory rate, whether the patient was admitted urgently, shock index, and whether palliative care is selected. The data processing module is used to process medical record text and numerical information; The processing of medical record text information includes: lowercase conversion of characters / letters, removal of special characters, sliding segmentation of sentences, word segmentation, and embedding representation of sentences. After processing, the medical record text information is represented as multiple vectors of predetermined length. The BERT-like model calculation module is based on a selected fine-tuned pre-trained clinical text-based model. It calculates the preICU_risk_score, a disease severity assessment before admission to the ICU, based on multiple pre-defined length encoding vectors of standard medical record text information. This allows the module to mimic a doctor's comprehensive and three-dimensional perception and assessment of a patient's current condition by inquiring about their medical history. The multivariate logistic regression model uses preICU_risk_score and GCS score, whether vasopressors are used, CCI index, whether the patient is in absolute bed rest, whether mechanical ventilation is performed, respiratory rate, whether the patient is admitted in an emergency, shock index, and whether palliative care is selected as inputs to calculate the patient's probability of in-hospital mortality and risk level. The Nomogram output module outputs a nomogram and disease severity score for the patient based on the input and parameters of the multivariate logistic regression model. For English medical records, the BERT-based model calculation module is one of Bio-ClinicalBERT, Clinical-Bigbird, Clinical-Longformer, or PubMedBERT; for Chinese medical records, the BERT-based model calculation module is one of PCL-MedBERT, ChineseBLUE, ChineseEHRBert, or Medbert. Disease severity scoring is based on the following formula: Total points = 10 preICU_risk_score-1.4116 GCS score + 3.5544 vasopressor +1.769 CCI score+1.0344 respiratory rate+2.8915 admission type+ 18.2479 shock index+4.4857 mechanical ventilation+9.9566 activity status+ 9.6055 code status; Among them, Total points is the severity score of the disease; GCS score is the GCS score; vasopressor indicates whether vasopressors are used (1 if yes, 0 if no); CCI score is the CCI index; respiratory rate is the respiratory rate; admission type indicates whether emergency admission is required (1 if yes, 0 if no); shock index is the shock index; mechanical ventilation indicates whether mechanical ventilation is required (1 if yes, 0 if no); activity status indicates whether absolute bed rest is required (1 if yes, 0 if no); code status indicates whether palliative care is selected (1 if yes, 0 if no).
2. The Nomogram ICU Geriatric Disease Risk Scoring Model integrating medical record text as described in claim 1, characterized in that: The data acquisition module extracts the patient's chief complaint, family history, present illness, medications used upon admission, past medical history, physical examination and social history from the electronic health record.
3. The Nomogram ICU Geriatric Disease Risk Scoring Model integrating medical record text as described in claim 1, characterized in that: Risk levels include low risk, medium risk, and high risk; Based on the calculated probability of hospitalized death, patients with a probability of 0-0.1 are considered low-risk, those with a probability of 0.1-0.35 are considered medium-risk, and those with a probability of 0.35-1 are considered high-risk.
4. The Nomogram ICU Geriatric Disease Risk Scoring Model that integrates medical record text according to claim 1, characterized in that: When segmenting sentences using a sliding window, a sliding window method is used for segmentation.
5. A Nomogram ICU geriatric disease risk scoring device that integrates medical record text, implemented by a computer, the device being configured to run the Nomogram ICU geriatric disease risk scoring model that integrates medical record text as described in any one of claims 1-4.
6. A method for establishing a Nomogram ICU geriatric disease risk scoring model that integrates medical record text, comprising: Data acquisition steps; Information for model building is obtained from the patient's electronic health record, including the patient's medical record text information before admission to the ICU and the numerical information collected on the first day of admission to the ICU; the medical record text information is extracted, including chief complaint, family history, present illness, medications at admission, past medical history, physical examination and social history; the numerical information is extracted, including basic information, cognitive function, activity tolerance, vital signs, laboratory tests, treatment interventions, fluid output (i.e., urine output) and commonly used clinical scores; Data processing steps; The process involves cleaning, processing, and feature construction of unstructured case text information and structured numerical information, and then splitting the dataset for subsequent model training and evaluation. Processing of unstructured data includes: lowercase conversion of characters / letters, removal of special characters, sentence sliding segmentation, word segmentation, and sentence embedding representation; processing of structured data includes: removal of outliers, data alignment, interpolation, construction of statistical features, and setting of factor variables; dataset segmentation includes the preparation of development sets, internal validation sets, and time-series validation sets; Model development steps: Obtain the preICU_risk_score using a pre-trained deep learning model; Train a machine learning model to select key risk factors; Integrate the medical record text into a nomogram. The ICU geriatric disease risk scoring model integrates the preICU_risk_score with selected key risk factors and is trained using a multivariate logistic regression model to obtain the patient's in-hospital mortality risk probability, risk level, nomogram, and disease severity score. A pre-trained clinical domain BERT model is used to perform downstream tasks assessing patient disease severity and in-hospital mortality risk, mimicking a physician's comprehensive and multi-dimensional perception and assessment of the patient's current condition through medical history taking. For English medical records, the pre-trained clinical domain BERT model is one of Bio-ClinicalBERT, Clinical-Bigbird, Clinical-Longformer, or PubMedBERT; for Chinese medical records, it is one of PCL-MedBERT, ChineseBLUE, ChineseEHRBert, or MedBERT. Through extensive medical record text analysis, the pre-trained clinical domain BERT model is fine-tuned to assess the patient's in-hospital mortality risk. The probability output of the pre-trained clinical domain BERT model serves as the preICU_risk_score, an assessment indicator of the patient's pre-ICU condition. The disease severity score is calculated using the following formula: Total points = 10 preICU_risk_score-1.4116 GCS score + 3.5544 vasopressor +1.769 CCI score+1.0344 respiratory rate+2.8915 admission type+ 18.2479 shock index+4.4857 mechanical ventilation+9.9566 activity status+ 9.6055 code status; Among them, Total points is the severity score of the disease; GCS score is the GCS score; vasopressor indicates whether vasopressors are used (1 for yes, 0 for no); CCI score is the CCI index; respiratory rate is the respiratory rate; admission type indicates whether emergency admission is required (1 for yes, 0 for no); shock index is the shock index; mechanical ventilation indicates whether mechanical ventilation is required (1 for yes, 0 for no); activity status indicates whether absolute bed rest is required (1 for yes, 0 for no); code status indicates whether palliative care is selected (1 for yes, 0 for no). Model evaluation steps: Under different validation methods, select clinically relevant performance indicators and compare multiple baseline models for different scenarios / needs; Baseline models include: preICU_risk_score obtained solely from medical record text, important risk factors selected solely from structured data modeling, and clinically commonly used disease severity scoring systems; Validation types / methods include: internal validation and time-series validation; Performance evaluation includes ROC curves, calibration curves, DCA curves, and seven related evaluation indicators, which include: area under the receiver operating procedure (ROC) curve, area under the curve enclosed by precision and recall, sensitivity, specificity, F1 score, accuracy and its corresponding 95% confidence interval, and Brier score.
7. The method for establishing a Nomogram ICU geriatric disease risk scoring model integrating medical record text according to claim 6, characterized in that: To obtain the optimal set of risk variables based on the entire training set: (1) Univariate logistic analysis was used to obtain the probability, 95% confidence interval and P value of each variable, and clinical variables with P value < 0.05 were selected; (2) The LASSO regression algorithm was used, and 5-fold cross-validation was performed to remove variables when selecting the lambda.1se parameter settings to achieve the most concise variable combination; (3) The forward and backward stepwise algorithms based on the Akaike information criterion were used to select important variables again; (4) The selected variables were input into the multivariate logistic regression model to obtain the OR value, 95% CI, P value and variable coefficients, and important variables were selected; (5) Variables with small coefficients or high missing rates were excluded. The optimal risk variables obtained included GCS score, use of vasopressors, CCI index, absolute bed rest, mechanical ventilation, respiratory rate, emergency admission, shock index, and palliative care. For the continuous variables, correlation analysis was performed to confirm whether they met the requirements for subsequent modeling.
8. The method according to claim 6, characterized in that: In the data processing step, the data processing module is used to simultaneously process the medical record text information of unstructured data and the numerical record information of structured data. The processed unstructured and structured data are matched and associated through the patient's unique identifier ID, so that they can be fused and analyzed. The unstructured data is embedded and represented as a vector of a predetermined length, representing the information contained in a piece of text. For structured data, the worst clinical value for each variable on the first day in the ICU was calculated; in addition, frailty index and geriatric nutritional risk index were constructed based on structured data.
9. The method according to claim 6, characterized in that: In the model development process, a pre-trained BERT model for the clinical domain, used as a deep learning model, and a multivariate logistic regression model, used as a machine learning model, are fused together to achieve the fusion analysis of unstructured and structured data. This allows for the integration and quantification of the patient's chronic / historical condition before admission to the ICU and the acute condition on the day of admission, thereby constructing a scoring model that is consistent with clinical behavior, interpretable, and transparent.
10. The method according to claim 6, characterized in that: The model evaluation step utilizes the model evaluation module to fully evaluate the performance of the obtained risk scoring model. The risk scoring model assesses the patient's prognosis based on the patient's condition before admission to the ICU and the condition on the first day of admission to the ICU. By combining actual clinical needs, patients using risk scoring models, and existing clinical assessment methods, three usage scenarios were designed to compare the performance differences between the selected scoring system and the model in the corresponding scenarios: (1) When a patient is admitted to the ICU, the severity of the patient’s disease and short-term prognosis are assessed based on the patient’s medical record text before admission to the ICU. That is, the preICU_risk_score obtained by the deep learning model is used to assess the patient’s prognosis. (2) On the first day of the patient’s admission to the ICU, very little information was obtained before the admission. Based on the patient information collected and measured on the day and the treatment interventions received by the patient, the severity of the patient’s disease and short-term prognosis were assessed. The multivariate logistic regression model of the machine learning model was used to calculate the patient’s in-hospital mortality risk probability and risk level based on the inputs of GCS score, whether vasopressors were used, CCI index, whether the patient was in bed, whether mechanical ventilation was performed, respiratory rate, whether the patient was admitted in an emergency, shock index, and whether palliative care was selected. This enabled the assessment of the patient’s prognosis. (3) Use commonly used clinical disease severity scoring systems such as SOFA and SAPS II scores to assess patient prognosis; Based on the comparison of the three types of use cases, and combined with the performance and indicators of clinical concern, the model performance is fully quantitatively evaluated, compared and presented intuitively.
Citation Information
Patent Citations
Inplausible fair early-stage death risk assessment model and device for severe elderly patients and establishment method of interpretable fair early-stage death risk assessment model and device
CN115101199A
Clinical endpoint adjudication system and method
WO2022161925A1