Prognosis risk prediction model construction system and method based on chronic disease real world data
By constructing a prognostic risk prediction model system based on real-world data of chronic diseases, the problems of data processing and model adaptability for chronic liver disease, diabetes, and hypertension were solved. This system achieved high-quality data collection, multi-model fusion, and individualized risk calculation, improving the robustness and clinical applicability of the model and supporting individualized intervention programs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI ZESHENXIN MEDICAL TECHNOLOGY CO LTD
- Filing Date
- 2026-04-28
- Publication Date
- 2026-07-03
AI Technical Summary
Existing technologies lack standardized solutions for real-world data processing of chronic liver disease, diabetes, and hypertension. Multi-source heterogeneous data are of low quality, models do not consider comorbidity features, prediction results have poor clinical adaptability, model training processes are not optimized for bias, validation dimensions are limited, individualized risk quantification and clinical interpretation are lacking, model deployment efficiency is low, and individualized intervention plans are not supported.
We construct a prognostic risk prediction model system based on real-world data of chronic diseases, including multi-source data collection and access, data standardization and quality control, multi-dimensional feature engineering and screening, comorbidity feature fusion and dataset construction, prognostic risk prediction model training and optimization, multi-dimensional model validation and evaluation, model deployment and risk stratification output, and adopting technical means such as clinical-data dual-dimensional feature screening, multi-model fusion, hyperparameter optimization, and lightweight deployment.
It improved data quality and the clinical suitability of the model, achieved accurate calculation and interpretability of individualized risk values, enhanced the robustness and generalization ability of the model, supported the clinical application of individualized intervention programs, and improved the translation efficiency of the model from the laboratory to the clinic.
Smart Images

Figure CN122337671A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical big data analysis and chronic disease prognosis prediction technology, specifically to a system and method for constructing a prognostic risk prediction model based on real-world data of chronic diseases. Background Technology
[0002] With an aging population and changing lifestyles, chronic liver disease, diabetes, and hypertension have become prevalent chronic non-communicable diseases in my country, and these three conditions often co-occur. Their insidious disease progression and complex prognostic factors pose significant challenges to clinical diagnosis and chronic disease management. Real-world data, due to its broad coverage and close relevance to actual clinical scenarios, has become a crucial data foundation for constructing chronic disease prognostic risk prediction models. However, current research and applications of chronic disease prognostic prediction based on real-world data still face numerous technical challenges.
[0003] Existing data processing systems lack standardized protocols specifically for chronic liver disease, diabetes, and hypertension. Multi-source, heterogeneous medical data suffers from inconsistent indicators, high rates of missing and outliers, and data redundancy, directly leading to low-quality data sources for model construction. Traditional prognostic models are often developed for single diseases, failing to consider the comorbidity among chronic liver disease, diabetes, and hypertension, thus becoming disconnected from the real-world scenario of patients with multiple diseases, resulting in poor clinical adaptability of model predictions. Feature selection often employs purely data-driven statistical methods, ignoring the guidance of clinical treatment guidelines, leading to features lacking clinical interpretability and making them difficult for clinicians to accept and apply. Model training does not specifically address biases in real-world data, and validation dimensions are limited to internal cross-validation, lacking external multi-center real-world cohort validation, resulting in insufficient robustness and generalization ability. After model construction, a standardized deployment and output system is lacking, failing to achieve individualized risk quantification and clinically interpretable risk stratification, resulting in low efficiency in translating models from the laboratory to clinical practice and failing to provide effective support for developing individualized intervention plans.
[0004] In addition, most existing model building systems are general-purpose big data analysis systems that do not have dedicated functional modules designed for the disease characteristics of chronic liver disease, diabetes, and hypertension. These systems suffer from problems such as general module functions, lack of technical details, and weak integration with clinical diagnosis and treatment processes. Summary of the Invention
[0005] The purpose of this invention is to provide a system and method for constructing a prognostic risk prediction model based on real-world data of chronic diseases, so as to solve the problems existing in the prior art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a prognostic risk prediction model construction system based on real-world data of chronic diseases, including a multi-source real-world data acquisition and access module, a data standardization and quality control module, a multi-dimensional feature engineering and screening module, a comorbidity feature fusion and dataset construction module, a prognostic risk prediction model training and optimization module, a multi-dimensional model verification and evaluation module, and a model deployment and risk stratification output module.
[0007] The multi-source real-world data acquisition and access module enables comprehensive acquisition and standardized protocol access of multi-source heterogeneous medical data related to chronic liver disease, diabetes, and hypertension, ensuring the integrity of the original data and the compatibility of the access.
[0008] The data standardization and quality control module is based on the clinical diagnosis and treatment guidelines for chronic diseases and medical data standards. It standardizes the collected raw data and improves data quality through multi-dimensional quality control methods.
[0009] The multi-dimensional feature engineering and screening module extracts multiple types of features from standardized data. After preprocessing, it uses a clinical-data dual-dimensional feature screening method to screen out core features that have both clinical significance and data distinguishability.
[0010] The comorbidity feature fusion and dataset construction module targets single diseases and comorbidity scenarios of chronic liver disease, diabetes, and hypertension. It constructs single-disease feature sets and mines the correlation between comorbidity features to complete multimodal comorbidity feature fusion. At the same time, it divides the training, validation, and test datasets according to stratified sampling.
[0011] The prognostic risk prediction model training and optimization module constructs multiple types of basic prediction models, uses a weighted fusion method for model training, improves model performance through hyperparameter optimization, and calculates individualized prognostic risk values by combining quantitative formulas.
[0012] The multi-dimensional model validation and evaluation module comprehensively evaluates the model from three dimensions: discrimination ability, calibration ability, and clinical applicability, through internal cross-validation and external real-world cohort validation.
[0013] The model deployment and risk stratification output module performs lightweight processing on the trained and optimized model and adapts it to the clinical information system, enabling real-time calculation of individualized prognostic risk values for patients, and completing risk stratification and visualization output based on clinical guidelines.
[0014] Furthermore, the multi-source real-world data acquisition and access module includes a data acquisition unit, a multi-protocol access unit, and a data caching unit;
[0015] The data acquisition unit collects multi-source real-world data on chronic liver disease, diabetes, and hypertension, covering demographic information, clinical symptoms, laboratory tests, treatment interventions, follow-up outcomes, and medical insurance records. Data sources include hospital electronic medical record systems, laboratory test systems, chronic disease follow-up management systems, and medical insurance information systems. The multi-protocol access unit supports common medical data protocols such as HL7, FHIR, DICOM, and SQL, enabling seamless data access from different heterogeneous medical information systems. The data caching unit uses a distributed caching mechanism to temporarily store raw real-world data and establishes a data acquisition log to record the data acquisition time, source, type, and completeness.
[0016] Furthermore, the data standardization and quality control module includes a data standardization unit, a missing value processing unit, an outlier detection and correction unit, and a data deduplication and integration unit;
[0017] The data standardization unit, based on the clinical diagnosis and treatment guidelines for chronic diseases and the HL7 international medical data standard, standardizes the indicator names, units of measurement, reference ranges, disease diagnosis codes, and surgical operation codes of the data, transforming unstructured clinical text data into structured data. The missing value processing unit uses interpolation based on data distribution characteristics to fill in continuous data, mode-based filling based on disease clinical characteristics to fill in categorized data, and mean-based filling based on the subgroups of patients with the same disease course and characteristics to fill in missing data for key clinical indicators. The outlier detection and correction unit uses an outlier detection algorithm combined with threshold values for chronic disease clinical indicators to detect outliers, and performs manual review and correction based on original medical records, marking and removing outliers that cannot be reviewed. The data deduplication and integration unit uses the patient's unique identifier as the core, combining medical insurance number, hospitalization number, and outpatient number for data association and matching, completing data deduplication and integrated integration of multi-source data for the same patient.
[0018] Furthermore, the multi-dimensional feature engineering and screening module includes a feature extraction unit, a feature preprocessing unit, and a clinical-data dual-dimensional feature screening unit;
[0019] The feature extraction unit extracts demographic features, clinical symptom features, laboratory test features, treatment intervention features, disease course features, and follow-up outcome features, including disease-specific features for chronic liver disease, diabetes, and hypertension. The feature preprocessing unit normalizes continuous features, performs one-hot encoding on categorical features, and divides time-series features into time windows and aggregates features. The clinical-data dual-dimensional feature screening unit first performs initial clinical screening based on chronic disease clinical treatment guidelines, then uses a feature comprehensive scoring formula for statistical screening, sets a feature comprehensive scoring threshold, and retains core features with scores higher than the threshold. The feature comprehensive scoring formula is as follows: ,in For feature comprehensive scoring, This is the clinical weighting coefficient. For statistical weighting coefficients, , and Based on the consensus of clinical experts on chronic diseases, it was determined that... Characteristic clinical scores ranging from 0 to 10. The feature statistics score is 0-10.
[0020] Furthermore, the comorbidity feature fusion and dataset construction module includes a single-disease feature set construction unit, a comorbidity association rule mining unit, a multimodal comorbidity feature fusion unit, and a training, validation, and test set partitioning unit;
[0021] The single-disease feature set construction unit classifies core features into three single-disease-specific feature sets: chronic liver disease, diabetes, and hypertension. The comorbidity association rule mining unit uses an association rule mining algorithm to mine the comorbidity feature associations of chronic liver disease-diabetes, hypertension-diabetes, and chronic liver disease-hypertension-diabetes, and extracts the comorbidity association feature set. The multimodal comorbidity feature fusion unit combines feature splicing and attention mechanisms to fuse the single-disease-specific feature set and the comorbidity association feature set, assigning differentiated weights to different features. The training, validation, and test set partitioning unit partitions the model training set, validation set, and test set based on stratified factors such as disease type, comorbidity status, age group, and disease stage, using stratified sampling.
[0022] Furthermore, the prognostic risk prediction model training and optimization module includes a basic model construction unit, a multi-model fusion training unit, a model hyperparameter optimization unit, and a prognostic risk core calculation unit.
[0023] The basic model building unit constructs four basic prognostic risk prediction models: logistic regression, XGBoost, random forest, and CNN-LSTM deep learning model. Each basic model takes the fused global feature set as input and the disease prognosis as output. The multi-model fusion training unit uses a weighted fusion method to train the basic models, allocating fusion weights based on the predictive performance of each basic model on the validation set. The model hyperparameter optimization unit uses a grid search combined with K-fold cross-validation to traverse and optimize the hyperparameters of the basic and fused models to determine the optimal hyperparameter combination. The prognostic risk core calculation unit calculates the individualized prognostic risk value of the patient based on the trained and optimized multi-model fusion prediction model using the prognostic risk quantification formula, which is: ,in This is the prognostic risk value, ranging from [0,1]. It is the Sigmoid activation function. The feature weight matrix, For global feature vectors, This is a bias term.
[0024] Furthermore, the multi-dimensional model verification and evaluation module includes an internal cross-validation unit, an external real-world queue verification unit, a model performance evaluation unit, and a model update triggering unit;
[0025] The internal cross-validation unit employs K-fold cross-validation, dividing the training set into K mutually exclusive subsets. One subset is selected sequentially as the validation set, and the rest are used as the training set for model training and validation. The average of the K validation results is taken as the internal validation performance index. The external real-world cohort validation unit collects real-world cohort data from multiple centers and different regions related to chronic diseases. After standardization and quality control, this data serves as the external validation dataset for external validation of the fusion model. The model performance evaluation unit assesses the model from three dimensions: discriminative ability, calibration ability, and clinical usability. Discriminative ability indicators include the area under the receiver operating characteristic (ROC) curve and the C-index. Calibration ability is assessed using calibration curve fitting and the Hosmer-Lemeshow test. Clinical usability is assessed using decision curve analysis. The model update triggering unit sets a model performance evaluation threshold based on industry standards and clinical expert consensus for chronic disease clinical prediction models. When the model performance index falls below the threshold, the model update process is triggered.
[0026] Furthermore, the model deployment and risk stratification output module includes a lightweight model deployment unit, an individualized prognostic risk calculation unit, a risk stratification output unit, and a clinical visualization display unit;
[0027] The lightweight model deployment unit compresses and lightens the fusion model to generate a lightweight model file adapted to clinical information systems and chronic disease management systems, supporting API calls and embedded deployment. The individualized prognostic risk calculation unit obtains standardized patient feature data through the clinical information system and inputs it into the lightweight model to calculate prognostic risk values in real time. The risk stratification output unit divides the prognostic risk values into three risk levels—low, medium, and high—according to the consensus of experts in chronic disease comorbidity management. The clinical visualization unit visualizes the prognostic risk values, risk stratification levels, and core feature contribution in chart form and generates a standardized prognostic risk prediction report.
[0028] The construction method of a prognostic risk prediction model system based on real-world data of chronic diseases includes the following steps:
[0029] Step 1: Start the multi-source real-world data acquisition and access module to collect multi-source real-world data on chronic diseases, access and cache it through multiple protocols, and record the acquisition log;
[0030] Step 2: Input the raw data into the data standardization and quality control module to complete standardization, missing value handling, outlier detection and correction, and data deduplication and integration to obtain standardized structured data;
[0031] Step 3: Input the standardized structured data into the multi-dimensional feature engineering and screening module, extract multi-dimensional features and preprocess them, and obtain the core features through clinical-data dual-dimensional screening;
[0032] Step 4: Input the core features into the comorbidity feature fusion and dataset construction module to build a single disease feature set, mine and fuse comorbidity-related features, and divide the training, validation and test sets by stratified sampling;
[0033] Step 5: Input the dataset into the prognostic risk prediction model training and optimization module, build the basic model and perform weighted fusion, optimize the hyperparameters, obtain the fusion prediction model, and calculate the risk value based on the quantitative formula;
[0034] Step 6: Input the fusion model into the model multi-dimensional validation and evaluation module to complete internal cross-validation and external multi-center validation. Evaluate the performance from three dimensions. If the performance does not meet the standard, return to step 5 to retrain.
[0035] Step 7: Deploy the validated fusion model to the clinical information system after it has been lightweighted, calculate individualized risk values for patients in real time, complete risk stratification and visualization, and generate prediction reports;
[0036] Step 8: Monitor the clinical application performance of the model in real time. When the performance is lower than the threshold, trigger the update process, collect new data, and repeat steps 1-7 to complete the model iteration update.
[0037] Furthermore, the weighted fusion method described in step 5 allocates differentiated fusion weights based on the prediction performance of each basic model on the validation set, the number of folds in the K-fold cross-validation described in step 6 is determined based on the data volume and distribution characteristics, and the model update process described in step 8 includes the entire process of new data collection, standardization, feature fusion, model retraining and validation.
[0038] Compared with the prior art, the beneficial effects of the present invention are:
[0039] A dedicated multi-source real-world data acquisition and quality control system for chronic liver disease, diabetes, and hypertension was constructed. Combined with clinical diagnosis and treatment guidelines for chronic diseases, data standardization processing was achieved. Through multi-dimensional missing value and outlier handling and data deduplication integration, the quality and usability of real-world data were effectively improved. This solved the core problem of messy and low-quality real-world data in existing technologies, and provided a high-quality data source foundation for model construction.
[0040] By introducing a comorbidity feature fusion mechanism, this study targets the comorbidity associations of chronic liver disease, diabetes, and hypertension, mines the associations of comorbidity features, and uses an attention mechanism to complete the fusion of multimodal features. This breaks through the limitations of traditional single-disease models, making the model more consistent with the real-world scenario of patients having multiple diseases in clinical practice, and significantly improving the clinical adaptability of the model's prediction results.
[0041] This approach employs a dual-dimensional feature selection method combining clinical and data analysis, integrating clinical guidelines for chronic disease diagnosis and treatment with statistical analysis. By using a comprehensive feature scoring formula to select core features, it ensures both statistical discriminative power and clear clinical interpretability of the features. This avoids the problem of purely data-driven feature selection being disconnected from clinical practice, making the model more easily recognized and accepted by clinicians.
[0042] A multi-model fusion prognostic risk prediction system was constructed, combining the advantages of four basic models: logistic regression, XGBoost, random forest, and CNN-LSTM. A weighted fusion method was adopted to improve the model's prediction performance, and an individualized risk value was accurately calculated through a quantitative prognostic risk calculation formula, making the model's prediction results more quantitative and interpretable.
[0043] A multi-dimensional model validation and evaluation system was established, integrating internal K-fold cross-validation and external multi-center real-world cohort validation. The model was comprehensively evaluated from three dimensions: discriminative ability, calibration ability, and clinical applicability, which effectively improved the robustness and generalization ability of the model and ensured its applicability in different medical institutions and regions.
[0044] It enables lightweight deployment of the model and clinically applicable risk stratification output. The trained and optimized model can be adapted to hospital clinical information systems and community chronic disease management systems, supports real-time individualized risk value calculation, and completes risk stratification based on clinical expert consensus. Combined with visualization and standardized report generation, it provides clinicians with a scientific and intuitive basis for developing individualized chronic disease diagnosis and intervention plans, effectively improving the efficiency of model translation from laboratory to clinical practice. Attached Figure Description
[0045] Figure 1 This is a system module diagram of the present invention;
[0046] Figure 2 This is a schematic diagram of the multi-source real-world data acquisition and access module of the present invention;
[0047] Figure 3 This is a schematic diagram of the data standardization and quality control module of the present invention;
[0048] Figure 4 This is a schematic diagram of the multi-dimensional feature engineering and screening module of the present invention;
[0049] Figure 5 This is a flowchart of the method of the present invention. Detailed Implementation
[0050] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0051] Please see Figure 1-5 This invention provides a prognostic risk prediction model construction system based on real-world data of chronic diseases, including a multi-source real-world data acquisition and access module, a data standardization and quality control module, a multi-dimensional feature engineering and screening module, a comorbidity feature fusion and dataset construction module, a prognostic risk prediction model training and optimization module, a multi-dimensional model verification and evaluation module, and a model deployment and risk stratification output module.
[0052] The multi-source real-world data acquisition and access module enables comprehensive acquisition and standardized protocol access of multi-source heterogeneous medical data related to chronic liver disease, diabetes, and hypertension, ensuring the integrity of the original data and the compatibility of the access.
[0053] The data standardization and quality control module is based on the clinical diagnosis and treatment guidelines for chronic diseases and medical data standards. It standardizes the collected raw data and improves data quality through multi-dimensional quality control methods, providing a high-quality data source for model construction.
[0054] The multi-dimensional feature engineering and screening module extracts multiple types of features from standardized data. After preprocessing, it uses a clinical-data dual-dimensional feature screening method to screen out core features that have both clinical significance and data discriminative power, providing a feature foundation for model construction.
[0055] The comorbidity feature fusion and dataset construction module targets single diseases and comorbidity scenarios of chronic liver disease, diabetes, and hypertension. It constructs single-disease feature sets and mines the correlation between comorbidity features to complete multimodal comorbidity feature fusion. At the same time, it divides the training, validation, and test datasets according to stratified sampling.
[0056] The prognostic risk prediction model training and optimization module constructs multiple types of basic prediction models, uses a weighted fusion method for model training, improves model performance through hyperparameter optimization, and calculates individualized prognostic risk values by combining quantitative formulas.
[0057] The multi-dimensional model validation and evaluation module comprehensively evaluates the model from three dimensions: discrimination ability, calibration ability, and clinical applicability, through internal cross-validation and external real-world cohort validation, to ensure the robustness and generalization ability of the model.
[0058] The model deployment and risk stratification output module lightweights the trained and optimized model and adapts it to the clinical information system, enabling real-time calculation of individualized prognostic risk values for patients and completing risk stratification and visualization output based on clinical guidelines.
[0059] The multi-source real-world data acquisition and access module specifically includes a data acquisition unit, a multi-protocol access unit, and a data caching unit;
[0060] Data Acquisition Unit: For chronic liver disease, diabetes, and hypertension, it collects multi-source real-world data covering demographic information, clinical symptoms, laboratory tests, treatment interventions, follow-up outcomes, and medical insurance records. Data sources include hospital electronic medical record systems, laboratory test systems, chronic disease follow-up management systems, and medical insurance information systems. Multi-protocol Access Unit: Supports common medical data protocols such as HL7, FHIR, DICOM, and SQL, enabling seamless data access from different heterogeneous medical information systems and ensuring data compatibility and universality. Data Caching Unit: Employs a distributed caching mechanism to temporarily store the collected raw real-world data in real time, establishing a data acquisition log to record the data acquisition time, source, type, and integrity, ensuring the lossless storage of raw data.
[0061] The data standardization and quality control module specifically includes a data standardization unit, a missing value handling unit, an outlier detection and correction unit, and a data deduplication and integration unit;
[0062] Data Standardization Unit: Based on the "Guidelines for the Diagnosis and Treatment of Chronic Hepatitis B," "Guidelines for the Prevention and Treatment of Type 2 Diabetes in China," "Guidelines for the Prevention and Treatment of Hypertension in China," and the HL7 international medical data standard, the indicator names, units of measurement, reference ranges, disease diagnosis codes, and surgical procedure codes in the data are standardized to transform unstructured clinical text data into structured data. Missing Value Handling Unit: For continuous data, interpolation based on data distribution characteristics is used for missing value imputation; for categorical data, mode imputation based on disease clinical characteristics is used; and for data with missing key clinical indicators, [further details needed]. The system employs a mean-filling method based on patients with the same disease course and characteristics in subgroups. An outlier detection and correction unit uses an outlier detection algorithm combined with threshold values for chronic disease clinical indicators to detect outliers in test results, follow-up time, and other data. Detected outliers are manually reviewed and corrected based on original medical records, and outliers that cannot be reviewed are marked and removed. A data deduplication and integration unit uses the patient's unique identifier as the core, combining medical insurance number, inpatient number, and outpatient number for data association and matching. Duplicate data collection and storage are deduplicated, achieving integrated integration of multi-source, multi-time-dimensional data for the same patient.
[0063] The multi-dimensional feature engineering and screening module specifically includes a feature extraction unit, a feature preprocessing unit, and a clinical-data dual-dimensional feature screening unit;
[0064] Feature Extraction Unit: Extracts multi-dimensional features from standardized data sources, including demographic features, clinical symptom features, laboratory test features, treatment intervention features, disease course features, and follow-up outcome features. Features specific to chronic liver disease include liver function indicators, liver fibrosis indicators, and virological indicators; features specific to diabetes include blood glucose indicators, glycated hemoglobin, and insulin levels; and features specific to hypertension include blood pressure monitoring values and cardiac function indicators. Feature Preprocessing Unit: Normalizes continuous features to eliminate dimensional differences; performs one-hot encoding on categorical features to convert unordered categorical features into computable numerical features; and performs time window segmentation and feature aggregation on time series features. Clinical-Data Dual-Dimensional Feature Screening Unit: First, performs initial clinical screening based on chronic disease clinical diagnosis and treatment guidelines to remove features without clinical significance. Then, uses a feature comprehensive scoring formula for statistical screening, sets a feature comprehensive scoring threshold, and retains core features with scores above the threshold. The feature comprehensive scoring formula is:
[0065]
[0066] in, For feature comprehensive scoring, This is the clinical weighting coefficient. For statistical weighting coefficients, satisfying , and Determined based on consensus among clinical experts on chronic diseases; The characteristic clinical score is a quantitative score of 0-10 assigned by clinical experts based on the degree to which the characteristic affects the prognosis of the disease. For feature statistical scoring, the feature importance calculated by the random forest algorithm is normalized and quantified into a score of 0-10.
[0067] The comorbidity feature fusion and dataset construction module specifically includes a single-disease feature set construction unit, a comorbidity association rule mining unit, a multimodal comorbidity feature fusion unit, and a training, validation, and test set partitioning unit;
[0068] The system comprises the following units: Single-disease feature set construction unit: After feature screening, core features are categorized into chronic liver disease, diabetes, and hypertension, and three dedicated feature sets are constructed for each disease. These feature sets contain the core prognostic features of the disease. Comorbidity association rule mining unit: Using association rule mining algorithms, the system mines the feature associations of three comorbidity combinations: chronic liver disease-diabetes, hypertension-diabetes, and chronic liver disease-hypertension-diabetes. Key features that synergistically influence prognosis are extracted, forming a comorbidity association feature set. Multimodal comorbidity feature fusion unit: Using a combination of feature concatenation and attention mechanisms, the system fuses the single-disease-specific feature sets with the comorbidity association feature sets. Differential weights are assigned to different features using the attention mechanism to highlight features with significant prognostic impact, forming a fused global feature set. Training, validation, and test set partitioning unit: Using stratified sampling, based on stratification factors such as disease type, comorbidity status, age group, and disease stage, the fused global feature set is divided into a model training set, a validation set, and a test set according to a preset ratio, ensuring consistency between the disease features and demographic features of each dataset.
[0069] The prognostic risk prediction model training and optimization module specifically includes a basic model building unit, a multi-model fusion training unit, a model hyperparameter optimization unit, and a prognostic risk core calculation unit;
[0070] The system comprises the following components: **Basic Model Construction Unit:** Based on the disease characteristics and data features of chronic liver disease, diabetes, and hypertension, four basic prognostic risk prediction models are constructed: logistic regression, XGBoost, random forest, and CNN-LSTM deep learning. Each basic model takes a fused global feature set as input and outputs the disease prognosis. **Multi-Model Fusion Training Unit:** A weighted fusion method is used to train the basic models. Differential fusion weights are assigned to each basic model based on its predictive performance on the validation set, constructing a multi-model fusion prognostic risk prediction model. **Model Hyperparameter Optimization Unit:** A grid search combined with K-fold cross-validation is used to traverse and optimize the hyperparameters of the basic and fusion models, determining the optimal hyperparameter combination for each model to improve predictive performance. **Core Prognostic Risk Calculation Unit:** Based on the trained and optimized multi-model fusion prediction model, an individualized prognostic risk value is calculated for each patient using a prognostic risk quantification formula. The prognostic risk quantification formula is as follows:
[0071]
[0072] in, This is a patient-specific prognostic risk value, ranging from [0,1]. A higher value indicates a higher prognostic risk for the patient. This is the Sigmoid activation function, used to map the model output to the [0,1] interval; This is the feature weight matrix obtained from model training, where each element represents the weight of the influence of each core feature on prognostic risk. This is the fused patient-wide feature vector; This refers to the bias term obtained during model training.
[0073] The multi-dimensional model validation and evaluation module specifically includes an internal cross-validation unit, an external real-world queue validation unit, a model performance evaluation unit, and a model update triggering unit;
[0074] Internal Cross-Validation Unit: Employing K-fold cross-validation, the model training set is divided into K mutually exclusive subsets. One subset is selected sequentially as the validation set, while the remaining subsets serve as the training set for model training and validation. The mean of the K validation results is used as the model's internal validation performance metric. External Real-World Cohort Validation Unit: Collects real-world cohort data from multiple centers and different regions related to chronic liver disease, diabetes, and hypertension. After standardization and quality control by this system, this data serves as the external validation dataset for external validation of the trained and optimized fusion model, verifying its generalization ability. Model Performance Evaluation Unit: The model is comprehensively evaluated from three dimensions: discriminant capability, calibration capability, and clinical usability. Discriminant capability evaluation metrics include the area under the receiver operating characteristic (ROC) curve and the C-index. Calibration capability is evaluated using calibration curve fitting and the Hosmer-Lemeshow test. Clinical usability is evaluated using decision curve analysis. Model Update Trigger Unit: A model performance evaluation threshold is set, determined based on industry standards and clinical expert consensus for chronic disease clinical prediction models. When the model's performance metric falls below the set threshold during external validation or clinical application, the model update process is triggered.
[0075] The model deployment and risk stratification output module specifically includes a lightweight model deployment unit, an individualized prognostic risk calculation unit, a risk stratification output unit, and a clinical visualization unit;
[0076] The model lightweight deployment unit performs model compression and lightweighting on the trained and optimized multi-model fusion prediction model, removes redundant parameters, and generates lightweight model files adapted to hospital clinical information systems and chronic disease management systems, supporting API calls and embedded deployment. The individualized prognostic risk calculation unit obtains standardized patient characteristic data from the clinical information system, inputs it into the lightweight model, and calculates the patient's individualized prognostic risk value in real time using a prognostic risk quantification formula. The risk stratification output unit outputs the patient's prognostic risk value based on the "Expert Consensus on the Management of Comorbidities of Chronic Liver Disease, Diabetes, and Hypertension." The risk levels are divided into three categories based on preset intervals: low risk, medium risk, and high risk. Low risk indicates that the patient's disease progresses slowly and the probability of adverse prognosis is low. Medium risk indicates that the patient has a certain probability of adverse prognosis and requires enhanced follow-up monitoring. High risk indicates that the patient has a high probability of adverse prognosis and requires individualized intensive intervention. The clinical visualization unit displays the patient's individualized prognostic risk value, risk stratification level, and the contribution of each core feature to prognostic risk in the form of bar charts, line charts, and radar charts. At the same time, it generates a standardized prognostic risk prediction report for clinicians to review and refer to.
[0077] This invention also provides a method for constructing a prognostic risk prediction model system based on real-world data of chronic diseases, comprising the following steps:
[0078] Step 1: Multi-source real-world data acquisition and access
[0079] The multi-source real-world data acquisition and access module is activated, and the multi-protocol access unit connects to various heterogeneous medical information systems. The data acquisition unit comprehensively collects multi-source real-world data on chronic liver disease, diabetes, and hypertension. The data is temporarily stored in real time and the acquisition log is recorded by the data caching unit to ensure the integrity of the original data.
[0080] Step 2: Data Standardization and Quality Control Processing
[0081] The cached raw data is input into the data standardization and quality control module. The data standardization unit performs standardization processing according to the chronic disease clinical diagnosis and treatment guidelines and medical data standards. Then, the missing value processing unit and the outlier detection and correction unit complete the data cleaning. Finally, the data deduplication and integration unit realizes the integrated integration of patient data to obtain high-quality standardized structured data.
[0082] Step 3: Multi-dimensional feature engineering and core feature selection
[0083] Standardized structured data is input into the multi-dimensional feature engineering and screening module, where the feature extraction unit extracts multi-dimensional features. The feature preprocessing unit completes feature normalization and encoding. Then, the clinical-data dual-dimensional feature screening unit first performs clinical screening and then statistical screening based on the feature comprehensive scoring formula to retain core features and obtain a single disease core feature set.
[0084] Step 4: Comorbidity Feature Fusion and Dataset Partitioning
[0085] The core feature set of a single disease is input into the comorbidity feature fusion and dataset construction module. First, the feature set of a single disease is constructed. Then, the comorbidity association rule mining unit extracts the comorbidity association features. The feature fusion is completed by the multimodal comorbidity feature fusion unit to obtain the global feature set. Finally, the training, validation and test set partitioning unit partitions the model training set, validation set and test set according to the stratified sampling method.
[0086] Step 5: Training and Optimization of Prognostic Risk Prediction Model
[0087] The divided dataset is input into the prognostic risk prediction model training and optimization module. The basic model building unit constructs four basic prediction models, which are then weighted and fused by the multi-model fusion training unit. The optimal hyperparameter combination is then determined by the model hyperparameter optimization unit to obtain the trained and optimized multi-model fusion prediction model. The prognostic risk core calculation unit calculates the risk value based on the prognostic risk quantification formula.
[0088] Step 6: Multi-dimensional Validation and Evaluation of the Model
[0089] The trained and optimized fusion model is input into the model multi-dimensional verification and evaluation module. First, the internal cross-validation unit completes the internal K-fold cross-validation, and then the external real-world queue verification unit completes the external multi-center verification. The model performance evaluation unit evaluates the performance from three dimensions. If the performance indicators meet the set threshold, proceed to the next step. If they do not meet the threshold, return to step 5 to retrain and optimize the model.
[0090] Step 7: Model Deployment and Clinical Risk Stratification Output
[0091] The validated fusion model is processed by the model lightweight deployment unit and then deployed to the clinical information system and chronic disease management system. The individualized prognostic risk calculation unit calculates the individualized prognostic risk value of the patient in real time. The risk level is classified by the risk stratification output unit. Finally, the clinical visualization display unit realizes the visualization display of the prediction results and generates reports.
[0092] Step 8: Dynamic Model Update
[0093] During clinical application, the model's performance indicators are monitored in real time by the model update triggering unit. When the performance indicators fall below the set threshold, the model update process is triggered to collect new real-world data and repeat steps 1-7 to complete the iterative update of the model, ensuring the long-term effectiveness of the model.
[0094] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A system for constructing prognostic risk prediction models based on real-world data of chronic diseases, characterized by: It includes modules for multi-source real-world data acquisition and access, data standardization and quality control, multi-dimensional feature engineering and screening, comorbidity feature fusion and dataset construction, prognostic risk prediction model training and optimization, multi-dimensional model validation and evaluation, and model deployment and risk stratification output. The multi-source real-world data acquisition and access module enables comprehensive acquisition and standardized protocol access of multi-source heterogeneous medical data related to chronic liver disease, diabetes, and hypertension, ensuring the integrity of the original data and the compatibility of the access. The data standardization and quality control module is based on the clinical diagnosis and treatment guidelines for chronic diseases and medical data standards. It standardizes the collected raw data and improves data quality through multi-dimensional quality control methods. The multi-dimensional feature engineering and screening module extracts multiple types of features from standardized data. After preprocessing, it uses a clinical-data dual-dimensional feature screening method to screen out core features that have both clinical significance and data distinguishability. The comorbidity feature fusion and dataset construction module targets single diseases and comorbidity scenarios of chronic liver disease, diabetes, and hypertension. It constructs single-disease feature sets and mines the correlation between comorbidity features to complete multimodal comorbidity feature fusion. At the same time, it divides the training, validation, and test datasets according to stratified sampling. The prognostic risk prediction model training and optimization module constructs multiple types of basic prediction models, uses a weighted fusion method for model training, improves model performance through hyperparameter optimization, and calculates individualized prognostic risk values by combining quantitative formulas. The multi-dimensional model validation and evaluation module comprehensively evaluates the model from three dimensions: discrimination ability, calibration ability, and clinical applicability, through internal cross-validation and external real-world cohort validation. The model deployment and risk stratification output module performs lightweight processing on the trained and optimized model and adapts it to the clinical information system, enabling real-time calculation of individualized prognostic risk values for patients, and completing risk stratification and visualization output based on clinical guidelines.
2. The prognostic risk prediction model construction system based on real-world data of chronic diseases according to claim 1, characterized in that: The multi-source real-world data acquisition and access module includes a data acquisition unit, a multi-protocol access unit, and a data caching unit; The data acquisition unit collects multi-source real-world data on chronic liver disease, diabetes, and hypertension, covering demographic information, clinical symptoms, laboratory tests, treatment interventions, follow-up outcomes, and medical insurance records. Data sources include hospital electronic medical record systems, laboratory test systems, chronic disease follow-up management systems, and medical insurance information systems. The multi-protocol access unit supports common medical data protocols such as HL7, FHIR, DICOM, and SQL, enabling seamless data access from different heterogeneous medical information systems. The data caching unit uses a distributed caching mechanism to temporarily store raw real-world data and establishes a data acquisition log to record the data acquisition time, source, type, and completeness.
3. The prognostic risk prediction model construction system based on real-world data of chronic diseases according to claim 1, characterized in that: The data standardization and quality control module includes a data standardization unit, a missing value processing unit, an outlier detection and correction unit, and a data deduplication and integration unit. The data standardization unit, based on the clinical diagnosis and treatment guidelines for chronic diseases and the HL7 international medical data standard, standardizes the indicator names, units of measurement, reference ranges, disease diagnosis codes, and surgical operation codes of the data, transforming unstructured clinical text data into structured data. The missing value processing unit uses interpolation based on data distribution characteristics to fill in continuous data, mode-based filling based on disease clinical characteristics to fill in categorized data, and mean-based filling based on the subgroups of patients with the same disease course and characteristics to fill in missing data for key clinical indicators. The outlier detection and correction unit uses an outlier detection algorithm combined with threshold values for chronic disease clinical indicators to detect outliers, and performs manual review and correction based on original medical records, marking and removing outliers that cannot be reviewed. The data deduplication and integration unit uses the patient's unique identifier as the core, combining medical insurance number, hospitalization number, and outpatient number for data association and matching, completing data deduplication and integrated integration of multi-source data for the same patient.
4. The prognostic risk prediction model construction system based on real-world data of chronic diseases according to claim 1, characterized in that: The multi-dimensional feature engineering and screening module includes a feature extraction unit, a feature preprocessing unit, and a clinical-data dual-dimensional feature screening unit; The feature extraction unit extracts demographic features, clinical symptom features, laboratory test features, treatment intervention features, disease course features, and follow-up outcome features, including disease-specific features for chronic liver disease, diabetes, and hypertension. The feature preprocessing unit normalizes continuous features, performs one-hot encoding on categorical features, and divides time-series features into time windows and aggregates features. The clinical-data dual-dimensional feature screening unit first performs initial clinical screening based on chronic disease clinical treatment guidelines, then uses a feature comprehensive scoring formula for statistical screening, sets a feature comprehensive scoring threshold, and retains core features with scores higher than the threshold. The feature comprehensive scoring formula is as follows: ,in For feature comprehensive scoring, This is the clinical weighting coefficient. For statistical weighting coefficients, , and Based on the consensus of clinical experts on chronic diseases, it was determined that... Characteristic clinical scores ranging from 0 to 10. The feature statistics score is 0-10.
5. The prognostic risk prediction model construction system based on real-world data of chronic diseases according to claim 1, characterized in that: The comorbidity feature fusion and dataset construction module includes a single-disease feature set construction unit, a comorbidity association rule mining unit, a multimodal comorbidity feature fusion unit, and a training, validation, and test set partitioning unit; The single-disease feature set construction unit classifies core features into three single-disease-specific feature sets: chronic liver disease, diabetes, and hypertension. The comorbidity association rule mining unit uses an association rule mining algorithm to mine the comorbidity feature associations of chronic liver disease-diabetes, hypertension-diabetes, and chronic liver disease-hypertension-diabetes, and extracts the comorbidity association feature set. The multimodal comorbidity feature fusion unit combines feature splicing and attention mechanisms to fuse the single-disease-specific feature set and the comorbidity association feature set, assigning differentiated weights to different features. The training, validation, and test set partitioning unit partitions the model training set, validation set, and test set based on stratified factors such as disease type, comorbidity status, age group, and disease stage, using stratified sampling.
6. The prognostic risk prediction model construction system based on real-world data of chronic diseases according to claim 1, characterized in that: The prognostic risk prediction model training and optimization module includes a basic model construction unit, a multi-model fusion training unit, a model hyperparameter optimization unit, and a prognostic risk core calculation unit. The basic model building unit constructs four basic prognostic risk prediction models: logistic regression, XGBoost, random forest, and CNN-LSTM deep learning model. Each basic model takes the fused global feature set as input and the disease prognosis as output. The multi-model fusion training unit uses a weighted fusion method to train the basic models, allocating fusion weights based on the predictive performance of each basic model on the validation set. The model hyperparameter optimization unit uses a grid search combined with K-fold cross-validation to traverse and optimize the hyperparameters of the basic and fused models to determine the optimal hyperparameter combination. The prognostic risk core calculation unit calculates the individualized prognostic risk value of the patient based on the trained and optimized multi-model fusion prediction model using the prognostic risk quantification formula, which is: ,in This is the prognostic risk value, ranging from [0,1]. It is the Sigmoid activation function. The feature weight matrix, For global feature vectors, This is a bias term.
7. The prognostic risk prediction model construction system based on real-world data of chronic diseases according to claim 1, characterized in that: The model multi-dimensional verification and evaluation module includes an internal cross-validation unit, an external real-world queue verification unit, a model performance evaluation unit, and a model update triggering unit. The internal cross-validation unit employs K-fold cross-validation, dividing the training set into K mutually exclusive subsets. One subset is selected sequentially as the validation set, and the rest are used as the training set for model training and validation. The average of the K validation results is taken as the internal validation performance index. The external real-world cohort validation unit collects real-world cohort data from multiple centers and different regions related to chronic diseases. After standardization and quality control, this data serves as the external validation dataset for external validation of the fusion model. The model performance evaluation unit assesses the model from three dimensions: discriminative ability, calibration ability, and clinical usability. Discriminative ability indicators include the area under the receiver operating characteristic (ROC) curve and the C-index. Calibration ability is assessed using calibration curve fitting and the Hosmer-Lemeshow test. Clinical usability is assessed using decision curve analysis. The model update triggering unit sets a model performance evaluation threshold based on industry standards and clinical expert consensus for chronic disease clinical prediction models. When the model performance index falls below the threshold, the model update process is triggered.
8. The prognostic risk prediction model construction system based on real-world data of chronic diseases according to claim 1, characterized in that: The model deployment and risk stratification output module includes a lightweight model deployment unit, an individualized prognostic risk calculation unit, a risk stratification output unit, and a clinical visualization unit. The lightweight model deployment unit compresses and lightens the fusion model to generate lightweight model files that are compatible with clinical information systems and chronic disease management systems, and supports API interface calls and embedded deployment. The individualized prognostic risk calculation unit obtains standardized patient characteristic data through the clinical information system and inputs it into the lightweight model to calculate the prognostic risk value in real time; the risk stratification output unit divides the prognostic risk value into three risk levels—low, medium, and high—according to the consensus of experts on chronic disease comorbidity management. The clinical visualization unit visualizes prognostic risk values, risk stratification levels, and the contribution of core features in chart form, and generates a standardized prognostic risk prediction report.
9. The method for constructing a prognostic risk prediction model construction system based on real-world data of chronic diseases according to any one of claims 1-8, characterized in that: Includes the following steps: Step 1: Start the multi-source real-world data acquisition and access module to collect multi-source real-world data on chronic diseases, access and cache it through multiple protocols, and record the acquisition log; Step 2: Input the raw data into the data standardization and quality control module to complete standardization, missing value handling, outlier detection and correction, and data deduplication and integration to obtain standardized structured data; Step 3: Input the standardized structured data into the multi-dimensional feature engineering and screening module, extract multi-dimensional features and preprocess them, and obtain the core features through clinical-data dual-dimensional screening; Step 4: Input the core features into the comorbidity feature fusion and dataset construction module to build a single disease feature set, mine and fuse comorbidity-related features, and divide the training, validation and test sets by stratified sampling; Step 5: Input the dataset into the prognostic risk prediction model training and optimization module, build the basic model and perform weighted fusion, optimize the hyperparameters, obtain the fusion prediction model, and calculate the risk value based on the quantitative formula; Step 6: Input the fusion model into the model multi-dimensional validation and evaluation module to complete internal cross-validation and external multi-center validation. Evaluate the performance from three dimensions. If the performance does not meet the standard, return to step 5 to retrain. Step 7: Deploy the validated fusion model to the clinical information system after it has been lightweighted, calculate individualized risk values for patients in real time, complete risk stratification and visualization, and generate prediction reports; Step 8: Monitor the clinical application performance of the model in real time. When the performance is lower than the threshold, trigger the update process, collect new data, and repeat steps 1-7 to complete the model iteration update.
10. The method for constructing a prognostic risk prediction model construction system based on real-world data of chronic diseases according to claim 9, characterized in that: The weighted fusion method described in step 5 allocates differentiated fusion weights based on the prediction performance of each basic model on the validation set. The number of folds in the K-fold cross-validation described in step 6 is determined based on the data volume and distribution characteristics. The model update process described in step 8 includes the entire process of new data collection, standardization, feature fusion, model retraining and validation.