Liver cancer occurrence risk prediction model and construction method

By providing a risk prediction model and construction method for liver cancer, including the construction of risk assessment and early warning models, using parameter tuning and statistical analysis to optimize the model, setting dynamic evaluation mechanisms and triggering conditions, the problem of inaccurate liver cancer risk prediction in the existing technology is solved, and more accurate risk judgment and automated early warning are achieved.

CN120048525AActive Publication Date: 2025-05-27SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202510166607.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-27
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The prior art has problems in the prediction of liver cancer risk that do not conduct targeted risk analysis, risk warning and intervention decision generation after data collection, resulting in inaccurate prediction; at the same time, there is a lack of model evaluation and optimization of risk judgment, and it is impossible to accurately obtain the degree of risk level; and it is also impossible to effectively formulate risk warning and intervention plans.

Method used

Provide a risk prediction model and construction method for liver cancer, including risk assessment models and risk warning models, and optimize model parameters through parameter tuning methods such as grid search and random search; use statistical tests and multiple regression analysis to evaluate the relationship between variables and liver cancer, perform feature selection and model evaluation; set dynamic evaluation mechanisms and trigger conditions, automatically trigger early warning mechanisms and generate intervention suggestions.

Benefits of technology

It improves the accuracy and flexibility of liver cancer risk prediction, reduces the work burden of medical staff, achieves more accurate risk judgment and model optimization, promptly detects and handles potential risks, and provides support for clinical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048525A_ABST
    Figure CN120048525A_ABST
Patent Text Reader

Abstract

The invention discloses a liver cancer occurrence risk prediction model and a construction method, relates to the technical field of liver cancer occurrence risk prediction, and aims to solve the problem that risk data of a patient cannot be accurately obtained. Diversified signal forms can comprehensively attract attention of medical staff, different signal forms can aim at different environments and situations, the flexibility and applicability of early warning are improved, intervention suggestions are automatically generated according to early warning intensity and early warning signals, the workload of the medical staff is greatly relieved, and the working efficiency of the medical staff is improved. By adopting parameter tuning methods such as grid search and random search, a parameter space can be systematically explored, and optimal model parameter configuration can be found, so that the prediction performance of the model is improved, features are reselected and transformed, redundant features can be removed, and the model is more stable and efficient; statistical testing and multiple regression analysis are used for evaluating the relation between the variables and the occurrence of the liver cancer, so that the variables having remarkable influence on the occurrence of the liver cancer can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of predicting the risk of liver cancer occurrence, and specifically to a model for predicting the risk of liver cancer occurrence and a construction method thereof. Background Art

[0002] Predicting the risk of liver cancer occurrence refers to estimating the likelihood of an individual developing liver cancer within a certain period in the future by analyzing various information of the individual, such as living habits, family medical history, biomarker levels, imaging examination results, etc.

[0003] The Chinese patent application document with the publication number CN116543898A discloses a method for constructing a risk prediction model for thyroid cancer in a target population. It mainly obtains the first sample data of the subjects in the target population and cleans the variables of the first sample data to obtain the first sample data. The influencing factor of the family history of hypertension in the target population is put into the process of constructing the risk prediction model, which is conducive to accurately and precisely predicting the risk of thyroid cancer in the target population; by constructing a sample information database according to the first sample data; based on statistical analysis, determining the characteristic variables for predicting the risk of thyroid cancer in the target population according to the sample information database; among them, the characteristic variables include the family history factor; constructing a risk prediction model for thyroid cancer in the target population according to the characteristic variables. Through this step, the effectiveness and reliability of the risk prediction model can be improved. Although the above patent document solves the problem of risk prediction, there are still the following problems in actual operation: 1. After collecting the patient's data, no targeted risk analysis is carried out, and no further risk warning and intervention decision generation are made according to the risk analysis results, resulting in inaccurate prediction of the patient's disease risk.

[0004] 2. No more accurate risk judgment is made according to the risk situation of the data, and no targeted model evaluation and optimization are carried out on the judged risk, resulting in the inability to accurately obtain the patient's risk level.

[0005] 3. No effective risk warning and intervention plan are formulated according to the patient's risk reasons and risk levels, resulting in the patient's inability to further diagnose the risk according to his own condition. Summary of the Invention

[0006] The object of the present invention is to provide a risk prediction model for liver cancer occurrence and a construction method. Diversified signal forms can attract the attention of medical staff in all aspects. Different signal forms can be targeted at different environments and situations, improving the flexibility and applicability of early warning. Intervention suggestions are automatically generated according to the early warning intensity and early warning signals, which greatly reduces the workload of medical staff. By using parameter tuning methods such as grid search and random search, the parameter space can be systematically explored to find the optimal model parameter configuration, thereby improving the prediction performance of the model. Re-selecting and transforming features can remove redundant features and make the model more stable and efficient. Using statistical tests and multiple regression analysis to evaluate the relationship between variables and liver cancer occurrence helps to accurately identify variables that have a significant impact on liver cancer occurrence, and can solve the problems in the prior art.

[0007] To achieve the above object, the present invention provides the following technical solutions: A risk prediction model for liver cancer occurrence, including a risk assessment model and a risk early warning model; The risk assessment model is used for: Collect the liver cancer-related data of the patient, and conduct a quantitative assessment of the liver cancer occurrence risk on the collected liver cancer-related data. Among them, the quantitative assessment includes risk factor analysis, risk assessment calculation and risk level division based on the liver cancer-related data of the patient. After the quantitative assessment is completed, a risk assessment model is obtained; The risk early warning model is used for: Conduct a risk early warning assessment of liver cancer occurrence for the patient according to the risk assessment model. The risk early warning assessment of liver cancer occurrence includes risk dynamic assessment, early warning signal triggering and intervention suggestion generation. After the risk early warning assessment of liver cancer occurrence is completed, a risk early warning model is obtained.

[0008] A construction method for a risk prediction model for liver cancer occurrence, and the construction method comprises the following steps: S1: Collect the liver cancer-related data of the patient from the database. After the collection of the liver cancer-related data is completed, data preprocessing is carried out. After the data preprocessing is completed, target collected data is obtained; S2: Conduct risk factor analysis, risk assessment calculation and risk level division on the target collected data respectively. After the risk factor analysis, risk assessment calculation and risk level division are completed, a risk assessment model is obtained; S3: Conduct model evaluation and optimization on the risk assessment model, and then deploy the risk assessment model after the model evaluation and optimization to the risk prediction port.

[0009] Preferably, for step S1 in the construction method of the risk assessment model, it includes: Collect the liver cancer-related data of the patient from the database. The database includes the databases of the hospital electronic medical record system, the health examination report system and the epidemiological investigation system; Hepatocellular carcinoma-related data includes clinical data, biomarker data, lifestyle information, and environmental exposure data; Perform data cleaning and data standardization on the hepatocellular carcinoma-related data; Data cleaning and data standardization include handling missing values, denoising the data, and integrating data from different sources into a unified format; After the data preprocessing is completed, the target collected data is obtained.

[0010] Preferably, for step S1 in the risk assessment model construction method, it further includes: Extract the data sets corresponding to different data sources; Retrieve the cleaning process data corresponding to the data cleaning of each data set, where the cleaning process data includes the number of missing values and their corresponding weight values and the format conversion duration; Obtain the data quality coefficients corresponding to different data sources by using the number of missing values and their corresponding weight values and the format conversion duration of the data sets corresponding to different data sources; Among them, the data quality coefficient is obtained through the following formula: Among them, J represents the data quality coefficient; n represents the number of missing values corresponding to the data source; w ni represents the data weight value corresponding to the i-th missing value; m represents the number of valid data corresponding to the data source; w mi represents the data weight value corresponding to the i-th valid data; T i represents the format conversion duration corresponding to the i-th valid data; T c represents the preset conversion duration reference value; z represents the total number of collected data corresponding to the data source; Use the data quality coefficients corresponding to different data sources to limit the collection data ratio of the hepatocellular carcinoma-related data collected by different data sources.

[0011] Preferably, using the data quality coefficients corresponding to different data sources to limit the collection data ratio of the hepatocellular carcinoma-related data collected by different data sources includes: Compare the data quality coefficients corresponding to different data sources with the preset data quality coefficient threshold; Take the data sources whose data quality coefficients are not lower than the preset data quality coefficient threshold as the target data sources; Extract the data ratio of the currently collected data of the target data source in the total collected data; Use the data quality coefficient corresponding to the target data source combined with the data ratio to obtain the upper limit value of the collection data ratio corresponding to the target data source; Among them, the upper limit value of the collected data ratio is obtained through the following formula: Among them, B represents the upper limit value of the collected data ratio corresponding to the target data source; B c represents the data ratio of the currently collected data of the standard data source in the total collected data; x represents the number of data sources other than the current target data source; B xi represents the data ratio of the collected data corresponding to the i-th data source other than the current target data source in the total collected data; J m represents the data quality coefficient corresponding to the target data source; J i represents the data quality coefficient corresponding to the i-th data source other than the current target data source; Control the data ratio of the collected data corresponding to the target data source according to the upper limit value of the collected data ratio as the standard.

[0012] Preferably, for the S2 step in the risk assessment model construction method, it includes: First, perform risk factor analysis on the target collected data. The risk factor analysis includes exploratory data analysis, variable screening, and feature selection; The exploratory analysis is to calculate the mean, median, standard deviation, minimum value, and maximum value in the target collected data, and visualize the calculation results in the form of a chart. After the visualization is generated, the distribution and abnormal conditions of the target collected data are obtained; Variable screening is to perform univariate analysis and multivariate analysis according to the abnormal conditions of the target collected data. Univariate analysis is to use statistical tests to evaluate the relationship between each variable and the occurrence of liver cancer. Multivariate analysis is to use multiple regression analysis to evaluate the combined effect of multiple variables on the occurrence of liver cancer when they exist simultaneously; Feature selection is to use the stepwise regression screening method to select the feature data that has a significant impact on the occurrence of liver cancer from the analysis results of the univariate analysis and multivariate analysis of the target collected data obtained in the variable screening; Label the data in the target collected data after the feature selection is completed as the risk factor analysis result.

[0013] Preferably, for the S2 step in the risk assessment model construction method, it also includes: Perform risk assessment calculation according to the risk factor analysis result. Before performing the risk assessment calculation, first construct an evaluation model according to the risk factor analysis result. The evaluation model is a Cox proportional hazards model, a random forest model, or a gradient boosting tree model; Train the obtained evaluation model, and perform the evaluation calculation of the risk factor analysis result on the trained evaluation model; After the evaluation calculation is completed, obtain the risk assessment model of the target collected data; Confirm the evaluation thresholds in the risk assessment model and divide the risk levels according to the confirmed thresholds; The risk level division is to assign risk levels to each evaluation threshold according to the threshold range of the evaluation threshold. The risk levels include low risk, medium risk, and high risk; After each evaluation threshold is assigned to the corresponding risk level, a risk assessment model for target data collection is obtained.

[0014] Preferably, for step S3 in the risk assessment model construction method, it includes: Evaluate the performance indicators of the risk assessment model using the cross-validation method. The performance indicators include the ROC curve, AUC value, precision, recall, F1 score, Brier score, and calibration curve; According to the performance indicator evaluation results, optimize the risk assessment model using the method of parameter tuning. Parameter tuning includes grid search and random search; Re-select the features for the risk assessment model after parameter tuning, and standardize, normalize, or perform logarithmic transformation on the selected features; Use the ensemble learning method to perform model fusion on the risk assessment model after re-selecting the features; After the model fusion is completed, a risk assessment model with model evaluation and optimization completed is obtained; Save the risk assessment model with model evaluation and optimization completed in a persistent format. The persistent format is PMML, ONNX, pickle file, or a custom binary format; Transfer the risk assessment model with model evaluation and optimization completed in the persistent format to the server. The server is signal-connected to the risk prediction port; The server transmits the risk assessment model with model evaluation and optimization completed in the persistent format to the risk prediction port.

[0015] Preferably, the construction method of the risk warning model includes the following steps: S4: Integrate the risk assessment results in the risk assessment model, and set a dynamic evaluation mechanism for the integrated risk assessment results. After the dynamic evaluation mechanism is set, patient risk prediction data is obtained; S5: Judge the warning intensity of the patient risk prediction data, generate warning signals according to the warning level, generate automated intervention suggestions according to the generated warning signals, and construct a risk warning model with the generated automated intervention suggestions; Regarding integrating the risk assessment results in the risk assessment model in step S4 and setting a dynamic evaluation mechanism for the integrated risk assessment results, it includes: Confirm the risk assessment results of the patients in the risk assessment model. After the confirmation is completed, perform dataset integration. After the dataset integration, obtain a risk dataset. Set a dynamic assessment mechanism for the risk dataset. The dynamic assessment mechanism is set to determine the risk assessment frequency of the patients. After the frequency confirmation is completed, set the trigger conditions. The trigger conditions are set to automatically trigger an alarm when the risk level is assigned according to the assessment threshold, based on the completed risk level assignment. After the dynamic assessment mechanism is set, obtain the patient risk prediction data.

[0016] Preferably, for step S5 of the method for constructing a risk warning model, it includes: Judge the warning intensity of the risk levels assigned to the patients in the patient risk prediction data. The higher the risk level, the higher the warning intensity. Distinguish different warning signals according to the warning intensity. Among them, the warning signals include visual signals, text descriptions, and sound signals. Generate automated intervention suggestions for the patients according to the warning intensity and warning signals. The automated intervention suggestions include treatment item guidelines, medical institution contact information, and emergency instructions. Perform visual conversion on the generated automated intervention suggestions, and transmit the visually converted automated intervention suggestions to the mobile terminals of medical staff and patients for display.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The liver cancer occurrence risk prediction model and construction method provided by the present invention use statistical tests and multiple regression analysis to evaluate the relationship between variables and liver cancer occurrence. This method helps to accurately identify the variables that have a significant impact on liver cancer occurrence, improve the accuracy of analysis. The feature selection process uses the stepwise regression screening method to select the feature data that has a significant impact on liver cancer occurrence from the variable screening results. This helps to reduce noise and redundant information and improve the prediction performance of the model.

[0018] 2. The liver cancer occurrence risk prediction model and construction method provided by the present invention adopt parameter tuning methods such as grid search and random search, which can systematically explore the parameter space to find the optimal model parameter configuration, thereby improving the prediction performance of the model. Re-selecting and transforming the features can remove redundant features, improve the feature quality, and make the model more stable and efficient.

[0019] 3. The liver cancer occurrence risk prediction model and its construction method provided by the present invention, through the setting of triggering conditions, when the evaluation threshold reaches a specific risk level, it can automatically trigger the early warning mechanism without manual intervention, improving the efficiency and accuracy of early warning. The automated early warning helps to detect and handle potential risks in a timely manner, providing strong support for clinical decision-making. The diverse signal forms can attract the attention of medical staff in all aspects. Different signal forms can be targeted at different environments and situations, improving the flexibility and applicability of early warning. Intervention suggestions are automatically generated according to the early warning intensity and early warning signals, which greatly reduces the workload of medical staff. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the steps for constructing a risk assessment model in the construction of the liver cancer occurrence risk prediction model of the present invention; Figure 2 It is a schematic diagram of the steps for constructing a risk early warning model in the construction of the liver cancer occurrence risk prediction model of the present invention; Figure 3 It is a schematic diagram of the construction process of the liver cancer occurrence risk prediction model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] In order to solve the problems in the prior art that after collecting the data of patients, no targeted risk analysis is carried out, and no further risk early warning and intervention decision generation are made according to the risk analysis results, resulting in inaccurate prediction of the disease risk of patients, please refer to Figures 1 - 3 , the following technical solutions are provided in this embodiment: A liver cancer occurrence risk prediction model, including a risk assessment model and a risk early warning model; The risk assessment model is used for: Collect the liver cancer-related data of patients, and conduct a quantitative assessment of the liver cancer occurrence risk on the collected liver cancer-related data. Among them, the quantitative assessment includes risk factor analysis, risk assessment calculation, and risk level division based on the liver cancer-related data of patients. After the quantitative assessment is completed, a risk assessment model is obtained; The risk early warning model is used for: Based on the risk assessment model, a risk early warning assessment for liver cancer occurrence in patients is carried out. The risk early warning assessment for liver cancer occurrence includes risk dynamic assessment, early warning signal triggering, and generation of intervention suggestions. After the risk early warning assessment for liver cancer occurrence is completed, a risk early warning model is obtained.

[0023] Specifically, the risk assessment model conducts personalized quantitative assessments by collecting liver cancer-related data of patients. This means that each patient will obtain an exclusive risk assessment result according to their unique situation, improving the accuracy and pertinence of the assessment. Through risk assessment calculation and risk level classification, patients can obtain a clear risk assessment result. The risk early warning model can conduct risk dynamic assessment and real-time monitor the changes in the risk of liver cancer occurrence in patients. This helps to promptly detect the trend of increasing risk, so as to take timely intervention measures. When the risk of liver cancer occurrence in patients reaches or exceeds the preset threshold, the risk early warning model will promptly trigger an early warning signal. The risk early warning model not only provides risk early warning, but also generates personalized intervention suggestions according to the specific situation of the patients.

[0024] To solve the problem in the prior art that after receiving the data of patients, more accurate risk judgments are not made according to the risk situation of the data, and targeted model evaluations and optimizations are not carried out on the judged risks, resulting in the inability to accurately obtain the risk degree of patients, please refer to Figures 1 - 3 , this embodiment provides the following technical solutions: A method for constructing a risk prediction model for liver cancer occurrence, including a method for constructing a risk assessment model. The construction method is as follows: S1: Collect the liver cancer-related data of patients from the database. After the collection of liver cancer-related data is completed, data preprocessing is carried out. After the data preprocessing is completed, target collected data is obtained; Among them, by collecting multi-dimensional information such as clinical data, biomarker data, lifestyle information, and environmental exposure data, a more comprehensive understanding of the liver cancer status of patients and its related factors can be obtained; S2: Conduct risk factor analysis, risk assessment calculation, and risk level classification on the target collected data respectively. After the risk factor analysis, risk assessment calculation, and risk level classification are completed, a risk assessment model is obtained; Among them, risk level classification according to the threshold range of the evaluation threshold helps to intuitively understand the risk degree of different evaluation results; S3: Conduct model evaluation and optimization on the risk assessment model, and then deploy the risk assessment model after the model evaluation and optimization are completed to the risk prediction port; Among them, the entire process from model development to deployment is systematically planned, ensuring the repeatability and verifiability of each step, which is conducive to subsequent model maintenance and update.

[0025] For step S1 in the method for constructing a risk assessment model, it includes: Collect the liver cancer-related data of patients from the database, where the database includes the databases of the hospital electronic medical record system, the health examination report system, and the epidemiological investigation system; The liver cancer-related data includes clinical data, biomarker data, lifestyle information, and environmental exposure data; Perform data cleaning and data standardization processing on the liver cancer-related data; The data cleaning and data standardization processing include handling missing values, denoising the data, and integrating data from different sources into a unified format; After the data preprocessing is completed, the target collected data is obtained.

[0026] Specifically, by collecting multi-dimensional information such as clinical data, biomarker data, lifestyle information, and environmental exposure data, the liver cancer condition of patients and its related factors can be understood more comprehensively. The data cleaning process can identify and handle problems such as missing values and outliers, improving the accuracy and reliability of the data. The data standardization processing helps to eliminate the format differences between data from different sources, ensuring the consistency and comparability of the data. The data after cleaning and standardization can be more conveniently used for subsequent data analysis, data mining, and machine learning tasks. The unified data format and representation method help to simplify the analysis process, improve the accuracy and efficiency of the analysis. Through the analysis and mining of these data, new biomarkers, treatment targets, etc. can be discovered, providing a scientific basis for the precise treatment and individualized treatment of liver cancer.

[0027] Specifically, for step S1 in the method for constructing a risk assessment model, it also includes: Extract the data sets corresponding to different data sources; Retrieve the cleaning process data corresponding to the data cleaning of each data set, where the cleaning process data includes the number of missing values and their corresponding weight values and the format conversion duration; Obtain the data quality coefficients corresponding to different data sources by using the number of missing values and their corresponding weight values and the format conversion duration of the data sets corresponding to different data sources; Among them, the data quality coefficient is obtained through the following formula: Among them, J represents the data quality coefficient; n represents the number of missing values corresponding to the data source; w ni represents the data weight value corresponding to the i-th missing value; m represents the number of valid data corresponding to the data source; w mi represents the data weight value corresponding to the i-th valid data; T i represents the format conversion duration corresponding to the i-th valid data; T crepresents a preset reference value for the conversion duration; z represents the total number of collected data corresponding to the data source; Use the data quality coefficients corresponding to the different data sources to limit the proportion of the collected data of the liver cancer-related data collected by the different data sources.

[0028] The technical effects of the above technical solution are as follows: By comprehensively considering the number of missing values and their weight values, the number of valid data and their weight values, and the format conversion duration, this technical solution can quantitatively evaluate the data quality of different data sources. This quantitative evaluation provides an important reference basis for subsequent data processing and analysis. Using the calculated data quality coefficients, the proportion of the collected data of the liver cancer-related data collected by the different data sources can be limited. This means that in the data collection process, data sources with higher quality can be given priority, and data collection from data sources with lower quality can be reduced or avoided, thereby improving the overall data quality and reliability. By quantitatively evaluating the data quality, high-quality data can be screened out before data processing and analysis, reducing the time for processing and analyzing low-quality data, thereby improving the efficiency of data processing. This technical solution not only considers traditional data quality indicators such as missing values and format conversion duration, but also introduces the factor of data weight values, making the evaluation of data quality more comprehensive and accurate. This helps to enhance the usability of the data and provide more reliable support for subsequent data analysis and decision-making. Through the evaluation of the format conversion duration, this technical solution encourages the standardization of the formats of data sources, making it easier to integrate and compare data between different data sources. This helps to promote the process of data standardization and improve the interoperability and comparability of data. The formula and the calculation method of the data quality coefficients in this technical solution have a certain degree of flexibility and scalability. As the number of data sources increases and data processing technologies continue to develop, the parameters and indicators in the formula can be further adjusted and optimized to meet new data processing requirements.

[0029] In summary, this technical solution helps to improve the efficiency of data processing, enhance the usability and standardization of data by quantitatively evaluating the data quality of different data sources and restricting the data collection ratio accordingly, and at the same time has a certain degree of flexibility and scalability. These technical effects are of great significance for improving the quality and reliability of liver cancer-related data.

[0030] Specifically, using the data quality coefficients corresponding to the different data sources to limit the proportion of the collected data of the liver cancer-related data collected by the different data sources includes: Compare the data quality coefficients corresponding to the different data sources with a preset data quality coefficient threshold; Take the data sources whose data quality coefficients are not lower than the preset data quality coefficient threshold as the target data sources; Extract the proportion of the currently collected data of the target data sources in the total collected data; Obtain the upper limit value of the collected data ratio corresponding to the target data source by combining the data quality coefficient corresponding to the target data source with the data ratio. Among them, the upper limit value of the collected data ratio is obtained through the following formula: Among them, B represents the upper limit value of the collected data ratio corresponding to the target data source; B c represents the data ratio of the currently collected data of the target data source in the total collected data; x represents the number of data sources other than the current target data source; B xi represents the data ratio of the collected data corresponding to the i-th data source other than the current target data source in the total collected data; J m represents the data quality coefficient corresponding to the target data source; J i represents the data quality coefficient corresponding to the i-th data source other than the current target data source; Control the data ratio of the collected data corresponding to the target data source according to the upper limit value of the collected data ratio as the standard.

[0031] The technical effects of the above technical solution are as follows: By comparing the data quality coefficients of different data sources with a preset threshold, this technical solution can screen out data sources with higher quality as target data sources. This data collection strategy based on data quality helps to ensure that the collected data has high reliability and accuracy. This technical solution not only considers the current data collection ratio of the target data source, but also combines the data collection ratios and data quality coefficients of other data sources, and calculates the upper limit value of the collected data ratio corresponding to the target data source through a formula. This dynamic adjustment mechanism can adjust the data collection ratio in real time according to the data quality of different data sources, thereby optimizing the data collection strategy. By restricting the collection ratio of low-quality data sources, this technical solution can reduce the collection of invalid or low-quality data, thereby improving the efficiency of data collection. At the same time, due to the high data quality of the target data source, subsequent data processing and analysis will also be more efficient. This technical solution restricts the data collection ratio through the data quality coefficient, and actually also plays an incentive role for the data source. In order to obtain more data collection opportunities, the data source will strive to improve the data quality to meet the requirements of the preset data quality coefficient threshold. This technical solution provides a scientific data management method. By quantitatively evaluating the data quality and adjusting the data collection ratio accordingly, the data management becomes more standardized and orderly. This helps to improve the overall level of data management and provide more reliable support for subsequent data analysis and decision-making. The formula and data quality coefficient calculation method in this technical solution have a certain degree of adaptability and can be adjusted and optimized according to the actual situation. With the increase of data sources and the continuous development of data processing technologies, this technical solution can continuously adapt to new data processing requirements.

[0032] In summary, through data collection strategies based on data quality, dynamically adjusting the data collection ratio, improving data collection efficiency, promoting data quality improvement, enhancing the scientific nature and strong adaptability of data management, etc., this technical solution provides an effective method for the collection and management of liver cancer-related data. These technical effects are of great significance for improving data quality, optimizing data collection strategies, and enhancing data management levels.

[0033] Preferably, for step S2 in the risk assessment model construction method, it includes: First, perform risk factor analysis on the target collected data. The risk factor analysis includes exploratory data analysis, variable screening, and feature selection; The exploratory analysis is to calculate the mean, median, standard deviation, minimum value, and maximum value in the target collected data, and generate the calculation results in the form of a chart for visualization. After the visualization is generated, the distribution and abnormal conditions of the target collected data are obtained; Variable screening is to perform univariate analysis and multivariate analysis based on the abnormal conditions of the target collected data. Univariate analysis is to use statistical tests to evaluate the relationship between each variable and the occurrence of liver cancer, and multivariate analysis is to use multiple regression analysis to evaluate the combined effect of multiple variables on the occurrence of liver cancer when they exist simultaneously; Feature selection is to use the stepwise regression screening method to select the feature data that has a significant impact on the occurrence of liver cancer from the analysis results of univariate analysis and multivariate analysis of the target collected data obtained in variable screening; Label the data after feature selection in the target collected data as the risk factor analysis result.

[0034] Perform risk assessment calculation based on the risk factor analysis result. Before performing the risk assessment calculation, first construct an evaluation model according to the risk factor analysis result. The evaluation model is a Cox proportional hazards model, a random forest model, or a gradient boosting tree model; Train the obtained evaluation model, and perform the evaluation calculation of the risk factor analysis result on the trained evaluation model; After the evaluation calculation is completed, the risk assessment model of the target collected data is obtained; Confirm the evaluation threshold in the risk assessment model, and perform risk level classification according to the confirmed threshold; The risk level classification is to assign a risk level to each evaluation threshold according to the threshold range of the evaluation threshold. The risk levels include low risk, medium risk, and high risk; After each evaluation threshold is assigned to the corresponding risk level, the risk assessment model of the target collected data is obtained.

[0035] Specifically, starting from exploratory data analysis, variable screening and feature selection are gradually carried out, ensuring the comprehensiveness and systematicness of the analysis. This step-by-step in-depth analysis method helps to identify key risk factors and provides a solid foundation for subsequent risk assessment. By calculating the mean, median, standard deviation, minimum, and maximum values and visualizing them in the form of charts, the distribution and anomalies of the data can be intuitively understood. This helps to discover potential problems and trends in the data and provides clues for subsequent analysis. The variable screening process includes univariate analysis and multivariate analysis, using statistical tests and multiple regression analysis to evaluate the relationship between variables and the occurrence of liver cancer. This method helps to accurately identify variables that have a significant impact on the occurrence of liver cancer and improve the accuracy of the analysis. The feature selection process uses stepwise regression screening to select feature data that has a significant impact on the occurrence of liver cancer from the variable screening results. This helps to reduce noise and redundant information and improve the prediction performance of the model. By performing risk assessment calculations using the trained evaluation model, an accurate risk assessment model can be obtained. This helps to conduct precise risk assessment on the target collected data and provides a basis for subsequent decision-making. According to the threshold range of the evaluation threshold, risk levels are divided into low risk, medium risk, and high risk. This clear division helps to intuitively understand the risk levels of different evaluation results and facilitates the adoption of corresponding risk management measures.

[0036] Regarding step S3 in the method for constructing a risk assessment model, it includes: Evaluate the performance metrics of the risk assessment model using the cross-validation method. The performance metrics include the ROC curve, AUC value, precision, recall, F1 score, Brier score, and calibration curve; According to the performance metric evaluation results, optimize the risk assessment model using the method of parameter tuning. Parameter tuning includes grid search and random search; Re-perform feature selection on the risk assessment model after parameter tuning, and standardize, normalize, or perform logarithmic transformation on the selected features; Use the ensemble learning method to perform model fusion on the risk assessment model after re-selecting features; After model fusion is completed, a risk assessment model with model evaluation and optimization completed is obtained; Save the risk assessment model with model evaluation and optimization completed in a persistent format. The persistent format is PMML, ONNX, pickle file, or a custom binary format; Migrate the risk assessment model with model evaluation and optimization completed in the persistent format to the server. The server is signal-connected to the risk prediction port; The server transmits the risk assessment model with model evaluation and optimization completed in the persistent format to the risk prediction port.

[0037] Specifically, various performance metrics such as ROC curve, AUC value, precision, recall, F1 score, Brier score, and calibration curve are used to comprehensively evaluate the performance of the model, ensuring that the model can perform well in different dimensions. Parameter tuning methods such as grid search and random search can be used to systematically explore the parameter space and find the optimal model parameter configuration, thereby improving the prediction performance of the model. Re-selecting and transforming features (such as standardization, normalization, logarithmic transformation) can remove redundant features, improve the quality of features, and make the model more stable and efficient. Using ensemble learning methods for model fusion can further improve the prediction accuracy and robustness of the model because ensemble learning can combine the advantages of multiple models and reduce the overfitting risk of a single model. Saving the model in PMML, ONNX, pickle file or custom binary format facilitates cross-platform and cross-language migration and use of the model, improving the flexibility and scalability of the model. Migrating the model in persistent format to the server and connecting it to the risk prediction port signal enables the model to be easily integrated into the existing business system to achieve real-time or batch risk prediction. The entire process from model development to deployment is systematically planned to ensure the repeatability and verifiability of each step, which is beneficial for subsequent model maintenance and update. By continuously optimizing and deploying the risk assessment model, enterprises can more accurately identify and manage risks, improve the scientificity and efficiency of business decisions, and thus bring greater business value.

[0038] To solve the problem in the prior art that there is no effective risk warning and intervention plan formulation based on the risk causes and risk levels of patients, resulting in patients being unable to conduct further risk diagnosis according to their own conditions, please refer to Figures 1 - 3 , the following technical solutions are provided in this embodiment: A method for constructing a risk warning model, comprising the following steps: S4: Integrate the risk assessment results in the risk assessment model, and set a dynamic assessment mechanism for the integrated risk assessment results. After the dynamic assessment mechanism is set, patient risk prediction data is obtained; Among them, automated warning helps to detect and handle potential risks in a timely manner, providing strong support for clinical decision-making; S5: Judge the warning intensity of the patient risk prediction data, generate warning signals according to the warning level, generate automated intervention suggestions according to the generated warning signals, and construct a risk warning model with the generated automated intervention suggestions; Among them, by transmitting the intervention suggestions after visualization conversion to the mobile terminals of medical staff and patients for display, the instant transmission and sharing of information are realized; Integrate the risk assessment results in the risk assessment model in step S4, and set up a dynamic assessment mechanism for the integrated risk assessment results, including: Confirm the risk assessment results of the patients in the risk assessment model. After the confirmation is completed, perform dataset integration. After the dataset integration, a risk dataset is obtained; Set up a dynamic assessment mechanism for the risk dataset. The dynamic assessment mechanism is set to determine the risk assessment frequency of the patients. After the frequency confirmation is completed, set up a trigger condition. The trigger condition is that when the assessment threshold is used for risk level allocation, according to the allocated risk level, an early warning is automatically triggered; After the dynamic assessment mechanism is set up, patient risk prediction data is obtained.

[0039] Specifically, by integrating the results in the risk assessment model, a more comprehensive risk dataset can be obtained, thus more accurately reflecting the risk status of the patients. The setting of the dynamic assessment mechanism enables the risk assessment to be adjusted in a timely manner as the patient's condition changes, improving the timeliness of the assessment. The setting of the trigger condition enables the early warning mechanism to be automatically triggered when the assessment threshold reaches a specific risk level without manual intervention, improving the efficiency and accuracy of the early warning. The automated early warning helps to detect and handle potential risks in a timely manner, providing strong support for clinical decision-making. According to the risk assessment results and risk levels, medical institutions can allocate resources more targeted, focus on the management and intervention of high-risk patients, and reduce potential losses. The dynamic assessment mechanism enables medical institutions to continuously monitor risk changes, adjust risk management strategies in a timely manner, and ensure the safety of patients. Based on the integrated risk dataset and dynamic assessment results, medical staff can make more scientific and reasonable medical decisions, improving the treatment effect and patient satisfaction. The risk prediction data provides an important reference basis for clinical pathway planning, treatment plan selection, etc. By implementing this solution, medical institutions can establish a complete set of risk assessment and management systems, improving the overall management level.

[0040] For step S5 of the risk early warning model construction method, it includes: Judge the early warning intensity of the risk levels assigned to the patients in the patient risk prediction data. The higher the risk level, the higher the early warning intensity; Distinguish different early warning signals according to the early warning intensity. Among them, the early warning signals include visual signals, text descriptions, and sound signals; Generate automated intervention suggestions for the patients according to the early warning intensity and early warning signals; The automated intervention suggestions include treatment item guidelines, medical institution contact information, and emergency instructions; Visualize the generated automated intervention suggestions and transmit the visualized automated intervention suggestions to the mobile terminals of medical staff and patients for display.

[0041] Specifically, by directly associating the risk level assigned to the patient with the warning intensity, the system can accurately reflect the patient's risk status. The higher the risk level, the higher the warning intensity, which helps medical staff quickly identify high-risk patients and thus take timely and effective intervention measures. The warning signals include visual signals, text descriptions, and sound signals. This diverse signal form can attract the attention of medical staff in all aspects. Different signal forms can be targeted at different environments and situations, improving the flexibility and applicability of the warning. Automatically generate intervention suggestions based on the warning intensity and warning signals, which greatly reduces the workload of medical staff. The intervention suggestions include treatment item guidelines, contact information of medical institutions, and emergency instructions, providing comprehensive reference and support for medical staff. Visualize the generated automated intervention suggestions to make the information more intuitive and understandable. Transmit the visualized intervention suggestions to the mobile terminals of medical staff and patients for display, realizing the instant transmission and sharing of information. This method improves the accessibility and usability of information, helps medical staff and patients quickly obtain key information and make responses. Precise risk warnings and timely intervention suggestions help reduce medical risks and improve patient safety and satisfaction.

[0042] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0043] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention.

Claims

1. A liver cancer risk prediction model, characterized in that: Including risk assessment model and risk early warning model; Risk assessment models for: Collecting the patient's liver cancer-related data, and using the collected liver cancer-related data to quantitatively assess the risk of liver cancer, wherein the quantitative assessment includes risk factor analysis, risk assessment calculation, and risk level classification based on the patient's liver cancer-related data, and obtaining a risk assessment model after the quantitative assessment is completed; Risk early warning models for: The patients are assessed for liver cancer risk warning based on the risk assessment model. The liver cancer risk warning includes dynamic risk assessment, warning signal triggering and intervention suggestion generation. After the liver cancer risk warning is completed, a risk warning model is obtained.

2. A method for constructing a liver cancer risk prediction model, applied to the liver cancer risk prediction model of claim 1, characterized in that: The method for constructing a risk assessment model includes the following steps: S1: Collecting the patient's liver cancer related data from the database, performing data preprocessing after the liver cancer related data are collected, and obtaining the target collected data after the data preprocessing is completed; S2: Perform risk factor analysis, risk assessment calculation and risk level classification on the target collected data, and obtain a risk assessment model after the risk factor analysis, risk assessment calculation and risk level classification are completed; S3: Evaluate and optimize the risk assessment model, and then deploy the risk assessment model to the risk prediction port after the model evaluation and optimization are completed.

3. The method for constructing a liver cancer risk prediction model according to claim 2, characterized in that: The S1 step in the risk assessment model construction method includes: Collecting patients' liver cancer-related data from databases, including the hospital's electronic medical record system, health examination report system, and epidemiological survey system databases; Liver cancer-related data include clinical data, biomarker data, lifestyle information, and environmental exposure data; Clean and standardize liver cancer-related data; Data cleaning and data standardization include handling missing values, data denoising, and integrating data from different sources into a unified format; After data preprocessing is completed, the target collected data is obtained.

4. The method for constructing a liver cancer risk prediction model according to claim 2, characterized in that: The S1 step in the risk assessment model construction method also includes: Extract data sets corresponding to different data sources; Retrieve the cleaning process data corresponding to the data cleaning of each data set, wherein the cleaning process data includes the number of missing values ​​and their corresponding weight values ​​and format conversion time; The data quality coefficients corresponding to different data sources are obtained by using the number of missing values ​​of the data sets corresponding to different data sources, their corresponding weight values, and format conversion time; The data quality coefficient is obtained by the following formula: Among them, J represents the data quality coefficient; n represents the number of missing values ​​corresponding to the data source; w ni represents the data weight value corresponding to the i-th missing value; m represents the number of valid data corresponding to the data source; w mi represents the data weight value corresponding to the i-th valid data; T i Indicates the format conversion duration corresponding to the i-th valid data; T c represents the preset conversion time reference value; z represents the total number of collected data corresponding to the data source; The data quality coefficients corresponding to the different data sources are used to limit the collection data ratios of the liver cancer related data collected by different data sources.

5. The method for constructing a liver cancer risk prediction model according to claim 4, characterized in that: The data quality coefficients corresponding to the different data sources are used to limit the proportion of liver cancer-related data collected by different data sources, including: Compare the data quality coefficients corresponding to different data sources with the preset data quality coefficient thresholds; A data source whose data quality coefficient is not lower than a preset data quality coefficient threshold is used as a target data source; Extract the proportion of the current collected data of the target data source to the total collected data; Obtaining an upper limit value of the collected data ratio corresponding to the target data source by using the data quality coefficient corresponding to the target data source in combination with the data ratio; The upper limit of the collected data ratio is obtained by the following formula: Where B represents the upper limit of the proportion of collected data corresponding to the target data source; B c Indicates the proportion of the current collected data of the target data source to the total collected data; x indicates the number of data sources other than the current target data source; B xi represents the proportion of the collected data corresponding to the i-th data source other than the current target data source in the total collected data; J m Indicates the data quality coefficient corresponding to the target data source; J i Indicates the data quality coefficient corresponding to the i-th data source other than the current target data source; The data ratio of the target data source corresponding to the collected data is controlled according to the upper limit of the collected data ratio as a standard.

6. The method for constructing a liver cancer risk prediction model according to claim 3, characterized in that: The S2 step in the risk assessment model construction method includes: First, collect the target data for risk factor analysis, which includes exploratory data analysis, variable screening, and feature selection; Exploratory analysis is to calculate the mean, median, standard deviation, minimum and maximum values ​​of the target collected data, and visualize the calculation results in the form of charts. After visualization, the distribution and abnormality of the target collected data can be obtained; Variable screening involves univariate analysis and multivariate analysis based on the abnormalities of the target collected data. Univariate analysis involves using statistical tests to evaluate the relationship between each variable and liver cancer occurrence. Multivariate analysis involves using multiple regression analysis to evaluate the combined effects of multiple variables on liver cancer occurrence. Feature selection is to select the feature data that have a significant impact on liver cancer occurrence from the analysis results of univariate analysis and multivariate analysis of the target collected data obtained in variable screening using stepwise regression screening method; The data after feature selection in the target collected data is labeled as the risk factor analysis results.

7. The method for constructing a liver cancer risk prediction model according to claim 6, characterized in that: The S2 step in the risk assessment model construction method also includes: Perform risk assessment calculations based on the risk factor analysis results. Before performing risk assessment calculations, an assessment model is constructed based on the risk factor analysis results. The assessment model is a Cox proportional hazard model, a random forest model, or a gradient boosting tree model; The acquired evaluation model is trained, and the trained evaluation model is used to perform evaluation calculations on the risk factor analysis results; After the evaluation calculation is completed, the risk assessment model of the target collection data is obtained; Confirm the assessment threshold in the risk assessment model and classify the risk level according to the confirmed threshold; The risk level is divided into a threshold range based on the assessment threshold, and a risk level is assigned to each assessment threshold, and the risk levels include low risk, medium risk and high risk; After each assessment threshold is assigned to the corresponding risk level, a risk assessment model for the target collection data is obtained.

8. The method for constructing a liver cancer risk prediction model according to claim 7, characterized in that: The S3 step in the risk assessment model construction method includes: The risk assessment model was evaluated by cross-validation method for performance indicators, including ROC curve, AUC value, precision, recall, F1 score, Brier score and calibration curve; According to the performance index evaluation results, the risk assessment model is optimized by using parameter tuning methods, including grid search and random search; Re-select features for the risk assessment model after parameter optimization, and standardize, normalize or logarithmically transform the selected features; The risk assessment model after feature reselection is integrated using the ensemble learning method for model fusion; After the model fusion is completed, the risk assessment model with model evaluation and optimization is obtained; Save the risk assessment model after model evaluation and optimization in a persistent format, which can be PMML, ONNX, pickle file or custom binary format; Migrate the model evaluation and optimized risk assessment model in a persistent format to the server, and connect the server to the risk prediction port signal; The server transmits the model evaluation and optimized risk assessment model in a persistent format to the risk prediction port.

9. The method for constructing a liver cancer risk prediction model according to claim 8, characterized in that: The method for constructing a risk warning model includes the following steps: S4: Integrate the risk assessment results in the risk assessment model, and set a dynamic assessment mechanism for the integrated risk assessment results. After the dynamic assessment mechanism is set, the patient risk prediction data is obtained; S5: Determine the warning intensity of the patient risk prediction data, generate a warning signal according to the warning degree, generate an automated intervention suggestion according to the generated warning signal, and construct a risk warning model with the generated automated intervention suggestion; In step S4, the risk assessment results in the risk assessment model are integrated, and a dynamic assessment mechanism is set for the integrated risk assessment results, including: Confirm the risk assessment results of the patients in the risk assessment model, integrate the data sets after confirmation, and obtain the risk data sets after data set integration; A dynamic assessment mechanism is set for the risk data set. The dynamic assessment mechanism is set to determine the frequency of risk assessment for the patient. After the frequency is confirmed, a trigger condition is set. The trigger condition is set to automatically trigger an early warning according to the risk level assigned when the assessment threshold is assigned a risk level. After the dynamic assessment mechanism is set up, patient risk prediction data can be obtained.

10. The method for constructing a liver cancer risk prediction model according to claim 9, characterized in that: The S5 step of the risk warning model construction method includes: The risk level assigned to the patient in the patient risk prediction data is used to determine the warning intensity. The higher the risk level, the higher the warning intensity. Different warning signals are distinguished according to the warning intensity, among which the warning signals include visual signals, text descriptions and sound signals; Generate automated intervention recommendations for patients based on warning intensity and warning signals; Automated intervention recommendations include guidance on treatment options, healthcare facility contacts, and emergency instructions; The generated automated intervention suggestions are converted into visualizations, and the automated intervention suggestions after the visualization conversion are transmitted to the mobile terminals of medical staff and patients for display.

Citation Information

Patent Citations

  • Target group thyroid cancer risk prediction model construction method

    CN116543898A

  • Data quality evaluation method and apparatus, computer readable storage medium, and terminal

    CN107633257A

  • Electric power system data quality assessment method and device, and storage medium

    CN108022046A

  • Data quality assessment method

    CN108334636A

  • A visualized modeling system and method for public security traffic management based on big data technology

    CN109189846A