A data filling model selection and health evaluation method and device

By constructing datasets with different missing proportions, evaluating multiple data imputation models, and selecting the optimal imputation model, the problem of missing medical data was solved, and the reliability of the data and the accuracy of health assessment were improved.

CN116483817BActive Publication Date: 2026-04-24SHANXI MEDICAL UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANXI MEDICAL UNIV
Filing Date
2023-04-25
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Due to individual differences in clinical treatment and inaccurate recording, medical data is often missing, affecting the accuracy of disease diagnosis and treatment. Existing technologies are insufficient to effectively fill in the missing data to obtain true, reliable, and complete data.

Method used

By constructing datasets with different missing proportions, evaluating the imputation results using multiple data imputation models, selecting the optimal imputation model, and combining it with a health assessment model for data imputation and evaluation, the contribution of missing feature data is determined.

Benefits of technology

It improves the accuracy and reliability of missing data filling, enhances the precision and completeness of health assessments, and ensures the credibility of assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116483817B_ABST
    Figure CN116483817B_ABST
Patent Text Reader

Abstract

The application provides a data filling model selection, health evaluation method and device, wherein the data filling model selection method comprises the following steps: obtaining a source data set, wherein the source data set comprises multiple groups of data, and each group of data comprises data of a preset type; preprocessing each type of data in the source data set to construct multiple groups of data sets with different missing ratios; inputting each group of data set with a missing ratio into multiple data filling models in sequence to obtain data filling results; and selecting a target data filling model corresponding to each group of data set with a missing ratio from the multiple data filling models according to the data filling results. The application can solve the technical problem of how to fill the data with missing values to obtain real, reliable and complete target data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data completion model selection, health assessment method and apparatus. Background Technology

[0002] Medical data resources are widely used to develop clinical decision support systems for the diagnosis and prediction of various diseases, and are of great value for disease diagnosis, treatment, and medical research. However, due to individual differences in clinical treatment or inaccurate recording and input of information, gaps in patients' clinical medical data are unavoidable. Incomplete medical data has low reference value for disease prediction, diagnosis, treatment, and medical research, and may lead to misdiagnosis, mistreatment, and misprediction. Therefore, how to fill in the missing data to obtain accurate, reliable, and complete target data has become an urgent technical problem to be solved. Summary of the Invention

[0003] Therefore, in order to solve the technical problem of how to fill in missing data to obtain real, reliable and complete target data, the present invention provides a data filling model selection, health assessment method and apparatus.

[0004] In a first aspect, embodiments of the present invention disclose a data imputation model selection method, comprising:

[0005] Obtain the source dataset, which includes multiple sets of data, each set containing data of a preset type; preprocess each type of data in the source dataset to construct multiple datasets with different missing proportions; input each dataset with a missing proportion into multiple data imputation models in sequence to obtain data imputation results; based on the data imputation results, select the target data imputation model corresponding to each dataset with a missing proportion from the multiple data imputation models.

[0006] The data imputation model construction method provided by this invention preprocesses each type of data in the source dataset to construct multiple datasets with different missing proportions. Each dataset with a missing proportion is then sequentially input into multiple data imputation models to obtain data imputation results. Based on the data imputation results, a target data imputation model corresponding to each dataset with a missing proportion is selected from the multiple data imputation models. This allows different data imputation models to impute preset types of data under corresponding missing proportions, obtaining the optimal imputation result for each type of data under the corresponding missing proportion condition. This makes the imputed data closer to the real data and improves the reliability of the imputed data.

[0007] Optionally, based on the data imputation results, a data imputation model corresponding to the missing proportion of each dataset is selected from multiple data imputation models, specifically including:

[0008] The data imputation model with the best imputation result corresponding to the first group of missing proportions in the first type is selected from multiple data imputation models and used as the target imputation model corresponding to the first group of missing proportions. Here, the first type is any type in the source dataset, and the first group of missing proportions is any set of datasets with different missing proportions.

[0009] Optionally, when the missing proportion of the first type of dataset is less than or equal to a preset proportion threshold, and the preset proportion threshold is a missing proportion set based on the missing proportion of the dataset with the first group of missing proportions, then the target imputation model corresponding to the first type of dataset is determined to be the data imputation model with the best imputation result corresponding to the dataset with the first group of missing proportions.

[0010] Alternatively, if the missing proportion of a dataset of the first type is greater than or equal to a preset proportion threshold, then the target imputation model corresponding to the dataset of the first type is determined to be the data imputation model with the best imputation result corresponding to the dataset of the second missing proportion in the first type, wherein the missing proportion of the dataset of the second missing proportion is greater than the first missing proportion.

[0011] Secondly, embodiments of the present invention disclose a health assessment method, including: obtaining the original feature dataset of the target object;

[0012] Missing feature data were obtained based on the analysis of the original feature dataset;

[0013] The missing feature data in the original feature dataset is filled in using any data imputation model selection method in the first aspect, resulting in the imputed feature dataset;

[0014] The imputed feature dataset is input into the pre-built health assessment model to obtain the health assessment results of the target object.

[0015] Optionally, after inputting the imputed feature dataset into the pre-built health assessment model to obtain the health assessment results for the target object, the following steps are also included:

[0016] Based on the imputed feature dataset, the health assessment results of the target object, and the health assessment model, the health assessment results of the target object are analyzed to determine the contribution of different feature data in the imputed feature dataset to the health assessment results.

[0017] Optionally, based on the imputed feature dataset, the health assessment results of the target object, and the health assessment model, the health assessment results of the target object are analyzed to determine the contribution of different feature data in the imputed feature dataset to the health assessment results, specifically including:

[0018] Arbitrarily arrange and combine different feature data in the imputed feature dataset to obtain multiple feature datasets with different arrangement orders;

[0019] The feature data in each set of feature datasets are input into the health assessment model in sequence, and the health assessment results of the target object are obtained after each input feature data.

[0020] Based on feature datasets with different arrangement orders and the corresponding health assessment results of target objects, the contribution of different feature data in the imputed feature dataset to the health assessment results is determined.

[0021] A third aspect of the present invention provides a data imputation model selection apparatus, comprising:

[0022] The first acquisition module is used to acquire the source dataset, which includes multiple sets of data, each set of data including data of a preset type;

[0023] The first processing module is used to preprocess each type of data in the source dataset to construct multiple datasets with different missing proportions.

[0024] The first input module is used to sequentially input the dataset of each missing proportion into multiple data imputation models to obtain the data imputation results;

[0025] The first selection module is used to select the target data imputation model corresponding to each set of missing proportions from multiple data imputation models based on the data imputation results.

[0026] The functions performed by each component in the data filling model selection device provided by the present invention have been applied in any of the method embodiments of the first aspect described above, and therefore will not be repeated here.

[0027] A fourth aspect of the present invention provides a health assessment device, comprising: a second acquisition module, used to acquire the original feature dataset of a target object;

[0028] The first analysis module is used to analyze the original feature dataset to obtain the missing feature data;

[0029] The first filling module is used to fill in the missing feature data in the original feature data of the target object using any of the data filling model selection methods in the second aspect, so as to obtain the filled target object feature data.

[0030] The second input module is used to input the filled feature dataset into the pre-built health assessment model to obtain the health assessment results of the target object.

[0031] The functions performed by each component in the health assessment device provided by the present invention have been applied in any of the method embodiments of the second aspect described above, and therefore will not be repeated here.

[0032] The fifth aspect of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to execute the program stored in the memory to implement the steps of the data filling model selection method of the first aspect or the steps of the health assessment method of the second aspect.

[0033] The sixth aspect of the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform a data imputation model selection method as provided in the first aspect of the present invention, or to perform a health assessment method as provided in the second aspect of the present invention. Attached Figure Description

[0034] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0035] Figure 1 This is a schematic diagram of a data imputation model selection method according to an embodiment of the present invention;

[0036] Figure 2 This is a schematic diagram of a health assessment method provided in an embodiment of the present invention;

[0037] Figure 3 This is a schematic diagram of a data imputation model selection device provided in an embodiment of the present invention;

[0038] Figure 4 This is a schematic diagram of a health assessment device provided in an embodiment of the present invention;

[0039] Figure 5 This is a schematic diagram of an electronic device structure provided in an embodiment of the present invention. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0041] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “a,” “an,” or “the,” as used in this disclosure, do not indicate a limitation of quantity, but rather indicate the presence of at least one. Terms such as “comprising” or “including” mean that an element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects.

[0042] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0043] To address the technical problems mentioned in the background section, embodiments of the present invention provide a data imputation model selection method, such as... Figure 1 As shown, the steps of this method include:

[0044] Step S110: Obtain the source dataset, which includes multiple sets of data, each set of data including data of a preset type.

[0045] Specifically, the source dataset can refer to a complete dataset without missing data. The source dataset is obtained by removing incomplete data types from the original dataset with missing data, thus extracting the complete dataset, i.e., the source dataset. This ensures that all data in the obtained source dataset is real data, thereby improving the accuracy of the data imputation model selection. The multiple sets of data included in the source dataset can refer to multiple target objects, each target object corresponding to a set of data. Predefined data types can refer to multiple variable factors corresponding to each target object, with each variable factor representing a different data type.

[0046] For example, as an optional embodiment, the original dataset may include 200 target objects, each target object including 30 variable factors. These 30 variable factors can be represented as X1, X2, X3, X4, X5, X6, X7, X8, ..., X30. The variable factors may include continuous variables and categorical variables. The original dataset can be structured with the target objects as rows and the variable factors of different types as columns. Removing rows or columns with missing data yields a complete dataset without missing data, which serves as the source dataset. Alternatively, the source dataset may include 100 target objects, each target object including 9 variable factors, which can be represented as X1, X2, X3, X4, X5, X6, X7, X8, and X9.

[0047] Step S120: Preprocess each type of data in the source dataset to construct multiple datasets with different missing proportions.

[0048] Specifically, preprocessing includes determining the missing percentage and missing mechanism of the source dataset. The missing percentage can be the percentage of missing data relative to the total data of the current data type, such as generating datasets with different missing percentages like 5%, 10%, 15%, and 30%. The missing mechanism can be the form of missing data, for example, it can be Missing Completely at Random (MCAR), Missing at Random (MAR), or Missing Not at Random (MNAR).

[0049] For example, consider a source dataset containing 100 target objects, each with 9 variable factors, such as X1, X2, X3, X4, X5, X6, X7, X8, and X9. The dataset is structured as 100 rows and 9 columns, with the 100 target objects as rows and the 9 variable factors as columns. Random missing values ​​are applied to the X1 variable factor, resulting in a first dataset with a 5% missing value, a second dataset with a 10% missing value, a third dataset with a 15% missing value, a fourth dataset with a 30% missing value, and so on. The same applies to other variable factors. It should be noted that even with the same missing value, different random missing values ​​exist, so there can be multiple datasets with missing values ​​of 5% / 10% / 15% / 30% for each column. This is not limited here, and those skilled in the art can determine the appropriate dataset based on the specific circumstances.

[0050] Step S130: Input the datasets of each missing proportion into multiple data imputation models in sequence to obtain the data imputation results.

[0051] Specifically, this embodiment provides a classifier-optimized hybrid imputation method for mixed datasets (HEIM), where the data imputation model can include, but is not limited to, Factorial Analysis for Mixed Data (FAMD), mRF, Multivariate Imputation by Chained Equations (MICE), and K Nearest Neighbors Imputation (KNN). It should be noted that in this embodiment, each data imputation model imputes different types of variable factors in each set of missing proportions. This approach considers both rows to account for the correlations between different types of variable factors within the same set and columns to account for statistical regularities and similarities among similar data, thereby improving the accuracy of data imputation.

[0052] For example, taking preset missing percentages of 5%, 10%, 15%, and 30% as examples, multiple sets of data with a missing percentage of 5% for variable factors X1-X9 are constructed respectively. The data with a missing percentage of 5% for X1-X9 are input into each data imputation model to impute the missing data, and the process is iterated a preset number of times to obtain the data imputation results. Other missing percentages are handled similarly, and will not be elaborated here.

[0053] Step S140: Based on the data imputation results, select the target data imputation model from multiple data imputation models that corresponds to the dataset with the missing proportion for each group.

[0054] Specifically, by recording the normalized root mean square error (NRMSE) of each imputation model for continuous variables and the proportion of falsely classified (PFC) for categorical variables, the optimal imputation method corresponding to each missing variable factor is selected.

[0055] For example, taking HEIM as an example, which includes FAMD, mRF, MICE, and KNN data imputation models, as shown in the HEIM algorithm:

[0056] Algorithm HEIM()

[0057] {

[0058] Input the original dataset DataA with missing data.

[0059] Perform column deletion to obtain the source dataset DataB

[0060] Generating random missing data (5 / 10 / 15 / 30%) based on DataB simData

[0061] The FAMD algorithm is used to predict missing values ​​in simData, resulting in simData1.

[0062] mRF is used to predict missing values ​​in simData, generating simData2.

[0063] The MICE algorithm is used to predict missing values ​​in simData, generating simData3.

[0064] Apply KNN to predict missing values ​​in simData, generating simData4.

[0065] Assign the dataset simDataN to DataC.

[0066] N = Number of variables in DataB

[0067] For (i = 1 to N)

[0068] {

[0069] Replace the i-th column of DataC with the i-th column of SimDta1, and use a classifier to calculate NRMSE / PFC.

[0070] Replace the i-th column of DataC with the i-th column of SimDta2, and use a classifier to calculate NRMSE / PFC.

[0071] Replace the i-th column of DataC with the i-th column of SimDta3, and use a classifier to calculate NRMSE / PFC.

[0072] Replace the i-th column of DataC with the i-th column of SimDta4, and use a classifier to calculate NRMSE / PFC.

[0073] Choose the dataset with the smallest NRMSE / PFC.

[0074] Replace the i-th column of dataset DataC with the dataset selected in the previous step.

[0075] }

[0076] Output the padded DataC}

[0077] The data imputation model selection method provided in this invention preprocesses each type of data in the source dataset to construct multiple datasets with different missing proportions. Each dataset with a missing proportion is sequentially input into multiple data imputation models to obtain data imputation results. Based on the data imputation results, a target data imputation model corresponding to each dataset with a missing proportion is selected from the multiple data imputation models. This leverages the different data imputation models to impute preset types of data under corresponding missing proportions, obtaining the optimal imputation result for each type of data under the corresponding missing proportion condition. This makes the imputed data closer to the real data and improves the reliability of the imputed data.

[0078] As an optional embodiment of the present invention, step S140 includes:

[0079] Step S210: Select the data imputation model with the best imputation result from multiple data imputation models that corresponds to the dataset with the first missing proportion in the first type, and use it as the target imputation model corresponding to the dataset with the first missing proportion. Here, the first type is any type in the source dataset, and the dataset with the first missing proportion is any set of datasets with different missing proportions.

[0080] Specifically, by recording the normalized root mean square error (NRMSE) of each imputation model for continuous variables and the proportion of falsely classified (PFC) for categorical variables, the optimal imputation method corresponding to each missing variable factor is selected. For example, if the first data type is X1 and the optimal imputation method for the first dataset with a missing percentage of 5% is FAMD, then the FAMD data imputation model is used as the target imputation model for the first data type X1 with a missing percentage of 5%. Similarly, if the first data type is X1 and the optimal imputation method for the first dataset with a missing percentage of 10% is MICE, then the MICE data imputation model is used as the target imputation model for the first data type X1 with a missing percentage of 10%. Likewise, if the first data type is X3 and the optimal imputation method for the first dataset with a missing percentage of 5% is mRF, then the mRF data imputation model is used as the target imputation model for the first data type X3 with a missing percentage of 5%. And so on. The same principle applies to determining the target imputation model for other data types and missing percentages, and will not be elaborated further here.

[0081] The data imputation model selection method provided in this embodiment of the invention selects the data imputation model with the best imputation result corresponding to the first group of missing proportions in the first type of dataset from multiple data imputation models, and uses it as the target imputation model corresponding to the first group of missing proportions of dataset. By selecting the target data imputation model corresponding to each group of missing proportions of dataset from multiple data imputation models, different data imputation models are used to imput the preset type of data under the corresponding missing proportions, so as to obtain the optimal imputation result of each type of data under the corresponding missing proportion conditions, making the imputed data closer to the real data and improving the reliability of the imputed data.

[0082] As an optional embodiment of the present invention, when the missing proportion of a dataset of the first type is less than or equal to a preset proportion threshold, and the preset proportion threshold is a missing proportion set based on the missing proportion of the dataset of the first group of missing proportions, then the target imputation model corresponding to the dataset of the first type is determined to be the data imputation model with the best imputation result corresponding to the dataset of the first group of missing proportions; or, when the missing proportion of a dataset of the first type is greater than or equal to the preset proportion threshold, then the target imputation model corresponding to the dataset of the first type is determined to be the data imputation model with the best imputation result corresponding to the dataset of the second group of missing proportions in the first type, wherein the missing proportion of the dataset of the second group of missing proportions is greater than the first missing proportion.

[0083] Specifically, each type of data in the source dataset is preprocessed to construct multiple datasets with different missing proportions. These multiple missing proportions are preset sets of different values. Values ​​that do not belong to the preset sets of different values ​​are categorized according to preset rules.

[0084] For example, consider a data imputation model selection method with four preset percentage thresholds: 5%, 10%, 15%, and 30%. These four preset percentage thresholds do not cover all percentages between 0-30% or 0-100%, so the uncovered values ​​need to be categorized using preset rules. After determining the target data imputation model for each set of missing percentages using the data imputation model selection method, as an optional embodiment, the preset rules can map 0-5% to 5%; 5%-10% to 10%; 10%-15% to 15%; and 15%-30% to 30%. When the missing percentage is greater than 30%, it can be considered that there is too much missing data, resulting in low accuracy of the imputation effect. Those skilled in the art can set these values ​​according to actual needs; no restrictions are imposed here. It should be noted that if the current missing percentage equals the preset percentage threshold, the preset percentage threshold corresponding to the current missing percentage can be determined by those skilled in the art based on the actual situation; no restrictions are imposed here. For example, if the optimal imputation method for a dataset with the first type being X1 and the first group of missing data having a missing percentage of 5% is FAMD, then the optimal imputation method for a dataset with the first type being X1 and the first group of missing data having a missing percentage of <5% can be set as FAMD. Alternatively, if the optimal imputation method for a dataset with the first type being X1 and the first group of missing data having a missing percentage of 10% is MICE, then the optimal imputation method for a dataset with the first type being X1 and the first group of missing data having a missing percentage of greater than 5% and less than 10% can be set as MICE.

[0085] For example, after determining the target data imputation model for each dataset with missing proportions using various data imputation model selection methods, the same type of data may correspond to different optimal data imputation models under different missing proportions. In this case, the optimal data imputation model for that type of data can be determined by the missing proportion of the current data type in the original data. For instance, when the first type is X1, the optimal imputation method for the first dataset with a missing proportion of 5% is FAMD; when the first type is X1, the optimal imputation method for the first dataset with a missing proportion of 10% is MICE. However, if the missing proportion of variable factor X1 in the original dataset is 3%, then the FAMD data imputation model can be selected as the optimal imputation model for variable factor X1. The same logic applies to other variable factors and missing proportions, and will not be elaborated further here.

[0086] The data imputation model selection method provided in this invention uses preset ratio thresholds to map different missing ratios to corresponding preset ratio thresholds. On the one hand, a limited number of preset ratio thresholds cover a relatively large range, allowing different data imputation models to impute preset types of data under corresponding missing ratios, obtaining the optimal imputation result for each type of data under the corresponding missing ratio condition. This makes the imputed data closer to the real data, improving the reliability of the imputed data while reducing the computational load of the data imputation model selection process and improving selection efficiency. On the other hand, it can effectively balance the relationship between the missing ratio range, the computational power of the data imputation model selection process, and the reliability of the imputed data.

[0087] This invention provides a health assessment method, such as... Figure 2 As shown, the steps of this method include:

[0088] Step S410: Obtain the original feature dataset of the target object.

[0089] Specifically, this embodiment uses patients with cardiovascular diseases, such as coronary heart disease, congestive heart failure, heart attack, stroke, and angina, as an example. However, this is not a limitation, and those skilled in the art can determine the target subject based on the actual situation. The original feature dataset corresponds to the target subject. This embodiment uses an original feature dataset including at least one of the following: vitamin exposure, age, alcohol consumption, smoking, glomerular filtration rate, γ-GT, creatinine, and total bilirubin. The above data can be obtained through medical device measurement and patient self-reporting, or it can be historical measurement data.

[0090] Step S420: Analyze the original feature dataset to obtain the missing feature data.

[0091] Specifically, a reference indicator dataset is obtained, which includes complete data types used in health assessments. The original feature dataset is then compared and analyzed with the complete data types to identify missing feature data in the original feature dataset.

[0092] For example, the complete data types include vitamin exposure, age, alcohol consumption, smoking, glomerular filtration rate, γ-GT, creatinine, and total bilirubin. The original feature dataset includes vitamin exposure, age, alcohol consumption, smoking, and glomerular filtration rate; analysis reveals that the missing feature data are γ-GT, creatinine, and total bilirubin.

[0093] Step S430: Use any of the above data imputation model selection methods to impute the missing feature data in the original feature dataset to obtain the imputed feature dataset.

[0094] Specifically, the optimal imputation model for different variable factors has been determined based on the data imputation model selection method. The missing feature data in the original feature dataset is then imputed using this method, resulting in the imputed feature dataset, i.e., the complete dataset.

[0095] Step S440: Input the filled feature dataset into the pre-built health assessment model to obtain the health assessment results of the target object.

[0096] Specifically, this embodiment can, but is not limited to, utilize tree models such as CatBoost (Categorical Features Gradient Boosting), Decision Tree (DT), Random Forest (RF), LightGBM (Light Gradient Boosting Machine), and eXtreme Gradient Boosting (XGBoost) to construct a health assessment model. The model is evaluated on a test set using one or more of the following metrics: Area Under Curve (AUC), Accuracy, Recall, Precision, F1 score, and Kappa score. Alternatively, different models can be visualized and evaluated by plotting the receiver operating characteristic curve (ROC), the model's calibration curve, and the Kolmogorov-Smirnov curve to determine the optimal model. The methods for constructing prediction models are relatively mature and will not be elaborated upon here.

[0097] For example, the original feature datasets of four groups of patients—male patient 1, male patient 2, female patient 1, and female patient 2—were obtained, and missing data were imputed to obtain four complete datasets. The imputed complete data were then input into a health prediction model to predict the patients' risk of all-cause mortality within 15 years, as shown in Table 1.

[0098] Table 1

[0099]

[0100] As shown in Table 1, compared with male patient 1, male patient 2, despite being under 62 years old, having high vitamin exposure, and having quit smoking for more than six months, had a 76.5% lower risk of all-cause mortality within 15 years, indicating that these three characteristics have a higher risk factor effect value for male patients. Compared with female patient 1, female patient 2, with being under 62 years old, having high vitamin exposure, having quit smoking for more than six months, and having lower total bilirubin, had a 58.89% lower risk.

[0101] The health assessment method provided in this invention uses a health prediction model to predict the health status of a target. The health prediction model can more effectively explore the correlation between variables and thus make more accurate predictions.

[0102] As an optional embodiment of the present invention, it further includes:

[0103] Step S510: Based on the imputed feature dataset, the health assessment results of the target object, and the health assessment model, analyze the health assessment results of the target object and determine the contribution of different feature data in the imputed feature dataset to the health assessment results.

[0104] Specifically, as an optional embodiment of the present invention, different feature data in the imputed feature dataset are arbitrarily arranged and combined to obtain multiple feature datasets with different arrangement orders. The feature data in each feature dataset are sequentially input into a health assessment model to obtain the health assessment result of the target object after each input feature data. Based on the multiple feature datasets with different arrangement orders and the corresponding health assessment results of the target objects, the contribution of different feature data in the imputed feature dataset to the health assessment result is determined.

[0105] For example, using the Catboost model as a health prediction model, the SHAP (SHapley Additive exPlanations) value calculated by the CatBoost model is applied to further calculate the Shapley value, and the feature attribution is visualized. SHAP is used to demonstrate the impact of different risk factors on all-cause mortality in cardiovascular patients. The importance and stability of eight risk factors—vitamin exposure, age, alcohol consumption, smoking, glomerular filtration rate, γ-GT, creatinine, and total bilirubin—are ranked in the general population, for men, and for women, respectively.

[0106] The SHAP value, based on the Shapley value, quantifies the contribution of each feature to the model's prediction and is a method of post-hoc explanation. SHAP constructs an additive explanatory model where all features are considered "contributors." For each predicted sample, the model generates a predicted value, and the SHAP value is the numerical value assigned to each feature in that sample.

[0107] It's important to note that Shapley value is a concept from game theory. It's a fair and quantitative assessment of the marginal contribution of a feature. It's a method to describe the "weight" or "importance" of a specific feature when a model predicts a particular data point. Positive or negative values ​​indicate the direction of the effect. The basic idea can be understood as calculating the marginal contribution of a feature when added to the model, then considering the different marginal contributions of that feature across all feature sequences, and finally averaging the results to obtain the Shapley value for that feature. In other words, it calculates the contribution value (Shapley Value) of each feature variable in each sample, and then sums the Shapley values ​​corresponding to each feature variable to explain how each feature variable affects the model's prediction.

[0108] This paper employs a permutation-based feature importance method for interpretation. This method measures feature importance by calculating the increase in model prediction error after rearranging the input feature order. If rearranging a feature's values ​​increases the model error, the feature is considered "important"; otherwise, rearranging the values ​​does not change the model error, the feature is considered "unimportant."

[0109] It should be noted that interpretability analysis based on predictive models is an existing and relatively mature technology, and will not be elaborated on here.

[0110] The health assessment method provided in this invention analyzes the health assessment results of the target object using the filled feature dataset, the health assessment results of the target object, and the health assessment model. It determines the contribution of different feature data in the filled feature dataset to the health assessment results, making the influence of each feature data used for prediction on the prediction results transparent.

[0111] Figure 3 A data imputation model selection device provided in one embodiment of the present invention includes:

[0112] The first acquisition module 710 is used to acquire a source dataset, wherein the source dataset includes multiple sets of data, each set of data including data of a preset type. For details, please refer to the description of the corresponding parts in the above embodiments, which will not be repeated here.

[0113] The first processing module 720 is used to preprocess each type of data in the source dataset to construct multiple datasets with different missing proportions. For details, please refer to the descriptions of the corresponding parts in the above embodiments, which will not be repeated here.

[0114] The first input module 730 is used to sequentially input the datasets of each missing proportion into multiple data imputation models to obtain the data imputation results. For details, please refer to the descriptions of the corresponding parts in the above embodiments, which will not be repeated here.

[0115] The first selection module 740 is used to select, based on the data imputation results, a target data imputation model corresponding to the dataset with the missing proportion from multiple data imputation models. For details, please refer to the description of the corresponding part in the above embodiments, which will not be repeated here.

[0116] As an optional embodiment of the present invention, the first selection module 740 includes:

[0117] The second selection module is used to select the data imputation model with the best imputation result from multiple data imputation models that corresponds to the dataset with the first missing proportion in the first type, as the target imputation model corresponding to the dataset with the first missing proportion. Here, the first type can be any type in the source dataset, and the dataset with the first missing proportion can be any set of datasets with different missing proportions from multiple sets of datasets. For details, please refer to the description of the corresponding part in the above embodiments, which will not be repeated here.

[0118] As an optional implementation of the present invention, the first determining module is used to determine the target imputation model corresponding to the first type of dataset as the data imputation model with the best imputation result corresponding to the dataset with the first group of missing proportions when the missing proportion of the dataset of the first type of dataset is less than or equal to a preset proportion threshold, and the preset proportion threshold is a missing proportion set based on the missing proportion of the dataset of the first group of missing proportions.

[0119] Alternatively, the second determining module is used to determine, when the missing proportion of a dataset of the first type is greater than or equal to a preset proportion threshold, the target imputation model corresponding to the dataset of the first type as the data imputation model with the optimal imputation result corresponding to the dataset of the second group of missing proportions in the first type, wherein the missing proportion of the dataset of the second group of missing proportions is greater than the first missing proportion. For details, please refer to the description of the corresponding parts in the above embodiments, which will not be repeated here.

[0120] Figure 4 A health assessment device provided in one embodiment of the present invention includes:

[0121] The second acquisition module 810 is used to acquire the original feature dataset of the target object. For details, please refer to the description of the corresponding part in the above embodiments, which will not be repeated here.

[0122] The first analysis module 820 is used to analyze the original feature dataset to obtain the missing feature data. For details, please refer to the description of the corresponding part in the above embodiments, which will not be repeated here.

[0123] The first filling module 830 is used to fill in the missing feature data in the original feature data of the target object using any of the above-mentioned data filling model selection methods, to obtain the filled target object feature data. For details, please refer to the description of the corresponding part in the above embodiments, which will not be repeated here.

[0124] The second input module 840 is used to input the imputed feature dataset into the pre-built health assessment model to obtain the health assessment results of the target object. For details, please refer to the description of the corresponding parts in the above embodiments, which will not be repeated here.

[0125] As an optional embodiment of the present invention, it further includes:

[0126] The third determining module is used to analyze the health assessment results of the target object based on the imputed feature dataset, the health assessment results of the target object, and the health assessment model, and to determine the contribution of different feature data in the imputed feature dataset to the health assessment results. For details, please refer to the corresponding descriptions in the above embodiments, which will not be repeated here.

[0127] As an optional embodiment of the present invention, the third determining module includes:

[0128] The first permutation module is used to arbitrarily arrange and combine different feature data in the imputed feature dataset to obtain multiple feature datasets with different permutation orders. For details, please refer to the corresponding descriptions in the above embodiments, which will not be repeated here.

[0129] The third input module is used to sequentially input the feature data from each set of feature datasets into the health assessment model, thereby obtaining the health assessment result of the target object after each input feature data. For details, please refer to the corresponding descriptions in the above embodiments, which will not be repeated here.

[0130] The fourth determining module is used to determine the contribution of different feature data in the imputed feature dataset to the health assessment results, based on feature datasets with different arrangement orders and the corresponding health assessment results of the target objects. For details, please refer to the descriptions of the corresponding parts in the above embodiments, which will not be repeated here.

[0131] This invention provides an electronic device, such as... Figure 5 As shown, the device includes one or more processors 3010 and a memory 3020, the memory 3020 including persistent memory, volatile memory, and a hard disk. Figure 5 Taking a processor 3010 as an example, the device may also include an input device 3030 and an output device 3040.

[0132] The processor 3010, memory 3020, input device 3030, and output device 3040 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0133] Processor 3010 may include, but is not limited to, a Central Processing Unit (CPU) and a Graphics Processing Unit (GPU). Processor 3010 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations thereof. The general-purpose processor may be a microprocessor or any conventional processor. Memory 3020 may include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created by the data filling model selection device or by the use of a health assessment device. Furthermore, memory 3020 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 3020 may optionally include memory remotely located relative to processor 3010, and this remote memory may be connected via a network to a data completion model selection device or a health assessment device. Input device 3030 may receive calculation requests (or other numeric or character information) input by the user, and generate key signal inputs related to the data completion model selection device or the health assessment device. Output device 3040 may include a display device such as a screen for outputting calculation results.

[0134] This invention provides a computer-readable storage medium that stores computer instructions. The computer-readable storage medium stores computer-executable instructions that can execute the data filling model selection method or the health assessment method described in any of the above method embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium may also include combinations of the above types of memory.

[0135] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable storage medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable storage media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), fiber optic devices, and compact disc read-only memory (CDROM). Furthermore, computer-readable storage media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0136] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0137] In the description of this specification, the references to terms such as "this embodiment," "an embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples, without contradiction. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0138] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for selecting a data imputation model, characterized in that, include: Obtain the source dataset, wherein the source dataset includes multiple sets of data, each set of data including data of a preset type; Preprocess each type of data in the source dataset to construct multiple datasets with different missing proportions; Each dataset with a missing proportion is sequentially input into multiple data imputation models to obtain the data imputation results; Based on the data imputation results, a target data imputation model corresponding to each set of missing proportions is selected from multiple data imputation models. The step of selecting a data imputation model corresponding to each set of missing proportions from multiple data imputation models based on the data imputation results specifically includes: The data imputation model with the best imputation result corresponding to the first group of missing proportions in the first type is selected from multiple data imputation models and used as the target data imputation model corresponding to the first group of missing proportions. The first type is any type in the source dataset, and the first group of missing proportions is any set of datasets with different missing proportions. When the missing proportion of a dataset of the first type is less than or equal to a preset proportion threshold, and the preset proportion threshold is a missing proportion set based on the missing proportion of the dataset of the first group of missing proportions, then the target imputation model corresponding to the dataset of the first type is determined to be the data imputation model with the best imputation result corresponding to the dataset of the first group of missing proportions. Alternatively, if the missing proportion of a dataset of the first type is greater than or equal to the preset proportion threshold, then the target imputation model corresponding to the dataset of the first type is determined to be the data imputation model with the best imputation result corresponding to the dataset of the second group of missing proportions in the first type, wherein the missing proportion of the dataset of the second group of missing proportions is greater than the missing proportion of the first group.

2. A health assessment method, characterized in that, include: Obtain the original feature dataset of the target object; Missing feature data were obtained based on the analysis of the original feature dataset; The missing feature data in the original feature dataset is filled in using the data imputation model selection method described in claim 1 to obtain the imputed feature dataset; The filled feature dataset is input into the pre-built health assessment model to obtain the health assessment results of the target object.

3. The method according to claim 2, characterized in that, After inputting the filled feature dataset into the pre-built health assessment model to obtain the health assessment result of the target object, the method further includes: Based on the augmented feature dataset, the health assessment results of the target object, and the health assessment model, the health assessment results of the target object are analyzed to determine the contribution of different feature data in the augmented feature dataset to the health assessment results.

4. The method according to claim 3, characterized in that, The step involves analyzing the health assessment results of the target object based on the augmented feature dataset, the health assessment results of the target object, and the health assessment model, to determine the contribution of different feature data in the augmented feature dataset to the health assessment results. Specifically, this includes: Arbitrarily arrange and combine different feature data in the filled feature dataset to obtain multiple feature datasets with different arrangement orders; The feature data in each set of feature datasets are sequentially input into the health assessment model to obtain the health assessment result of the target object after each input of the feature data. Based on the feature datasets with different arrangement orders and the corresponding health assessment results of the target object, the contribution of different feature data in the filled feature dataset to the health assessment results is determined.

5. A data imputation model selection device, characterized in that, include: The first acquisition module is used to acquire a source dataset, wherein the source dataset includes multiple sets of data, and each set of data includes data of a preset type; The first processing module is used to preprocess each type of data in the source dataset to construct multiple datasets with different missing proportions. The first input module is used to sequentially input the dataset of each missing proportion into multiple data imputation models to obtain the data imputation results; The first selection module is used to select, from multiple data imputation models, a target data imputation model corresponding to each set of missing proportions of the dataset, based on the data imputation results. The second selection module is used to select the data imputation model with the best imputation result from multiple data imputation models that corresponds to the dataset with the first missing proportion in the first type, as the target imputation model corresponding to the dataset with the first missing proportion. Here, the first type is any type in the source dataset, and the dataset with the first missing proportion is any set of datasets with different missing proportions from multiple sets of datasets. The first determining module is used to determine the target imputation model corresponding to the first type of dataset as the data imputation model with the best imputation result corresponding to the dataset with the first group of missing proportions when the missing proportion of the dataset of the first type of dataset is less than or equal to a preset proportion threshold, and the preset proportion threshold is a missing proportion set based on the missing proportion of the dataset of the first group of missing proportions. Alternatively, the second determining module is used to determine, when the missing proportion of a dataset of the first type is greater than or equal to a preset proportion threshold, the target imputation model corresponding to the dataset of the first type is the data imputation model with the best imputation result corresponding to the dataset of the second group of missing proportions in the first type, wherein the missing proportion of the dataset of the second group of missing proportions is greater than the missing proportion of the first group.

6. A health assessment device, characterized in that, include: The second acquisition module is used to acquire the original feature dataset of the target object; The first analysis module is used to analyze the original feature dataset to obtain missing feature data; The first filling module is used to fill in the missing feature data in the original feature data of the target object using the data filling model selection method described in claim 1, so as to obtain the filled target object feature data; The second input module is used to input the filled feature dataset into the pre-built health assessment model to obtain the health assessment result of the target object.

7. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory is coupled to the processor; The memory stores computer-readable program instructions, which, when executed by the processor, implement the data imputation model selection method as described in claim 1, or the health assessment method as described in any one of claims 2-4.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data imputation model selection method as described in claim 1, or the health assessment method as described in any one of claims 2-4.

Citation Information

Patent Citations

  • Health data missing value prediction method and device, computer equipment and storage medium

    CN114171150A