Data Missing Value Filling Method, Device, Equipment, Medium and Program

By dividing the data into discrete time periods, obtaining key correlation characteristics and using interpretation models to fill missing data, the problem of insufficient accuracy and reliability of filling missing data in the prior art is solved, and more efficient data processing and user experience improvement is achieved.

CN119943244BActive Publication Date: 2025-07-22SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510078831.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-22
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

In the prior art, the data missing value filling method uses mean, median or mode to cause poor data accuracy and reliability, ignores the correlation between variables and dynamic changes in data, lacks targeted processing strategies, affecting the accuracy and reliability of data processing.

Method used

The target data is divided into multiple discrete time data, the key correlation characteristics of each patient in the target database are obtained, the prediction results are output by interpreting the model, and the follow-up data corresponding to the prediction results are filled into the missing database, and the nonlinear relationship is captured using the efficient maximum information coefficient estimate and the multi-layer perceptron layer to improve the accuracy and reliability of the filling.

Benefits of technology

It improves the accuracy and reliability of missing value filling, improves processing efficiency, improves the follow-up system, enhances user experience, and improves the prediction accuracy and data quality of postoperative survival time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943244B_ABST
    Figure CN119943244B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and particularly relates to a method, device, equipment, medium and program for filling missing data values. The method includes: dividing target data into multiple discrete-time data; obtaining the key associated features of each patient in the target database, where the key associated features represent the features related to the target data; inputting the key associated features into an interpretation model, and the interpretation model outputs corresponding prediction results, where the prediction results are the results of the target data falling into the corresponding discrete time; associating the prediction results with the corresponding follow-up data and filling them into the missing database. Thus, the problems in the related art that the filling process of missing data values uses the mean, median or mode to fill, resulting in poor data accuracy and reliability, etc., are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a method, device, equipment, medium and program for filling missing data values. Background Art

[0002] In related technologies, the filling of missing data values mostly uses the mean, median or mode. Although this method is simple to calculate, it ignores the correlation between variables and the dynamic changes of data, which may lead to information loss and deviation; it has poor processing effects on complex non-linear relationships; it lacks targeted processing strategies for different types of data missing, and outlier detection is often ignored, resulting in poor accuracy and reliability of data processing. Summary of the Invention

[0003] This application provides a method, device, equipment, medium and program for filling missing data values to solve the problems such as poor data accuracy and reliability caused by using the mean, median or mode to fill missing data values in related technologies.

[0004] The first aspect embodiment of this application provides a method for filling missing data values, including the following steps: dividing the target data into multiple discrete time data; obtaining the key associated features of each patient in the target database, where the key associated features represent the relevant features of the target data; inputting the key associated features into an interpretation model, and the interpretation model outputs corresponding prediction results, where the prediction result is that the target data falls into the corresponding discrete time result; associating the prediction result with the corresponding follow-up data and filling it into the missing database.

[0005] Optionally, the obtaining of the key associated features of each patient in the target database includes: identifying the follow-up data of each missing target key content in the target database; performing format preprocessing on the follow-up data to convert it into sequence integers to generate numerical features, and filtering the numerical features to select key associated features with a similarity greater than the first target threshold in the historical database.

[0006] Optionally, the training process of the interpretation model: obtaining historical follow-up data; performing format preprocessing on the discrete feature data format in the historical follow-up data to convert it into sequence integers to generate numerical features; screening the key associated features in the numerical features, and generating a training data set according to the key associated features and the corresponding survival time to train the interpretation model.

[0007] Optionally, the key associated features in the screened numerical features include: calculating the efficient maximum information coefficient estimate between each numerical feature and the postoperative survival time of the patient; screening the numerical features with the efficient maximum information coefficient estimate greater than the second target threshold to generate associated features; calculating the efficient maximum information coefficient estimate between the associated features, and removing any one of the associated features with the efficient maximum information coefficient estimate greater than the third target threshold to generate key associated features.

[0008] Optionally, the key associated features in the screened numerical features further include: performing low-variance filtering processing on the numerical features to generate target features, and pollenizing the class labels of the postoperative survival time for the target features; calculating the similarity value between each target feature and the class label, sorting the target features and the corresponding similarity values in descending order and storing them in a target container; sequentially extracting target features from the target container and adding them to a subset linked list container, and determining whether the conditions for adding to the subset linked list container are met until the number of target features in the subset linked list container reaches the target number to generate corresponding key associated features.

[0009] Optionally, the structure of the interpretation model includes: an input layer, a multi-layer perceptron layer, an output layer, and a loss function. Among them, the input of the input layer is the key associated feature set of the patient; the multi-layer perceptron layer includes multiple fully connected layers, and each layer uses an objective function to capture the non-linear relationship and linear relationship between the key associated features and convert them into the predicted probability value of the postoperative survival time of the patient; the output layer is used to output the predicted probability value of the postoperative survival time corresponding to each key associated feature; the loss function is used to calculate the difference value according to the predicted probability value of the postoperative survival time corresponding to each key associated feature and the actual postoperative survival time corresponding to each key associated feature, and optimize the parameters of the interpretation model with the aim of minimizing the loss function.

[0010] An embodiment of the second aspect of the present application provides a data missing value filling device, including: a dividing module for dividing the continuous postoperative survival time of the patient into multiple discrete time periods; an obtaining module for obtaining the key associated features of each patient in the target database, where the key associated features represent the relevant features of the target data; an input module for inputting the key associated features into an interpretation model, and the interpretation model outputs a corresponding prediction result, where the prediction result is the result that the target data falls into the corresponding discrete time; a filling module for filling the follow-up data associated with the prediction result into the missing database.

[0011] An embodiment of the third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to perform the data missing value filling method as described in the above embodiment.

[0012] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. The program is executed by a processor to perform the data missing value filling method as described in the above embodiments.

[0013] In the fifth aspect of the embodiments of the present application, a computer program product is provided, including a computer program or instruction, characterized in that when the computer program or instruction is executed, it realizes the data missing value filling method as described in the above embodiments.

[0014] Therefore, the present application has at least the following beneficial effects:

[0015] In the embodiments of the present application, the target data is divided into multiple discrete time data; the key associated features of each patient in the target database are obtained, where the key associated features represent the relevant features related to the target data; the key associated features are input into an interpretation model, and the interpretation model outputs a corresponding prediction result, where the prediction result is that the target data falls into the corresponding discrete time result; the prediction result is associated with the corresponding follow-up data and filled into the missing database, thereby improving the accuracy and reliability of missing value filling, improving the processing efficiency, facilitating the improvement of the follow-up system, and enhancing the user experience.

[0016] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0018] Figure 1 is a flowchart of a data missing value filling method according to an embodiment of the present application;

[0019] Figure 2 is an example diagram of a follow-up time period according to an embodiment of the present application;

[0020] Figure 3 is a schematic diagram of the data format preprocessing process according to an embodiment of the present application;

[0021] Figure 4 is a schematic diagram of the selection of key associated features according to an embodiment of the present application;

[0022] Figure 5 is a schematic diagram of the interpretation of missing values according to an embodiment of the present application;

[0023] Figure 6 is a schematic diagram of the process of the data missing value filling method according to an embodiment of the present application;

[0024] Figure 7 It is a block diagram example of a data missing value filling device provided according to an embodiment of the present application;

[0025] Figure 8 It is a schematic structural diagram of an electronic device provided according to an embodiment of the present application. Detailed implementation manners

[0026] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as a limitation to the present application.

[0027] In clinical research, the follow-up of postoperative patients is a crucial task. However, in actual operation, some difficult-to-overcome difficulties are often encountered, and one of the most prominent problems is the interruption of follow-up caused by the death of patients. The follow-up of postoperative patients usually needs to last for a period of time so that doctors can comprehensively understand the recovery of patients; however, some patients unfortunately die or do not cooperate with the follow-up during the follow-up period, which leads to the interruption of follow-up and affects the subsequent judgment of doctors on the condition of patients. Due to the uncertainty of the death time, it is impossible to accurately understand the survival time of patients after surgery, resulting in the lack of database content and making the data analysis of the postoperative survival period more difficult, and the follow-up system is not perfect.

[0028] First of all, the evaluation of the surgical effect usually needs to consider the survival time of patients, and the lack of the death time makes it impossible for us to accurately judge whether the surgery is successful. Secondly, the lack of the death time also affects the analysis of the postoperative survival period. The analysis of the postoperative survival period requires a large amount of data support, and the lack of the death time will directly affect the accuracy of the analysis results. In addition, the lack of the death time also affects the prevention of postoperative complications. The prevention of postoperative complications needs to formulate corresponding preventive measures according to the survival time of patients, and the lack of the death time makes it impossible to accurately judge whether the patient is in a high-risk period of complications, so that effective preventive measures cannot be taken.

[0029] In short, the problem of the lack of the death time in the follow-up of postoperative patients is a problem that needs to be highly valued. Effective measures need to be taken to solve this problem so that postoperative complications can be more effectively prevented, thereby improving the survival quality of patients.

[0030] The data missing value filling method, device, equipment, storage medium and program of the embodiments of the present application will be described below with reference to the accompanying drawings.

[0031] Specifically, Figure 1Schematic flowchart of a method for filling missing data values provided by an embodiment of the present application.

[0032] As Figure 1 shown, the method for filling missing data values includes the following steps:

[0033] In step S101, the target data is divided into multiple discrete time data.

[0034] Among them, the target data can be the postoperative survival time data of each patient, without specific limitation.

[0035] It can be understood that the embodiments of the present application can divide the target data into multiple discrete time data to facilitate subsequent improvement of the efficiency of data analysis and prediction.

[0036] For example, as Figure 2 shown, assume that the survival time y to be interpreted is divided into 3 time periods. The first is 1 week, and the second and third are 1 month and 3 months respectively. Then the corresponding time periods are 7 days, 7 days to 37 days, and 37 to 127 days.

[0037] In step S102, the key associated features of each patient in the target database are obtained, where the key associated features represent the features related to the target data.

[0038] It can be understood that the embodiments of the present application can effectively extract the most valuable key associated features from a large amount of raw data by obtaining the key associated features of each patient in the target database, improving the accuracy of postoperative survival time prediction.

[0039] It should be noted that after the selection of key associated features, the key associated features of the patient's postoperative survival time are x = [x1, x2,..., x m .

[0040] In the embodiments of the present application, obtaining the key associated features of each patient in the target database includes: identifying the follow-up data of each missing key content in the target database; performing format preprocessing on the follow-up data to convert it into sequence integers to generate numerical features, and filtering the numerical features to select the key associated features with a similarity greater than the first target threshold in the historical database.

[0041] Among them, the first target threshold can be set according to actual needs, without specific limitation.

[0042] It can be understood that the embodiments of the present application identify the follow-up data of each missing target key content in the target database; perform format preprocessing on the follow-up data to convert it into sequence integers to generate numerical features, and perform feature filtering on the numerical features to screen out the key association features with a similarity greater than the first target threshold in the historical database, which can effectively extract the most valuable key association features from the original data, not only improving the accuracy of postoperative survival time prediction, but also improving the quality and usability of the data, and enhancing the interpretability and practicality of the model.

[0043] Specifically, as Figure 3 shown, the categorical feature data, that is, the values taken by the categorical (discrete) features, are converted into ordinal integers, so that each feature has only one column of integers. The data format processing solves the problem of digital representation of information such as race and gender in the data; data dimensionality reduction mainly aims at the problem of redundant features, extracts the key association features related to the postoperative survival time, eliminates the redundant features, and reduces the data dimension.

[0044] In step S103, the key association features are input into the interpretation model, and the interpretation model outputs the corresponding prediction result, where the prediction result is that the target data falls into the corresponding discrete time result.

[0045] It can be understood that the embodiments of the present application input the key association features into the interpretation model, and the interpretation model outputs the corresponding prediction result, thereby improving the accuracy and reliability of missing value filling, improving the processing efficiency, facilitating the improvement of the follow-up system, and enhancing the user experience.

[0046] It should be noted that the missing value interpretation obtains the survival time interpretation model through learning the past case data, and interprets and fills the missing death time information with the data model.

[0047] In the embodiments of the present application, the training process of the interpretation model: obtain historical follow-up data; preprocess the data format of the discrete feature data in the historical follow-up data to convert it into sequence integers to generate numerical features; screen out the key association features in the numerical features, and generate a training data set according to the key association features and the corresponding survival time to train the interpretation model.

[0048] It can be understood that the embodiments of the present application obtain historical follow-up data; preprocess the data format of the discrete feature data in the historical follow-up data to convert it into sequence integers to generate numerical features; screen out the key association features in the numerical features, and generate a training data set according to the key association features and the corresponding survival time to train the interpretation model, which can effectively improve the performance and prediction accuracy of the interpretation model.

[0049] In the embodiments of the present application, screening key associated features among numerical features includes: calculating the efficient maximum information coefficient estimate of each numerical feature and the postoperative survival time of the patient; screening numerical features with efficient maximum information coefficient estimates greater than the second target threshold to generate associated features; calculating the efficient maximum information coefficient estimates among the associated features, and removing any one of the associated features with efficient maximum information coefficient estimates greater than the third target threshold to generate key associated features.

[0050] Among them, the second target threshold and the third target threshold can be set according to actual needs, and no specific limitations are made.

[0051] It can be understood that the embodiments of the present application can calculate the efficient maximum information coefficient estimate of each numerical feature and the postoperative survival time of the patient; screen numerical features with efficient maximum information coefficient estimates greater than the second target threshold to generate associated features; calculate the efficient maximum information coefficient estimates among the associated features, and remove any one of the associated features with efficient maximum information coefficient estimates greater than the third target threshold to generate key associated features, which can effectively extract the most valuable key associated features from a large amount of original data, not only improving the quality and relevance of the features, but also enhancing the interpretability and practicability of the model.

[0052] Specifically, as Figure 4 shown, the difference between the efficient maximum information coefficient estimation method MICe and the maximum information coefficient MIC method lies in the strategy of grid search.

[0053] MICe reduces the grid search process of dynamic partitioning to a great extent and significantly reduces the time cost by more finely dividing the evenly divided axis, and only when the number of divisions of the evenly divided axis is greater than the number of divisions of the dynamic divided axis.

[0054] 1) Given the binary variable data set D as the sample data set of the binary random variables (X, Y), the maximum information of evenly dividing the y-axis is:

[0055]

[0056] In the formula, k is the number of columns for dividing the x-axis in any form; [l] is the number of rows for evenly dividing the y-axis; G1(k, [l]) represents the grid set of dividing the x-axis into k columns in any form and evenly dividing the y-axis into l rows; G1 ∈ G1(k, [l]) represents the grid element in G1(k, [l]) that maximizes the mutual information; (D)|G1 represents the division of D within the grid G1; I((D)|G1) represents the mutual information calculated under the division of D in the grid G1.

[0057] 2) The maximum mutual information of evenly dividing the x-axis:

[0058]

[0059] In the formula, [k] represents the number of columns evenly dividing the x-axis; l represents the number of rows dividing the y-axis in any form; G2([k], l) represents the grid set with the x-axis evenly divided into k columns and the y-axis divided into l rows in any form; G2 ∈ G2([k], l) represents the grid element in G2([k], l) that maximizes the mutual information; D|G2 represents the partition of the random variables (X, Y) within the grid G2; I(D|G2) represents the mutual information calculated under the partition of D in the grid G2.

[0060] 3) The evenly divided maximum mutual information is:

[0061]

[0062] In the formula, l represents the number of rows dividing the y-axis, and k represents the number of columns dividing the x-axis; G2 ∈ G2([k], l) represents that the grid G2 is an element in the grid set with the x-axis evenly divided into k columns and the y-axis divided into l rows in any form; G1 ∈ G1(k, [l]) represents that the grid G1 is an element in the grid set with the x-axis divided into k columns in any form and the y-axis evenly divided into l rows.

[0063] 4) The evenly divided feature matrix is:

[0064]

[0065] In the formula, represents the evenly divided feature matrix; D represents the binary variable data set; (D) k,l represents the maximum normalized mutual information obtained by partitioning the binary variable data set into a grid of l rows and k columns; I * (D, k, l) represents the evenly divided maximum mutual information.

[0066] 5) Given a binary variable data set D with a sample size of n and an upper limit condition kl < B(n) satisfied by the number k of partitions on the y-axis and the number l of partitions on the x-axis, the efficient maximum information coefficient estimate MICe is:

[0067]

[0068] In the formula, kl represents the product of the positive integers k and l in l rows and k columns; n represents the number of samples; the expression of B(n) is B(n) = n α , 0 < α < 1.

[0069] In the embodiments of the present application, screening key associated features in numerical features further includes: performing low-variance filtering on numerical features to generate target features, and generating class labels of postoperative survival time for the target features; calculating the similarity values between each target feature and the class labels, sorting the target features and the corresponding similarity values in descending order and storing them in a target container; sequentially extracting target features from the target container and adding them to a subset linked list container, and determining whether the conditions for adding to the subset linked list container are met until the number of target features in the subset linked list container reaches the target quantity to generate corresponding key associated features.

[0070] It can be understood that in the embodiments of the present application, by performing low-variance filtering on numerical features to generate target features, and generating class labels of postoperative survival time for the target features; calculating the similarity values between each target feature and the class labels, sorting the target features and the corresponding similarity values in descending order and storing them in a target container; sequentially extracting target features from the target container and adding them to a subset linked list container, and determining whether the conditions for adding to the subset linked list container are met until the number of target features in the subset linked list container reaches the target quantity to generate corresponding key associated features, the most valuable key associated features can be effectively extracted from a large amount of raw data, which not only improves the quality and relevance of the features, but also enhances the interpretability and practicality of the model.

[0071] Specifically, as Figure 4 shown, the feature selection based on MICe first calculates the association between physical features and survival time, measures the correlation between each physical feature and survival time, and saves the correlation between each feature and survival time, and then eliminates the redundant features related between the features.

[0072] 1) The correlation between features and survival time. Let F=(f1, f2, …, f m ) represent the m features after low-variance filtering in the first stage, and C represent the classification label. The correlation between the i-th feature f i and the label C is expressed as MICe(f i , C), and the results of the features and their correlations with the label are sorted in descending order and stored in the linked list container FC_Corr.

[0073] 2) Redundancy elimination between features. Redundancy analysis is the core of selecting k best features. The correlation between the i-th feature f i and the j-th feature f j in F is expressed as MICe(f i , f j ). The larger the value of MICe(f i , f j ), the more it indicates that f i and the j-th feature f jThe stronger the substitutability between them, that is, the stronger the redundancy, the screening of the subset of body features in this stage is carried out cyclically. First, the first feature is taken out from FC_Corr and added to the subset linked list container SF, and then the next feature is taken out from FC_Corr in turn to determine whether it can be added to the subset linked list container SF, and so on until the number of features in the subset linked list container SF reaches k.

[0074] When selecting the s + 1-th feature, for the s features (s < K < m) stored in the linked list container SF that have been selected, for the s + 1-th feature f selected from FC_Corr and added to SF y (s + 1 ≤ K) and any one of the remaining m - s - 1 features f selected from FC_Corr y after f x satisfies the fractional inequality as

[0075]

[0076] In the embodiment of the present application, the structure of the interpretation model includes: an input layer, a multi-layer perceptron layer, an output layer, and a loss function. Among them, the input of the input layer is the key associated feature set of the patient; the multi-layer perceptron layer includes multiple fully connected layers, and each layer uses the objective function to capture the non-linear relationship and linear relationship between the key associated features and convert them into the predicted probability value of the patient's postoperative survival time; the output layer is used to output the predicted probability value of the patient's postoperative survival time corresponding to each key associated feature; the loss function is used to calculate the difference value according to the predicted probability value of the patient's postoperative survival time corresponding to each key associated feature and the actual postoperative survival time of the patient corresponding to each key associated feature, and optimize the parameters of the interpretation model with the aim of minimizing the loss function.

[0077] It can be understood that the interpretation model of the embodiment of the present application can effectively extract valuable information from the key associated features of the patient and convert it into the predicted probability value of the postoperative survival time. The powerful expression ability of the multi-layer perceptron layer and the application of the activation function enable the model to accurately capture complex non-linear relationships, significantly improving the prediction accuracy. The clear classification results and probability distribution provided by the output layer enhance the interpretability of the model. The introduction of the loss function ensures the stability and generalization ability of the model during the training process, reduces the risk of overfitting, has strong robustness, and the simplified model structure and efficient activation function selection accelerate the training speed and save computing resources.

[0078] Specifically, as Figure 5 shown, after the key associated features are selected, the key associated features of the patient's postoperative survival time are x = [x1, x2,..., xm], and the relationship between the survival time and the key associated features can be y = F(x) = α1x1 + α2x2 +,..., α mx m is represented by +η, where α = [α1, α2, …, α m , and η is a parameter to be determined; find the optimal estimated values of the parameters of the survival time and y = F(x). For the given (historical data) observed data, design the objective function L(y, F(x)) = ∑[y - F(x)] 2 , and transform it into minimizing the objective function minL(y, F(x)) to obtain the relationship of y = F(x).

[0079] In the interpretation, to avoid overfitting of y = F(x), an artificial neural network (Neural Network) is introduced. For the given (historical data) observed data, train the artificial neural network to obtain the neural network model y = N(x) for predicting the survival time. Given the key correlation feature x, the final interpretation value is the average value integrated by the functional relationship y = N(x) and the artificial neural network y = N(x).

[0080] In step S104, fill the predicted results associated with the corresponding follow-up data into the missing database.

[0081] It can be understood that in the embodiment of the present application, the predicted results are associated with the corresponding follow-up data and filled into the missing database, thereby improving the accuracy and reliability of the missing value filling, improving the processing efficiency, facilitating the improvement of the follow-up system, enhancing the user experience, and facilitating subsequent applications.

[0082] According to the data missing value filling method proposed in the embodiment of the present application, the target data is divided into multiple discrete time data; obtain the key correlation features of each patient in the target database, where the key correlation features represent the relevant features of the target data; input the key correlation features into the interpretation model, and the interpretation model outputs the corresponding predicted results, where the predicted results are the results of the target data falling into the corresponding discrete time; fill the predicted results associated with the corresponding follow-up data into the missing database, thereby improving the accuracy and reliability of the missing value filling, improving the processing efficiency, facilitating the improvement of the follow-up system, and enhancing the user experience.

[0083] Next, it will be combined with Figure 6 For the data missing value filling method of the present application, the process of the data missing value filling method includes the preprocessing of historical data, key feature extraction, and learning of the model based on key correlation feature data, as follows:

[0084] Step 1, data format preprocessing

[0085] The categorical feature data represents the values taken by categorical (discrete) features. These features are converted into ordinal integers so that each feature has only one column of integers.

[0086] Step 2, key correlation feature selection includes efficient maximum correlation coefficient and MICe-based feature selection.

[0087] (1) The difference between the efficient maximum information coefficient estimation method MICe and the maximum information coefficient MIC method lies in the strategy of grid search.

[0088] MICe reduces the time cost to a great extent by dividing the equal-axis more carefully. Only when the number of divisions of the equal-axis is greater than that of the dynamic division axis, the dynamic division of the grid search process is reduced.

[0089] 1) Given the binary variable dataset D as the sample dataset of the binary random variables (X, Y), the maximum information of the equally divided y-axis is:

[0090]

[0091] In the formula, k is the number of columns for dividing the x-axis in any form; [l] is the number of rows for equally dividing the y-axis; G1(k, [l]) represents the grid set of dividing the x-axis into k columns in any form and equally dividing the y-axis into l rows; G1 ∈ G1(k, [l]) represents the grid element in G1(k, [l]) that maximizes the mutual information; (D)|G1 represents the division of D within the grid G1; I((D)|G1) represents the mutual information calculated under the division of D in the grid G1.

[0092] 2) The maximum mutual information of the equally divided x-axis:

[0093]

[0094] In the formula, [k] represents the number of columns for equally dividing the x-axis; l represents the number of rows for dividing the y-axis in any form; G2([k], l) represents the grid set of equally dividing the x-axis into k columns and dividing the y-axis into l rows in any form; G2 ∈ G2([k], l) represents the grid element in G2([k], l) that maximizes the mutual information; D|G2 represents the division of the random variables (X, Y) within the grid G2; I(D|G2) represents the mutual information calculated under the division of D in the grid G2.

[0095] 3) The equally divided maximum mutual information is:

[0096]

[0097] In the formula, l represents the number of rows for dividing the y-axis, and k represents the number of columns for dividing the x-axis; G2 ∈ G2([k], l) represents that the grid G2 is an element in the grid set of equally dividing the x-axis into k columns and dividing the y-axis into l rows in any form; G1 ∈ G1(k, [l]) represents that the grid G1 is an element in the grid set of dividing the x-axis into k columns in any form and equally dividing the y-axis into l rows.

[0098] 4) The evenly divided feature matrix is as follows:

[0099]

[0100] In the formula, represents the evenly divided feature matrix; D represents the binary variable dataset; (D) k,l represents the maximum normalized mutual information obtained by dividing the binary variable dataset into a grid of l rows and k columns; I * (D, k, l) represents the evenly divided maximum mutual information.

[0101] 5) Given a binary variable dataset D with a sample size of n, and the upper limit condition kl < B(n) satisfied by the number of divisions k on the y-axis and the number of divisions l on the x-axis, the efficient maximum information coefficient estimate MICe is:

[0102]

[0103] In the formula, kl represents the product of the positive integers k and l in l rows and k columns; n represents the number of samples; the expression of B(n) is B(n) = n α , 0 < α < 1

[0104] (2) Feature selection based on the characteristics of MICe. First, calculate the correlation between physical features and survival time, measure the correlation between each physical feature and survival time, and save the correlation between each feature and survival time. Then, remove the redundant features that are correlated between features.

[0105] 1) The correlation between features and survival time. Let F = (f1, f2,..., f m ) represent the m features after low-variance filtering in the first stage, and C represent the classification label. The correlation between the i-th feature f i and the label C is expressed as MICe(f i , C). The results of the features and their correlations with the label are sorted in descending order and stored in the linked list container FC_Corr.

[0106] 2) Redundancy removal between features. Redundancy analysis is the core of selecting k best features. Among them, the correlation between the i-th feature f i and the j-th feature f j in F is expressed as MICe(f i , f j ). The larger the value of MICe(f i , f j ), the more it indicates that f i and the j-th feature f jThe stronger the substitutability, that is, the stronger the redundancy, the screening of the subset of body features in this stage is carried out cyclically. First, the first feature is taken out from FC_Corr and added to the subset linked list container SF, and then the next feature is taken out from FC_Corr in turn to determine whether it can be added to the subset linked list container SF, and so on until the number of features in the subset linked list container SF reaches k.

[0107] When selecting the s + 1th feature, for the s features (s < K < m) stored in the linked list container SF that have been selected, for the s + 1th feature f selected from FC_Corr and added to SF y (s + 1 ≤ K) and any one feature f among the remaining m - s - 1 features selected by FC_Corr y after selecting f x satisfies the fractional inequality

[0108]

[0109]

[0110] Step 3, missing value judgment and filling

[0111] Suppose the survival time y to be judged is divided into 3 time periods. The first is 1 week, and the second and third are 1 month and 3 months respectively. Then the corresponding time periods are 7 days, 7 days to 37 days, and 37 to 127 days.

[0112] After the key associated features are selected, the key associated features of the patient's postoperative survival time are x = [x1, x2,..., x m , then the relationship between the survival time and the key associated features can be expressed as y = F(x) = α1x1 + α2x2 +,..., α m x m + η, where α = [α1, α2,..., α m , and η is a parameter to be determined.

[0113] Find the optimal estimated values of the parameters of the survival time and y = F(x). For the given (historical data) observed data, design the objective function L(y, F(x)) = ∑[y - F(x)] 2 , and transform it into minimizing the objective function minL(y, F(x)) to obtain the relationship of y = F(x).

[0114] In the interpretation, in order to avoid overfitting of y = F(x), an artificial neural network is introduced. For the given (historical data) observed data, the artificial neural network is trained to obtain a neural network model y = N(x) for predicting the survival time. Given the key correlation feature x, the final interpretation value is the average value integrated by the functional relationship y = N(x) and the artificial neural network y = N(x).

[0115] In summary, in view of the redundant feature problem, the present application designs an efficient method for extracting key correlation features for the postoperative survival time, eliminates redundant features, and improves the interpretation filling efficiency; integrates the interpretation of missing values, obtains two survival time interpretation models through learning from past case data, interprets and fills the missing death time information with a data model, and takes the average value to avoid the problem of overfitting of a single model.

[0116] Next, a data missing value filling device according to an embodiment of the present application will be described with reference to the accompanying drawings.

[0117] Figure 7 It is a block diagram of the data missing value filling device according to an embodiment of the present application.

[0118] As Figure 7 shown, the data missing value filling device 10 includes: a division module 100, an acquisition module 200, an input module 300, and a filling module 400.

[0119] Among them, the division module 100 is used to divide the continuous postoperative survival time of patients into multiple discrete time periods; the acquisition module 200 is used to acquire the key correlation features of each patient in the target database, where the key correlation features represent the relevant features of the target data; the input module 300 is used to input the key correlation features into the interpretation model, and the interpretation model outputs the corresponding prediction result, where the prediction result is the result that the target data falls into the corresponding discrete time; the filling module 400 is used to fill the prediction result associated with the corresponding follow-up data into the missing database.

[0120] It should be noted that the foregoing explanation of the data missing value filling method embodiment also applies to the data missing value filling device of this embodiment, and will not be repeated here.

[0121] The missing value filling device for data according to the embodiments of the present application divides the target data into multiple discrete-time data; obtains the key associated features of each patient in the target database, where the key associated features represent the features related to the target data; inputs the key associated features into an interpretation model, and the interpretation model outputs corresponding prediction results, where the prediction results are the results of the target data falling into the corresponding discrete time; associates the prediction results with the corresponding follow-up data and fills them into the missing database, thereby improving the accuracy and reliability of missing value filling, improving the processing efficiency, facilitating the improvement of the follow-up system, and enhancing the user experience.

[0122] Figure 8 The structural schematic diagram of the electronic device provided by the embodiments of the present application. The electronic device may include:

[0123] A memory 801, a processor 802, and a computer program stored on the memory 801 and executable on the processor 802.

[0124] When the processor 802 executes the program, it implements the data missing value filling method provided in the above embodiments.

[0125] Further, the electronic device further includes:

[0126] A communication interface 803 for communication between the memory 801 and the processor 802.

[0127] The memory 801 is used to store a computer program executable on the processor 802.

[0128] The memory 801 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0129] If the memory 801, the processor 802, and the communication interface 803 are implemented independently, the communication interface 803, the memory 801, and the processor 802 may be connected to each other through a bus and complete communication with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0130] Optionally, in a specific implementation, if the memory 801, the processor 802, and the communication interface 803 are integrated on a single chip, the memory 801, the processor 802, and the communication interface 803 can communicate with each other through an internal interface.

[0131] The processor 802 may be a central processing unit (CPU for short), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present application.

[0132] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program or instruction is stored. When the computer program or instruction is executed by a processor, the data missing value filling method as described above is implemented.

[0133] The embodiments of the present application further provide a computer program product, including a computer program or instruction, characterized in that when the computer program or instruction is executed, the data missing value filling method as described above is implemented.

[0134] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0135] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0136] Any process or method description depicted in the flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of this application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application pertain.

[0137] It should be understood that the various parts of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented by a combination of any one or more of the following techniques known in the art: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0138] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried out in implementing the above-described embodiment methods can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium, and when executed, includes one or a combination of the steps of the method embodiments.

Claims

1. A method for filling missing data values, characterized in that, Including the following steps: Dividing the target data into multiple discrete time data; Obtaining the key associated features of each patient in the target database, where the key associated features represent the relevant features of the target data; Inputting the key associated features into an interpretation model, and the interpretation model outputs corresponding prediction results, where the prediction result is that the target data falls into the corresponding discrete time result. The training process of the interpretation model: obtaining historical follow-up data; preprocessing the discrete feature data format in the historical follow-up data and converting it into sequence integers to generate numerical features; screening the key associated features in the numerical features, and generating a training dataset according to the key associated features and the corresponding survival time to train the interpretation model; Among them, screening the key associated features in the numerical features includes: calculating the efficient maximum information coefficient estimate value of each numerical feature and the postoperative survival time of the patient; screening the numerical features with the efficient maximum information coefficient estimate value greater than the second target threshold to generate associated features; calculating the efficient maximum information coefficient estimate value between the associated features, and removing any one of the associated features with the efficient maximum information coefficient estimate value greater than the third target threshold to generate key associated features; Among them, the structure of the interpretation model includes: an input layer, a multi-layer perceptron layer, an output layer, and a loss function. The input of the input layer is the key associated feature set of the patient; the multi-layer perceptron layer contains multiple fully connected layers, and each layer uses an objective function to capture the non-linear relationship and linear relationship between the key associated features and convert them into the predicted probability value of the postoperative survival time of the patient; the output layer is used to output the predicted probability value of the postoperative survival time of the patient corresponding to each key associated feature; the loss function is used to calculate the difference value according to the predicted probability value of the postoperative survival time of the patient corresponding to each key associated feature and the actual postoperative survival time of the patient corresponding to each key associated feature, and optimize the parameters of the interpretation model with the goal of minimizing the loss function; Associating the prediction results with the corresponding follow-up data and filling them into the missing database.

2. The method for filling missing data values according to claim 1, wherein The obtaining of the key associated features of each patient in the target database includes: Identifying the follow-up data of each missing target key content in the target database; Preprocessing the format of the follow-up data and converting it into sequence integers to generate numerical features, and filtering the numerical features to select the key associated features with a similarity greater than the first target threshold in the historical database.

3. The method for filling missing data values according to claim 1, wherein The screening of the key associated features in the numerical features also includes: Performing low-variance filtering processing on the numerical features to generate target features, and dividing the target features into category labels of postoperative survival time; Calculating the similarity value of each target feature and the category label, and sorting the target features and the corresponding similarity values in descending order and storing them in a target container; Sequentially extracting target features from the target container and adding them to the subset linked list container, and judging whether the conditions for adding to the subset linked list container are met until the number of target features in the subset linked list container reaches the target number to generate the corresponding key associated features.

4. A data missing value filling device, characterized in that Including: A division module for dividing the target data into multiple discrete time data; An acquisition module, configured to acquire the key associated features of each patient in a target database, where the key associated features represent the features related to the target data; An input module, configured to input the key associated features into an interpretation model, and the interpretation model outputs a corresponding prediction result, where the prediction result is that the target data falls into the corresponding discrete time result. The training process of the interpretation model is as follows: acquiring historical follow-up data; preprocessing and converting the discrete feature data format in the historical follow-up data into sequence integers to generate numerical features; screening the key associated features in the numerical features, and generating a training dataset based on the key associated features and the corresponding survival time to train the interpretation model; Among them, the screening of the key associated features in the numerical features includes: calculating the efficient maximum information coefficient estimation value between each numerical feature and the postoperative survival time of the patient; screening the numerical features with the efficient maximum information coefficient estimation value greater than a second target threshold to generate associated features; calculating the efficient maximum information coefficient estimation value between the associated features, and removing any one of the associated features with the efficient maximum information coefficient estimation value greater than a third target threshold to generate key associated features; Among them, the structure of the interpretation model includes: an input layer, a multi-layer perceptron layer, an output layer, and a loss function. The input of the input layer is the key associated feature set of the patient; the multi-layer perceptron layer includes multiple fully connected layers, and each layer uses an objective function to capture the non-linear relationship and linear relationship between the key associated features and convert them into the predicted probability value of the postoperative survival time of the patient; the output layer is used to output the predicted probability value of the postoperative survival time of the patient corresponding to each key associated feature; the loss function is used to calculate the difference value based on the predicted probability value of the postoperative survival time of the patient corresponding to each key associated feature and the actual postoperative survival time of the patient corresponding to each key associated feature, and optimize the parameters of the interpretation model with the aim of minimizing the loss function; A filling module, configured to fill the missing database with the follow-up data associated with the prediction result.

5. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the data missing value filling method according to any one of claims 1-3.

6. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by the processor, it is used to implement the data missing value filling method according to any one of claims 1-3.

7. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed, it implements the data missing value filling method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Medical data missing value filling method and system based on partition data division

    CN117668474A

  • Method and system for predicting hemoglobin concentration of postoperative patient

    CN119028473A