Data missing value filling method and device, equipment, medium and program

By dividing the target data into multiple discrete time data, obtaining key correlation characteristics and inputting the interpretation model, the problem of missing value filling in the prior art ignores correlation and dynamic changes, and a more efficient and reliable missing value filling effect is achieved.

CN119943244AActive Publication Date: 2025-05-06SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510078831.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-05-06
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

When filling in missing values ​​in data processing, the prior art ignores the correlation between variables and dynamic changes in data, resulting in information loss and deviation, especially for complex nonlinear relationship processing.

Method used

By dividing the target data into multiple discrete time data, the key correlation characteristics of each patient are obtained, and these characteristics are entered into the interpretation model to obtain the prediction results, and finally the prediction results are associated with the follow-up data to fill in the missing values.

Benefits of technology

It improves the accuracy and reliability of missing value filling, improves processing efficiency, improves the follow-up system, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943244A_ABST
    Figure CN119943244A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a data missing value filling method and device, equipment, a medium and a program, and the method comprises the steps: dividing target data into a plurality of discrete time data; key associated features of each patient in the target database are obtained, and the key associated features represent features related to the target data; the key correlation features are input into an interpretation model, the interpretation model outputs a corresponding prediction result, and the prediction result is a result that the target data falls into the corresponding discrete time; and filling the missing database with the follow-up data associated and corresponding to the prediction result. Therefore, the problems of poor data accuracy and reliability and the like caused by the fact that data missing value filling processing adopts mean value, median or mode filling in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, device, equipment, medium and program for filling missing data values. Background Art

[0002] In related technologies, data processing mostly uses mean, median or mode to fill missing values. Although this method is simple to calculate, it ignores the correlation between variables and the dynamic changes of data, which may lead to information loss and deviation; it is not effective for complex nonlinear relationships; there is a lack of targeted processing strategies for different types of missing data, and outlier detection is often ignored, resulting in poor accuracy and reliability of data processing. Summary of the invention

[0003] The present application provides a method, device, equipment, medium and program for filling missing data values ​​to solve the problem that the data missing value filling process in the related art uses the mean, median or mode to fill in the missing data, resulting in poor data accuracy and reliability.

[0004] The first aspect of the present application provides a method for filling missing data values, comprising the following steps: dividing the target data into multiple discrete time data; obtaining key correlation features of each patient in the target database, wherein the key correlation features represent related features with the target data; inputting the key correlation features into a judgment model, and the judgment model outputs a corresponding prediction result, wherein the prediction result is the target data falling into the corresponding discrete time result; and associating the prediction result with the corresponding follow-up data to fill in the missing database.

[0005] Optionally, the method of obtaining the key associated features of each patient in the target database includes: identifying follow-up data of each missing target key content in the target database; performing format preprocessing on the follow-up data to convert it into a sequence integer to generate numerical features, and performing feature filtering on the numerical features to select key associated features with a similarity greater than a first target threshold in the historical database.

[0006] Optionally, the training process of the interpretation model includes: obtaining historical follow-up data; pre-processing the discrete feature data format in the historical follow-up data and converting it into a sequence integer to generate numerical features; screening key associated features in the numerical features, and generating a training data set based on the key associated features and the corresponding survival time to train the interpretation model.

[0007] Optionally, the screening of key associated features in the numerical features includes: calculating the efficient maximum information coefficient estimate of each numerical feature and the patient's postoperative survival time; screening the numerical features whose efficient maximum information coefficient estimate is greater than a second target threshold to generate associated features; calculating the efficient maximum information coefficient estimate between associated features, and eliminating any associated features whose efficient maximum information coefficient estimate is greater than a third target threshold to generate key associated features.

[0008] Optionally, the screening of key associated features in the numerical features also includes: performing low variance filtering processing on the numerical features to generate target features, and assigning category labels for postoperative survival time to the target features; calculating similarity values ​​between each target feature and the category labels, and sorting the target features and the corresponding similarity values ​​in descending order and storing them in a target container; extracting target features from the target container in turn and adding them to a subset linked list container, and determining whether the conditions for adding the subset linked list container are met until the target number of target features in the subset linked list container reaches the target number to generate corresponding key associated features.

[0009] Optionally, the structure of the interpretation model includes: an input layer, a multi-layer perceptron layer, an output layer and a loss function, wherein the input layer input is a set of key correlation features of the patient; the multi-layer perceptron layer includes multiple fully connected layers, each layer uses an objective function to capture the nonlinear relationship and linear relationship between the key correlation features and converts them into a predicted probability value of the patient's postoperative survival time; the output layer is used to output the predicted probability value of the patient's postoperative survival time corresponding to each key correlation feature; the loss function is used to calculate the difference value based on the predicted probability value of the patient's postoperative survival time corresponding to each key correlation feature and the actual postoperative survival time of the patient corresponding to each key correlation feature, and optimize the parameters of the interpretation model with the purpose of minimizing the loss function.

[0010] The second aspect of the present application provides a device for filling missing data values, including: a division module, used to divide the continuous postoperative survival time of patients into multiple discrete time periods; an acquisition module, used to obtain key correlation features of each patient in the target database, wherein the key correlation features represent related features with the target data; an input module, used to input the key correlation features into a judgment model, and the judgment model outputs a corresponding prediction result, wherein the prediction result is the result that the target data falls into the corresponding discrete time; a filling module, used to associate the prediction result with the corresponding follow-up data and fill it into the missing database.

[0011] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to perform the method for filling missing data values ​​as described in the above embodiment.

[0012] The fourth aspect of the present application provides a computer-readable storage medium on which a computer program is stored. The program is executed by a processor to perform the method for filling missing data values ​​as described in the above embodiments.

[0013] The fifth aspect of the present application provides a computer program product, including a computer program or instructions, characterized in that when the computer program or instructions are executed, the method for filling missing data values ​​as described in the above embodiments is implemented.

[0014] Therefore, this application has at least the following beneficial effects: The embodiment of the present application divides the target data into multiple discrete time data; obtains the key correlation features of each patient in the target database, wherein the key correlation features represent the related features with the target data; inputs the key correlation features into the judgment model, and the judgment model outputs the corresponding prediction results, wherein the prediction results are the target data falling into the corresponding discrete time results; associates the prediction results with the corresponding follow-up data and fills them into the missing database, thereby improving the accuracy and reliability of missing value filling and improving processing efficiency, so as to improve the follow-up system and enhance the user experience.

[0015] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart of a method for filling missing data values ​​provided according to an embodiment of the present application; Figure 2 This is an example diagram of a follow-up period provided according to an embodiment of the present application; Figure 3 A schematic diagram of a data format preprocessing process according to an embodiment of the present application; Figure 4 A schematic diagram of the selection of key associated features according to an embodiment of the present application; Figure 5 A schematic diagram of missing value interpretation according to an embodiment of the present application; Figure 6 A schematic diagram of a missing value filling method according to an embodiment of the present application; Figure 7 This is a block diagram of an apparatus for filling missing data values ​​according to an embodiment of the present application; Figure 8 It is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0017] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0018] In clinical research, follow-up of postoperative patients is a crucial task. However, in actual operation, some insurmountable difficulties are often encountered. One of the most prominent problems is the interruption of follow-up caused by patient death. Postoperative follow-up of patients usually needs to continue for a period of time so that doctors can fully understand the patient's recovery; however, some patients unfortunately die during the follow-up period or do not cooperate with the follow-up, which leads to the interruption of follow-up and affects the subsequent doctor's judgment of the patient's condition. Due to the uncertainty of the time of death, it is impossible to accurately understand the patient's survival time after surgery, resulting in the loss of database content, making the data analysis of the postoperative survival period more difficult, and the follow-up system is incomplete.

[0019] First, the evaluation of surgical results usually needs to consider the patient's survival time, and the lack of death time makes it impossible for us to accurately judge whether the operation is successful. Secondly, the lack of death time will also affect the analysis of the postoperative survival period. The analysis of the postoperative survival period requires a large amount of data support, and the lack of death time will directly affect the accuracy of the analysis results. In addition, the lack of death time will also affect the prevention of postoperative complications. The prevention of postoperative complications requires the formulation of corresponding preventive measures based on the patient's survival time, and the lack of death time makes it impossible to accurately judge whether the patient is at a high risk of complications, and thus it is impossible to take effective preventive measures.

[0020] In conclusion, the problem of missing death time in postoperative patient follow-up is an issue that needs to be taken seriously. Effective measures need to be taken to solve this problem so that postoperative complications can be more effectively prevented, thereby improving the patient's quality of survival.

[0021] The following describes the data missing value filling method, device, equipment, storage medium and program of the embodiments of the present application with reference to the accompanying drawings.

[0022] Specifically, Figure 1 A flowchart of a method for filling missing data values ​​provided in an embodiment of the present application.

[0023] like Figure 1 As shown, the method for filling missing values ​​in data includes the following steps: In step S101, target data is divided into a plurality of discrete time data.

[0024] The target data may be the postoperative survival time data of each patient, without any specific limitation.

[0025] It is understandable that the embodiment of the present application can divide the target data into multiple discrete time data to facilitate subsequent improvement of the efficiency of data analysis and prediction.

[0026] For example, Figure 2 As shown, suppose the survival time y to be judged is divided into three time periods, the first one is 1 week, the second and the third are 1 month and 3 months respectively, then the corresponding time periods are 7 days, 7 days to 37 days, and 37 to 127 days.

[0027] In step S102, key correlation features of each patient in the target database are obtained, wherein the key correlation features represent features related to the target data.

[0028] It can be understood that the embodiment of the present application can effectively extract the most valuable key correlation features from a large amount of original data by acquiring the key correlation features of each patient in the target database, thereby improving the accuracy of postoperative survival time prediction.

[0029] It should be noted that after the key correlation features are selected, the key correlation features of the patient's postoperative survival time are x =[ x 1 , x 2 , … , x m ].

[0030] In an embodiment of the present application, key associated features of each patient in the target database are obtained, including: identifying follow-up data of each missing target key content in the target database; performing format preprocessing on the follow-up data and converting it into a sequence integer to generate numerical features, and performing feature filtering on the numerical features to select key associated features with a similarity greater than a first target threshold in the historical database.

[0031] Among them, the first target threshold can be set according to actual needs without specific limitation.

[0032] It can be understood that the embodiment of the present application identifies the follow-up data of each missing target key content in the target database; performs format preprocessing on the follow-up data to convert it into a sequence integer to generate numerical features, and performs feature filtering on the numerical features to screen out key associated features with a similarity greater than the first target threshold in the historical database. This can effectively extract the most valuable key associated features from the original data, which not only improves the accuracy of postoperative survival time prediction, but also improves the quality and availability of the data, and also enhances the interpretability and practicality of the model.

[0033] Specifically, if Figure 3 As shown in the figure, the categorical feature data, that is, the values ​​of the categorical (discrete) features, are converted into ordinal integers so that each feature has only one column of integers. The data format processing solves the problem of digital representation of information such as race and gender in the data. Data dimensionality reduction mainly targets the problem of redundant features, extracts key features related to postoperative survival time, eliminates redundant features, and reduces data dimensions.

[0034] In step S103, the key correlation features are input into the judgment model, and the judgment model outputs the corresponding prediction results, wherein the prediction results are the target data falling into the corresponding discrete time results.

[0035] It can be understood that the embodiment of the present application inputs key correlation features into the judgment model, and the judgment model outputs the corresponding prediction results, thereby improving the accuracy and reliability of missing value filling and improving processing efficiency, so as to improve the follow-up system and enhance the user experience.

[0036] It should be noted that the missing value interpretation is achieved by studying previous case data to obtain a survival time interpretation model, and the data model is used to interpret and fill in the missing death time information.

[0037] In an embodiment of the present application, the training process of the interpretation model includes: obtaining historical follow-up data; pre-processing the discrete feature data format in the historical follow-up data and converting it into a sequence integer to generate numerical features; screening the key associated features in the numerical features, and generating a training data set based on the key associated features and the corresponding survival time to train the interpretation model.

[0038] It can be understood that the embodiments of the present application obtain historical follow-up data; pre-process the discrete feature data format in the historical follow-up data and convert it into a sequence integer to generate numerical features; screen the key associated features in the numerical features, and generate a training data set based on the key associated features and the corresponding survival time to train the interpretation model, which can effectively improve the performance and prediction accuracy of the interpretation model.

[0039] In an embodiment of the present application, key associated features in numerical features are screened, including: calculating the efficient maximum information coefficient estimate of each numerical feature and the patient's postoperative survival time; screening the numerical features whose efficient maximum information coefficient estimate is greater than a second target threshold to generate associated features; calculating the efficient maximum information coefficient estimate between associated features, and eliminating any associated features whose efficient maximum information coefficient estimate is greater than a third target threshold to generate key associated features.

[0040] Among them, the second target threshold and the third target threshold can be set according to actual needs without specific limitation.

[0041] It can be understood that the embodiments of the present application can calculate the efficient maximum information coefficient estimate of each numerical feature and the patient's postoperative survival time; screen the numerical features whose efficient maximum information coefficient estimate is greater than the second target threshold to generate associated features; calculate the efficient maximum information coefficient estimate between associated features, and eliminate any associated features whose efficient maximum information coefficient estimate is greater than the third target threshold to generate key associated features, which can effectively extract the most valuable key associated features from a large amount of raw data, which not only improves the quality and relevance of the features, but also enhances the interpretability and practicality of the model.

[0042] Specifically, if Figure 4 As shown, the efficient maximum information coefficient estimation method MICe differs from the maximum information coefficient MIC method in the grid search strategy.

[0043] MICe divides the equal-division axis more finely, and only when the number of equal-division axis divisions is greater than the number of dynamic division axis divisions exists, it reduces the grid search process caused by dynamic division and greatly reduces the time cost.

[0044] 1) Given a binary variable dataset D is a binary random variable The sample data set is divided equally y The maximum information of the axis is: In the formula, k To divide in any form x The number of columns of the axis; For equal distribution y The number of rows of the axis; Represents any form of partitioning x Axis k Columns, evenly divided y axis l A grid collection of rows; express The grid element that maximizes the mutual information; express D In the grid G1 division within; express D In the grid G 1 The mutual information calculated under the partition.

[0045] 2) Equal distribution x Maximum mutual information of axes: In the formula, Means evenly divided x The number of columns of the axis; l Indicates that the division is in any form y The number of rows of the axis; Means evenly divided x Axis k Columns, arbitrary division y Axis l A grid collection of rows; express The grid element that maximizes the mutual information; Represents a random variable In the grid G 2 division within; express D In the grid G 2 The mutual information calculated under the partition.

[0046] 3) The average maximum mutual information is: In the formula, l Representation Division y The number of rows of the axis, k Representation Division x The number of columns of the axis; Representation Grid G 2 Belong to the average x Axis k Columns, arbitrary division y Axis l An element in the grid collection of rows; Representation Grid G 1 Any form of division x Axis k Column, evenly divided y Axis l An element in the grid collection for a row.

[0047] 4) The evenly divided feature matrix is: In the formula, represents the equally partitioned feature matrix; represents a binary variable data set; Indicates that the binary variable data set is divided into l OK, k The maximum normalized mutual information obtained by the column grid; represents the average maximum mutual information.

[0048] 5) Given a sample size of n A binary variable dataset D and in y Number of axis divisions k with x Number of axis divisions l Upper limit conditions met , the efficient maximum information coefficient estimate MICe is: In the formula, kl express l OK, k Positive integers in the column k and l The product of n Indicates the number of samples; B ( n ) is expressed as , .

[0049] In an embodiment of the present application, screening key associated features in numerical features also includes: performing low variance filtering on the numerical features to generate target features, and performing category labels on the target features for postoperative survival time; calculating similarity values ​​between each target feature and the category label, and sorting the target features and the corresponding similarity values ​​in descending order and storing them in a target container; extracting target features from the target container in turn and adding them to a subset linked list container, and determining whether the conditions for adding the subset linked list container are met until the target features in the subset linked list container reach the target number to generate the corresponding key associated features.

[0050] It can be understood that the embodiment of the present application generates target features by performing low-variance filtering processing on numerical features, and classifies the target features for postoperative survival time; calculates the similarity value between each target feature and the category label, and sorts the target features and the corresponding similarity values ​​in descending order and stores them in a target container; extracts target features from the target container in turn and adds them to a subset linked list container, and determines whether the conditions for adding the subset linked list container are met until the target number of target features in the subset linked list container reaches the target number to generate the corresponding key correlation features, which can effectively extract the most valuable key correlation features from a large amount of raw data, which not only improves the quality and relevance of the features, but also enhances the interpretability and practicality of the model.

[0051] Specifically, if Figure 4 As shown in the figure, the feature selection based on MICe first calculates the association between physical features and survival time, measures the correlation between each physical feature and survival time, saves the correlation between each feature and survival time, and then eliminates the redundant features related to each feature.

[0052] 1) Correlation between features and survival time, F =( f 1 , f 2 ,…, f m ) indicates that the first stage is processed by low variance filtering. m Features, C Represents a classification label. i Features f i With label C The correlation between them is expressed as MICe( f i , C ), the features and the results of their correlation with the labels are sorted in descending order and stored in the linked list container FC_Corr.

[0053] 2) Eliminate redundancy between features. Redundancy analysis is to select k The core of the best features, F Middle i Features f i With j Features f j The correlation between them is expressed as MICe( f i , f j ), MICe( f i , f j ) value is larger, indicating f i With j Features f j The stronger the substitutability between them, that is, the stronger the redundancy, the screening of the body feature subset in this stage is cyclical. First, the first feature is taken out from FC_Corr and added to the subset linked list container SF. Then, the next feature is taken out from FC_Corr in turn to determine whether it can be added to the subset linked list container SF. This process is repeated until the number of features in the subset linked list container SF reaches k indivual.

[0054] In selectings +1 feature, the selected one stored in the linked list container SF s Features ( s < K < m ), for the selected join in FC_Corr SF The s +1 feature f y ( s +1≤ K ) and FC_Corr select f y After the remaining m - s - Any one of 1 features f x The relationship satisfies the inequality In an embodiment of the present application, the structure of the interpretation model includes: an input layer, a multi-layer perceptron layer, an output layer and a loss function, wherein the input layer inputs a set of key correlation features of the patient; the multi-layer perceptron layer includes multiple fully connected layers, each layer uses an objective function to capture the nonlinear relationship and linear relationship between the key correlation features and converts them into a predicted probability value of the patient's postoperative survival time; the output layer is used to output the predicted probability value of the patient's postoperative survival time corresponding to each key correlation feature; the loss function is used to calculate the difference value based on the predicted probability value of the patient's postoperative survival time corresponding to each key correlation feature and the actual postoperative survival time of the patient corresponding to each key correlation feature, and optimize the parameters of the interpretation model with the purpose of minimizing the loss function.

[0055] It can be understood that the interpretation model of the embodiment of the present application can effectively extract valuable information from the key correlation features of the patient and convert it into a predicted probability value of postoperative survival time. The powerful expression ability of the multi-layer perceptron layer and the application of the activation function enable the model to accurately capture complex nonlinear relationships and significantly improve the accuracy of the prediction. The clear classification results and probability distribution provided by the output layer enhance the interpretability of the model. The introduction of the loss function ensures the stability and generalization ability of the model during the training process, reduces the risk of overfitting, and has strong robustness. The simplified model structure and efficient activation function selection speed up the training speed and save computing resources.

[0056] Specifically, if Figure 5 As shown in the figure, after the key correlation features are selected, the key correlation feature of the patient's postoperative survival time is x= [x1, x2, …, xm]. The relationship between the survival time and the key correlation features can be To indicate that α =[ α1 , α 2 , … , α m ], is a parameter to be determined; find the survival time and y = F ( x ) is the optimal estimate of the parameters of the design objective function for a given (historical) observation data. , which is transformed into the minimization objective function ,get y = F ( x )relation.

[0057] In order to avoid y = F ( x ) overfitting, introduce artificial neural network (Neural Network), train the artificial neural network for given (historical data) observation data, and obtain the neural network model y= N ( x ). Given the key correlation feature x, the final judgment value is the functional relationship y = N ( x ) and artificial neural network y = N ( x ) The average value of the integration .

[0058] In step S104, the follow-up data corresponding to the predicted result is filled into the missing database.

[0059] It can be understood that the embodiment of the present application fills the follow-up data corresponding to the prediction results into the missing database, thereby improving the accuracy and reliability of missing value filling and improving processing efficiency, so as to improve the follow-up system and enhance the user experience for subsequent applications.

[0060] According to the data missing value filling method proposed in the embodiment of the present application, the target data is divided into multiple discrete time data; the key correlation features of each patient in the target database are obtained, wherein the key correlation features represent the related features with the target data; the key correlation features are input into the judgment model, and the judgment model outputs the corresponding prediction results, wherein the prediction results are the target data falling into the corresponding discrete time results; the prediction results are associated with the corresponding follow-up data and filled into the missing database, thereby improving the accuracy and reliability of the missing value filling and improving the processing efficiency, so as to improve the follow-up system and enhance the user experience.

[0061] The following will be combined Figure 6 The data missing value filling method of the present application includes the preprocessing of historical data, key feature extraction, and learning of the model based on key associated feature data, as follows: Step 1: Data format preprocessing Categorical feature data is the value that a categorical (discrete) feature takes. These features are converted to ordinal integers so that each feature has only one column of integers.

[0062] Step 2, key correlation feature selection includes efficient maximum correlation coefficient and MICe-based feature selection.

[0063] (1) The efficient maximum information coefficient estimation method MICe differs from the maximum information coefficient MIC method in the grid search strategy.

[0064] MICe divides the equal-division axis more finely, and only when the number of equal-division axis divisions is greater than the number of dynamic division axis divisions exists, it reduces the grid search process caused by dynamic division and greatly reduces the time cost.

[0065] 1) Given a binary variable dataset D is a binary random variable The sample data set is divided equally y The maximum information of the axis is: ; In the formula, k To divide in any form x The number of columns of the axis; For equal distribution y The number of rows of the axis; Represents any form of partitioning x Axis k Columns, evenly divided y axis l A grid collection of rows; express The grid element that maximizes the mutual information; express D In the grid G 1 division within; express D In the grid G 1 The mutual information calculated under the partition.

[0066] 2) Equal distribution x Maximum mutual information of axes: ; In the formula, Means evenly divided x The number of columns of the axis;l Indicates that the division is in any form y The number of rows of the axis; Means evenly divided x Axis k Columns, arbitrary division y Axis l A grid collection of rows; express The grid element that maximizes the mutual information; Represents a random variable In the grid G 2 division within; express D In the grid G 2 The mutual information calculated under the partition.

[0067] 3) The average maximum mutual information is: In the formula, l Representation Division y The number of rows of the axis, k Representation Division x The number of columns of the axis; Representation Grid G 2 Belong to the average x Axis k Columns, arbitrary division y Axis l An element in the grid collection of rows; Representation Grid G 1 Any form of division x Axis k Column, evenly divided y Axis l An element in the grid collection for a row.

[0068] 4) The evenly divided feature matrix is: In the formula, represents the equally partitioned feature matrix; represents a binary variable data set; Indicates that the binary variable data set is divided into l OK, k The maximum normalized mutual information obtained by the column grid; represents the average maximum mutual information.

[0069] 5) Given a sample size of n A binary variable dataset D and in yNumber of axis divisions k with x Number of axis divisions l Upper limit conditions met The efficient maximum information coefficient estimate MICe is: In the formula, kl express l OK, k Positive integers in the column k and l The product of n Indicates the number of samples; B ( n ) is expressed as ,

[0070] (2) The feature selection based on MICe first calculates the correlation between physical features and survival time, measures the correlation between each physical feature and survival time, and saves the correlation between each feature and survival time, and then eliminates the redundant features related to each feature.

[0071] 1) Correlation between features and survival time, F =( f 1 , f 2 ,…, f m ) indicates that the first stage is processed by low variance filtering. m Features, C Represents a classification label. i Features f i With label C The correlation between them is expressed as MICe( f i , C ), the features and the results of their correlation with the labels are sorted in descending order and stored in the linked list container FC_Corr.

[0072] 2) Eliminate redundancy between features. Redundancy analysis is to select k The core of the best features, F Middle i Features f i With j Features f j The correlation between them is expressed as MICe( f i , f j ), MICe(f i , f j ) value is larger, indicating f i With j Features f j The stronger the substitutability between them, that is, the stronger the redundancy, the screening of the body feature subset in this stage is cyclical. First, the first feature is taken out from FC_Corr and added to the subset linked list container SF. Then, the next feature is taken out from FC_Corr in turn to determine whether it can be added to the subset linked list container SF. This process is repeated until the number of features in the subset linked list container SF reaches k indivual.

[0073] In selecting s +1 feature, the selected one stored in the linked list container SF s Features ( s < K < m ), for the selected join in FC_Corr SF The s +1 feature f y ( s +1≤ K ) and FC_Corr select f y After the remaining m - s - Any one of 1 features f x The relationship satisfies the inequality

[0074] Step 3: Missing value interpretation and filling Assume that the survival time to be judged y It is divided into 3 time periods, the first one is 1 week, the second and the third are 1 month and 3 months respectively, so the corresponding time periods are 7 days, 7 days to 37 days, and 37 to 127 days.

[0075] After the key correlation features are selected, the key correlation features of the patient's postoperative survival time are x = [ x 1 , x 2 , … , x m ], then the relationship between survival time and key associated features can be To indicate that α =[ α1 , α 2 , … , α m ], To be determined parameter.

[0076] Find the survival time and y = F ( x ) is the optimal estimate of the parameters of the design objective function for a given (historical) observation data. , which is transformed into the minimization objective function ,get y = F ( x )relation.

[0077] In order to avoid y = F ( x ) overfitting, introduce artificial neural network (Neural Network), train the artificial neural network for given (historical data) observation data, and obtain the neural network model y= N ( x ). Given the key correlation feature x, the final judgment value is the functional relationship y = N ( x ) and artificial neural network y = N ( x ) The average value of the integration .

[0078] In summary, in response to the problem of redundant features, this application has designed an efficient method for extracting key related features for postoperative survival time, eliminating redundant features and improving the efficiency of interpretation and filling; integrating the interpretation of missing values, and obtaining two survival time interpretation models through learning from previous case data, using the data model to interpret and fill in the missing death time information, and taking the average to avoid the problem of overfitting of a single model.

[0079] Next, a device for filling missing data values ​​according to an embodiment of the present application will be described with reference to the accompanying drawings.

[0080] Figure 7 It is a block diagram of a device for filling missing data values ​​according to an embodiment of the present application.

[0081] like Figure 7 As shown, the data missing value filling device 10 includes: a division module 100, an acquisition module 200, an input module 300 and a filling module 400.

[0082] Among them, the division module 100 is used to divide the continuous postoperative survival time of patients into multiple discrete time periods; the acquisition module 200 is used to obtain the key correlation features of each patient in the target database, wherein the key correlation features represent the related features with the target data; the input module 300 is used to input the key correlation features into the judgment model, and the judgment model outputs the corresponding prediction results, wherein the prediction results are the target data falling into the corresponding discrete time results; the filling module 400 is used to associate the prediction results with the corresponding follow-up data to fill in the missing database.

[0083] It should be noted that the aforementioned explanation of the embodiment of the method for filling missing values ​​in data is also applicable to the device for filling missing values ​​in data of this embodiment, and will not be repeated here.

[0084] According to the data missing value filling device proposed in the embodiment of the present application, the target data is divided into multiple discrete time data; the key correlation features of each patient in the target database are obtained, wherein the key correlation features represent the related features with the target data; the key correlation features are input into the judgment model, and the judgment model outputs the corresponding prediction results, wherein the prediction results are the target data falling into the corresponding discrete time results; the prediction results are associated with the corresponding follow-up data and filled into the missing database, thereby improving the accuracy and reliability of the missing value filling and improving the processing efficiency, so as to improve the follow-up system and enhance the user experience.

[0085] FIG8 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include: A memory 801 , a processor 802 , and a computer program stored in the memory 801 and executable on the processor 802 .

[0086] When the processor 802 executes the program, the data missing value filling method provided in the above embodiment is implemented.

[0087] Furthermore, the electronic device further comprises: The communication interface 803 is used for communication between the memory 801 and the processor 802 .

[0088] The memory 801 is used to store computer programs that can be executed on the processor 802 .

[0089] The memory 801 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0090] If the memory 801, the processor 802 and the communication interface 803 are implemented independently, the communication interface 803, the memory 801 and the processor 802 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0091] Optionally, in a specific implementation, if the memory 801, the processor 802 and the communication interface 803 are integrated on a chip, the memory 801, the processor 802 and the communication interface 803 can communicate with each other through an internal interface.

[0092] The processor 802 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0093] An embodiment of the present application also provides a computer-readable storage medium having a computer program or instruction stored thereon. When the computer program or instruction is executed by a processor, the method for filling missing data values ​​as described above is implemented.

[0094] An embodiment of the present application also provides a computer program product, including a computer program or instructions, characterized in that when the computer program or instructions are executed, the method for filling missing data values ​​as described above is implemented.

[0095] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0096] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0097] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0098] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one or a combination of multiple of the following technologies known in the art: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0099] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

Claims

1. A method for filling missing data values, characterized in that: The following steps are involved: Divide the target data into multiple discrete time data; Obtain key correlation features of each patient in the target database, wherein the key correlation features represent features related to the target data; The key correlation feature is input into the interpretation model, and the interpretation model outputs the corresponding prediction result, wherein the prediction result is the result that the target data falls into the corresponding discrete time; The follow-up data corresponding to the predicted results are associated with the missing data.

2. The method for filling missing data values ​​according to claim 1, characterized in that: The step of obtaining key associated features of each patient in the target database includes: Identify follow-up data for each missing target key element in the target database; The follow-up data is formatted and converted into a sequence integer to generate a numerical feature, and the numerical feature is feature filtered to select key associated features with a similarity greater than a first target threshold in a historical database.

3. The method for filling missing data values ​​according to claim 1, characterized in that: The training process of the interpretation model: Obtain historical follow-up data; Preprocessing the discrete feature data format in the historical follow-up data into a sequence integer to generate numerical features; The key associated features in the numerical features are screened, and a training data set is generated according to the key associated features and the corresponding survival time to train the interpretation model.

4. The method for filling missing data values ​​according to claim 3, characterized in that: The key associated features in the screening numerical features include: Calculate the efficient maximum information coefficient estimate between each numerical feature and the patient's postoperative survival time; Screening the numerical features whose estimated values ​​of the efficient maximum information coefficient are greater than a second target threshold to generate associated features; Calculate the efficient maximum information coefficient estimation value between the associated features, and eliminate any associated features whose efficient maximum information coefficient estimation value is greater than the third target threshold to generate the key associated features.

5. The method for filling missing data values ​​according to claim 3, characterized in that: The screening of key related features in the numerical features also includes: Performing low variance filtering processing on the numerical features to generate target features, and performing category labels for postoperative survival time on the target features; Calculate the similarity value between each target feature and the category label, sort the target features and the corresponding similarity values ​​in descending order and store them in the target container; Target features are extracted from the target container in sequence and added to the subset linked list container, and it is determined whether the conditions for adding to the subset linked list container are met until the target number of target features in the subset linked list container reaches the target number to generate corresponding key associated features.

6. The method for filling missing data values ​​according to claim 1, characterized in that: The structure of the interpretation model includes: an input layer, a multi-layer perceptron layer, an output layer and a loss function, wherein: The input layer input is a set of key associated features of the patient; The multi-layer perceptron layer includes multiple fully connected layers, each of which uses an objective function to capture the nonlinear relationship and linear relationship between key related features and converts them into a predicted probability value of the patient's postoperative survival time; The output layer is used to output the predicted probability value of the patient's postoperative survival time corresponding to each key correlation feature; The loss function is used to calculate the difference value based on the predicted probability value of the patient's postoperative survival time corresponding to each key correlation feature and the patient's actual postoperative survival time corresponding to each key correlation feature, and optimize the parameters of the interpretation model with the purpose of minimizing the loss function.

7. A device for filling missing data values, characterized in that: include: A partitioning module, used for partitioning the target data into a plurality of discrete time data; An acquisition module, used to acquire key associated features of each patient in the target database, wherein the key associated features represent features related to the target data; An input module, used to input the key correlation features into a judgment model, and the judgment model outputs a corresponding prediction result, wherein the prediction result is a result that the target data falls into a corresponding discrete time; A filling module is used to fill the follow-up data corresponding to the predicted result into the missing database.

8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for filling missing data values ​​as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, it is used to implement the method for filling missing data values ​​as described in any one of claims 1-6.

10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed, the method for filling missing data values ​​as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Medical data missing value filling method and system based on partition data division

    CN117668474A

  • Method and system for predicting hemoglobin concentration of postoperative patient

    CN119028473A