A method for completing missing data of medical indicators based on information bottleneck graph compression
Through the information bottleneck graph compression method, the graph neural network is used to generate sub-graphs and perform feature causal analysis, which solves the problem of missing data in electronic health records, and achieves streamlining of information and improving the accuracy of diagnostic prediction.
Patent Information
- Application Number
- CN202310623896.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2043-05-30
AI Technical Summary
The prior art is difficult to effectively deal with the missing data problems in electronic health records, especially in the case of uneven data distribution and short time intervals, which lead to difficult data analysis and unsatisfactory results.
The information bottleneck graph compression method is used to model all attributes of patients of the same type as undirected complete graphs, and a sub-graph is generated through the graph neural network, and the cross entropy loss and connectivity loss are compressed. Combined with causal reasoning and a learnable mutual information estimator, the amount of information is minimized to streamline the features and generate a characteristic causal relationship diagram.
On the premise of ensuring the accuracy of diagnostic prediction, the amount of information required for missing data processing is effectively reduced, the data processing process is simplified, and the effectiveness and commercial value of data analysis are improved.
Smart Images

Figure CN116630777B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method for completing missing data of medical indicators by using information bottleneck graph compression under the guidance of information theory. Background Art
[0002] In recent years, interest in analyzing patient electronic health record (EHR) data has surged. The rapid adoption of EHRs in healthcare systems worldwide has resulted in the accumulation of vast amounts of EHR data. Consequently, data-driven healthcare has emerged, leveraging existing, large-scale healthcare data to provide the best, most personalized healthcare possible. Because patients' EHRs are the primary vehicle for data-driven healthcare research, understanding the information contained in EHRs, including demographic information, diagnoses, laboratory test results, prescriptions, radiology images, clinical notes, and more, is crucial.
[0003] However, the widespread problem of missing data in electronic health record data significantly limits the modeling of electronic health record data by many mainstream data analysis methods. Furthermore, electronic health record data contains physical measurements of hospitalized patients. Due to objective factors, these data contain a high number of missing values and their distribution is uneven. Deleting all data samples with missing values would result in the deletion of many samples, resulting in a significant loss of information. Without processing this data, it would be difficult to use it in traditional machine learning algorithms. Therefore, it is necessary to employ predictive methods to handle missing data.
[0004] In existing research, Tang et al. used a set of rules derived from clinical experience to fill missing values using the median or mean of the patient's normal range, achieving good results. However, this method becomes more difficult to implement when the data interval is short and the number of missing values is large. The expectation maximization method (EM) studied by Dempster et al. builds a model based on the observed data and uses the marginal distribution of the observed data to estimate unknown parameters. However, this method is prone to falling into local minima and has a slow convergence rate.
[0005] With the rapid development of artificial intelligence, the medical field has also begun to incorporate machine learning methods to process data. However, because medical data contains a vast amount of patient information, including medical history, test results, and medications, the dimensionality of the training data is extremely high and contains many noisy features, making the final experimental results unsatisfactory. Therefore, it is important to adopt a new method to handle missing data in electronic health records. Summary of the Invention
[0006] The present invention aims to address the shortcomings of the existing technology by providing a method for completing missing data of medical indicators using information bottleneck graph compression. The method adopts a flexible and scalable interpolation framework to model all attributes of patients of the same type into an undirected complete graph and input it into a multi-layer feedforward graph neural network. The obtained node features are screened and a subgraph is generated. The obtained subgraph participates in the calculation of cross entropy loss and a connectivity loss. The subgraph is continuously pruned and compressed according to the diagnostic requirements and the information bottleneck principle to ensure that the subgraph is consistent with the downstream task objectives while being as concise as possible. The interpolation framework has a learnable estimator. The input is the features corresponding to the subgraph and its corresponding complete graph, and the output is the mutual information estimate between the two. By minimizing this mutual information, the total amount of information contained in the subgraph can be compressed, so that it filters out minor attributes that are not helpful for downstream tasks, and the total amount of information used for missing interpolation and diagnosis prediction can be simplified as much as possible. The present invention introduces a causal reasoning method to mine the causal relationship between features, find the optimal feature subset, and simultaneously complete the visualization of the subset to generate a feature causal relationship graph. The method is simple and has good application prospects and commercial value.
[0007] The present invention achieves its objectives by providing a method for completing missing data for medical indicators using information bottleneck graph compression. The method is characterized by modeling all attributes of patients of the same type as an undirected complete graph, which is continuously pruned and compressed according to diagnostic requirements and the information bottleneck principle. This ensures that the total amount of information used for missing data interpolation and diagnostic prediction is minimized while ensuring that the subgraph is consistent with downstream task objectives. Specifically, the missing data completion method includes the following steps:
[0008] Step 1: Preprocess the input data
[0009] Step 2: Build an undirected complete graph based on all attributes of patients of the same type as nodes.
[0010] Step 3: Feed the initial complete graph forward to the graph neural network to obtain a series of node features X, and pass it to the subgraph generator for pruning and compression to obtain the subgraph and its contained node feature X sub .
[0011] Step 4: Use the READOUT function and the trainable estimator to calculate and minimize the subgraph With the complete graph Mutual information between
[0012] Step 5: According to the subgraph where the missing number is located, the features of its adjacent points are integrated and interpolated to obtain the completed subgraph Its node characteristics
[0013] Step 6, and Connectivity constraints and classification losses are imposed separately to ensure that the subgraph generation process meets the minimum amount of information required for the task.
[0014] The step 1 specifically includes:
[0015] In step S101, important features in different types of elderly chronic disease data are selected by deleting irrelevant features and redundant features with less effective information, and one-hot encoding is performed on features containing multiple categories.
[0016] In step S102 , to address the common but serious imbalance problem in medical data, undersampling is performed on the input samples based on Tomek's links.
[0017] Step S103: Initialize and fill in the missing data using the MissForest interpolation strategy.
[0018] The step 3 specifically includes:
[0019] Step S301: Initialize the graph neural network θ1 with the graph attention network as the skeleton and the multilayer perceptron θ2 with hidden units of [32, 64, 128].
[0020] Step S302: Feed forward to θ1 to obtain the corresponding node feature X, and then use θ2 to filter out the subgraph according to its corresponding adjacency matrix A. and the node features X it contains sub .
[0021] The step 4 specifically includes:
[0022] Step S4, use the READOUT function to get and The global features of and The relative entropy between the distributions, using the estimator φ2 to estimate the mutual information between the two Its specific goal is defined by the following formula (a):
[0023]
[0024] in, represents a trainable mutual information estimator, with input Its corresponding subgraph and Respectively represent and The joint distribution between them and their respective marginal distributions.
[0025] The step 5 specifically includes: using the component θ3 of the bottle mouth structure to fuse The feature information of the adjacent points of the missing data contained in is interpolated and obtained Corresponding node features
[0026] The step 6 specifically includes:
[0027] Step S601: To ensure the consistency of subgraph screening and interpolation strategies with downstream tasks, a classification loss function is used. constraint
[0028] constraint where p φ represents the classifier, y gt Represents the true label, and y represents the feature of the completed node Classification predictions made.
[0029] Step S602: To improve the effectiveness of subgraph partitioning, connectivity loss is used. Constraint X sub structure, where Norm(·) represents row-by-row normalization; represent The probability distribution of the nodes being divided into subgraphs, N represents Total number of midpoints; is the identity matrix, ||·|| F stands for the Frobenius norm. The essence of is to promote the probability of each node being divided into a subgraph or not to present a one-hot distribution, reducing ambiguous judgments and potential redundant nodes.
[0030] Compared with the existing technology, the present invention has the advantage of streamlining the total amount of information used for missing data interpolation and diagnostic prediction as much as possible while ensuring consistency with the task objectives. The method is simple and effectively solves the problem of processing missing data in electronic health records. The experimental results are satisfactory and it has good application prospects and commercial value. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0032] The present invention addresses the common problem of missing data in medical diagnosis and proposes a flexible and scalable interpolation framework. First, all the patient's attributes are modeled into an undirected complete graph and input into a multi-layer feedforward graph neural network. Then, the obtained node features are screened and a subgraph is generated. The obtained subgraph participates in the calculation of cross entropy loss and a connectivity loss to ensure that the subgraph is consistent with the downstream task objectives while being as concise as possible. In addition, the framework also includes a learnable estimator, which takes as input the features corresponding to the subgraph and its corresponding complete graph, and outputs an estimate of the mutual information between the two. By minimizing this mutual information, the total amount of information contained in the subgraph can also be compressed, so that it can filter out minor attributes that are not very helpful for downstream tasks. The main advantages are as follows: introducing causal reasoning methods, exploring the causal relationship between features, finding the optimal feature subset, and simultaneously completing the visualization of the subset to generate a feature causal relationship graph.
[0033] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It is obvious that the described embodiments are part of the embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0034] Example 1
[0035] See Figure 1 According to a preferred embodiment of the present invention, the interpolation and diagnosis prediction of medical missing data specifically include the following steps:
[0036] Step 1: Preprocess the input data. The case dataset to be processed contains subsets at different time nodes and has very serious imbalance and missing problems. Therefore, each subset needs to be cleaned and initialized first, and then the subsets are fused according to the time sequence.
[0037] To address the severe class imbalance issue, the algorithm undersamples the class with the majority of training examples: Assume that samples A and B come from the majority and minority classes, respectively. Sample A is sampled only if A and B are each other's k-nearest neighbors. To facilitate subsequent interpolation model training, preprocessing also requires initializing and filling missing data. The strategy used is the MissForest method based on random forests. The resulting interpolation initial values are updated iteratively and momentum-based during subsequent training.
[0038] For the pathology dataset after class balancing and initialization imputation, the static data is fused with the dynamic data in the time dimension. At this time, each patient in the dataset corresponds to multiple samples, and the total number is the time series length * the number of patients. The data is reshaped into [sample id, time series, feature] and fused using a time series transformer.
[0039] Step 2: Based on step 1, take all attributes of patients of the same type as nodes and build an undirected complete graph based on them At the same time, initialize the corresponding adjacency matrix (The superscript i∈{1,…c} represents the category index).
[0040] Step 3: Complete the initial graph for each category Feed forward to the corresponding node features in the graph neural network θ1, that is Then, the subgraph generator θ2 is used to perform pruning and compression, filter key nodes and structures in each complete graph, and generate subgraphs corresponding to each category. and
[0041] The graph neural network θ1 structure can be flexibly selected. In the present invention, the graph attention network ( P., Cucurull, G., Casanova, A., Romero, A., Liò, P., & Bengio, Y. Graph Attention Networks. In International Conference on Learning Representations.), and a multilayer perceptron with a hidden unit structure of [32, 64, 128] was selected as the subgraph generator.
[0042] Step 4: Use the READOUT function and the trainable estimator to calculate and minimize the subgraph With the complete graph Mutual information between The estimator is a mutual information estimator, and there are multiple options, such as: Belghazi, MI, Baratin, A., Rajeshwar, S., Ozair, S., Bengio, Y., Courville, A., & Hjelm, D. (2018, July). Mutual information neural estimation. In International conference on machine learning (pp. 531-540). PMLR, Lehner, J., Alkin, B., Fürst, A., Rumetshofer, E., Miklautz, L., & Hochreiter, S. (2023). Contrastive Tuning: A Little Help to MakeMasked Autoencoders Forget. arXiv preprint arXiv: 2304.10520, etc.
[0043] The subgraph generator θ2 essentially outputs a distribution matrix Therefore, it is necessary to avoid redundant nodes from being ambiguously assigned to subgraphs as much as possible. The specific approach is to optimize For example, suppose there are two nodes in the subgraph corresponding to the i-th category, and I2 is The identity matrix of ij for The element in row i and column j in the graph, if the first node is included in the subgraph Then minimize will make and If it is not in the subgraph, the result is the opposite, and the situation is similar for other nodes.
[0044] The above embodiments are only for further explanation of the present invention and are not intended to limit the present invention. Any equivalent implementation of the present invention should be included in the scope of the claims of the present invention.
Claims
1. A method for completing missing data of medical indicators using information bottleneck graph compression, characterized by: We model all attributes of patients of the same type as an undirected complete graph. We continuously prune and compress it based on diagnostic requirements and the information bottleneck principle to ensure that the subgraph is consistent with downstream task objectives and to reduce the total amount of information used for missing data interpolation and diagnostic prediction. The specific steps for missing data completion include: Step 1: Preprocess the input raw data; Step 2: Use all attributes of the same type of patients as nodes and build an undirected complete graph on top of the nodes Step 3: Feed the initial complete graph forward to the graph neural network to obtain a series of node features, which are passed to the subgraph generator for pruning and compression to obtain the subgraph and the node features it contains; Step 4: Use the READOUT function and the trainable estimator to calculate and minimize the subgraph With the complete graph Mutual information between Step 5: According to the subgraph where the missing number is located, fuse its adjacent point features and interpolate to obtain the completed subgraph Its node characteristics Step 6: Minimize the subgraph and complete subgraph Connectivity constraints and classification losses are imposed separately to ensure that the subgraph generation process meets the minimum amount of information required for the task.
2. The method for completing missing data of medical indicators using information bottleneck graph compression according to claim 1, characterized in that: The step 1 specifically includes: Step S101: By deleting irrelevant features and redundant features with less effective information, important features in different types of elderly chronic disease data are selected, and one-hot encoding is performed on features containing multiple categories; Step S102: performing undersampling on the input samples based on Tomek's links; Step S103: Initialize and fill in the missing data using the MissForest interpolation method.
3. The method for completing missing data of medical indicators by information bottleneck graph compression according to claim 1, characterized in that: Step 2 models all attributes of the sample into nodes and builds an undirected complete graph by combining the attributes of patients of the same category.
4. The method for completing missing data of medical indicators using information bottleneck graph compression according to claim 1, characterized in that: The step 3 specifically includes: Step S301: Initialize a graph neural network with a graph attention network as the skeleton and a multi-layer perceptron with hidden units of [32, 64, 128]; Step S302: Undirected complete graph Feed forward to the graph neural network to obtain the corresponding node features, and use the multi-layer perceptron to filter out the subgraph based on the node features and their corresponding adjacency matrix and the node features it contains.
5. The method for completing missing data of medical indicators by information bottleneck graph compression according to claim 1, characterized in that: The step 4 specifically includes: Step S4: Use the READOUT function to obtain the undirected complete graph and subgraphs The global features are modeled using the DONSKER-VARADHAN feature and The relative entropy between the distributions, using an estimator to estimate the mutual information between the two Its specific goal is defined by the following formula (a): in, represents a trainable mutual information estimator, with input Its corresponding subgraph and Respectively represent and The joint distribution between them and their respective marginal distributions.
6. The method for completing missing data of medical indicators using information bottleneck graph compression according to claim 1, characterized in that: The step 5 specifically includes: using the components of the bottle mouth structure to fuse the subgraphs The feature information of the adjacent points of the missing data contained in the graph is interpolated to obtain the completed subgraph. The corresponding completion node features 7. The method for completing missing medical indicator data by compressing information bottleneck graph according to claim 1, characterized in that: The step 6 specifically includes: Step S601: Using classification loss function Constrained completion node features structure to ensure the consistency of subgraph screening and interpolation strategies with downstream tasks, the classification loss function It is expressed by the following formula (b): Among them, represents the classifier, y gt Represents the real label; y represents the feature of the completed node The classification predictions made; Step S602: Using connectivity loss function Constraining the node feature structure, the connectivity loss function It is expressed by the following formula (c): Among them, N+rm· represents row-by-row normalization; represent The probability distribution of the nodes being divided into subgraphs, N represents Total number of midpoints; is the identity matrix, ||·|| F stands for Frobenius norm.
Citation Information
Patent Citations
Android malicious software detection method and detection system based on traffic fingerprints and graph data features
CN114640502A
Medical diagnosis missing data completion method and completion apparatus, and electronic device and medium
WO2022222026A1