Gaussian Mixture Model for Missing Medical Data Imputation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for handling missing medical data in patient monitoring systems, particularly in telemedicine, are limited as they either ignore valuable data or generate biased estimates that fail to capture short-term changes in patient conditions.
Innovation Solution
A method and system for estimating missing medical data using a telehealth analysis system that models mean and covariance structures with polynomial curves and Gaussian mixture models, accounting for changes over time to provide accurate imputations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If correlation methods are used to estimate missing data, then values can be obtained for missing variables, but the estimates are biased and do not accurately represent individual patients
Solution Approach 1:
The patent segments the population into latent subgroups based on similar longitudinal patterns, then estimates missing data separately for each subgroup. This allows individualized estimation that captures patient-specific patterns rather than applying a single population-level correlation model to all patients, thereby improving both accuracy and representativeness for individual patients.
Solution Approach 2:
The patent uses dynamic clustering that adapts to changing patterns in the data over time. The latent subgroup assignments are not static but are determined by the actual longitudinal patterns observed for each patient, allowing the estimation method to dynamically adjust to individual patient trajectories and capture short-term changes in condition.
2Measurement precision
If regression techniques are used to impute missing data, then values can be generated based on historical patient data, but the estimates reflect long-term averages and ignore short-term changes
Solution Approach 1:
The patent employs dynamic clustering that captures time-varying patterns in patient data. By identifying latent subgroups based on longitudinal patterns and estimating missing data within these dynamic groups, the method preserves short-term fluctuations and individual trajectories rather than smoothing them into long-term averages, thereby reliably capturing short-term condition changes.
Solution Approach 2:
The patent performs preliminary clustering analysis to identify the appropriate latent subgroup for each patient before estimating missing data. This preliminary action of grouping patients by similar patterns ensures that subsequent imputation respects individual patient trajectories and short-term variations, rather than applying generic long-term averages.
3Measurement precision
If complete medical records are required for analysis, then data quality is high, but valuable partial data is discarded
Solution Approach 1:
The patent converts the previously harmful effect of missing data into a beneficial feature by using the pattern of observed and missing values themselves to identify latent subgroups. Rather than discarding partial records, the method uses the specific pattern of which variables are observed versus missing to inform the clustering and imputation process, thereby benefiting from all available data including previously unusable partial records.
Solution Approach 2:
The patent replaces the mechanical filtering approach (discarding incomplete records) with a statistical imputation approach. Instead of mechanically removing partial data, the system uses statistical methods to estimate missing values, thereby substituting a data loss mechanism with a data recovery mechanism that preserves and utilizes all available information.
Data Source
Figure 1
Figure 2
Figure 2
AI summary
A method for estimating values of missing data in partial sets of medical data includes generating a Gaussian mixture with a time-varying mean and time and lag varying covariances. The method generates the estimate for a missing datum with the Gaussian distribution having a selected mean and covariance corresponding to the time of the missing datum. An estimate of the missing datum is generated with reference to the mean of the Gaussian distribution conditioned on other medical data that are observed at the time of the missing datum.