RPCA and SPCA Anomaly Removal for Healthcare Metric Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for reducing data metrics in healthcare systems are hindered by redundancies and inaccuracies, which increase the data burden and processing costs.

Innovation Solution

The use of Robust Principal Component Analysis (RPCA) and Sparse Principal Component Analysis (SPCA) to identify and remove anomalies in data sets, thereby recommending a smaller set of metrics that reduce the data processing burden without compromising information accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If a large set of metrics is collected to ensure comprehensive data coverage, then information completeness is improved, but data processing burden increases

Engineering Contradiction:
Improveinformation completenessVSAvoiddata processing burden
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts and removes redundant metrics from the dataset by analyzing correlations between metrics. The system identifies metrics that can be inferred from other metrics and removes them, keeping only the essential subset of metrics that maintains information completeness while reducing processing burden.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the large set of metrics into independent groups based on correlation analysis. By dividing the metrics into independent segments, the system reduces the overall processing burden while ensuring that each segment contributes unique information not redundant with other segments.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If standard correlation methods are used to identify redundant metrics, then analysis simplicity is improved, but measurement accuracy deteriorates due to anomalies

Engineering Contradiction:
Improveanalysis simplicityVSAvoidcorrelation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary anomaly detection and removal before conducting correlation analysis. By preprocessing the data to eliminate anomalies, the system ensures that subsequent correlation calculations are based on clean, accurate data, thereby improving measurement precision while maintaining analysis simplicity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary preprocessing step that mediates between raw data and correlation analysis. This intermediary layer detects and corrects anomalies, transcription errors, and data quality issues before the main analysis, protecting the correlation calculations from being distorted by poor quality data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If PCA is used for dimensionality reduction, then mathematical optimization is improved, but metric interpretability deteriorates due to linear combinations

Engineering Contradiction:
Improvedimensionality reduction efficiencyVSAvoidmetric interpretability
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

Instead of using PCA to create linear combinations of metrics (which loses interpretability), the patent inverts the approach by using correlation analysis to identify and remove redundant metrics, keeping only the original, interpretable metrics. This maintains metric interpretability while achieving dimensionality reduction through selective elimination rather than transformation.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS12282464B2Systems and methods for reducing data collection burden
Publication Date: 2025.04.22 THE MITRE CORPORATION
  • US12282464B2 patent drawing
  • US12282464B2 patent drawing
  • US12282464B2 patent drawing

AI summary

A system for reducing data collection burden, comprising: one or more programs including instructions for: receiving a first set of metrics for a plurality of facilities; receiving data associated with the first set of metrics from one or more facilities of the plurality of facilities; determining one or more anomalies in the received data; removing the determined one or more anomalies from the received data; selecting a second set of metrics from the first set of metrics, wherein a number of metrics of the second set is less than a number of metrics of the first set of metrics; and outputting a recommendation applicable to the plurality of facilities based on the second set of metrics.