Health Data Mining Using Machine Learning for Unclean Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Collecting and analyzing clean data is complex, time-consuming, and costly, limiting its availability for drawing statistically-valid conclusions, especially in healthcare where large sample sizes are needed for statistical significance.

Innovation Solution

Utilizing machine-learning and data-mining algorithms to process and analyze large amounts of unclean data from non-selected populations, facilitating the extraction of valuable information and providing health-related services by correlating genetic data with health history and behaviors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If clean data from randomized controlled trials is used, then measurement precision and reliability are improved, but productivity and time consumption worsen

Engineering Contradiction:
Improvedata qualityVSAvoiddata collection efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent employs inexpensive, easily deployable data collection methods such as mobile health applications, wearable sensors, and electronic health records that can rapidly gather large volumes of data without the lengthy recruitment and protocol adherence requirements of traditional clinical trials. This enables fast data accumulation while maintaining sufficient quality for statistical analysis.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Solution Approach 2:

The patent transforms the parameters of data collection by shifting from small sample sizes with strict inclusion criteria to large sample sizes with more flexible data capture methods. This parameter change allows real-world data to be collected at scale, improving productivity while the large sample size compensates for individual data point variability through statistical power.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If large sample sizes are obtained, then statistical significance is improved, but device complexity and data processing requirements worsen

Engineering Contradiction:
Improvestatistical significanceVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the complex task of analyzing large datasets into manageable components through standardized data collection protocols, modular analysis pipelines, and distributed computing architectures. This segmentation reduces processing complexity by breaking down the overall analytical challenge into smaller, independently manageable tasks that can be executed efficiently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary layers including data standardization frameworks, preprocessing algorithms, and intermediate storage systems that mediate between raw data collection and final analysis. These intermediaries simplify the processing complexity by preparing and organizing data before it reaches the analytical engine, making large-scale data handling more manageable.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If unclean data from non-selected populations is analyzed, then productivity and sample size are improved, but measurement precision worsens

Engineering Contradiction:
Improvedata collection speedVSAvoiddata cleanliness
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent converts the previously harmful characteristic of data 'dirtiness' into a benefit by embracing real-world variability and heterogeneity as valuable features rather than noise to be eliminated. This approach allows analysis of authentic population diversity, and through appropriate statistical methods, extracts meaningful patterns that reflect actual clinical practice and patient populations.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent changes the acceptance parameters for data quality by lowering thresholds for inclusion criteria and embracing greater variability in data sources, collection methods, and population characteristics. This parameter change enables rapid data collection from diverse sources while statistical techniques account for the increased heterogeneity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7647285B2Tools for health and wellness
Publication Date: 2010.01.12 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7647285B2 patent drawing
  • US7647285B2 patent drawing
  • US7647285B2 patent drawing

AI summary

A tool for providing health and/or wellness services is described herein. Not necessarily clean or unclean data about a plurality of self-selected or non-selected or unselected subjects is received. The data can be aggregated and mined at least in part by employing a statistical algorithm, a data-mining algorithm and/or a machine-learning algorithm. The data can be further employed to provide health and/or wellness services to participants.