Map Representations of Categorical Data for Classification Predictions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing predictive data analysis systems are ill-equipped to perform accurate, efficient, and reliable predictive data analysis in domains with high-dimensional categorical feature spaces and high cardinality, particularly in areas with scarce data such as rare diseases, due to inefficiencies in computational resource usage and inability to effectively communicate temporal information.

Innovation Solution

The use of map representations of categorical data through machine learning models that generate cumulative clinical and provider history maps without temporal signals, allowing for efficient and reliable predictive analysis by transforming clinical and provider data into vector representations that can be processed using machine learning frameworks, enabling accurate predictions and insights.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing predictive data analysis systems are used to process high-dimensional categorical data with high cardinality, then computational resources are consumed, but predictive accuracy and efficiency deteriorate

Engineering Contradiction:
Improvepredictive accuracyVSAvoidcomputational resource usage
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent transforms the parameter representation of categorical data by encoding high-cardinality features into compressed vector representations (e.g., using hashing or embedding techniques). This parameter transformation reduces the dimensionality of the input space while preserving predictive information, enabling machine learning models to achieve accurate predictions with lower computational resource consumption.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent extracts and separates temporal information from the categorical data processing pipeline by using cumulative histograms that inherently capture temporal patterns. This extraction allows the system to convey temporal information without requiring complex temporal processing, thereby improving predictive accuracy while reducing computational overhead.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If existing systems process high-dimensional categorical data, then data coverage is attempted, but efficiency and reliability deteriorate due to inability to effectively communicate temporal information

Engineering Contradiction:
Improvepredictive reliabilityVSAvoidtemporal information communication
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent introduces a new dimensional representation for categorical data by using cumulative histograms that add a temporal aggregation dimension. This dimensional transformation allows temporal information to be communicated effectively through the structure of the histogram data, improving predictive reliability without time loss.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If map representations with multiple coding standards are generated, then predictive accuracy improves, but device complexity increases

Engineering Contradiction:
Improveclassification prediction accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the categorical data processing by applying different coding standards to different subsets of features or data sources. Each coding standard can be optimized for specific types of categorical data, improving overall prediction accuracy while keeping individual processing modules relatively simple and manageable.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240428088A1Machine learning using map representations of categorical data to provide classification predictions
Publication Date: 2024.12.26 UNITEDHEALTH GROUP INC
  • US20240428088A1 patent drawing
  • US20240428088A1 patent drawing
  • US20240428088A1 patent drawing

AI summary

Various embodiments of the present disclosure provide machine learning using map representations of categorical data to provide classification predictions. In one example, an embodiment provides for generating a first map representation of a first categorical input feature set for categorical data based on a first coding standard. A second map representation of a second categorical input feature set for the categorical data may also be generated based on a second coding standard. Additionally, at least one machine learning model may be applied to the first map representation and the second map representation to generate the prediction output. Based on the prediction output one or more prediction-based actions may also be performed.