Map Representations of Categorical Data for Classification Predictions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing predictive data analysis systems are ill-equipped to perform accurate, efficient, and reliable predictive data analysis in domains with high-dimensional categorical feature spaces and high cardinality, particularly in areas with scarce data such as rare diseases, due to inefficiencies in computational resource usage and inability to effectively communicate temporal information.
Innovation Solution
The use of map representations of categorical data through machine learning models that generate cumulative clinical and provider history maps without temporal signals, allowing for efficient and reliable predictive analysis by transforming clinical and provider data into vector representations that can be processed using machine learning frameworks, enabling accurate predictions and insights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing predictive data analysis systems are used to process high-dimensional categorical data with high cardinality, then computational resources are consumed, but predictive accuracy and efficiency deteriorate
Solution Approach 1:
The patent transforms the parameter representation of categorical data by encoding high-cardinality features into compressed vector representations (e.g., using hashing or embedding techniques). This parameter transformation reduces the dimensionality of the input space while preserving predictive information, enabling machine learning models to achieve accurate predictions with lower computational resource consumption.
Solution Approach 2:
The patent extracts and separates temporal information from the categorical data processing pipeline by using cumulative histograms that inherently capture temporal patterns. This extraction allows the system to convey temporal information without requiring complex temporal processing, thereby improving predictive accuracy while reducing computational overhead.
2Reliability
If existing systems process high-dimensional categorical data, then data coverage is attempted, but efficiency and reliability deteriorate due to inability to effectively communicate temporal information
Solution Approach 1:
The patent introduces a new dimensional representation for categorical data by using cumulative histograms that add a temporal aggregation dimension. This dimensional transformation allows temporal information to be communicated effectively through the structure of the histogram data, improving predictive reliability without time loss.
3Measurement precision
If map representations with multiple coding standards are generated, then predictive accuracy improves, but device complexity increases
Solution Approach 1:
The patent segments the categorical data processing by applying different coding standards to different subsets of features or data sources. Each coding standard can be optimized for specific types of categorical data, improving overall prediction accuracy while keeping individual processing modules relatively simple and manageable.
Data Source
AI summary
Various embodiments of the present disclosure provide machine learning using map representations of categorical data to provide classification predictions. In one example, an embodiment provides for generating a first map representation of a first categorical input feature set for categorical data based on a first coding standard. A second map representation of a second categorical input feature set for the categorical data may also be generated based on a second coding standard. Additionally, at least one machine learning model may be applied to the first map representation and the second map representation to generate the prediction output. Based on the prediction output one or more prediction-based actions may also be performed.


