Categorical Input Models With Level Merging for Reliable Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing predictive data analysis systems face challenges in handling categorical input data, leading to inaccurate and unreliable predictions due to the imposition of ordinal ordering, which introduces noise and overfitting, and lack of interpretability, especially in high-impact business contexts where explanations are required.

Innovation Solution

The use of categorical level merging, mutual-information-based feature filtering, and feature-correlation-based feature filtering to preprocess training data, followed by training categorical input machine learning models that generate interpretable predictions by leveraging statistical relationships between feature values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing predictive data analysis systems use traditional machine learning models on categorical data, then they can process the data, but they introduce noise and overfitting due to imposition of ordinal ordering leading to inaccurate predictions

Engineering Contradiction:
Improveprediction accuracyVSAvoidnoise and overfitting
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent changes the fundamental parameter representation by treating categorical data as non-ordinal entities rather than imposing numerical ordering. This is achieved through one-hot encoding and using tree-based models that naturally handle categorical variables without requiring ordinal assumptions, thereby eliminating the noise and overfitting introduced by traditional ordinal-based approaches

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent substitutes traditional mechanical preprocessing steps (such as manual feature engineering and ordinal encoding) with automated machine learning techniques including tree-based models and automatic feature selection algorithms that can directly process categorical data without requiring artificial ordering

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If existing predictive data analysis systems use traditional models, then they can generate predictions, but they lack interpretability which is required in high-impact business contexts

Engineering Contradiction:
ImproveinterpretabilityVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the model into interpretable components by using tree-based models where each decision path can be individually analyzed and explained. This allows the complex predictive task to be broken down into understandable decision rules that stakeholders can interpret and validate

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces feature importance metrics and decision path analysis as intermediary tools that bridge the gap between complex model operations and human interpretation, providing explanatory outputs that make the model's reasoning transparent without simplifying the underlying complex calculations

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If categorical data is processed using traditional methods, then computation can be performed, but processing efficiency is reduced due to noise and overfitting requiring extensive preprocessing

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidpreprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by implementing automated preprocessing steps including automatic feature selection, one-hot encoding, and data validation before model training. This upfront preparation reduces the need for iterative tuning and reprocessing, thereby improving overall processing efficiency despite the initial computational investment

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12423611B2Categorical input machine learning models
Publication Date: 2025.09.23 OPTUM SERVICES IRELAND LTD
  • US12423611B2 patent drawing
  • US12423611B2 patent drawing
  • US12423611B2 patent drawing

AI summary

There is a need for more effective and efficient predictive data analysis based at least in part on categorical input data. This need can be addressed by, for example, solutions for performing predictive data analysis that utilize at least one of categorical level merging, mutual-information-based feature filtering, feature-correlation-based feature filtering to generate training data feature value arrangements, as well as training and using categorical input machine learning models trained using the training data feature value arrangements.