Heterogeneous Kernel Embedding for Outlier-Rich Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning methods fail to accurately identify and model order in large datasets, especially when dealing with outliers and heterogeneous data, due to computational intractability and the need for homogeneous data structures.

Innovation Solution

A computer-implemented method and system for generating an interpretable kernel embedding for heterogeneous data by identifying base kernels, applying unique composition rules, fitting them into a stochastic process model, and standardizing the results to create an interpretable representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If known machine learning methods model all or most of the data in the presence of outliers, then they can handle large datasets, but they fail to accurately identify and model order present in the data

Engineering Contradiction:
Improveaccuracy of identifying orderVSAvoidability to handle heterogeneous data
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the data into multiple subsets and applies different kernel components to different subsets. This allows the model to identify order in specific subsets without being overwhelmed by outliers in the entire dataset, thereby improving measurement precision while maintaining adaptability to heterogeneous data structures.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by allowing different kernel components to have different properties for different data subsets. Each subset can be modeled with appropriate kernel characteristics, enabling accurate order identification in local regions while accommodating heterogeneity across the entire dataset.

Inventive Principle:
Principle #3Local quality

2Device complexity

If a single shared kernel component is used to model multiple time series, then computational complexity is reduced, but the multiple time series must be somewhat homogeneous which is problematic when outliers exist or data is heterogeneous

Engineering Contradiction:
Improvecomputational complexityVSAvoidability to handle heterogeneous data
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent segments the kernel into multiple components, each responsible for different subsets of time series. This segmentation allows the model to handle heterogeneous data with outliers while maintaining computational feasibility through modular structure and efficient computation strategies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a composite kernel structure by combining multiple kernel components. Each component has specific properties suited for different data characteristics, and their combination enables the model to handle heterogeneous data while maintaining computational efficiency through the structured composite architecture.

Inventive Principle:
Principle #40Composite materials

3Extent of automation

If compositional kernel search builds explanation from simple concepts and combines them iteratively, then automatic description of data characteristics is achieved, but searching through all possible structure-sharing combinations results in explosion in complexity

Engineering Contradiction:
Improveautomatic hypothesis generationVSAvoidsearch space complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent segments the search space by assigning different kernel components to different subsets of time series. This segmentation dramatically reduces the search space complexity while maintaining automatic hypothesis generation capability, as each subset can be searched independently rather than searching through all possible combinations for the entire dataset.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by focusing the search on relevant subsets of data rather than attempting to find a single kernel that explains all data. This partial approach to kernel selection for each subset reduces computational complexity while still achieving automatic hypothesis generation for the overall heterogeneous dataset.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11354600B2System and method for heterogeneous relational kernel learning
Publication Date: 2022.06.07 BOOZ ALLEN HAMILTON INC
  • US11354600B2 patent drawing
  • US11354600B2 patent drawing
  • US11354600B2 patent drawing

AI summary

A computer-implemented method for generating an interpretable kernel embedding for heterogeneous data. The method can include identifying a set of base kernels in the heterogeneous data; and creating multiple sets of transformed kernels by applying a unique composition rule or a unique combination of multiple composition rules to the set of base kernels. The method can include fitting the multiple sets into a stochastic process model to generate fitting scores that respectively indicate a degree of the fitting for each of the multiple sets; storing the fitting scores in a matrix; and standardizing the matrix to generate the interpretable kernel embedding for the heterogeneous data.