Heterogeneous Kernel Embedding for Outlier-Rich Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning methods fail to accurately identify and model order in large datasets, especially when dealing with outliers and heterogeneous data, due to computational intractability and the need for homogeneous data structures.
Innovation Solution
A computer-implemented method and system for generating an interpretable kernel embedding for heterogeneous data by identifying base kernels, applying unique composition rules, fitting them into a stochastic process model, and standardizing the results to create an interpretable representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If known machine learning methods model all or most of the data in the presence of outliers, then they can handle large datasets, but they fail to accurately identify and model order present in the data
Solution Approach 1:
The patent segments the data into multiple subsets and applies different kernel components to different subsets. This allows the model to identify order in specific subsets without being overwhelmed by outliers in the entire dataset, thereby improving measurement precision while maintaining adaptability to heterogeneous data structures.
Solution Approach 2:
The patent applies local quality by allowing different kernel components to have different properties for different data subsets. Each subset can be modeled with appropriate kernel characteristics, enabling accurate order identification in local regions while accommodating heterogeneity across the entire dataset.
2Device complexity
If a single shared kernel component is used to model multiple time series, then computational complexity is reduced, but the multiple time series must be somewhat homogeneous which is problematic when outliers exist or data is heterogeneous
Solution Approach 1:
The patent segments the kernel into multiple components, each responsible for different subsets of time series. This segmentation allows the model to handle heterogeneous data with outliers while maintaining computational feasibility through modular structure and efficient computation strategies.
Solution Approach 2:
The patent creates a composite kernel structure by combining multiple kernel components. Each component has specific properties suited for different data characteristics, and their combination enables the model to handle heterogeneous data while maintaining computational efficiency through the structured composite architecture.
3Extent of automation
If compositional kernel search builds explanation from simple concepts and combines them iteratively, then automatic description of data characteristics is achieved, but searching through all possible structure-sharing combinations results in explosion in complexity
Solution Approach 1:
The patent segments the search space by assigning different kernel components to different subsets of time series. This segmentation dramatically reduces the search space complexity while maintaining automatic hypothesis generation capability, as each subset can be searched independently rather than searching through all possible combinations for the entire dataset.
Solution Approach 2:
The patent applies partial action by focusing the search on relevant subsets of data rather than attempting to find a single kernel that explains all data. This partial approach to kernel selection for each subset reduces computational complexity while still achieving automatic hypothesis generation for the overall heterogeneous dataset.
Data Source
AI summary
A computer-implemented method for generating an interpretable kernel embedding for heterogeneous data. The method can include identifying a set of base kernels in the heterogeneous data; and creating multiple sets of transformed kernels by applying a unique composition rule or a unique combination of multiple composition rules to the set of base kernels. The method can include fitting the multiple sets into a stochastic process model to generate fitting scores that respectively indicate a degree of the fitting for each of the multiple sets; storing the fitting scores in a matrix; and standardizing the matrix to generate the interpretable kernel embedding for the heterogeneous data.


