Landmark Feature Selection for Multidimensional Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for analyzing large multidimensional datasets are inefficient and fail to identify important relationships, often breaking relationships and being too sensitive to large scale distances, requiring sophisticated experts and lacking interactive visualization capabilities for exploratory data analysis.
Innovation Solution
A method involving the selection of landmark features, where distances between features are calculated using a metric, and the closest non-selected features are added to the set of landmark features, allowing for interactive visualization and clustering of data points in a mathematical reference space.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If clustering methods are used to process large datasets, then data can be grouped for analysis, but important relationships are broken and the method is too blunt to identify important relationships
Solution Approach 1:
The patent segments the dataset by selecting a subset of landmark features rather than applying clustering to all features. This segmentation approach identifies relationships among selected landmarks first, then uses those relationships to guide analysis of the remaining data, preserving important relationships while enabling scalable processing.
Solution Approach 2:
The patent transforms the high-dimensional feature space into a lower-dimensional landmark space by selecting a subset of representative features. This dimensional reduction allows efficient processing while maintaining relationship information through the landmark-to-full-feature mapping approach.
2Measurement precision
If linear algebraic and analytic methods are used, then data can be analyzed, but the methods are too sensitive to large scale distances and lose detail
Solution Approach 1:
The patent segments the feature space by identifying landmark features that represent different regions or clusters of the data. By analyzing relationships among these segmented landmarks rather than all features uniformly, the method reduces sensitivity to large-scale distances while maintaining precision in identifying important relationships.
Solution Approach 2:
The patent applies local quality by focusing computational analysis on local relationships among landmark features rather than global relationships across all features. This localized approach preserves detail in relationship identification while reducing the harmful effect of large-scale distance sensitivity.
3Loss of information
If previous analysis methods are used, then some relationships can be depicted in graphs, but the graphs are not interactive and require considerable time for experts to understand
Solution Approach 1:
The patent implements self-service by providing an interactive visualization system that allows users to directly explore and modify the analysis parameters, rather than requiring expert interpretation of static graphs. Users can interactively adjust landmark selections and immediately see updated relationship depictions, eliminating the time-consuming expert interpretation step.
Solution Approach 2:
The patent transforms static relationship graphs into dynamic, interactive visualizations where users can modify parameters and immediately observe changes in relationship depictions. This dynamic approach reduces expert interpretation time by allowing users to self-guide the analysis process.
4Adaptability or versatility
If previous analysis methods are used, then data can be analyzed, but the output does not allow for exploratory data analysis where analysis can be quickly modified to discover new relationships
Solution Approach 1:
The patent implements a dynamic analysis framework where the landmark feature set can be interactively modified during the analysis process. Users can add or remove landmarks and immediately see updated relationship depictions, enabling rapid exploratory analysis without sacrificing processing efficiency through the landmark subset approach.
Solution Approach 2:
The patent creates a universal analysis framework that handles both initial analysis and exploratory modifications through the same landmark-based approach. This multi-functional system supports various analysis scenarios (initial processing, parameter adjustment, relationship exploration) without requiring different methodologies, maintaining productivity across all use cases.
Data Source
AI summary
An example method comprises receiving a multidimensional data set, receiving a predetermined number of features for a set of landmark features, when a current number of features of the set is less than the predetermined number: for each landmark feature of the set of landmark features, calculate a distance between that particular landmark feature and each non-selected feature that is not within the set, identify a closest non-selected feature to that particular landmark feature, identify a particular closest non-selected feature related to a largest distance among the distances, and adding the particular non-selected feature to the set of landmark features, and if the current number of features of the set of landmark features is equal to or greater than the predetermined number of features for the set of landmark features, then providing identification of at least a subset of features of the set of landmark features.


