Kernel Density Shape Interpolation for Data Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data classification methods are inadequate as they fail to accurately represent cluster shapes, leading to improper classification of data points based solely on proximity to centroids or support vectors, neglecting the shape of clusters.
Innovation Solution
A method involving unsupervised non-parametric clustering followed by supervised partitional clustering, using kernel density estimation to generate a dense representation of cluster shapes, allowing for shape interpolation and improved classification by considering the overall shape of clusters rather than just proximity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If classification is based solely on proximity to centroids or support vectors, then the classification process is simple and computationally efficient, but the classification accuracy deteriorates because the shape of clusters is not considered
Solution Approach 1:
The patent segments the cluster representation into multiple boundary points that define the cluster shape, rather than using a single centroid. These boundary points are obtained through clustering algorithms and are used to create a more accurate geometric representation of cluster boundaries, improving classification accuracy while maintaining computational feasibility.
Solution Approach 2:
The patent transitions from zero-dimensional centroid representation to one-dimensional boundary curves by connecting boundary points in sequence. This dimensional expansion allows the classification system to consider the actual shape and extent of clusters, not just their central location, thereby improving classification accuracy.
2Measurement precision
If cluster shapes are represented using detailed boundary points and shape interpolation, then classification accuracy is improved, but the computational complexity and data processing requirements increase
Solution Approach 1:
The patent uses a limited number of boundary points to represent cluster shapes rather than attempting to model every detail of the cluster boundary. This partial representation captures the essential shape characteristics needed for accurate classification while avoiding the excessive computational complexity of complete boundary modeling.
Solution Approach 2:
The patent creates simplified geometric copies of cluster shapes using straight line segments connecting boundary points. These linear approximations serve as computationally efficient representations that capture the essential shape characteristics without requiring complex curve fitting or detailed geometric modeling.
3Productivity
If traditional clustering methods are used, then the clustering process is straightforward and computationally efficient, but the representation of cluster shapes is inadequate for accurate classification
Solution Approach 1:
The patent performs preliminary clustering to identify boundary points that define cluster shapes before proceeding to classification. By pre-processing the data to extract shape-defining points and constructing boundary representations, the system preserves cluster shape information that would otherwise be lost in traditional centroid-based approaches.
Solution Approach 2:
The patent creates a composite representation of clusters by combining multiple boundary points into a unified shape model. This composite structure integrates information from multiple data points to form a coherent geometric representation that captures the overall cluster shape while maintaining the efficiency benefits of structured data organization.
Data Source
AI summary
A method for representing a dataset comprises clustering the dataset using an unsupervised, non-parametric clustering method to generate a set of clusters each comprising a set of data points in an image; clustering the data points of each cluster using a supervised, partitional clustering method to partition each cluster into a specified number of sub-clusters; generating a density estimate value of each grid point of a set of grid points sampled from the image at a specified resolution for each sub-cluster using a kernel density function; identifying a maximum density estimate value and a sub-cluster associated with the maximum density estimate value for the grid point; adding each grid point for which the maximum density estimate value exceeds a specified threshold to the sub-cluster associated with the maximum density estimate value; and, for each cluster, merging the sub-clusters of the cluster into a corresponding cluster region in the image.


