Incremental Agglomerative Clustering for Digital Image Fingerprints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data clustering techniques, such as Hierarchical Agglomerative Clustering (HAC), are time- and computationally expensive when dealing with high-dimensional data sets, particularly in digital image processing, as they require re-clustering all existing data with new data points, making them impractical for large datasets with frequent additions.
Innovation Solution
The implementation of incremental agglomerative clustering methods that sample existing data points to cluster new data points efficiently, mapping new clusters to existing clusters based on similarity, reducing the need to re-cluster all existing data, thereby improving computational efficiency and speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional hierarchical agglomerative clustering is used to cluster new data points with existing data, then clustering accuracy is maintained, but computational time and processing cost increase significantly
Solution Approach 1:
The patent segments the clustering process into two distinct phases: (1) clustering new data points only with sampled existing data points to form preliminary clusters, and (2) mapping these preliminary clusters to existing clusters in the hierarchy. This segmentation avoids the computationally expensive operation of clustering new points with all existing points, thereby reducing computational time while maintaining clustering accuracy through the mapping phase.
Solution Approach 2:
The patent applies partial action by sampling a subset of existing data points rather than using all existing data points for clustering operations. This partial approach reduces the computational burden significantly while the subsequent mapping step ensures that the clustering results remain accurate by referencing the full existing cluster structure.
2Reliability
If all existing data is re-clustered with new data points, then up-to-date clustering results are achieved, but computational resources are excessively consumed
Solution Approach 1:
The patent divides the clustering operation into separate stages: new data points are clustered only with sampled existing points, and then the resulting clusters are mapped to the existing hierarchical structure. This segmentation eliminates the need to re-cluster all existing data, thereby reducing computational resource consumption while maintaining reliable clustering results through the mapping mechanism.
Solution Approach 2:
The patent performs preliminary clustering of new data points with a sampled subset of existing points before integrating them into the full hierarchical structure. This preliminary action reduces the immediate computational burden, and the subsequent mapping step ensures the results are reliably integrated into the complete dataset structure without requiring full re-clustering.
3Productivity
If incremental clustering with sampling is used, then computational efficiency is improved, but the complexity of cluster mapping increases
Solution Approach 1:
The patent introduces cluster centroids as intermediaries to facilitate the mapping process. Instead of directly comparing all new clusters with all existing clusters, the centroids serve as representative mediators that simplify the comparison and mapping operation. This intermediary approach manages the algorithmic complexity while maintaining clustering speed by reducing the dimensionality of the mapping problem.
Solution Approach 2:
The patent transforms the complex cluster-to-cluster mapping problem into a simpler centroid-to-centroid comparison by changing the parameter of comparison from full cluster data to condensed centroid representations. This parameter change significantly reduces the computational complexity of the mapping step while preserving the essential clustering information needed for accurate integration.
Data Source
AI summary
Techniques are disclosed for incremental agglomerative clustering of data, including but not limited to digital image data. Fewer than all of a plurality of existing digital image fingerprints are sampled from a first hierarchical data cluster of digital image fingerprints stored in a data storage device, the first hierarchical data cluster excluding a new digital image fingerprint. The new digital image fingerprint and the existing digital image fingerprints sampled from the first hierarchical data cluster are clustered to produce a second hierarchical data cluster of digital image fingerprints, the second hierarchical data cluster including the new digital image fingerprint. If a majority of the existing digital image fingerprints in the first hierarchical data cluster match the new digital image fingerprint, then the second hierarchical data cluster is mapped to the first hierarchical data cluster based on the determination.


