Topological Coding for Visual Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current mobile visual search systems face high computational costs and communication expenses due to the large size of visual descriptors, which hinder real-time operations and compromise matching accuracy when searching large image repositories.
Innovation Solution
A visual search system generates a topologically encoded vector from a point set of an image, using a graph Laplacian matrix and affinity matrix, which is invariant to rotation and scaling, and can be compressed using methods like discrete cosine transform, facilitating efficient image search and identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If visual descriptors (feature points with locations) are used for image search, then matching accuracy is improved, but communication cost and data size increase significantly
Solution Approach 1:
The patent extracts only the essential topological relationships and geometric constraints from complete visual descriptors, representing images through a compact set of topological codes that capture the essential structure without transmitting full feature point data. This extraction reduces data size while preserving matching capability.
Solution Approach 2:
Instead of transmitting actual feature point coordinates and descriptors, the patent creates a simplified topological copy that represents the essential spatial relationships. This topological code serves as a compact representation that can be used for image search without requiring the original detailed descriptor data.
2Measurement precision
If complete images are sent for visual search, then matching accuracy is maintained, but computational cost and communication expense become prohibitively high
Solution Approach 1:
The patent extracts essential topological and geometric information from complete images, creating a simplified representation that retains matching accuracy while enabling real-time processing. The topological code captures the essential structure needed for search operations.
Solution Approach 2:
The patent transforms image data from high-dimensional pixel values to a compact topological parameter representation. This parameter transformation reduces computational complexity while preserving the essential features needed for accurate image matching and search operations.
3Quantity of substance
If feature point size is reduced to lower communication cost, then data transmission efficiency is improved, but searching performance and matching accuracy are compromised
Solution Approach 1:
The patent changes the parameter representation from detailed feature point coordinates to topological codes that encode spatial relationships. This parameter transformation maintains searching performance by preserving essential geometric constraints while significantly reducing feature point size for efficient transmission.
4Manufacturing precision
If high-resolution image data is transmitted, then image quality is maintained, but communication cost and processing time increase
Solution Approach 1:
The patent extracts essential topological and geometric information from high-resolution images, creating a compact representation that maintains image quality for search purposes while dramatically reducing processing time. The topological code preserves the essential structure needed for accurate matching without requiring full high-resolution data processing.
Data Source
Figure 1(a)~1(b)
Figure 2
Figure 3
AI summary
A method and an apparatus for processing images generate a first vector of a first number dimension for the image from a first number of points of the image based on topological information of the first number of points, and the first vector for the image is invariant to rotation and scaling in creating the image. The first number of points can be locations of a set of rotation and scaling invariant feature points of the image, and the generated first vector can be a graph spectrum of a pair-wise distance matrix generated from the first number of points of the image.