Hyperdimensional Network for Edge Object Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current industry practices for object detection and visual grounding, especially few-shot and zero-shot scenarios, are not well-suited for edge devices due to their reliance on large annotated datasets and high computational requirements.
Innovation Solution
A method utilizing a hyperdimensional network trained on edge devices for object detection and visual grounding, which involves determining regions of interest, generating numerical and hyperdimensional vector representations for images and text, and embedding them into a common hyperdimensional embedding space to preserve similarity, enabling identification and visual grounding of objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current industry practice of training object detectors on large annotated datasets is used, then detection accuracy is improved, but device complexity and computational requirements increase making it unsuitable for edge devices
Solution Approach 1:
The patent extracts only the essential features and characteristics of objects from large annotated datasets, rather than processing entire images. By focusing on key discriminative features and creating compact feature representations, the system achieves accurate detection while significantly reducing computational requirements for edge devices.
Solution Approach 2:
The patent transforms the problem from processing high-dimensional image data to working with optimized feature vectors and probability distributions. By changing the parameter representation from raw pixels to extracted features, and from deterministic labels to probability distributions, the system achieves better efficiency on edge devices while maintaining detection accuracy.
2Measurement precision
If object detectors are trained on specific object categories, then detection precision for those categories is improved, but adaptability to new object categories deteriorates
Solution Approach 1:
The patent creates a universal object detection framework that can handle multiple object categories through a single trained model. By using unsupervised feature extraction and probability distribution matching, the system achieves both precision for trained categories and adaptability to new categories without requiring retraining, making the detector multi-functional across diverse object types.
Solution Approach 2:
The patent performs preliminary unsupervised feature extraction and clustering to identify object characteristics before actual detection. By pre-processing images to extract universal features and organize them into probability distributions, the system prepares the data in advance, enabling both accurate detection of known categories and flexible adaptation to new categories.
3Reliability
If extensive training data for specific object categories is used, then detection reliability is improved, but loss of time for training and data processing increases
Solution Approach 1:
The patent performs unsupervised feature extraction and clustering as a preliminary step that does not require extensive labeled training data. By pre-processing images to extract features and organize them into probability distributions without supervision, the system reduces the time-consuming labeled data collection and training process while maintaining detection reliability through robust feature representations.
Solution Approach 2:
The patent enables the system to automatically extract features and learn object characteristics from unlabeled images through unsupervised clustering. The algorithm self-organizes the feature space and creates probability distributions without human intervention or extensive labeled data, reducing training time and enabling the system to serve itself by learning from available data.
Data Source
AI summary
A method, apparatus, and system for object detection on an edge device include projecting a hyperdimensional vector of a query request for an image received at the edge device into a hyperdimensional embedding space to identify at least one exemplar in the hyperdimensional embedding space having a predetermined measure of similarity to the query request using a network trained to: generate a respective hyperdimensional image vector and a respective hyperdimensional text vector for the image and received text descriptions of the image, generate a hyperdimensional query text vector of the query request, combine and embed respective ones of the hyperdimensional image vectors and the hyperdimensional text vectors into a hyperdimensional embedding space to generate respective exemplars, project the hyperdimensional query text vector into the hyperdimensional embedding space, and determine a similarity measure between the hyperdimensional query text vector and at least one of the respective exemplars.


