Head-Mounted Device Collaboration for Object Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing head-mounted devices face limitations in processing power, accuracy, and efficiency due to individual device constraints such as component cost, size, weight, heat generation, and occlusions, which hinder effective sensory perception and object recognition.
Innovation Solution
Multiple head-mounted devices can operate in concert to share sensory input, distribute processing workload, and enhance data accuracy by leveraging combined sensory input and computing power, as well as external devices, to improve sensory perception and processing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple head-mounted devices share processing workload, then object recognition accuracy and speed improve, but device complexity and communication requirements increase
Solution Approach 1:
The system segments the object recognition task across multiple head-mounted devices, where each device captures images from its own perspective and the server divides the overall processing workload. This segmentation allows each device to operate independently while contributing to the collective accuracy of object recognition through multiple viewpoints.
Solution Approach 2:
A server acts as an intermediary between multiple head-mounted devices, coordinating image collection, distributing processing tasks, and aggregating results. This intermediary manages the complexity of inter-device communication and data fusion, allowing individual devices to remain relatively simple while achieving enhanced collective performance.
2Measurement precision
If multiple head-mounted devices collaborate, then sensory perception accuracy improves, but power consumption increases
Solution Approach 1:
Instead of having all devices continuously capture and process images, the system uses partial action by selectively assigning image capture tasks to specific devices based on their current field of view and the server's processing needs. This reduces overall power consumption while maintaining the accuracy benefits of multiple perspectives.
Solution Approach 2:
Each head-mounted device autonomously determines whether it has useful images to contribute based on its current view and the requested object characteristics, eliminating the need for constant polling or coordination overhead. This self-service approach reduces communication power consumption while maintaining collaborative efficiency.
3Speed
If head-mounted devices process images locally, then response speed improves, but device heat generation and processing limitations worsen
Solution Approach 1:
The processing workload is segmented between local device processing for immediate response-critical tasks and server-based processing for computationally intensive image analysis. This segmentation allows fast local responses while distributing heat-generating processing tasks to the server with better cooling capabilities.
Solution Approach 2:
The server acts as an intermediary that receives images from devices, performs computationally intensive processing, and returns results. This intermediary handles the heat-generating processing tasks remotely, allowing client devices to maintain lower temperatures while still achieving fast overall response times through optimized task distribution.
Data Source
AI summary
A system can include head-mounted devices that collaborate to process views from cameras of the respective head-mounted devices and identify objects from different perspectives and/or objects that are within the view of only one of the head-mounted devices. Sharing sensory input between multiple head-mounted devices can complement and enhance individual units by interpreting and reconstructing objects, surfaces, and/or an external environment with perceptive data from multiple angles and positions, which also reduces occlusions and inaccuracies. As more detailed information is available at a specific moment in time, the speed and accuracy of object recognition, hand and body tracking, surface mapping, and/or digital reconstruction can be improved. Such collaboration can provide more effective and efficient mapping of space, surfaces, objects, gestures and users.


