Interactive Apparatus for Gaze-Directed Physical Object Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current interactive apparatuses for virtual, augmented, or mixed reality lack the ability to interact with physical objects in the real world.
Innovation Solution
An interactive apparatus equipped with a first image capture apparatus for capturing a user's face image, a second image capture apparatus for capturing a scene image, and a processor to determine gaze status and identify scene objects using a recognition model, enabling control requests to be transmitted to these objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If interactive apparatuses are used for virtual reality, augmented reality, or mixed reality, then users can interact with virtual objects, but the ability to interact with physical objects in the real world is lacking
Solution Approach 1:
The patent merges virtual reality interaction capabilities with augmented reality features by integrating a recognition model that identifies both virtual and physical objects. The system combines image capture from multiple cameras, gaze status detection through face image analysis, and scene object recognition to enable unified interaction with both virtual and real-world objects through a single interactive apparatus.
Solution Approach 2:
The interactive apparatus is designed with multi-functionality to handle diverse interaction scenarios. It can capture face images for gaze detection, capture scene images for object identification, process both virtual and physical objects through the recognition model, and transmit control requests to various types of objects, making it universally applicable across different interaction contexts.
2Ease of operation
If gaze status detection and scene object recognition are implemented, then users can control physical objects, but the system complexity increases
Solution Approach 1:
The patent segments the interaction system into distinct functional modules: a first image capture apparatus for face images, a second image capture apparatus for scene images, a recognition model for object identification, and a control module for transmitting requests. This segmentation allows each component to specialize in specific tasks, improving ease of operation while managing complexity through modular design.
Solution Approach 2:
The recognition model serves as an intermediary between the image capture apparatuses and the control system. It processes scene images to identify physical objects and determines which objects should receive control requests based on gaze status, thereby simplifying the overall system architecture by introducing a dedicated intermediate processing layer.
3Measurement precision
If multiple image capture apparatuses and recognition models are used, then scene object identification accuracy improves, but energy consumption increases
Solution Approach 1:
The system applies partial action by selectively activating the recognition model only when gaze status is detected, rather than continuously processing all scene images. The first image capture apparatus captures face images for gaze detection, and only when gaze is confirmed does the system proceed to identify scene objects using the second image capture apparatus and recognition model, thereby reducing unnecessary energy consumption while maintaining high recognition accuracy.
Solution Approach 2:
The patent implements periodic action through continuous gaze monitoring using the first image capture apparatus. The system periodically checks face images to determine gaze status, and only triggers the more energy-intensive scene object recognition process when gaze is detected, creating a rhythmic pattern of low-power and high-power operation that balances accuracy with energy efficiency.
Data Source
AI summary
An interactive apparatus includes a first image capture apparatus, a second image capture apparatus, a communication interface, and a processor. The first image capture apparatus is configured to capture a face image of a user. The second image capture apparatus is configured to capture a scene image in a scene. The processor is configured to determine whether the user is in a gaze status based on the face image. In response to the user being in the gaze status, the processor is configured to identify a scene object in the scene image by using a recognition model. In response to the processor identifying the scene object in the scene image, the communication interface is configured to transmit a control request to the scene object to control the scene object.


