Air Gesture Video Wall Control via Depth Sensing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for controlling video wall content in public venues using touch gestures are inefficient and cumbersome, especially when multiple users need to interact with the content simultaneously.
Innovation Solution
Implementing a gesture analysis system that captures and analyzes air gestures using computer vision and machine learning, allowing multiple users to control different areas of a video wall by detecting and tracking gestures in real-time, with the help of Multi-access Edge Computing (MEC) networks and 5G NR connections for low latency and high processing capacity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If touch gestures are used to control video wall content, then users can interact with the device, but multiple users cannot interact simultaneously and the interaction becomes cumbersome
Solution Approach 1:
The patent replaces physical touch gestures with air gesture recognition technology. The system uses depth sensors and computer vision algorithms to detect and interpret hand movements in the air without requiring physical contact with the display surface. This substitution enables multiple users to interact simultaneously with the video wall content through their respective air gestures, resolving the limitation of single-user touch interaction while maintaining ease of operation.
2Adaptability or versatility
If air gesture recognition is implemented, then multiple users can interact simultaneously, but processing complexity and computational requirements increase
Solution Approach 1:
The patent segments the gesture recognition process into distinct functional modules: depth image acquisition from sensors, hand detection algorithms that identify hand regions in the depth images, gesture classification systems that interpret specific gesture patterns, and control signal generation that translates gestures into video wall commands. This modular segmentation distributes computational complexity across separate processing stages, making the system more manageable and efficient while supporting multiple simultaneous users.
Solution Approach 2:
The patent introduces depth images as an intermediary representation between the physical air gestures and the digital control commands. The depth sensors capture three-dimensional hand position data, which is then processed through intermediate algorithms to identify gesture patterns before generating final control signals. This intermediary layer simplifies the overall processing by working with structured depth data rather than raw sensor signals, reducing computational complexity while enabling multi-user interaction.
3Ease of operation
If real-time gesture analysis is performed, then user interaction responsiveness improves, but processing speed requirements and latency increase
Solution Approach 1:
The patent implements preliminary action by continuously capturing depth images and pre-processing sensor data even before a complete gesture is formed. The system maintains a buffer of recent depth frames and pre-identifies hand positions and movement trends, so that when a recognizable gesture pattern emerges, the classification and command generation can occur rapidly with minimal additional processing time. This reduces perceived latency while maintaining real-time responsiveness.
Solution Approach 2:
The patent ensures continuous useful action by maintaining constant depth image acquisition and background processing of gesture data streams. The system does not wait for complete gestures to begin processing but continuously analyzes hand positions and gesture progression in real-time. This continuous processing pipeline eliminates idle time between gestures and maintains responsive interaction while optimizing processing efficiency through sustained computational activity.
Data Source
AI summary
A device may include a memory storing instructions and processor configured to execute the instructions to receive a video image stream of a scene. The device may be further configured to perform person detection on the received video image stream to identify a plurality of persons; perform gesture detection on the received video image stream to identify a gesture; select a person from the plurality of persons to associate with the identified gesture; and use the identified gesture to enable the selected person to control an object or area displayed on a video screen associated with the scene.


