Multimedia Device Gesture Recognition Using User Position and Viewing Direction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In complex multimedia environments with multiple devices, existing gesture recognition systems face challenges in reliably distinguishing between intended control gestures and non-control gestures, especially when users are not directly facing the screen or when multiple users are present.

Innovation Solution

A method that involves detecting a wake-up event, determining the user's position and viewing direction using audio and video information, and performing filtered gesture recognition based on this information to ensure only relevant gestures from the primary user are recognized and processed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If gesture recognition is used to control multimedia devices in a shared display environment, then user control capability is improved, but control accuracy deteriorates due to inability to distinguish between gestures from different users or unintended gestures

Engineering Contradiction:
Improveuser control capabilityVSAvoidcontrol accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system performs preliminary actions by detecting wake-up events and determining user position and viewing direction before processing gesture recognition. This preliminary identification of the primary user and their orientation establishes a reference framework that filters out gestures from other users or unintended targets, thereby improving control accuracy while maintaining ease of operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies local quality by focusing gesture recognition on specific spatial regions and temporal windows corresponding to the identified primary user's position and viewing direction. This localized approach ensures that only gestures from the relevant user are processed, eliminating false triggers from other users or unintended gestures while preserving the user-friendly gesture control interface.

Inventive Principle:
Principle #3Local quality

2Speed

If gesture recognition processes all detected gestures without filtering, then responsiveness is improved, but reliability deteriorates due to inclusion of non-essential gestures from passive users

Engineering Contradiction:
Improveresponse speedVSAvoidgesture recognition reliability
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The system performs preliminary detection of wake-up events and determination of user position and viewing direction before processing gestures. This preliminary action creates a filter framework that ensures only gestures from the identified primary user are processed, maintaining fast response to legitimate commands while eliminating unreliable gestures from passive users.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary filtering mechanism that mediates between raw gesture detection and final command execution. This intermediary layer uses user position and viewing direction information to filter out non-essential gestures from passive users, ensuring that only reliable gestures from the primary user trigger device commands, thus improving reliability without significantly impacting response speed.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If the system processes gestures from all users present in the environment, then user interaction coverage is improved, but control precision deteriorates due to inability to identify the primary user's intent

Engineering Contradiction:
Improveuser interaction coverageVSAvoidprimary user intent identification
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary detection of wake-up events and determination of user position and viewing direction before processing gestures. This preliminary action identifies the primary user and their orientation, creating a precise filter that ensures only gestures from the identified primary user are processed for device control. This maintains broad user interaction coverage while achieving precise intent identification through spatial and temporal filtering.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies local quality by focusing gesture recognition on specific spatial regions and temporal windows corresponding to the identified primary user's position and viewing direction. This localized approach maintains the ability to interact with multiple users in the environment while achieving precise intent identification by exclusively processing gestures from the currently active primary user.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP2595401A1Multimedia device, multimedia environment and method for controlling a multimedia device in a multimedia environment
Publication Date: 2013.05.22 THOMSON LICENSING SA
  • EP2595401A1 patent drawing
  • EP2595401A1 patent drawing
  • EP2595401A1 patent drawing

AI summary

There is a multimedia device 4, a multimedia environment 2 and a method for controlling the multimedia device 4. The multimedia environment 2 further comprises a sensor 6 for acquisition of audio and/or video information. The multimedia device 4 is configured to perform gesture and or speech recognition based on the acquired audio and/or video information. A wake-up event that is assigned to activation of the multimedia device 4 and is initiated by a user 18 of the multimedia environment 2 is detected based on acquired audio and/or video information. The multimedia device 4 is set to an active state upon the detection of the wake-up event. Further, a position 24 of the user 22 in the multimedia environment 2 is determined. A viewing direction 26 of the user 22 may further be detected. The filtered gesture recognition is performed based on subsequently acquired video information wherein the step of filtering takes into account the determined position 24 and the determined viewing direction 26 of the user 22.