Operation Terminal Voice Input Gesture Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-operated terminals require cumbersome gestures or constant voice collection, leading to user discomfort and privacy concerns, as they often necessitate specific arm directions or phrases for initiating voice recognition.

Innovation Solution

An operation terminal that uses an imaging part to detect a user, a human detecting part to identify the user, a voice inputting part to receive spoken voice, and a condition determining part to bring the terminal into a voice inputting state based on the positional relationship between upper limb and upper body coordinates, allowing for simple gestures like raising the arm without considering arm direction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If constant voice collection is implemented, then voice inputting responsiveness is improved, but user privacy security deteriorates

Engineering Contradiction:
Improvevoice inputting responsivenessVSAvoiduser privacy security
Core Design Contradiction:
SpeedVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary detection of user presence and gesture recognition before activating voice collection mode. The imaging part detects user approach and specific gestures in advance, triggering the voice inputting part only when needed, thus maintaining responsiveness while protecting privacy during non-usage periods.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The imaging part serves as an intermediary between the user and the voice collection system. It detects user presence and gestures, acting as a gatekeeper that controls when the voice inputting part should be activated, thereby balancing responsiveness with privacy protection.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If specific gesture directions are required for voice inputting activation, then operation precision is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvegesture detection accuracyVSAvoidgesture performance simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system focuses detection resources on specific local regions where gestures are most likely to occur. By analyzing the imaging data in targeted areas around the terminal, the system achieves high detection accuracy for simple gestures without requiring users to perform complex directional movements.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system analyzes more imaging data than strictly necessary for simple gesture detection. By processing comprehensive spatial information from the imaging part, the system can accurately recognize simple gestures while maintaining robustness against false detections, thus achieving both precision and ease of operation.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11195525B2Operation terminal, voice inputting method, and computer-readable recording medium
Publication Date: 2021.12.07 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US11195525B2 patent drawing
  • US11195525B2 patent drawing
  • US11195525B2 patent drawing

AI summary

An operation terminal includes: an imaging part configured to image a space; a human detecting part configured to detect a user based on information on the space imaged; a voice inputting part configured to receive inputting of the spoken voice of the user; a coordinates detecting part configured to detect a first coordinate of a predetermined first part of an upper limb of the user and a second coordinate of a predetermined second part of an upper half body excluding the upper limb of the user based on information acquired by a predetermined unit when the user is detected by the human detecting part; and a condition determining part configured to compare a positional relationship between the first coordinate and the second coordinate, and configured to bring the voice inputting part into a voice inputting receivable state when the positional relationship satisfies a predetermined first condition at least one time.