Device-Facing HCI via Visual Recognition for Natural Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current human-computer interaction technologies are single and unnatural, leading to inconvenient user operations, as they often require specific gestures or voice commands, limiting natural interaction and flexibility.

Innovation Solution

A human-computer interaction method and system based on direct view, utilizing image acquisition devices to determine the user-device interaction state through visual recognition technologies like face, gesture, and speech recognition, enabling natural and versatile interaction by recognizing user behavior and intentions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional key-pressing mode, voice word activation mode, or gesture mode is used for human-computer interaction, then the device can recognize user input, but the interaction process becomes unnatural and operation becomes inconvenient due to single mode and pre-set specific gestures

Engineering Contradiction:
Improveuser operation convenienceVSAvoidinteraction mode diversity
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system integrates multiple recognition technologies (face recognition, speech recognition, gesture recognition, lip recognition, voiceprint recognition, expression recognition, card recognition, pupil recognition, and iris recognition) into a single human-computer interaction framework, allowing the device to respond to various user inputs through different modalities without requiring pre-set specific gestures, thereby achieving universal adaptability while maintaining ease of operation

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adapts the interaction mode based on the user's direct view state and real-time behavior recognition, transitioning between different recognition technologies as needed rather than relying on a fixed pre-set gesture system, making the interaction process more natural and convenient

Inventive Principle:
Principle #15Dynamics

2Reliability

If specific gesture actions are pre-set for interaction, then the device can identify user commands, but the interaction process becomes unnatural and limits user flexibility

Engineering Contradiction:
Improvecommand recognition accuracyVSAvoidinteraction flexibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system uses dynamic behavior recognition through multiple technologies (face, speech, gesture, lip, voiceprint, expression, card, pupil, and iris recognition) to identify user commands in real-time based on the user's direct view state, replacing static pre-set gesture requirements with flexible, natural, and reliable command recognition that adapts to various user behaviors

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system enables users to interact naturally without needing to learn or perform specific pre-set gestures, allowing users to communicate with the device through their natural behaviors (speaking, gesturing, facial expressions) which the system automatically recognizes and processes, thereby maintaining reliability while enhancing flexibility

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11163356B2Device-facing human-computer interaction method and system
Publication Date: 2021.11.02 LIU GUOHUA
  • US11163356B2 patent drawing
  • US11163356B2 patent drawing
  • US11163356B2 patent drawing

AI summary

A device-facing human-computer interaction method and system is provided. The method includes acquiring device-facing image data acquired by an image acquisition device when a user is in a device-facing state relative to the device; acquiring current image data of the user and comparing the currently acquired image data with the device-facing image data; if the currently acquired image data is consistent with the device-facing image data, identifying a user behavior and intention by means of a device-facing recognition technique and a voice recognition technique of a computer; and according to a preset correspondence between the user behaviors and intention and an operation, controlling the device to perform an operation corresponding to the current user behavior and intention.