Specifying Unit Combining Gaze and Speech for Editing Position

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for specifying a character editing position using gaze recognition are unreliable due to high requirements for gaze recognition accuracy and instability, leading to frequent changes in the editing position.

Innovation Solution

An information processing device and method that specify a selected spot intended by a user from visual information using both non-verbal and verbal actions, such as gaze and speech, to enhance accuracy and stability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If gaze recognition technique is used to specify editing position, then input operation without keyboard or mouse is enabled, but gaze recognition accuracy must be significantly high and editing position changes frequently due to instability

Engineering Contradiction:
Improveinput operation convenienceVSAvoidediting position stability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent combines multiple input modalities (gaze recognition, voice recognition, and body movement recognition) into a unified input system. By merging these different recognition techniques, the system achieves more reliable and stable editing position specification compared to using gaze recognition alone, while maintaining hands-free operation capability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces voice commands and body movements as intermediary inputs to confirm and stabilize the editing position. These intermediaries help resolve the instability of pure gaze recognition by providing additional verification layers through alternative input channels that are more reliable.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If only gaze recognition is used to specify selected spot, then system complexity is reduced, but specification accuracy is insufficient

Engineering Contradiction:
Improvesystem complexityVSAvoidselected spot specification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges multiple recognition systems (gaze, voice, and body movement recognition) to achieve higher specification accuracy. By combining these modalities, the system overcomes the limitations of gaze-only recognition and achieves more precise selected spot identification.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent segments the input recognition process into multiple independent components (gaze recognition module, voice recognition module, body movement recognition module). Each module handles a specific aspect of input, and their results are integrated to achieve high accuracy while keeping individual module complexity manageable.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If multiple input modalities are combined to improve accuracy, then specification precision is improved, but system complexity increases

Engineering Contradiction:
Improveselected spot specification precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal recognition framework that handles multiple input modalities (gaze, voice, body movements) through a common processing architecture. This multi-functional system can process different types of inputs using unified algorithms, reducing the overall system complexity compared to having separate dedicated systems for each modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs dynamic input selection where the system can adaptively choose which recognition modalities to use based on the current context and user behavior. This dynamic approach allows the system to achieve high precision when needed while reducing complexity by not always activating all recognition systems simultaneously.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11513768B2Information processing device and information processing method
Publication Date: 2022.11.29 SONY GROUP CORP
  • US11513768B2 patent drawing
  • US11513768B2 patent drawing
  • US11513768B2 patent drawing

AI summary

An information processing device including a specifying unit configured to, based on a speech of a user, specify a selected spot that is intended by the user from visual information that is displayed, wherein the specifying unit is configured to specify the selected spot based on a non-verbal action and a verbal action of the user, is provided. Furthermore, an information processing method including, by a processor, based on a speech of a user, specifying a selected spot that is intended by the user from visual information that is displayed, wherein the specifying includes specifying the selected spot based on a non-verbal action and a verbal action of the user, is provided.