Specifying Unit Combining Gaze and Speech for Editing Position
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for specifying a character editing position using gaze recognition are unreliable due to high requirements for gaze recognition accuracy and instability, leading to frequent changes in the editing position.
Innovation Solution
An information processing device and method that specify a selected spot intended by a user from visual information using both non-verbal and verbal actions, such as gaze and speech, to enhance accuracy and stability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If gaze recognition technique is used to specify editing position, then input operation without keyboard or mouse is enabled, but gaze recognition accuracy must be significantly high and editing position changes frequently due to instability
Solution Approach 1:
The patent combines multiple input modalities (gaze recognition, voice recognition, and body movement recognition) into a unified input system. By merging these different recognition techniques, the system achieves more reliable and stable editing position specification compared to using gaze recognition alone, while maintaining hands-free operation capability.
Solution Approach 2:
The patent introduces voice commands and body movements as intermediary inputs to confirm and stabilize the editing position. These intermediaries help resolve the instability of pure gaze recognition by providing additional verification layers through alternative input channels that are more reliable.
2Device complexity
If only gaze recognition is used to specify selected spot, then system complexity is reduced, but specification accuracy is insufficient
Solution Approach 1:
The patent merges multiple recognition systems (gaze, voice, and body movement recognition) to achieve higher specification accuracy. By combining these modalities, the system overcomes the limitations of gaze-only recognition and achieves more precise selected spot identification.
Solution Approach 2:
The patent segments the input recognition process into multiple independent components (gaze recognition module, voice recognition module, body movement recognition module). Each module handles a specific aspect of input, and their results are integrated to achieve high accuracy while keeping individual module complexity manageable.
3Measurement precision
If multiple input modalities are combined to improve accuracy, then specification precision is improved, but system complexity increases
Solution Approach 1:
The patent implements a universal recognition framework that handles multiple input modalities (gaze, voice, body movements) through a common processing architecture. This multi-functional system can process different types of inputs using unified algorithms, reducing the overall system complexity compared to having separate dedicated systems for each modality.
Solution Approach 2:
The patent employs dynamic input selection where the system can adaptively choose which recognition modalities to use based on the current context and user behavior. This dynamic approach allows the system to achieve high precision when needed while reducing complexity by not always activating all recognition systems simultaneously.
Data Source
AI summary
An information processing device including a specifying unit configured to, based on a speech of a user, specify a selected spot that is intended by the user from visual information that is displayed, wherein the specifying unit is configured to specify the selected spot based on a non-verbal action and a verbal action of the user, is provided. Furthermore, an information processing method including, by a processor, based on a speech of a user, specifying a selected spot that is intended by the user from visual information that is displayed, wherein the specifying includes specifying the selected spot based on a non-verbal action and a verbal action of the user, is provided.


