Voice and Motion Text Input for HMD Immersion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text input is challenging when using head-mounted displays, as traditional keyboards obstruct the field of vision and virtual keyboards are difficult to handle, prone to erroneous recognition, and can disrupt the immersive experience.

Innovation Solution

An image processing apparatus with voice recognition, motion recognition, and text object control sections that displays text as an object in a three-dimensional virtual space, allowing users to input text by voice and interact with it using gestures, enabling efficient and intuitive text input without physical keyboards.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a virtual keyboard is displayed on screen, then text input becomes possible without mechanical input apparatus, but the virtual keyboard is difficult to handle and prone to erroneous recognition

Engineering Contradiction:
Improvetext input capabilityVSAvoidhandling difficulty
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent replaces mechanical keyboards with voice recognition technology. The voice recognition section captures spoken input and converts it to text, eliminating the need for physical or virtual keyboard interaction. This substitution resolves the handling difficulty by using natural speech instead of precise finger movements on a virtual interface.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces motion recognition as an intermediary between voice input and text output. The motion recognition section detects hand gestures or body movements that modify or correct the voice-recognized text. This intermediary layer provides intuitive control without requiring direct manipulation of virtual keyboard elements, reducing erroneous recognition through contextual gesture-based correction.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If a virtual keyboard is displayed on screen, then text input becomes possible, but the mechanical appearance of the virtual keyboard may spoil an originally-presented world view of content

Engineering Contradiction:
Improvetext input capabilityVSAvoidimmersive experience disruption
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts the text input function from the visual display domain. Instead of showing a virtual keyboard on the screen that would disrupt the immersive world view, the system processes voice and motion inputs behind the scenes. The text appears directly in the content without any keyboard interface visible, maintaining the integrity of the presented world view while enabling text input capability.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If voice recognition is used for text input, then text input efficiency is improved, but accuracy of text recognition may be insufficient

Engineering Contradiction:
Improvetext input efficiencyVSAvoidvoice recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements a feedback loop where motion recognition provides corrective input to voice recognition results. The motion recognition section detects gestures that indicate correction intentions, and this feedback modifies the voice-recognized text accordingly. This multi-stage verification process improves recognition accuracy while maintaining the efficiency of voice-based input.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent merges voice recognition and motion recognition systems into a unified text input process. The voice recognition section handles primary text generation while the motion recognition section handles correction and modification. By combining these two recognition methods, the system achieves both high efficiency (from voice input) and high accuracy (from motion-based verification).

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11308694B2Image processing apparatus and image processing method
Publication Date: 2022.04.19 SONY INTERACTIVE ENTERTAINMENT LLC
  • US11308694B2 patent drawing
  • US11308694B2 patent drawing
  • US11308694B2 patent drawing

AI summary

There is provided an image processing apparatus which includes a voice recognition section that recognizes a voice uttered by a user, a motion recognition section that recognizes a motion of the user, a text object control section that disposes an object of text representative of the contents of the voice in a three-dimensional virtual space, and varies text by implementing interaction based on the motion, and an image generation section that displays an image with the three-dimensional virtual space projected thereon.