Adaptive Voice Recognition Presentation Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition systems require multiple voice inputs for the same sentence, increasing user burden as sentence length increases, and fail to adapt effectively to varying noise environments and recognition use cases.

Innovation Solution

An information processing device with a presentation control unit that controls the separation of voice recognition results based on context, including noise environment and recognition use, allowing for adaptive presentation modes such as one-character, word, article/possessive connection, and clause/phrase connection, to facilitate easier modification and improve recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple voice inputs are required for the same sentence to improve recognition accuracy, then recognition reliability is improved, but user burden increases and ease of operation deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoiduser burden
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent applies dynamics by making the presentation granularity adaptive rather than fixed. The system dynamically adjusts the separation level of recognition results based on context parameters such as noise environment and recognition use case. In high-noise environments or for critical recognitions, finer granularity (character-level) is used to allow precise correction, while in low-noise environments, coarser granularity (phrase-level) is used to reduce user burden. This dynamic adaptation resolves the contradiction between ensuring recognition accuracy and maintaining ease of operation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of presentation granularity based on context conditions. By introducing context parameters (noise level, recognition use) and adjusting the separation granularity accordingly, the system optimizes the balance between recognition reliability and user burden. The control information transmitted to the presentation device includes instructions on the appropriate separation level, enabling parameter-based adaptation to different operational contexts.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If uniform presentation mode is used for all voice recognition results, then device complexity is reduced, but adaptability to different noise environments and use cases deteriorates

Engineering Contradiction:
Improvepresentation mode uniformityVSAvoidcontext adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements universality by designing a multi-functional presentation system that can handle multiple presentation granularities (character-level, word-level, phrase-level) within a single unified framework. The voice recognition device transmits both the recognition results and control information specifying the appropriate separation level, enabling the presentation device to adapt to different noise environments and use cases without requiring multiple specialized systems. This universal approach maintains relatively simple device complexity while achieving high context adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If fine-grained separation of recognition result is used to allow precise modification, then ease of correction is improved, but information processing complexity increases

Engineering Contradiction:
Improveease of correctionVSAvoidinformation processing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the voice recognition result into different granularity levels (characters, words, phrases) based on context requirements. The voice recognition device segments the recognition result and transmits it along with control information indicating the appropriate separation level. This allows the presentation device to display and allow modification at the appropriate granularity without requiring complex processing on the presentation side, as the segmentation logic is centralized in the voice recognition device.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10950240B2Information processing device and information processing method
Publication Date: 2021.03.16 SONY GROUP CORP
  • US10950240B2 patent drawing
  • US10950240B2 patent drawing
  • US10950240B2 patent drawing

AI summary

There is provided an information processing device and an information processing method that enable a desired voice recognition result to be easily obtained. The information processing device includes a presentation control unit that controls a separation at a time of presenting a recognition result of voice recognition on the basis of context relating to voice recognition. The present technology can be applied, for example, to an information processing device that independently performs voice recognition, a server that performs voice recognition in response to a call from a client and transmits the recognition result to the client, or the client that requests voice recognition to the server, receives the recognition result from the server, and presents the recognition result.