Adaptive Voice Recognition Presentation Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems require multiple voice inputs for the same sentence, increasing user burden as sentence length increases, and fail to adapt effectively to varying noise environments and recognition use cases.
Innovation Solution
An information processing device with a presentation control unit that controls the separation of voice recognition results based on context, including noise environment and recognition use, allowing for adaptive presentation modes such as one-character, word, article/possessive connection, and clause/phrase connection, to facilitate easier modification and improve recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple voice inputs are required for the same sentence to improve recognition accuracy, then recognition reliability is improved, but user burden increases and ease of operation deteriorates
Solution Approach 1:
The patent applies dynamics by making the presentation granularity adaptive rather than fixed. The system dynamically adjusts the separation level of recognition results based on context parameters such as noise environment and recognition use case. In high-noise environments or for critical recognitions, finer granularity (character-level) is used to allow precise correction, while in low-noise environments, coarser granularity (phrase-level) is used to reduce user burden. This dynamic adaptation resolves the contradiction between ensuring recognition accuracy and maintaining ease of operation.
Solution Approach 2:
The patent changes the parameter of presentation granularity based on context conditions. By introducing context parameters (noise level, recognition use) and adjusting the separation granularity accordingly, the system optimizes the balance between recognition reliability and user burden. The control information transmitted to the presentation device includes instructions on the appropriate separation level, enabling parameter-based adaptation to different operational contexts.
2Device complexity
If uniform presentation mode is used for all voice recognition results, then device complexity is reduced, but adaptability to different noise environments and use cases deteriorates
Solution Approach 1:
The patent implements universality by designing a multi-functional presentation system that can handle multiple presentation granularities (character-level, word-level, phrase-level) within a single unified framework. The voice recognition device transmits both the recognition results and control information specifying the appropriate separation level, enabling the presentation device to adapt to different noise environments and use cases without requiring multiple specialized systems. This universal approach maintains relatively simple device complexity while achieving high context adaptability.
3Ease of operation
If fine-grained separation of recognition result is used to allow precise modification, then ease of correction is improved, but information processing complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the voice recognition result into different granularity levels (characters, words, phrases) based on context requirements. The voice recognition device segments the recognition result and transmits it along with control information indicating the appropriate separation level. This allows the presentation device to display and allow modification at the appropriate granularity without requiring complex processing on the presentation side, as the segmentation logic is centralized in the voice recognition device.
Data Source
AI summary
There is provided an information processing device and an information processing method that enable a desired voice recognition result to be easily obtained. The information processing device includes a presentation control unit that controls a separation at a time of presenting a recognition result of voice recognition on the basis of context relating to voice recognition. The present technology can be applied, for example, to an information processing device that independently performs voice recognition, a server that performs voice recognition in response to a call from a client and transmits the recognition result to the client, or the client that requests voice recognition to the server, receives the recognition result from the server, and presents the recognition result.


