Speech Input Support Device Offline Recognition Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In maintenance and inspection settings, speech input devices face challenges in improving speech recognition accuracy, especially when not connected to a network, due to limited terminal specifications, and difficulties in correctly associating speech input with corresponding items in entry fields without explicit information.
Innovation Solution
A speech input support device that employs a dual speech recognition approach, using a first speech recognition engine for initial input and a second, cloud-based engine for comparison, to generate and compare recording contents, with guidance to ensure accurate item association and error detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition is performed on a terminal with limited specifications (offline), then the device can be used at maintenance and inspection sites without network connection, but the accuracy of speech recognition deteriorates
Solution Approach 1:
The system performs preliminary speech recognition using a first speech recognition engine on the terminal before network connection is available. This preliminary recognition allows the device to function offline while maintaining the ability to improve accuracy later through comparison with a second speech recognition engine when connected to the network.
Solution Approach 2:
The system compares the recognition result from the first speech recognition engine with the recognition result from the second speech recognition engine (when available). This feedback mechanism allows the terminal to identify and correct recognition errors, thereby improving speech recognition accuracy while maintaining offline operational capability.
2Measurement precision
If a dual speech recognition approach is used to improve accuracy, then speech recognition accuracy improves, but the device complexity increases
Solution Approach 1:
The speech recognition system is segmented into two distinct engines: a first speech recognition engine that operates on the terminal with limited specifications, and a second speech recognition engine that operates on a more powerful server or cloud system. This segmentation allows each engine to be optimized for its specific environment while working together to improve overall accuracy.
Solution Approach 2:
The system introduces an intermediary comparison mechanism that takes the recognition results from both speech recognition engines and identifies differences. This intermediary process resolves the complexity by providing a clear, systematic method for improving accuracy without requiring complete redesign of the speech recognition architecture.
3Ease of operation
If speech input does not include explicit item information, then the input process remains flexible and natural, but the ability to correctly associate speech input with corresponding items deteriorates
Solution Approach 1:
The system performs preliminary processing to determine which item corresponds to the speech input before performing speech recognition. By establishing the item context in advance, the system can naturally process speech input without requiring explicit item information while still maintaining accurate item association through the comparison mechanism.
Data Source
AI summary
According to one embodiment, a speech input support device includes a recording unit and a processor. The recording unit records speech of a user using a speech input device. The processor includes hardware. The processor recognizes the recorded speech separately from speech recognition for input of a first recording content by the speech input device. The processor generates a second recording content based on a result of the separately recognized speech and a next operation for the user for the input using the speech input device. The processor compares the first recording content with the second recording content.


