Input Support Apparatus for Speech Recognition Form Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies, particularly dictation type engines, face challenges in accurately determining item values for form data slots due to recognition errors, making it difficult to input correct values without complex rule creation and expert knowledge.
Innovation Solution
An input support apparatus that uses a form template with predefined item names and alternatives, along with a determination unit that processes recognition result data to accurately determine item values for alternative and free description type slots, even with errors, without requiring detailed speech recognition parameter settings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If dictation type speech recognition engine is used for versatile recognition of arbitrary utterance, then versatility is improved, but recognition accuracy deteriorates
Solution Approach 1:
The patent segments the speech recognition process into two distinct stages: first, a dictation-type engine performs initial recognition of arbitrary utterance to ensure versatility; second, the system segments and analyzes recognition results by comparing them against predefined form slots and item values. This segmentation allows the system to handle diverse utterance types while improving accuracy through structured validation and error correction in the second stage.
2Measurement precision
If grammar type speech recognition is used for high recognition accuracy, then recognition accuracy is improved, but ease of operation deteriorates due to complex rule creation requirements
Solution Approach 1:
The patent applies preliminary action by pre-defining form templates with slots and valid item values before the speech recognition process. This preliminary structure enables the system to automatically validate and correct recognition results without requiring complex grammar rules during operation. The form template acts as a pre-prepared framework that guides the recognition and validation process, eliminating the need for expert knowledge in rule creation while maintaining high accuracy.
3Ease of operation
If dictation type speech recognition is used to handle flexible user utterance, then ease of operation is improved, but reliability deteriorates due to recognition errors
Solution Approach 1:
The patent implements feedback mechanisms where recognition results are continuously validated against the predefined form template. When discrepancies or errors are detected between the recognized utterance and the expected slot values, the system provides feedback by identifying the mismatch and selecting the most appropriate item value from the predefined options. This feedback loop maintains reliability by ensuring that final form data entries are accurate even when initial recognition contains errors.
4Device complexity
If traditional speech recognition techniques are used without form template validation, then device complexity is reduced, but manufacturing precision deteriorates in determining correct item values
Solution Approach 1:
The patent applies preliminary action by pre-defining form templates with slots and valid item values before the speech recognition process. This preliminary structure enables the system to automatically validate and correct recognition results without requiring complex grammar rules during operation. The form template acts as a pre-prepared framework that guides the recognition and validation process, eliminating the need for expert knowledge in rule creation while maintaining high accuracy.
Data Source
AI summary
An input support apparatus of an embodiment includes a template storage unit configured to store a form template that is a template for form data having one or more slots to which item values are input in correspondence with item names, the form template describing item names of the respective slots and alternatives of an alternative type slot in which an item value is selected from a plurality of alternatives together with respective readings thereof; an acquisition unit configured to acquire recognition result data obtained by speech recognition performed on utterance of a user, the recognition result data containing a transcription and a reading; and a determination unit configured to determine the item values to be input to the slots of the form data based on the reading of the recognition result data and the readings of the item names and the alternatives described in the form template.


