Speech Recognition Normalization for Image Display
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in accurately executing device operations due to differences in vocabulary formats and languages, leading to recognition results that match function or content titles but fail to perform the intended action, especially with spontaneous speech recognition engines.
Innovation Solution
An image display apparatus and method that normalize speech recognition results by generating phonetic symbols for comparison with pre-stored command sets, using exceptional phonetic symbol databases to handle linguistic differences and convert recognition results into compatible formats for device operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If spontaneous speech recognition engine is used to recognize various vocabularies, then the recognition coverage is improved, but the format compatibility with device functions deteriorates
Solution Approach 1:
The patent introduces a normalization unit as an intermediary component between the speech recognition engine and the device function execution system. This normalization unit converts diverse speech recognition results into a standardized format that matches device function titles, enabling reliable execution while maintaining broad recognition coverage. The normalization unit acts as a mediator that bridges the format gap between spontaneous speech recognition outputs and structured device commands.
Solution Approach 2:
The patent changes the parameter of speech recognition result format from diverse/unstructured to standardized/structured through the normalization process. By transforming the output format parameter of the speech recognition engine, the system maintains versatility in recognizing various vocabularies while ensuring compatibility with device functions through standardized formatting.
2Speed
If speech recognition result is directly matched with function title, then the execution speed is improved, but the recognition accuracy deteriorates
Solution Approach 1:
The patent applies preliminary action by performing normalization of speech recognition results before the matching and execution process. The normalization unit pre-processes the recognition results to convert them into the correct format, ensuring that subsequent matching operations can proceed efficiently with high accuracy. This preliminary formatting step prevents mismatches that would require costly re-processing.
3Measurement precision
If post-processing technologies using correlations and parallel corpora are applied, then the recognition error rate is improved, but the format normalization problem persists
Solution Approach 1:
The patent extracts the format normalization function as a separate, dedicated component (normalization unit) from the complex post-processing chain involving correlations and parallel corpora. By isolating the normalization task, the system maintains the benefits of advanced post-processing for error reduction while adding a simple, efficient normalization step that doesn't increase overall system complexity significantly.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An image display apparatus, a method for driving the same, and a computer readable recording medium are provided. The image display apparatus includes a speech acquirer configured to acquire a speech command created by a user, a speech recognition executor configured to acquire text information which has a phonetic symbol that is the same as or similar to a phonetic symbol of a text-based recognition result corresponding to the acquired speech command and is expressed in a form that is different from a form of the text-based recognition result, and an operation performer configured to perform an operation corresponding to the acquired text information.