On-Device Speech Recognition Using Content-Based Expected Words
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current electronic devices face limitations in providing accurate subtitles for deaf individuals due to the need for massive databases and high human resource costs, resulting in only a small portion of content being subtitled, and they often rely on servers for speech recognition, which limits accuracy and requires significant resources.
Innovation Solution
An electronic device equipped with a speech recognition module that uses expected words based on content type, broadcast timing, viewing history, or user input to perform on-device speech recognition, providing both text and sign language images without relying on large databases, and allows for error correction and updates, enabling improved accuracy and user assistance in noisy environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a speech recognition module with a massive database is embedded in the TV to increase recognition accuracy, then speech recognition accuracy is improved, but device complexity and storage requirements worsen
Solution Approach 1:
The patent extracts only the essential linguistic rules and common word patterns from the massive database, separating them from the full speech recognition system. This allows the TV to perform basic speech recognition with minimal stored data, while more complex recognition can be performed by external servers or through incremental learning from user corrections.
Solution Approach 2:
The system performs preliminary speech recognition using the limited onboard database, then refines accuracy through subsequent steps including server verification and user feedback. This preliminary action allows the system to function immediately with minimal resources while progressively improving accuracy over time.
2Measurement precision
If subtitles are pre-produced by broadcasting companies to provide accurate subtitles for deaf persons, then subtitle accuracy is improved, but human resource costs and production time worsen
Solution Approach 1:
The system enables self-service subtitle generation through automated speech recognition. The TV automatically transcribes speech from content into subtitles without requiring external human intervention, allowing immediate subtitle provision for any content while maintaining reasonable accuracy through iterative refinement.
Solution Approach 2:
The system incorporates feedback mechanisms where user corrections to automatically generated subtitles are captured and used to improve future recognition accuracy. This feedback loop allows the system to progressively improve subtitle quality without requiring manual pre-production for each piece of content.
3Measurement precision
If the TV transmits speech data to a server for recognition to obtain subtitles, then speech recognition capability is improved, but response time and network dependency worsen
Solution Approach 1:
The speech recognition system is segmented into multiple components: basic recognition performed locally on the TV using the embedded database, intermediate verification performed on external servers, and final refinement through user feedback. This segmentation allows the system to provide immediate subtitles for common speech patterns while using server resources only when needed.
Solution Approach 2:
The system performs partial speech recognition locally handling the most common and straightforward cases, then transfers only the uncertain or complex cases to external servers. This partial action approach maintains fast response times for the majority of content while still achieving high overall accuracy.
Data Source
AI summary
An electronic device for providing content including an image and a voice is disclosed. The electronic device comprises: a display configured to display an image; a memory in which a voice recognition module including various executable instructions is stored; and a processor configured to acquire expected words that will possibly be included in a voice, based on information about content, using the expected words to perform voice recognition for the voice through the voice recognition module, and displaying, on the display, text converted from the voice based on the voice recognition.


