Universal Language Model for Accurate Speech Captioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices face challenges in accurately recognizing and displaying speech data as captions without relying on specific language models, especially when dealing with diverse domains and content types.
Innovation Solution
An electronic apparatus equipped with a communication interface, memory, and processor that extracts objects and characters from image data, generates a bias keyword list based on identified names and relationships, and converts speech data to text using a language contextual model, allowing for accurate captioning without requiring individual language models for each domain.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If individual language models are created for each domain to improve speech recognition accuracy, then recognition precision improves, but device complexity and development time increase significantly
Solution Approach 1:
The patent creates a single universal language model that can handle multiple domains (news, weather, sports, etc.) by training on diverse datasets. This model serves all domains simultaneously through domain identification and adaptive processing, eliminating the need for separate models for each domain while maintaining high recognition accuracy across all areas.
Solution Approach 2:
The system performs preliminary domain identification and keyword extraction before speech recognition. By pre-processing the input to identify the domain and extract relevant keywords, the system prepares the universal language model in advance to handle domain-specific terminology accurately, improving recognition precision without requiring domain-specific models.
2Device complexity
If a universal language model is used to reduce complexity, then device complexity decreases, but speech recognition accuracy deteriorates due to inability to handle domain-specific terminology
Solution Approach 1:
The system extracts domain-specific keywords and identifies the domain before processing speech recognition. This preliminary action allows the universal language model to be adaptively tuned to domain-specific terminology, maintaining high accuracy without requiring separate models for each domain.
Solution Approach 2:
The patent applies different processing strategies to different parts of the speech input based on domain identification. Domain-specific keywords receive specialized handling with adjusted recognition parameters, while general speech uses standard processing. This localized adaptation maintains high accuracy across diverse domains using a single universal model.
3Productivity
If speech data is processed without domain-specific optimization, then processing speed increases, but caption accuracy decreases due to misrecognition of domain-specific terms
Solution Approach 1:
The system performs rapid domain identification and keyword extraction as preliminary steps before speech recognition. This quick pre-processing enables the universal language model to be efficiently adapted to domain-specific terms, achieving both high processing speed and accurate recognition of domain-specific terminology in the final caption.
Data Source
AI summary
An electronic apparatus and a control method thereof are provided. The electronic apparatus includes a communication interface configured to receive content comprising image data and speech data; a memory configured to store a language contextual model trained with relevance between words; a display; and a processor configured to: extract an object and a character included in the image data, identify an object name of the object and the character, generate a bias keyword list comprising an image-related word that is associated with the image data, based on the identified object name and the identified character, convert the speech data to a text based on the bias keyword list and the language contextual model, and control the display to display the text that is converted from the speech data, as a caption.


