Display Voice Control for Multilingual Text Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing display apparatuses face limitations in voice recognition control when the system language differs from the language used in hyperlink text, preventing accurate selection of hyperlinks or execution of commands.
Innovation Solution
A display apparatus and method that allows voice recognition across multiple languages by displaying text objects in a language different from the preset language with accompanying symbols or numbers, enabling user selection through voice commands, and utilizing a server for multi-language voice recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a voice recognition engine is determined in advance based on the system language, then the voice recognition system is simple to implement, but it cannot recognize voices in languages different from the system language
Solution Approach 1:
The patent implements dynamic language switching by detecting the language of displayed content and automatically selecting the appropriate voice recognition engine. The system transitions from a static, pre-determined engine selection to a dynamic adaptation mechanism that changes the voice recognition language based on the current display content language, thereby achieving multi-language support without requiring manual configuration.
Solution Approach 2:
The system changes the language parameter of the voice recognition engine based on the detected language of the displayed content. By monitoring language parameters of on-screen text and adjusting the voice recognition engine's language parameter accordingly, the system achieves adaptability to different languages while maintaining a relatively simple architecture.
2Measurement precision
If the system language is used for voice recognition, then the voice recognition engine works correctly for the system language, but it fails to recognize voices corresponding to hyperlink text in different languages
Solution Approach 1:
The patent implements a feedback mechanism where the system continuously monitors the language of displayed content and uses this information to adjust the voice recognition engine's language setting. This closed-loop feedback ensures that the voice recognition accuracy is maintained for the current display language by automatically aligning the recognition engine's language with the language of the content being displayed.
Solution Approach 2:
The system performs preliminary detection of the display content language and proactively switches the voice recognition engine to the appropriate language before voice input is received. This preliminary action ensures that when the user speaks, the voice recognition engine is already configured to recognize the correct language, thereby maintaining high accuracy without requiring real-time language switching during voice input.
3Measurement precision
If multiple voice recognition engines for different languages are supported, then voice recognition accuracy improves for various languages, but the system complexity increases
Solution Approach 1:
The patent implements a universal voice recognition system that can handle multiple languages through a single integrated architecture. By creating a language-adaptive voice recognition engine that automatically selects and switches between different language models based on detected content language, the system achieves multi-language support without requiring separate, independent voice recognition systems for each language, thereby reducing overall system complexity.
Data Source
AI summary
A display apparatus is provided. The display apparatus according to an embodiment includes a display, and a processor configured to control the display to display a UI screen including a plurality of text objects, control the display to display a text object in a different language from a preset language among the plurality of text objects, along with a preset number, and in response to a recognition result of a voice uttered by a user including the displayed number, perform an operation relating to a text object corresponding to the displayed number.


