Multi-Language Voice Recognition via AI Server Mediator
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems in digital TVs are limited to performing intent analysis only in a preset language, failing to accurately analyze voices that include multiple languages.
Innovation Solution
A display device and an artificial intelligence server collaborate to receive and convert voice commands into text data, determine the languages present, and provide intent analysis in the appropriate language for voice recognition services, allowing for multi-language support.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If intent analysis is performed only through a preset language in the TV, then the device complexity is reduced and processing is simplified, but the adaptability to handle multi-language voice inputs deteriorates
Solution Approach 1:
The patent introduces an artificial intelligence server as an intermediary component that handles the complex task of multi-language voice recognition and intent analysis. The display device communicates with this external server, which possesses advanced language processing capabilities beyond what the preset TV system can provide. This mediator approach allows the display device to gain enhanced adaptability without significantly increasing its own complexity.
Solution Approach 2:
The artificial intelligence server provides universal language processing capabilities that can handle multiple languages and dialects. By using this external service, the display device gains the ability to process various languages through a single integrated system, rather than requiring separate processing mechanisms for each language, thus improving versatility while maintaining relatively simple device architecture.
2Measurement precision
If the voice recognition system uses only the preset language setting, then the ease of operation is maintained, but the measurement precision of language identification deteriorates
Solution Approach 1:
The system automatically detects and identifies the language being spoken by the user without requiring manual language selection or configuration. The artificial intelligence server analyzes the voice input and determines the appropriate language, then performs intent analysis in that language. This self-service approach improves language identification accuracy while maintaining ease of operation, as users simply speak naturally without needing to adjust any settings.
3Adaptability or versatility
If multi-language processing is implemented in the display device itself, then the adaptability improves, but the device complexity and processing requirements increase significantly
Solution Approach 1:
The patent employs an external artificial intelligence server as a mediator to handle the computationally intensive multi-language processing tasks. The display device maintains a relatively simple architecture by delegating complex language analysis to this external service. The server receives voice data, performs comprehensive multi-language recognition and intent analysis, then returns results to the display device, thus achieving high adaptability without significantly increasing device complexity.
Solution Approach 2:
The solution moves the complex language processing functionality from the display device dimension to an external server dimension. By distributing the processing workload across different system layers (local device and remote server), the patent achieves multi-language support capability while keeping the display device itself relatively simple. This dimensional separation allows advanced features to be accessed without burdening the main device architecture.
Data Source
AI summary
A display device according to an embodiment of the present disclosure includes an output unit, a communication unit configured to perform communication with an artificial intelligence server, and a control unit configured to receive a voice command, convert the received voice command into text data, determine whether the converted text data is composed of a plurality of languages, when the text data is composed of the plurality of languages, determine a language for a voice recognition service among the plurality of languages based on the text data, and output an intent analysis result of the voice command in the determined language.


