Voice-Controlled Media Selection Using Phonetic Transcription
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controlled multimedia systems face challenges in recognizing speech commands when media file names are in different languages, as the language of the entry is often unknown, leading to low recognition rates and difficulties in processing variable vocabulary.
Innovation Solution
Incorporating phonetic data into file identification metadata, which includes different pronunciations of file names across languages, allowing the speech recognition unit to compare user inputs more accurately and improve selection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech recognition is used to select media files with names in different languages, then voice-controlled operation is enabled, but recognition rate decreases when the language of the entry is unknown
Solution Approach 1:
The system performs preliminary action by storing multiple phonetic transcriptions of media file names in different languages before the speech recognition process. When a user speaks a media file name in their native language, the system can match it against pre-stored transcriptions in various languages, enabling successful recognition even when the user's language differs from the original media file language.
Solution Approach 2:
The system achieves universality by implementing a multi-language phonetic transcription system that can handle speech recognition across multiple languages. The phonetic data structure stores transcriptions in various languages for each media file, allowing the speech recognition unit to universally process user inputs regardless of the user's native language or the original media file language.
2Productivity
If phonetic transcriptions are generated automatically, then processing speed is improved, but recognition accuracy decreases
Solution Approach 1:
The system performs preliminary action by pre-storing multiple phonetic transcriptions of media file names in different languages before the speech recognition process. When a user speaks a media file name in their native language, the system can match it against pre-stored transcriptions in various languages, enabling successful recognition even when the user's language differs from the original media file language.
Solution Approach 2:
The system achieves universality by implementing a multi-language phonetic transcription system that can handle speech recognition across multiple languages. The phonetic data structure stores transcriptions in various languages for each media file, allowing the speech recognition unit to universally process user inputs regardless of the user's native language or the original media file language.
Data Source
AI summary
A voice-controlled data system is providing that has a storage medium for storing media files, the media files having associated file identification data for allowing the identification of the media files, the file identification data including phonetic data having phonetic information corresponding to the file identification data. The phonetic data is supplied to a speech recognition unit that compares the phonetic data to a speech command input into the speech recognition unit. The data system further includes a file selecting unit that selects one of the media files based on the comparison result.


