User-Specific Grammar for Speech Recognition Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Managing large user media libraries with tens of thousands of items and providing a rewarding user experience through easy media item selection is challenging due to the complexity of existing systems, which often require trade-offs between accuracy and intuitive speech recognition.
Innovation Solution
A system that uses a speech recognizer trained with a user-specific grammar library to receive digital representations of spoken commands, providing confidence ratings for media items, and automatically plays back the item with the highest rating, while allowing for intuitive and flexible speech utterances and disambiguation when necessary.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a speech recognizer uses a general grammar for media item selection, then it can handle a wider variety of speech inputs, but accuracy decreases when multiple media items have similar names
Solution Approach 1:
The patent applies local quality by creating a user-specific grammar library that is customized for each individual user based on their media library contents. Instead of using a uniform general grammar for all users, the system generates personalized grammar rules that reflect the specific media items, artists, and albums in each user's collection. This localized approach ensures high accuracy for each user's specific context while maintaining overall system versatility.
Solution Approach 2:
The system performs preliminary action by generating and storing a user-specific grammar library in advance, before speech recognition is needed. The grammar library is pre-populated with media item names, artists, and albums from the user's media library, allowing the speech recognizer to efficiently and accurately process speech inputs without having to analyze the entire media catalog in real-time.
2Measurement precision
If the system requires multiple steps for media item selection, then accuracy can be improved through confirmation, but user experience deteriorates due to complexity
Solution Approach 1:
The patent replaces the mechanical multi-step confirmation system with an acoustic-based single-step speech recognition approach. Instead of requiring users to navigate through menus, click through interfaces, or confirm selections via multiple interactions, the system uses a customized speech recognizer with user-specific grammar to accurately identify media items from a single speech input, thereby maintaining accuracy while dramatically simplifying user interaction.
3Adaptability or versatility
If the speech recognizer is trained with a large grammar library covering all media items, then coverage is improved, but processing time and computational resources increase
Solution Approach 1:
The patent extracts only the relevant portion of the media library needed for speech recognition by creating a user-specific grammar library that contains only the media items, artists, and albums present in the user's personal collection. This extracted subset provides complete coverage for the user's needs while being significantly smaller and more efficient than a comprehensive library covering all possible media items, thereby reducing processing time and computational resources.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A storage machine holds instructions executable by a logic machine to receive a digital representation of a spoken command. The digital representation is provided to a speech recognizer trained with a user-specific grammar library. The logic machine then receives from the speech recognizer a confidence rating for each of a plurality of different media items. The confidence rating indicates the likelihood that the media item is named in the spoken command. The logic machine then automatically plays back the media item with a greatest confidence rating.