Voice Search Manager with Parallel Text and Voice Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Open voice searching is more challenging than local searches due to a larger speech recognition space and the diversity of users' information needs, leading to less than satisfactory search results, even with improved speech recognition performance.
Innovation Solution
The system integrates a textual index with a voice index to execute simultaneous text and voice matching processes, using speech and metadata to classify users and re-order queries, thereby enhancing search results by predicting user interests and preferences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If open voice searching is implemented to allow users to search any information available over a communication network, then the versatility and coverage of search is improved, but the speech recognition accuracy and search result quality deteriorate due to the larger speech recognition space and diversity of user information needs
Solution Approach 1:
The patent segments the search system into two parallel processes: a text matching process that handles speech recognition and text-based search, and a voice matching process that handles acoustic feature extraction and voiceprint matching. This segmentation allows each process to specialize, improving overall accuracy while maintaining open-domain search capability
Solution Approach 2:
The patent introduces voiceprints as an intermediary element that bridges the user and the search system. Voiceprints capture acoustic characteristics and are used to retrieve relevant search results, acting as a mediator that enhances search accuracy without restricting the open-domain nature of the search
2Device complexity
If a pipeline system architecture is used where speech recognition is performed first followed by conventional web search, then the system simplicity is maintained, but the search result quality deteriorates when speech recognition generates incorrect results
Solution Approach 1:
The patent merges the text matching process and voice matching process into a unified search system that operates in parallel. The text matching process handles conventional speech-to-text search, while the voice matching process simultaneously performs acoustic feature extraction and voiceprint-based search, combining results to improve reliability
Solution Approach 2:
The system uses feedback from the voice matching process to correct and refine results from the text matching process. When speech recognition generates incorrect results, the voiceprint matching provides feedback that helps retrieve more accurate search results by comparing acoustic features against stored voiceprints
3Device complexity
If conventional voice search systems are used that rely solely on speech recognition, then the system complexity is reduced, but the ability to handle diverse user information needs deteriorates
Solution Approach 1:
The patent implements dynamic adaptability by using voiceprints to identify user characteristics and preferences. The system dynamically adjusts search results based on acoustic features that reveal information about the kind of person calling and what they might be interested in, making the system adaptable to diverse user needs without increasing apparent complexity
Data Source
AI summary
Techniques disclosed herein include systems and methods for open-domain voice-enabled searching that is speaker sensitive. Techniques include using speech information, speaker information, and information associated with a spoken query to enhance open voice search results. This includes integrating a textual index with a voice index to support the entire search cycle. Given a voice query, the system can execute two matching processes simultaneously. This can include a text matching process based on the output of speech recognition, as well as a voice matching process based on characteristics of a caller or user voicing a query. Characteristics of the caller can include output of voice feature extraction and metadata about the call. The system clusters callers according to these characteristics. The system can use specific voice and text clusters to modify speech recognition results, as well as modifying search results.


