Two-Stream Audio Indexing for Spoken Web Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing technologies face challenges in effectively indexing and searching audio content on the Spoken Web, particularly due to the spontaneous and noisy nature of voice data in diverse accents and languages, with limited precision recall and failure to incorporate contextual information.
Innovation Solution
The proposed solution involves a two-stream processing method, where a keyword stream identifies specific keywords through metadata analysis and an ontology, and a large vocabulary stream performs standard speech-to-text recognition, with search results from both streams being independently ranked and fused to create a comprehensive index.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If standard speech recognition analysis is used on audio content, then a comprehensive text index can be generated, but the precision and recall remain limited due to noisy voice data in diverse accents and languages
Solution Approach 1:
The patent divides the audio indexing process into two separate streams: a keyword stream that extracts specific keywords from audio content, and a large vocabulary stream that performs comprehensive speech-to-text recognition. Each stream processes the audio data differently and independently, allowing the system to leverage the strengths of both approaches while mitigating their individual weaknesses regarding noisy voice data
Solution Approach 2:
The patent changes the processing parameters by applying different recognition thresholds and methodologies to the same audio content. The keyword stream uses metadata analysis and ontology-based keyword extraction with higher precision thresholds, while the large vocabulary stream uses more permissive recognition to capture broader content, thereby improving overall precision recall despite noise
2Adaptability or versatility
If only standard speech recognition is used, then the system can handle diverse accents and languages, but contextual information is not incorporated leading to lower search accuracy
Solution Approach 1:
The patent separates the processing into two streams where the keyword stream specifically incorporates contextual information through metadata analysis and ontology-based keyword extraction, while the large vocabulary stream handles diverse accents and languages with comprehensive speech-to-text recognition. This segmentation allows each stream to optimize for its specific strength
Solution Approach 2:
The patent introduces metadata and ontology as intermediary elements that bridge the gap between diverse audio inputs and accurate search results. The metadata provides contextual information about the audio content, and the ontology structures keywords semantically, allowing the system to maintain high adaptability to diverse accents while improving search accuracy through contextual grounding
3Device complexity
If a single indexing method is used, then the system complexity is lower, but the precision and comprehensiveness of search results are insufficient
Solution Approach 1:
The patent divides the indexing system into two independent but complementary streams: keyword stream processing and large vocabulary stream processing. Each stream has its own processing pipeline and optimization criteria, allowing the system to achieve high precision results through multiple specialized pathways rather than a single general-purpose method
Solution Approach 2:
The patent merges the results from the keyword stream and large vocabulary stream into a unified search result set. By combining the outputs of both streams and ranking them together, the system achieves comprehensive and precise search results that leverage the strengths of both indexing methods while managing complexity through modular architecture
Data Source
AI summary
Systems and methods provide for indexing audio content by fusing the indexes derived from a keyword stream and a large vocabulary stream search. For example, systems and methods provide for two stream searching of Spoken Web VoiceSites, wherein metadata is extracted from the VoiceSite and is used to determine a set of keywords for high precision search while a traditional standard vocabulary set is used to perform a high results, low precision search. The results of the keyword search and the standard vocabulary search are fused together to form a comprehensive, ranked list of results.


