Voice Search Prosody Analysis for Query Disambiguation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice-enabled search systems face challenges in accurately recognizing user queries that are curt, incomplete, or inconsistent, leading to user inconvenience and inefficiency when requiring repetition or follow-up questions to clarify the meaning.
Innovation Solution
The system processes spoken queries through an automatic speech recognizer and prosodic analyzer, generating a word lattice that incorporates prosody information, additional contextual data, and behavioral history to approximate relevant responses, enhancing query understanding and response relevance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the system asks the user to repeat the query, then the accuracy of query recognition is improved, but the user convenience deteriorates
Solution Approach 1:
The system performs preliminary analysis of the user's original query using prosodic features (stress, pitch, duration) before asking for repetition. This allows the system to identify likely intended meanings and formulate targeted clarification questions, rather than requiring complete repetition of the query.
Solution Approach 2:
The patent introduces prosodic analysis as an intermediary between the user's spoken query and the search system. By analyzing stress patterns, pitch contours, and duration of syllables, the system extracts semantic information that helps disambiguate curt or incomplete queries without requiring user repetition.
2Measurement precision
If the system asks follow-up questions to clarify the query, then the query understanding is improved, but the time required to obtain information increases
Solution Approach 1:
The system performs preliminary prosodic analysis of the user's query to pre-identify ambiguities and likely intended meanings. This allows the system to either directly interpret curt queries accurately or ask minimal follow-up questions, rather than engaging in extended clarification dialogues.
Solution Approach 2:
The patent applies partial action by selectively analyzing only the most critical prosodic features (stress patterns on key words, pitch contours) rather than attempting to analyze every aspect of the speech. This provides sufficient information to resolve ambiguities in curt queries without requiring excessive processing time or user interaction.
3Productivity
If the system processes only the literal content of the query, then the processing speed is improved, but the relevance of responses deteriorates
Solution Approach 1:
The system performs preliminary extraction of prosodic features (stress, pitch, duration) from the user's speech signal before processing the literal query content. This parallel preliminary processing allows the system to enrich the query interpretation without significantly increasing overall processing time, as prosodic analysis occurs concurrently with or before text processing.
Solution Approach 2:
The patent adds a prosodic dimension to the traditional text-based query processing. By incorporating stress patterns, pitch contours, and duration information as an additional dimension of analysis, the system achieves more relevant responses to curt or ambiguous queries without sacrificing processing efficiency.
Data Source
AI summary
Disclosed herein are systems, methods, and non-transitory computer-readable storage media for approximating relevant responses to a user query with voice-enabled search. A system practicing the method receives a word lattice generated by an automatic speech recognizer based on a user speech and a prosodic analysis of the user speech, generates a reweighted word lattice based on the word lattice and the prosodic analysis, approximates based on the reweighted word lattice one or more relevant responses to the query, and presents to a user the responses to the query. The prosodic analysis examines metalinguistic information of the user speech and can identify the most salient subject matter of the speech, assess how confident a speaker is in the content of his or her speech, and identify the attitude, mood, emotion, sentiment, etc. of the speaker. Other information not described in the content of the speech can also be used.


