Utterance Subject Identification via Audio Markers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech processing systems face difficulties in resolving ambiguity and identifying the subject of anaphors, such as pronouns, without prompting the user for additional information, especially when spoken commands do not follow a predetermined format.
Innovation Solution
The use of identifiers or markers is introduced to associate specific elements or portions of audio presentations, allowing speech processing systems to determine the subject of spoken commands during playback, even when multiple audio programs are active.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech processing systems use natural language commands without predetermined formats, then ease of operation is improved, but reliability deteriorates due to ambiguity in identifying utterance subjects
Solution Approach 1:
The patent introduces an intermediary mechanism (context tracking module) that mediates between the user's natural language utterance and the speech processing system. This intermediary maintains context information about previously presented audio content and uses it to disambiguate pronouns and anaphors, allowing users to speak naturally while ensuring reliable subject identification.
2Reliability
If speech processing systems prompt users for additional information to resolve ambiguity, then reliability is improved, but productivity deteriorates due to additional interaction steps
Solution Approach 1:
The system performs preliminary actions by pre-processing and storing context information about audio content before the user issues commands. The context tracking module maintains a record of presented content, speakers, and temporal relationships in advance, so when a user utterance is received, the system can immediately resolve ambiguity using pre-established context without requiring additional user prompts.
3Adaptability or versatility
If speech processing systems track context information from multiple audio programs, then adaptability is improved, but device complexity increases
Solution Approach 1:
The patent segments the context tracking function into distinct modular components: an audio content analyzer that processes individual audio streams, a context database that stores structured information about presented content, and a subject identification module that queries the context database. This segmentation allows the system to handle multiple audio programs independently while maintaining manageable system complexity through clear separation of concerns.
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
Features are disclosed for generating markers for elements or other portions of an audio presentation so that a speech processing system may determine which portion of the audio presentation a user utterance refers to. For example, an utterance may include a pronoun with no explicit antecedent. The marker may be used to associate the utterance with the corresponding content portion for processing. The markers can be provided to a client device with a text-to-speech ("TTS") presentation. The markers may then be provided to a speech processing system along with a user utterance captured by the client device. The speech processing system, which may include automatic speech recognition ("ASR") modules and/or natural language understanding ("NLU") modules, can generate hints based on the marker. The hints can be provided to the ASR and/or NLU modules in order to aid in processing the meaning or intent of a user utterance.