Utterance Subject Identification via Audio Markers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech processing systems face difficulties in resolving ambiguity and identifying the subject of anaphors, such as pronouns, without prompting the user for additional information, especially when spoken commands do not follow a predetermined format.

Innovation Solution

The use of identifiers or markers is introduced to associate specific elements or portions of audio presentations, allowing speech processing systems to determine the subject of spoken commands during playback, even when multiple audio programs are active.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech processing systems use natural language commands without predetermined formats, then ease of operation is improved, but reliability deteriorates due to ambiguity in identifying utterance subjects

Engineering Contradiction:
Improveease of operationVSAvoidreliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces an intermediary mechanism (context tracking module) that mediates between the user's natural language utterance and the speech processing system. This intermediary maintains context information about previously presented audio content and uses it to disambiguate pronouns and anaphors, allowing users to speak naturally while ensuring reliable subject identification.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If speech processing systems prompt users for additional information to resolve ambiguity, then reliability is improved, but productivity deteriorates due to additional interaction steps

Engineering Contradiction:
ImprovereliabilityVSAvoidproductivity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary actions by pre-processing and storing context information about audio content before the user issues commands. The context tracking module maintains a record of presented content, speakers, and temporal relationships in advance, so when a user utterance is received, the system can immediately resolve ambiguity using pre-established context without requiring additional user prompts.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If speech processing systems track context information from multiple audio programs, then adaptability is improved, but device complexity increases

Engineering Contradiction:
ImproveadaptabilityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the context tracking function into distinct modular components: an audio content analyzer that processes individual audio streams, a context database that stores structured information about presented content, and a subject identification module that queries the context database. This segmentation allows the system to handle multiple audio programs independently while maintaining manageable system complexity through clear separation of concerns.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP2936482B1Identification of utterance subjects
Publication Date: 2021.04.07 AMAZON TECH INC
  • EP2936482B1 patent drawingFigure 1
  • EP2936482B1 patent drawingFigure 2A
  • EP2936482B1 patent drawingFigure 2B

AI summary

Features are disclosed for generating markers for elements or other portions of an audio presentation so that a speech processing system may determine which portion of the audio presentation a user utterance refers to. For example, an utterance may include a pronoun with no explicit antecedent. The marker may be used to associate the utterance with the corresponding content portion for processing. The markers can be provided to a client device with a text-to-speech ("TTS") presentation. The markers may then be provided to a speech processing system along with a user utterance captured by the client device. The speech processing system, which may include automatic speech recognition ("ASR") modules and/or natural language understanding ("NLU") modules, can generate hints based on the marker. The hints can be provided to the ASR and/or NLU modules in order to aid in processing the meaning or intent of a user utterance.