Multi-tier Speech Routing via Contextual Data Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech processing systems face challenges in accurately determining the appropriate actions to perform and applications to use for ambiguous spoken utterances due to limited access to relevant contextual data, leading to inefficiencies in user interaction and inability to integrate new actions or data without modifying underlying natural language understanding (NLU) processing.

Innovation Solution

A system that employs contextual data management and intra-domain routing to dynamically route utterances to appropriate domains and applications by using confidence providers and contextual data, including location, usage patterns, and displayed content, to enhance routing decisions and support integration of new commands and queries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech processing systems use limited contextual data for routing decisions, then the system complexity remains low, but the accuracy of determining appropriate actions and applications deteriorates

Engineering Contradiction:
Improveaccuracy of routing decisionsVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech processing system into multiple independent domains, each handling specific types of utterances and contextual data. This allows the system to manage complexity through modular domain-specific processors while improving routing accuracy by directing utterances to the most appropriate domain based on contextual analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary routing layer that sits between the utterance input and domain-specific processors. This intermediary analyzes contextual data and makes routing decisions, separating the complexity of contextual analysis from the domain processing logic, thereby improving accuracy without proportionally increasing overall system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the system integrates new actions or contextual data sources, then the versatility of the speech processing system improves, but the complexity of underlying NLU processing increases

Engineering Contradiction:
Improveability to integrate new actionsVSAvoidcomplexity of NLU processing
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent divides the speech processing system into separate domains, where each domain handles specific actions and contextual data types. This segmentation allows new actions to be integrated by adding new domains or extending existing ones without modifying the core NLU processing, thereby improving versatility while maintaining manageable complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal domain framework where each domain can handle multiple types of utterances and contextual data through standardized interfaces. This multi-functionality allows the system to accommodate new actions and data sources through configuration rather than structural changes, enhancing adaptability without increasing NLU processing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If the system processes ambiguous utterances without sufficient contextual information, then the processing speed remains high, but the success rate of interaction goals deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidsuccess rate of interaction goals
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary contextual analysis and routing decisions before utterance processing begins. By pre-fetching and analyzing relevant contextual data and determining the appropriate domain in advance, the system minimizes processing delays while ensuring that ambiguous utterances are handled by the most capable domain, thereby maintaining high speed with improved success rates.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where routing decisions and processing results are continuously refined based on contextual data and interaction outcomes. This feedback loop allows the system to learn from ambiguous utterance handling and improve future routing decisions, enhancing the success rate of interaction goals while maintaining efficient processing through optimized decision-making.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11657807B2Multi-tier speech processing and content operations
Publication Date: 2023.05.23 AMAZON TECH INC
  • US11657807B2 patent drawing
  • US11657807B2 patent drawing
  • US11657807B2 patent drawing

AI summary

A multi-tier architecture is provided for processing user voice queries and making routing decisions for generating responses, including responses to book browsing requests and other content requests. When an utterance is associated with multiple applications in a given domain, the applications may be organized into a subdomain and a tier of routing decisions may be added to the inter-domain and intra-domain routing decision system. The system uses contextual signals to make subdomain routing decisions, including signals regarding content items that are already in a user's content catalog, consumption status of individual content items in the user's catalog, and the like.