Speech System Answer Prediction via Entity Disambiguation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition systems face challenges in disambiguating entities spoken or implied in user requests, particularly in actions involving communication, where multiple contacts with similar names or implied relationships can lead to ambiguity, necessitating follow-up questions to ensure correct action initiation while maintaining security and privacy.

Innovation Solution

A system that combines automatic speech recognition (ASR), natural language understanding (NLU), user-specific entity libraries, and historical data models to predict answers with high confidence, thereby skipping unnecessary questions and ensuring secure action initiation by disambiguating parameters like target contacts, sources, and networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the system asks follow-up questions to disambiguate entities, then the accuracy of action initiation is improved, but the user experience and interaction time deteriorate

Engineering Contradiction:
Improveentity disambiguation accuracyVSAvoidinteraction time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary entity disambiguation using historical data models and entity libraries before requiring user input. By predicting the most likely intended entity based on context and history, the system can skip unnecessary follow-up questions and proceed directly to action initiation when confidence is sufficient.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from historical data and user interaction patterns to refine entity predictions. By continuously learning from past actions and adjusting the historical data model, the system improves its prediction accuracy over time, reducing the need for follow-up questions while maintaining high disambiguation accuracy.

Inventive Principle:
Principle #23Feedback

2Loss of time

If the system skips follow-up questions to improve user experience, then interaction time is reduced, but the reliability of action initiation deteriorates

Engineering Contradiction:
Improveinteraction timeVSAvoidaction initiation reliability
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The system replaces the mechanical approach of asking explicit follow-up questions with an automated prediction system based on historical data models and entity libraries. This substitution allows the system to infer entity identities automatically, reducing interaction time while maintaining reliability through confidence thresholding and prediction accuracy optimization.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If the system uses historical data models to predict entities, then the productivity of action initiation is improved, but the device complexity increases

Engineering Contradiction:
Improveaction initiation speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The historical data model and entity library serve multiple functions: they predict entity identities, determine action parameters, and assess confidence levels. By making this single system component multi-functional, the patent reduces overall system complexity while improving productivity through automated entity resolution across different action types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11798538B1Answer prediction in a speech processing system
Publication Date: 2023.10.24 AMAZON TECH INC
  • US11798538B1 patent drawing
  • US11798538B1 patent drawing
  • US11798538B1 patent drawing

AI summary

This disclosure relates to answer prediction in a speech processing system. The system may disambiguate entities spoken or implied in a request to initiate an action with respect to a target user. To initiate the action, the system may determine one or more parameters; for example, the target (e.g., a contact/recipient), a source (e.g., a caller/requesting user), and a network (voice over internet protocol (VOIP), cellular, video chat, etc.). Due to the privacy implications of initiating actions involving data transfers between parties, the system may apply a high threshold for a confidence associated with each parameter. Rather than ask multiple follow-up questions, which may frustrate the requesting user, the system may attempt to disambiguate or determine a parameter, and skip a question regarding the parameter if it can predict an answer with high confidence. The system can improve the customer experience while maintaining security for actions involving, for example, communications.