Contextual Voice Output System Using NLP and Session State

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech processing systems lack the ability to dynamically determine when and what information to output to users, often providing irrelevant information, leading to a suboptimal user experience.

Innovation Solution

A system that incorporates user permissions and context data to determine the relevance of information output, using natural language processing to associate information with skill identifiers and session points, ensuring information is provided only when most relevant, such as during skill sessions or based on user interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If information is continuously output to users during skill sessions, then user engagement may be maintained, but users receive irrelevant information leading to suboptimal user experience

Engineering Contradiction:
Improveuser experienceVSAvoidinformation relevance
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The system performs preliminary analysis of context data including skill session state, user profile, and interaction history before determining what information to output. This advance preparation enables the system to anticipate user needs and provide relevant information proactively, rather than continuously outputting information and hoping it is useful.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The information output system dynamically adjusts its behavior based on real-time context data. The determination of what information to output changes continuously based on skill session state, user interactions, and relevance assessments, making the system adaptive rather than static in its information delivery approach.

Inventive Principle:
Principle #15Dynamics

2Reliability

If the system outputs all available information, then comprehensive user support is provided, but unnecessary interactions increase reducing efficiency

Engineering Contradiction:
Improveuser support completenessVSAvoidinteraction efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system applies different information output strategies to different contexts and user states. Rather than using a uniform approach, it tailors information delivery to specific skill sessions, user profiles, and interaction contexts, providing comprehensive support only where and when needed based on local conditions.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system uses user interactions and skill session outcomes as feedback to refine its information output decisions. By monitoring whether provided information leads to successful user tasks, the system learns to adjust its information delivery to maintain completeness while reducing unnecessary interactions that do not contribute to user goals.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If context data processing is performed to determine information relevance, then information output accuracy improves, but system complexity increases

Engineering Contradiction:
Improveinformation relevance accuracyVSAvoidsystem processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The context data processing system is divided into modular components that handle different aspects of relevance determination separately. Each module processes specific types of context data (skill session state, user profile, interaction history) independently, then combines results to make final determination, reducing overall complexity through functional segmentation.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11699441B2Contextual content for voice user interfaces
Publication Date: 2023.07.11 AMAZON TECH INC
  • US11699441B2 patent drawing
  • US11699441B2 patent drawing
  • US11699441B2 patent drawing

AI summary

The present disclosure describes techniques for dynamically determining when information is to be output to a user, as well as what information is to be output to a user. A natural language processing system may receive, from a first device, first data representing information to be output at a first point during a skill session. The natural language processing system may also receive, from a second device, second data representing a natural language input. The natural language processing system may determine a skill component is to execute with respect to the natural language input. The natural language processing system may send, to the skill component, second data representing the natural language input. The natural language processing system may receive, from the skill component, an indication that an ongoing first skill session with the second device has reached the first point. After receiving the indication and based at least in part on system usage data associated with at least one user, the natural language processing system may determine third data representing a prompt corresponding to the information and send, to the second device, the third data for output.