Configurable Speech Output Data Formats

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech processing systems require users to speak multiple utterances to obtain additional information, leading to a segmented and time-consuming user experience, as they are limited to default content outputs from core domains and lack seamless integration with add-on domains for customized responses.

Innovation Solution

A speech processing system that configures core domains with customizable content portions, allowing users to select add-on domains for enhanced outputs, combining data from multiple sources to provide comprehensive responses to a single user command, such as incorporating air quality and wind speed with weather information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the speech processing system uses default content outputs from core domains only, then the system structure remains simple and manageable, but the user experience becomes segmented and time-consuming as users must speak multiple utterances to obtain additional information

Engineering Contradiction:
Improveuser experience continuityVSAvoidsystem integration complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The speech processing system is designed to handle multiple types of information requests (core domains and add-on domains) through a unified processing framework. The system can process both default weather information and additional add-on information (air quality, wind speed, etc.) using the same speech recognition and natural language understanding pipelines, eliminating the need for separate processing paths and improving user experience continuity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines results from multiple domains (core and add-on) into a single integrated response. The system merges weather data, air quality data, wind speed data, and other add-on information into one comprehensive output that is presented to the user in a unified manner, avoiding segmented interactions and reducing the number of utterances required.

Inventive Principle:
Principle #5Merging (Combining)

2Loss of information

If the system provides comprehensive responses by integrating multiple domains, then the information completeness improves, but the processing time and system complexity increase

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-configuring available domains and their associated parameters before user interaction. The speech processing system has already established connections to multiple domains (weather, air quality, wind speed) and their data structures are prepared in advance, allowing rapid retrieval and integration of information when a user speaks a single command, thus reducing processing time while maintaining information completeness.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If users can select add-on domains for customized outputs, then the system versatility improves, but the configuration complexity and data integration difficulty increase

Engineering Contradiction:
Improveoutput customization capabilityVSAvoiddomain integration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements dynamic configuration where users can select and customize which add-on domains they want to include in their speech queries. The system adapts its behavior based on user preferences, dynamically adjusting which domains are activated and integrated. This dynamic approach allows the system to maintain versatility while managing complexity by only activating the necessary domain integrations for each user.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12190885B2Configurable output data formats
Publication Date: 2025.01.07 AMAZON TECH INC
  • US12190885B2 patent drawing
  • US12190885B2 patent drawing
  • US12190885B2 patent drawing

AI summary

Configurable core domains of a speech processing system are described. A core domain output data format for a given command is originally configured with default content portions. When a user indicates additional content should be output for the command, the speech processing system creates a new output data format for the core domain. The new output data format is user specific and includes both default content portions as well as user preferred content portions.