Configurable Speech Output Data Formats
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech processing systems require users to speak multiple utterances to obtain additional information, leading to a segmented and time-consuming user experience, as they are limited to default content outputs from core domains and lack seamless integration with add-on domains for customized responses.
Innovation Solution
A speech processing system that configures core domains with customizable content portions, allowing users to select add-on domains for enhanced outputs, combining data from multiple sources to provide comprehensive responses to a single user command, such as incorporating air quality and wind speed with weather information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the speech processing system uses default content outputs from core domains only, then the system structure remains simple and manageable, but the user experience becomes segmented and time-consuming as users must speak multiple utterances to obtain additional information
Solution Approach 1:
The speech processing system is designed to handle multiple types of information requests (core domains and add-on domains) through a unified processing framework. The system can process both default weather information and additional add-on information (air quality, wind speed, etc.) using the same speech recognition and natural language understanding pipelines, eliminating the need for separate processing paths and improving user experience continuity.
Solution Approach 2:
The patent combines results from multiple domains (core and add-on) into a single integrated response. The system merges weather data, air quality data, wind speed data, and other add-on information into one comprehensive output that is presented to the user in a unified manner, avoiding segmented interactions and reducing the number of utterances required.
2Loss of information
If the system provides comprehensive responses by integrating multiple domains, then the information completeness improves, but the processing time and system complexity increase
Solution Approach 1:
The system performs preliminary actions by pre-configuring available domains and their associated parameters before user interaction. The speech processing system has already established connections to multiple domains (weather, air quality, wind speed) and their data structures are prepared in advance, allowing rapid retrieval and integration of information when a user speaks a single command, thus reducing processing time while maintaining information completeness.
3Adaptability or versatility
If users can select add-on domains for customized outputs, then the system versatility improves, but the configuration complexity and data integration difficulty increase
Solution Approach 1:
The system implements dynamic configuration where users can select and customize which add-on domains they want to include in their speech queries. The system adapts its behavior based on user preferences, dynamically adjusting which domains are activated and integrated. This dynamic approach allows the system to maintain versatility while managing complexity by only activating the necessary domain integrations for each user.
Data Source
AI summary
Configurable core domains of a speech processing system are described. A core domain output data format for a given command is originally configured with default content portions. When a user indicates additional content should be output for the command, the speech processing system creates a new output data format for the core domain. The new output data format is user specific and includes both default content portions as well as user preferred content portions.


