Multi-portion Spoken Command Framework for Speech System Content Integration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech processing systems face difficulties in integrating content from various services due to different content formats, requiring customized infrastructure and software for each type of content, which complicates the integration process.
Innovation Solution
A multi-portion spoken command framework is introduced, where metadata is associated with incoming content to map it to existing spoken commands, allowing for easier integration and processing of content from different services within a single speech processing system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If customized infrastructure and software are implemented for each type of content, then content integration capability is improved, but system complexity increases
Solution Approach 1:
The patent implements a universal metadata schema that can represent multiple content types (audio, video, text, articles, books) using a common structure. This allows the speech processing system to handle diverse content formats through a single standardized interface, eliminating the need for separate customized infrastructure for each content type while maintaining full adaptability.
Solution Approach 2:
The system uses configurable parameters within the metadata schema to adapt to different content types without changing the underlying infrastructure. By modifying metadata parameters rather than system architecture, the system achieves content-type specificity while maintaining overall simplicity and reusability across different services.
2Ease of operation
If a single standardized spoken command framework is used, then ease of operation is improved, but adaptability to different content formats decreases
Solution Approach 1:
The metadata schema is designed to be multi-functional, supporting audio, video, text, articles, and books through a unified structure. This allows standardized spoken commands to work across all content types without requiring format-specific command variations, achieving both simplicity and adaptability simultaneously.
Solution Approach 2:
The metadata acts as an intermediary layer between the standardized spoken commands and the diverse content formats. The metadata schema translates various content types into a common representation that the speech processing system can uniformly interpret and execute, enabling format-agnostic command processing.
Data Source
AI summary
A framework for efficiently importing content into a speech-controlled system in a manner that makes the content easily accessible using voice commands. A speech-controlled system that can be controlled using a variety of commands, including a command to retrieve audio content, can be configured using a framework of content organization that allows new content to be ingested using the framework, thus making the new content accessible to users of the system without manually adjusting the system to recognize when incoming commands call for the new content. The framework can include configured content demarcations (such as information demarcations that divide content into articles, or other sized portions), labels for those demarcations (such as topic descriptors or the like), etc. Various language processing components such as intent classification and named entity recognition can be configured to recognize words that relate to the new content such as recognizing intents to access the new content using the topic descriptors or other labels. In this manner data related to new content may be more efficiently incorporated into a speech processing system.


