Multi-portion Spoken Command Framework for Speech System Content Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing systems face difficulties in integrating content from various services due to different content formats, requiring customized infrastructure and software for each type of content, which complicates the integration process.

Innovation Solution

A multi-portion spoken command framework is introduced, where metadata is associated with incoming content to map it to existing spoken commands, allowing for easier integration and processing of content from different services within a single speech processing system.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If customized infrastructure and software are implemented for each type of content, then content integration capability is improved, but system complexity increases

Engineering Contradiction:
Improvecontent integration capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal metadata schema that can represent multiple content types (audio, video, text, articles, books) using a common structure. This allows the speech processing system to handle diverse content formats through a single standardized interface, eliminating the need for separate customized infrastructure for each content type while maintaining full adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses configurable parameters within the metadata schema to adapt to different content types without changing the underlying infrastructure. By modifying metadata parameters rather than system architecture, the system achieves content-type specificity while maintaining overall simplicity and reusability across different services.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If a single standardized spoken command framework is used, then ease of operation is improved, but adaptability to different content formats decreases

Engineering Contradiction:
Improvecommand usage simplicityVSAvoidcontent format compatibility
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The metadata schema is designed to be multi-functional, supporting audio, video, text, articles, and books through a unified structure. This allows standardized spoken commands to work across all content types without requiring format-specific command variations, achieving both simplicity and adaptability simultaneously.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The metadata acts as an intermediary layer between the standardized spoken commands and the diverse content formats. The metadata schema translates various content types into a common representation that the speech processing system can uniformly interpret and execute, enabling format-agnostic command processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11837225B1Multi-portion spoken command framework
Publication Date: 2023.12.05 AMAZON TECH INC
  • US11837225B1 patent drawing
  • US11837225B1 patent drawing
  • US11837225B1 patent drawing

AI summary

A framework for efficiently importing content into a speech-controlled system in a manner that makes the content easily accessible using voice commands. A speech-controlled system that can be controlled using a variety of commands, including a command to retrieve audio content, can be configured using a framework of content organization that allows new content to be ingested using the framework, thus making the new content accessible to users of the system without manually adjusting the system to recognize when incoming commands call for the new content. The framework can include configured content demarcations (such as information demarcations that divide content into articles, or other sized portions), labels for those demarcations (such as topic descriptors or the like), etc. Various language processing components such as intent classification and named entity recognition can be configured to recognize words that relate to the new content such as recognizing intents to access the new content using the topic descriptors or other labels. In this manner data related to new content may be more efficiently incorporated into a speech processing system.