Voice-Controlled Entertainment System with Cloud-Based Audio Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies do not effectively enable individuals to immerse themselves in broadcast entertainment by interacting with voice-controlled devices, as they lack the capability to process and analyze spoken words in real-time, providing a seamless and engaging experience.

Innovation Solution

A system that utilizes voice-controlled electronic devices and natural language processing to allow individuals to interact with broadcast content by analyzing spoken words, comparing them to broadcast data, and providing feedback, scores, and interactive elements, enabling participation in games and events as if they were on the show.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice-controlled electronic devices are used to interact with broadcast content, then user engagement and immersion are improved, but real-time processing capability and system response speed deteriorate

Engineering Contradiction:
Improveuser engagementVSAvoidreal-time processing capability
Core Design Contradiction:
Ease of operationVSSpeed

Solution Approach 1:

A cloud-based processing system is introduced as an intermediary between the voice-controlled electronic device and the broadcast content. The cloud system receives audio streams from the broadcast, processes spoken words, compares them to broadcast data, and provides feedback through the electronic device. This architecture allows complex real-time processing to occur in the cloud rather than on the device itself, resolving the contradiction between user engagement and real-time processing capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If spoken words are analyzed and compared to broadcast data in real-time, then interactive feedback is improved, but system complexity and processing requirements worsen

Engineering Contradiction:
Improveinteractive feedbackVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system is segmented into distinct functional components: a cloud-based processing system that handles audio stream reception, spoken word analysis, and data comparison; and a voice-controlled electronic device that handles user interaction and feedback delivery. This segmentation allows each component to be optimized independently, reducing the complexity burden on any single device while enabling comprehensive interactive feedback capabilities.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If comprehensive analysis of spoken words is performed, then accuracy of interaction is improved, but processing time and response delay worsen

Engineering Contradiction:
Improveaccuracy of interactionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The cloud-based processing system continuously monitors and processes audio streams from broadcast content in advance, preparing analysis results before user interactions occur. By maintaining an ongoing analysis of spoken words and comparing them to broadcast data in real-time, the system can provide immediate feedback responses without significant processing delays, thus improving accuracy while minimizing time loss.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10950228B1Interactive voice controlled entertainment
Publication Date: 2021.03.16 AMAZON TECH INC
  • US10950228B1 patent drawing
  • US10950228B1 patent drawing
  • US10950228B1 patent drawing

AI summary

Methods and systems for receiving shouted-out user responses to broadcast entertainment content, and for determining the responsiveness of those responses in relation to the broadcast content. In particular, entertainment broadcasts can be accompanied by mark-up data that represents various events within a given broadcast, which can be compared to the shouted-out responses to determine their accuracy. For example, if a game show was broadcast and an individual started shouting out answers during the broadcast, embodiments disclosed herein could utilize a voice-controlled electronic device that captures the shouted-out answers and passes them on to a language processing system that determines whether they are correct by comparing the answers to the mark-up data. The voice-controlled electronic device can also “listen” to background sounds to capture the broadcast of the entertainment content, and send that content to the language processing system, which can use that captured data to synchronize the actual broadcast with the analysis of the shouted-out answers to provide individuals with an immersive entertainment experience.