Speech-Controlled Communication Requests with Server-Mediated Call Setup

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition and natural language understanding systems are computationally expensive and not configured for efficient communication establishment between speech-controlled devices, lacking integration with traditional communication methods.

Innovation Solution

A system that utilizes distributed computing to process speech inputs, enabling speech-controlled devices to establish communications by converting audio signals into text, determining recipients and communication topics, and providing visual or audible indications before connection, allowing recipients to accept or decline calls.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition and natural language understanding processing techniques are used to enable voice-controlled communication requests, then user interaction capability is improved, but computational cost and system complexity increase

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces a server as an intermediary component that handles speech recognition and natural language understanding processing. The server receives audio inputs from speech-controlled devices, processes them through speech recognition and NLU techniques, and manages communication establishment. This intermediary approach allows complex processing to be centralized while keeping individual devices simpler.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system is segmented into distinct functional components: speech-controlled devices for audio capture and basic processing, a server for comprehensive speech recognition and NLU processing, and communication infrastructure for call establishment. This segmentation distributes computational burden and allows each component to be optimized independently.

Inventive Principle:
Principle #1Segmentation

2Productivity

If distributed computing is used to process speech inputs and establish communications, then processing efficiency is improved, but system complexity increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges speech recognition and natural language understanding processing into a unified server-based system. Instead of implementing separate processing chains for each function, the server combines both capabilities to handle speech inputs end-to-end, improving processing efficiency by eliminating redundant steps and data transformations.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If visual or audible indications are provided before connection establishment, then communication reliability is improved, but communication time increases

Engineering Contradiction:
Improvecommunication reliabilityVSAvoidcommunication time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by providing visual or audible indications to users before establishing communication connections. This allows users to verify that the correct recipient is being contacted and to prepare for the incoming call, improving communication reliability. The indication phase acts as a preliminary step that confirms communication intent before actual connection is made.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12424223B2Voice-controlled communication requests and responses
Publication Date: 2025.09.23 AMAZON TECH INC
  • US12424223B2 patent drawing
  • US12424223B2 patent drawing
  • US12424223B2 patent drawing

AI summary

Systems and methods for establishing communication connections using speech, such as establishing calls between speech-controlled devices, are described. A first speech-controlled device receives a communication request in the form of audio and sends audio data corresponding to the captured audio to a server. The server performs speech processing on the audio data to determine a recipient, a subject for the call, and a device associated with the recipient. The server then sends a message indicating the communication request and audio data corresponding to the communication topic to the recipient's speech-controlled device. The recipient device outputs audio to the recipient requesting whether the recipient accepts the communication request. The recipient audibly refuses or accepts the communication request, and the recipient's speech-controlled device sends an indication of the recipient's audible decision to the server. If the recipient accepted the communication request, the server causes a communication connection be established between the two speech-controlled devices.