Speech-Controlled Communication Requests with Server-Mediated Call Setup
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition and natural language understanding systems are computationally expensive and not configured for efficient communication establishment between speech-controlled devices, lacking integration with traditional communication methods.
Innovation Solution
A system that utilizes distributed computing to process speech inputs, enabling speech-controlled devices to establish communications by converting audio signals into text, determining recipients and communication topics, and providing visual or audible indications before connection, allowing recipients to accept or decline calls.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech recognition and natural language understanding processing techniques are used to enable voice-controlled communication requests, then user interaction capability is improved, but computational cost and system complexity increase
Solution Approach 1:
The patent introduces a server as an intermediary component that handles speech recognition and natural language understanding processing. The server receives audio inputs from speech-controlled devices, processes them through speech recognition and NLU techniques, and manages communication establishment. This intermediary approach allows complex processing to be centralized while keeping individual devices simpler.
Solution Approach 2:
The system is segmented into distinct functional components: speech-controlled devices for audio capture and basic processing, a server for comprehensive speech recognition and NLU processing, and communication infrastructure for call establishment. This segmentation distributes computational burden and allows each component to be optimized independently.
2Productivity
If distributed computing is used to process speech inputs and establish communications, then processing efficiency is improved, but system complexity increases
Solution Approach 1:
The patent merges speech recognition and natural language understanding processing into a unified server-based system. Instead of implementing separate processing chains for each function, the server combines both capabilities to handle speech inputs end-to-end, improving processing efficiency by eliminating redundant steps and data transformations.
3Reliability
If visual or audible indications are provided before connection establishment, then communication reliability is improved, but communication time increases
Solution Approach 1:
The system performs preliminary actions by providing visual or audible indications to users before establishing communication connections. This allows users to verify that the correct recipient is being contacted and to prepare for the incoming call, improving communication reliability. The indication phase acts as a preliminary step that confirms communication intent before actual connection is made.
Data Source
AI summary
Systems and methods for establishing communication connections using speech, such as establishing calls between speech-controlled devices, are described. A first speech-controlled device receives a communication request in the form of audio and sends audio data corresponding to the captured audio to a server. The server performs speech processing on the audio data to determine a recipient, a subject for the call, and a device associated with the recipient. The server then sends a message indicating the communication request and audio data corresponding to the communication topic to the recipient's speech-controlled device. The recipient device outputs audio to the recipient requesting whether the recipient accepts the communication request. The recipient audibly refuses or accepts the communication request, and the recipient's speech-controlled device sends an indication of the recipient's audible decision to the server. If the recipient accepted the communication request, the server causes a communication connection be established between the two speech-controlled devices.


