Voice Session Routing Using Centralized Action Parsing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Disparate computing resources face challenges in efficiently processing and accurately parsing audio-based instructions due to inconsistent or outdated voice models, leading to inefficient bandwidth utilization and degraded response quality.
Innovation Solution
A data processing system that routes packetized actions via a computer network, using voice models trained on aggregate voice to parse voice-based instructions and generate an action data structure, which is transmitted to third-party provider devices for processing, thereby improving efficiency and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If disparate computing resources process audio-based instructions independently using their own voice models, then each device can operate autonomously, but processing accuracy and consistency deteriorate due to unsynchronized voice models
Solution Approach 1:
A centralized data processing system acts as an intermediary between audio input and third-party provider devices. This mediator receives audio signals, processes them through unified voice models to generate action data structures, and transmits these structures to provider devices. This eliminates the need for each device to independently process audio, ensuring consistent and accurate parsing across all disparate computing resources while maintaining autonomous operation of individual devices.
Solution Approach 2:
The patent merges the voice processing functionality into a centralized system that consolidates multiple voice models into a unified processing architecture. By combining the processing capabilities of disparate devices into a single coordinated system, the patent achieves consistent action data structure generation across all third-party provider devices, resolving the accuracy inconsistency problem.
2Adaptability or versatility
If multiple third party provider devices process the same audio input independently, then device autonomy is maintained, but network traffic and processing time increase
Solution Approach 1:
The centralized data processing system performs preliminary processing of audio signals by generating action data structures before transmitting them to third-party provider devices. This preliminary action converts raw audio into structured data that provider devices can process efficiently, eliminating the need for each device to independently perform computationally intensive audio processing, thus reducing network traffic and improving overall processing efficiency.
Solution Approach 2:
The patent extracts the computationally intensive voice processing function from individual third-party provider devices and relocates it to a centralized data processing system. This extraction allows provider devices to focus on their core functions while the centralized system handles audio processing, reducing the processing burden on each device and improving overall system productivity.
3Adaptability or versatility
If voice models are updated independently on each device, then local customization is possible, but synchronization and consistency become difficult to maintain
Solution Approach 1:
The centralized data processing system provides universal voice processing functionality that serves all third-party provider devices. By implementing a single unified voice model architecture that can handle multiple types of audio inputs and generate standardized action data structures, the system achieves both consistency across devices and the ability to adapt to different service requirements, eliminating synchronization issues while maintaining versatility.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Routing packetized actions in a voice activated data packet based computer network environment is provided. A system can receive audio signals detected by a microphone of a device. The system can parse the audio signal to identify trigger keyword and request, and generate an action data structure. The system can transmit the action data structure to a third party provider device. The system can receive an indication from the third party provider device that a communication session was established with the device.