Voice Session Routing Using Centralized Action Parsing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Disparate computing resources face challenges in efficiently processing and accurately parsing audio-based instructions due to inconsistent or outdated voice models, leading to inefficient bandwidth utilization and degraded response quality.

Innovation Solution

A data processing system that routes packetized actions via a computer network, using voice models trained on aggregate voice to parse voice-based instructions and generate an action data structure, which is transmitted to third-party provider devices for processing, thereby improving efficiency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If disparate computing resources process audio-based instructions independently using their own voice models, then each device can operate autonomously, but processing accuracy and consistency deteriorate due to unsynchronized voice models

Engineering Contradiction:
Improveautonomous operationVSAvoidprocessing accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

A centralized data processing system acts as an intermediary between audio input and third-party provider devices. This mediator receives audio signals, processes them through unified voice models to generate action data structures, and transmits these structures to provider devices. This eliminates the need for each device to independently process audio, ensuring consistent and accurate parsing across all disparate computing resources while maintaining autonomous operation of individual devices.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent merges the voice processing functionality into a centralized system that consolidates multiple voice models into a unified processing architecture. By combining the processing capabilities of disparate devices into a single coordinated system, the patent achieves consistent action data structure generation across all third-party provider devices, resolving the accuracy inconsistency problem.

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If multiple third party provider devices process the same audio input independently, then device autonomy is maintained, but network traffic and processing time increase

Engineering Contradiction:
Improvedevice autonomyVSAvoidprocessing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The centralized data processing system performs preliminary processing of audio signals by generating action data structures before transmitting them to third-party provider devices. This preliminary action converts raw audio into structured data that provider devices can process efficiently, eliminating the need for each device to independently perform computationally intensive audio processing, thus reducing network traffic and improving overall processing efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts the computationally intensive voice processing function from individual third-party provider devices and relocates it to a centralized data processing system. This extraction allows provider devices to focus on their core functions while the centralized system handles audio processing, reducing the processing burden on each device and improving overall system productivity.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If voice models are updated independently on each device, then local customization is possible, but synchronization and consistency become difficult to maintain

Engineering Contradiction:
Improvelocal customizationVSAvoidmodel synchronization
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The centralized data processing system provides universal voice processing functionality that serves all third-party provider devices. By implementing a single unified voice model architecture that can handle multiple types of audio inputs and generate standardized action data structures, the system achieves both consistency across devices and the ability to adapt to different service requirements, eliminating synchronization issues while maintaining versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3905653B1Natural language processing for session establishment with service providers
Publication Date: 2026.05.20 GOOGLE LLC
  • EP3905653B1 patent drawingFigure 1
  • EP3905653B1 patent drawingFigure 2
  • EP3905653B1 patent drawingFigure 3

AI summary

Routing packetized actions in a voice activated data packet based computer network environment is provided. A system can receive audio signals detected by a microphone of a device. The system can parse the audio signal to identify trigger keyword and request, and generate an action data structure. The system can transmit the action data structure to a third party provider device. The system can receive an indication from the third party provider device that a communication session was established with the device.