Voice Command Routing via Centralized Aggregation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Disparate computing resources face challenges in accurately and consistently processing audio-based instructions in voice-based computing environments due to differences in voice models, leading to inefficiencies and inaccuracies in information transmission and processing.

Innovation Solution

A data processing system that uses aggregate voice-trained models to parse voice-based inputs, construct action data structures, and route these to third-party providers, improving the reliability, efficiency, and accuracy of voice-based instruction processing by bypassing the need for real-time voice processing at multiple devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple disparate computing resources process voice-based inputs independently, then each device can autonomously handle voice commands, but processing accuracy and consistency deteriorate due to differences in voice models

Engineering Contradiction:
Improveautonomous voice processing capabilityVSAvoidvoice processing accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces a centralized voice processing server as an intermediary between client devices and voice commands. The server aggregates voice model capabilities across multiple devices, receiving audio inputs from various sources and processing them through a unified model. This mediator ensures consistent and accurate voice processing while allowing individual devices to maintain their autonomous processing capability for simpler tasks.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If real-time voice processing is performed at multiple devices, then response time can be fast, but resource consumption and bandwidth usage increase

Engineering Contradiction:
Improvevoice command response timeVSAvoidprocessor and battery efficiency
Core Design Contradiction:
SpeedVSLoss of energy

Solution Approach 1:

The patent divides voice processing tasks into two segments: local preprocessing and centralized processing. Client devices perform initial audio capture and basic preprocessing locally to maintain fast response times for simple commands. More complex processing, including advanced voice recognition and model aggregation, is segmented and performed at the centralized server, reducing the computational burden on individual devices and their energy consumption.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If voice models are updated across multiple disparate computing resources, then system adaptability improves, but processing consistency deteriorates during transition periods

Engineering Contradiction:
Improvevoice model update capabilityVSAvoidprocessing consistency
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent merges voice model management across multiple computing resources into a unified system. The centralized server maintains a single, aggregated voice model that combines capabilities from various source models. When updates are needed, the server updates the unified model in one location and distributes it to all client devices, ensuring that all resources process voice inputs consistently while still benefiting from the combined adaptability of multiple original models.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11030239B2Audio based entity-action pair based selection
Publication Date: 2021.06.08 GOOGLE LLC
  • US11030239B2 patent drawing
  • US11030239B2 patent drawing
  • US11030239B2 patent drawing

AI summary

Routing packetized actions in a voice activated data packet based computer network environment is provided. A system can receive audio signals detected by a microphone of a device. The system can parse the audio signal to identify trigger keyword and request, and generate an action data structure. The action data structure can include digital components and entity-action pairs.