Audio Signal Transmission with Distributed Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition in computing devices is affected by noise, volume, and orientation, and existing techniques for real-time processing are computationally expensive and resource-intensive.

Innovation Solution

The implementation of a distributed computing platform that reduces audio signal data by determining time offsets and signal differences between multiple audio transducers, allowing for efficient data transmission and processing, including beamforming, echo cancellation, and automatic speech recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If advanced speech extraction techniques are applied in near real-time, then speech recognition accuracy is improved, but computational cost and processing time increase significantly

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the audio signal processing into two distinct stages: (1) local device processing that performs computationally intensive operations like beamforming and echo cancellation to extract and clean the audio signal, and (2) cloud-based speech recognition that processes the pre-processed signal. This segmentation allows complex operations to be performed locally with optimized resources while maintaining near real-time performance, resolving the contradiction between accuracy and processing speed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by performing audio signal processing operations (beamforming, echo cancellation, noise reduction) at the local device before transmitting the signal to the cloud for speech recognition. This preliminary processing cleans and optimizes the audio data in advance, reducing the computational burden on remote servers and enabling faster overall processing while maintaining high accuracy.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If multiple audio transducers are used to capture audio signals, then speech recognition accuracy in noisy environments is improved, but data transmission volume and processing complexity increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and transmits only the essential information from the audio signals captured by multiple transducers. By performing local processing to determine time offsets and signal differences, the system extracts only the necessary audio data for speech recognition, reducing the volume of data that needs to be transmitted and processed in the cloud while maintaining the accuracy benefits of multi-transducer capture.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by performing computationally intensive signal processing operations (beamforming, echo cancellation) at the local device where the audio is captured, rather than transmitting all raw data to the cloud. This allows the system to leverage local computational resources to handle the complexity of multi-transducer data processing, reducing the burden on remote systems while maintaining high processing accuracy.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9570071B1Audio signal transmission techniques
Publication Date: 2017.02.14 AMAZON TECH INC
  • US9570071B1 patent drawing
  • US9570071B1 patent drawing
  • US9570071B1 patent drawing

AI summary

A voice interaction architecture that compiles multiple audio signals captured at different locations within an environment, determines a time offset between a primary audio signal and other captured audio signals and identifies differences between the primary signal and the other signal(s). Thereafter, the architecture may provide the primary audio signal, an indication of the determined time offset(s) and the identified differences to remote computing resources for further processing. For instance, the architecture may send this information to a network-accessible distributed computing platform that performs beamforming and/or automatic speech recognition (ASR) on the received audio. The distributed computing platform may in turn determine a response to provide based upon the beamforming and/or ASR.