Volume-Triggered Voice Link Activation for Low-Bandwidth Speech Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems require significant bandwidth and computing resources, and raise privacy concerns due to continuous audio transmission, with inefficiencies in processing and resource utilization when no commands are being issued.

Innovation Solution

Implementing a distributed speech processing system that activates only upon a user-defined waking command, using volume-initiated communications to determine the intended recipient and transmit audio data selectively, reducing unnecessary processing and bandwidth usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If continuous audio transmission is implemented for speech recognition, then speech processing capability is improved, but bandwidth consumption and computing resource utilization increase significantly

Engineering Contradiction:
Improvespeech processing capabilityVSAvoidbandwidth consumption
Core Design Contradiction:
Ease of operationVSLoss of energy

Solution Approach 1:

The system performs preliminary volume assessment of audio input before initiating full speech processing and transmission. By evaluating the volume threshold in advance, the system determines whether further processing is necessary, thereby avoiding unnecessary bandwidth consumption and computing resource usage when no valid command is present

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of continuously transmitting all audio data for complete speech processing, the system applies partial action by only transmitting audio data that meets the volume threshold criteria. This selective transmission reduces bandwidth consumption while maintaining effective speech recognition capability when needed

Inventive Principle:
Principle #16Partial or excessive action

2Ease of operation

If continuous audio transmission is implemented for speech recognition, then speech processing capability is improved, but privacy concerns increase due to continuous data transmission

Engineering Contradiction:
Improvespeech processing capabilityVSAvoidprivacy concerns
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary volume assessment before transmitting audio data to the server. By evaluating whether the audio volume exceeds the threshold in advance, the system determines whether transmission is necessary, thereby minimizing privacy concerns by only transmitting data when a valid command is detected

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial transmission by sending only those audio segments that meet the volume threshold criteria. This selective approach reduces the amount of personal audio data transmitted to external servers, thereby addressing privacy concerns while maintaining speech recognition functionality

Inventive Principle:
Principle #16Partial or excessive action

3Loss of energy

If volume threshold detection is added to activate communications, then bandwidth utilization is reduced, but device complexity increases

Engineering Contradiction:
Improvebandwidth utilizationVSAvoidsystem complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The speech processing system is segmented into distinct functional stages: volume detection stage, threshold evaluation stage, and speech processing stage. This segmentation allows the system to perform simple volume checks without committing to full speech processing, thereby reducing bandwidth utilization while maintaining manageable device complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11348579B1Volume initiated communications
Publication Date: 2022.05.31 AMAZON TECH INC
  • US11348579B1 patent drawing
  • US11348579B1 patent drawing
  • US11348579B1 patent drawing

AI summary

A system incorporating volume activated communications. A system with local devices linked to a remote server that is capable of processing voice commands is capable of linking local devices in a communications link where the link is initiated upon high volume speech being detected by a receiving device. A target device is determined based on an estimated connection between users at the respective devices. Audio is then sent from the receiving device to the target device resulting in a communication link between the two devices, allowing communication between the users.