Audio Volume Adjustment via User and Noise Location Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech processing systems face challenges in ensuring that output audio is at a sufficient volume for users in noisy environments, as they struggle to accurately determine the user's location and noise levels, leading to inconsistent audio quality.

Innovation Solution

The system employs a combination of microphone arrays, beamforming techniques, and environmental noise estimation to calculate a gain for output audio, ensuring it reaches the user at an appropriate volume by determining the user's and noise sources' locations and sound pressure levels, and adjusts based on user feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the system uses standard audio output without environmental adaptation, then the device complexity is low, but the audio quality and user experience deteriorate in noisy environments

Engineering Contradiction:
Improveaudio quality consistencyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by determining user location and noise levels before audio output is generated. The speech processing system calculates the appropriate gain for output audio based on predicted user position and environmental noise characteristics, ensuring the audio is optimized for the specific listening conditions before the user hears it.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system incorporates feedback mechanisms where user responses (such as volume adjustments or explicit feedback) are used to refine and update the volume model. The speech processing system continuously improves its audio output by learning from user interactions and adapting to individual preferences and environmental conditions.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If the system accurately determines user location and noise levels, then the audio output quality improves, but the measurement precision requirements and device complexity increase

Engineering Contradiction:
Improveuser location and noise level detection accuracyVSAvoiddetection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The speech processing system performs multiple functions using the same audio input hardware. The microphone array that captures speech for recognition also serves to determine user location through beamforming and to estimate noise levels. This multi-functional approach allows accurate measurement without adding separate dedicated sensors or detection systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses its own audio output and the user's speech input to automatically determine listening conditions. By analyzing the audio environment and user responses, the system self-adjusts the output volume and characteristics without requiring external measurement devices or manual user configuration.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10147439B1Volume adjustment for listening environment
Publication Date: 2018.12.04 AMAZON TECH INC
  • US10147439B1 patent drawing
  • US10147439B1 patent drawing
  • US10147439B1 patent drawing

AI summary

A speech-capturing device that can modulate its output audio data volume based on environmental sound conditions at the location of a user speaking to the device. The device detects the sound pressure of a spoken utterance at the device location and determines the distance of the user from the device. The device also detects the sound pressure of noise at the device and uses information about the location of the noise source and user to determine the sound pressure of noise at the location of the talker. The device can then adjust the gain for output audio (such as a spoken response to the utterance) to ensure that the output audio is at a certain desired sound pressure when it reaches the location of the user.