Smart Audio Component for Noisy Environment Voice Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current social networking systems face challenges in effectively managing audio inputs during audio-visual communications, particularly in noisy environments where multiple sound sources are present, leading to difficulties in isolating and amplifying the most relevant human voices over background noise.

Innovation Solution

An intelligent communication device equipped with a 'smart audio' component that utilizes a microphone array, sound localization, classification, and engagement metric calculation to distinguish between sound sources, amplify the most interesting conversation, and attenuate less relevant sounds, based on features like human voice recognition, movement, and user engagement metrics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional audio processing is used in noisy environments, then all sound sources are captured equally, but relevant human voices are lost in background noise

Engineering Contradiction:
Improveaudio clarityVSAvoidnoise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system extracts and isolates human voice signals from the mixed audio input by using sound classification to identify and separate relevant speech from background noise and other sound sources, thereby improving audio clarity in noisy environments

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies different audio processing qualities to different sound sources based on their classification - human voices receive amplification and enhancement while background noise and less relevant sounds are attenuated or suppressed, creating localized quality differences in the audio output

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If multiple sound sources are amplified equally, then all conversations are heard, but user engagement with relevant content decreases

Engineering Contradiction:
Improveaudio source selectionVSAvoiduser engagement
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system uses engagement metrics as feedback to continuously adjust audio amplification - by monitoring user interaction with different sound sources and calculating engagement levels, the system dynamically adjusts which conversations are amplified to maximize user engagement

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically adjusts audio amplification levels based on real-time engagement metrics and user behavior, allowing the audio processing to adapt and change during the communication session rather than maintaining static amplification levels

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If smart audio processing is implemented, then relevant voices are amplified, but device complexity increases

Engineering Contradiction:
Improvesound source discriminationVSAvoidaudio processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio processing system is divided into separate functional modules - sound classification, engagement metric calculation, and selective amplification - allowing each component to be optimized independently and simplifying the overall system architecture while maintaining high measurement precision

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10838689B2Audio selection based on user engagement
Publication Date: 2020.11.17 META PLATFORMS INC
  • US10838689B2 patent drawing
  • US10838689B2 patent drawing
  • US10838689B2 patent drawing

AI summary

In one embodiment, a method includes receiving audio input during an audio-video communication session. The audio input is generated by a first sound source within an environment and a second sound source within the environment. The method includes receiving video input depicting the first sound source and the second sound source in the environment. The method includes identifying the first sound source and the second sound source using the audio input and the video input. The method includes predicting a first engagement metric for the first sound source and a second engagement metric for the second sound source based on the identifying. The method includes processing the audio input to generate an audio output signal based on a comparison of the first engagement metric and the second engagement metric. The method includes providing the audio output signal to a computing device associated with the audio-video communication session.