Computer Vision Audio Control for Conferencing Echo Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Stand-alone telephone conferencing devices often produce sub-optimal sound and fail to effectively detect speaker input due to environmental factors, with current technologies lacking adequate solutions to address these issues.

Innovation Solution

A device equipped with a processor and storage that uses computer vision and augmented reality to identify the location of a conferencing device and objects within the environment, adjusting the operation of its speakers and microphones to optimize audio performance by reducing echo and ambient noise through commands transmitted to the device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If stand-alone telephone conferencing devices are used in various environments, then the devices can be deployed flexibly, but the audio quality deteriorates due to echo and ambient noise

Engineering Contradiction:
Improvedeployment flexibilityVSAvoidecho and ambient noise
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary environmental scanning using computer vision to identify objects and surfaces before audio processing. By pre-mapping the environment and identifying potential echo sources (walls, furniture, other devices), the system can proactively adjust audio parameters to prevent echo and noise issues before they degrade call quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors the environment using camera input and adjusts speaker and microphone operations in real-time based on detected objects. The feedback loop involves capturing visual data, identifying environmental features, analyzing audio characteristics, and dynamically adjusting device parameters to maintain optimal audio quality despite environmental variations

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If the conferencing device operates in environments with objects nearby, then the device can function in diverse settings, but sound detection accuracy deteriorates

Engineering Contradiction:
Improveenvironmental adaptabilityVSAvoidspeaker input detection accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system applies different operational characteristics to different microphones and speakers based on their local environmental context. By identifying which microphones are oriented toward objects or potential noise sources, the system can selectively adjust sensitivity levels for specific microphones while maintaining optimal performance for others, preserving overall detection accuracy in diverse settings

Inventive Principle:
Principle #3Local quality

3Area of stationary object

If the conferencing device outputs sound in all directions, then audio coverage is maximized, but echo is generated from surrounding objects

Engineering Contradiction:
Improveaudio coverage areaVSAvoidecho
Core Design Contradiction:
Area of stationary objectVSObject-generated harmful factors

Solution Approach 1:

The system enables different speakers to operate with different characteristics based on their orientation and environmental context. Speakers facing open spaces or areas without reflective surfaces operate at full power for maximum coverage, while speakers oriented toward walls or furniture have their output reduced or deactivated to prevent echo generation

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system pre-identifies potential echo sources using computer vision before audio output begins. By mapping the environment and identifying reflective surfaces and objects, the system can proactively configure speaker operations to direct audio away from potential echo sources while maintaining coverage in appropriate directions

Inventive Principle:
Principle #10Preliminary action

4Object-affected harmful factors

If computer vision and AR processing are used to identify objects and adjust audio, then audio quality is improved, but device complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidsystem complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The system uses a single camera device to perform multiple functions: environmental mapping, object identification, and audio quality monitoring. By making the camera serve multiple purposes rather than requiring separate sensors for each function, the system reduces overall device complexity while still achieving sophisticated environmental awareness and audio optimization

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The conferencing device performs its own environmental assessment and self-adjustment without requiring external configuration or manual intervention. The device autonomously captures visual data, identifies its environment, analyzes audio characteristics, and adjusts its own speaker and microphone operations, eliminating the need for separate setup systems or expert configuration

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11258417B2Techniques for using computer vision to alter operation of speaker(s) and/or microphone(s) of device
Publication Date: 2022.02.22 LENOVO SWITZERLAND INTERNATIONAL GMBH
  • US11258417B2 patent drawing
  • US11258417B2 patent drawing
  • US11258417B2 patent drawing

AI summary

In one aspect, a first device includes at least one processor and storage accessible to the at least one processor. The storage includes instructions that may be executable by the processor to receive input from a camera and identify a second device based on the input from the camera. The second device may include at least one speaker and at least one microphone. The instructions may also be executable to identify a current location of the second device within an environment based on the input from the camera and to identify a current location of an object within the environment that is different from the second device. The instructions may then be executable to provide a command to alter operation of the at least one speaker and/or the at least one microphone based on the current location of the second device and the current location of the object.