Smart Speaker Microphone Array Calibration via Optical Sensors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Smart speaker systems face challenges in accurately recognizing voice commands due to acoustic reflections from hard and soft surfaces in environments, which are not typically omni-directional and can create undesirable reverberation, affecting voice recognition accuracy.

Innovation Solution

The integration of onboard image sensors, such as self-lit infrared cameras, to detect room conditions and calibrate the microphone array by adjusting the far-field microphone algorithm and potentially turning off microphones to minimize voice reflections, allowing for more accurate voice recognition by accounting for proximity to surfaces and adjusting beamforming techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Length of moving object

If the microphone array is placed in a circular pattern for optimized omni-directional long-range voice pickup, then the voice recognition range is improved, but acoustic reflections from hard and soft surfaces create reverberation that degrades voice recognition accuracy

Engineering Contradiction:
Improvevoice pickup rangeVSAvoidvoice recognition accuracy
Core Design Contradiction:
Length of moving objectVSMeasurement precision

Solution Approach 1:

The system performs preliminary room calibration before normal operation by playing test tones and using image sensors to map the environment. This advance preparation allows the system to pre-identify reflective surfaces and pre-adjust microphone weights to compensate for expected reverberation, thereby maintaining voice recognition accuracy while preserving omni-directional pickup capability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically changes the parameters of the beamforming algorithm based on detected room conditions. By adjusting microphone weights and beamforming coefficients according to the calibrated room model, the system adapts to different acoustic environments, eliminating the need to compromise between pickup range and recognition accuracy

Inventive Principle:
Principle #35Parameter changes

2Stability of the object's composition

If the smart speaker is placed against a hard wall for space efficiency and stability, then device stability is improved, but voice reflections from the wall create indeterminate reflections that degrade voice recognition

Engineering Contradiction:
Improvedevice stabilityVSAvoidvoice recognition accuracy
Core Design Contradiction:
Stability of the object's compositionVSMeasurement precision

Solution Approach 1:

During initial setup, the system performs room calibration that specifically maps walls and reflective surfaces relative to the speaker's position. By预先 identifying the hard wall placement scenario, the system can pre-compensate for the expected strong reflections, allowing the device to maintain its stable wall-mounted position without sacrificing voice recognition accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors the acoustic environment using the microphone array and compares received signals against the calibrated room model. When wall reflections are detected, the system provides feedback to adjust the beamforming algorithm in real-time, dynamically compensating for the fixed placement against hard surfaces

Inventive Principle:
Principle #23Feedback

3Measurement precision

If onboard image sensors are added to detect room conditions and calibrate the microphone array, then voice recognition accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The image sensors serve multiple functions: they map the room geometry, identify reflective surfaces, detect the presence of objects, and track user position. By consolidating these environmental sensing tasks into a single sensor type, the system achieves improved voice recognition without proportionally increasing complexity, as the same hardware infrastructure supports multiple calibration and operational functions

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution enhances voice recognition accuracy by reducing unwanted acoustic reflections, improving the ability of smart speakers to process voice signals effectively even in environments with strong reflective surfaces, thereby enhancing user interaction and command recognition.

Implementation Method 1

a processor of the speaker system. The optical signal can be generated by an optical source of the one or more optical sensors and the reflected optical signal can be detected by an optical detector of the one or more optical sensors

Methodology Applied
Scientific EffectOptical reflection: Reflection

Data Source

PatentUS11095980B2Smart speaker system with microphone room calibration
Publication Date: 2021.08.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11095980B2 patent drawing
  • US11095980B2 patent drawing
  • US11095980B2 patent drawing

AI summary

Systems and methods can be implemented to include a speaker system with microphone room calibration in a variety of applications. The speaker system can be implemented as a smart speaker. The speaker system can include a microphone array having multiple microphones, one or more optical sensors, one or more processors, and a storage device comprising instructions. The one or more optical sensors can be used to determine distances of one or more surfaces to the speaker system. Based on the determined distances, an algorithm to manage beamforming of an incoming voice signal to the speaker system can be adjusted or selected one or more microphones of the microphone array can be turned off, with an adjustment of an evaluation of the voice signal to the microphone array to account for the one or more microphones turned off. Additional systems and methods are disclosed.