Multi-modal Audio Processing for Voice Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice-controlled devices struggle with accurately recognizing voice commands in noisy environments due to interference from background sounds generated by electronic devices, which are not effectively addressed by traditional noise cancellation techniques.

Innovation Solution

The implementation of multi-modal audio processing in voice-controlled devices, which involves receiving audio signals through both microphones and electromagnetic signals, allowing for the reduction of noise interference by correlating and subtracting noise signals from the primary audio signal.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional noise cancellation techniques are used, then the device complexity is reduced, but the speech recognition accuracy deteriorates in noisy environments

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing pipeline complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces electromagnetic signal reception as a new dimension alongside traditional acoustic microphone reception. By receiving the same audio content through two different physical dimensions (acoustic waves and electromagnetic waves), the system creates redundant information channels that enable more effective noise cancellation and improve speech recognition accuracy in noisy environments.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent uses electromagnetic signals as an intermediary to obtain a reference version of the audio content being played by electronic devices. This intermediary signal serves as a template for identifying and subtracting noise components from the acoustic microphone signal, effectively separating desired speech from background electronic device noise.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multi-modal audio processing is implemented, then the speech recognition accuracy improves, but the device complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing pipeline complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the noise cancellation process into distinct modular steps: (1) receiving acoustic signal through microphone, (2) receiving electromagnetic signal, (3) correlating the two signals to identify noise components, (4) subtracting identified noise from acoustic signal. This segmentation makes the complex multi-modal processing more manageable and implementable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary correlation analysis between acoustic and electromagnetic signals to identify noise components before final speech recognition processing. By pre-identifying and removing noise components in advance, the system simplifies the subsequent speech recognition stage and improves overall accuracy.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If noise signals are subtracted from primary audio signal, then the purity of speech signal improves, but the processing time increases

Engineering Contradiction:
Improvesignal purityVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by selectively subtracting only the noise components that are correlated between electromagnetic and acoustic signals, rather than processing the entire signal spectrum. This selective approach removes sufficient noise to improve speech purity while avoiding excessive processing of all signal components, thus reducing processing time.

Inventive Principle:
Principle #16Partial or excessive action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach significantly improves the accuracy of speech recognition in voice-controlled devices by effectively filtering out noise from electronic devices, enhancing the responsiveness and reliability of voice-controlled systems in diverse environments.

Implementation Method 1

A microphone converts the sound waves into an audio signal

Methodology Applied
Scientific EffectAcoustic transduction:

Implementation Method 2

A typical home may have dozens of such devices, many with stereo or other multi-channel output, such as televisions, radios, smart speakers, telephones, computers, and portable 'boom-boxes' just to name a few. Each of these devices may obtain audio signals and use the audio signal to generate sound waves

Methodology Applied
Scientific EffectElectroacoustic transduction:

Data Source

PatentUS12273679B2Multi-modal audio processing
Publication Date: 2025.04.08 SOUNDHOUND AI IP LLC
  • US12273679B2 patent drawing
  • US12273679B2 patent drawing
  • US12273679B2 patent drawing

AI summary

A method for processing an audio signal involves receiving sound waves at a microphone, converting them into a first audio signal, and extracting a second audio signal from an electromagnetic signal received at a receiver. The first audio signal is correlated with the second audio signal to calculate a correlation value. If the correlation value exceeds a threshold, the first audio signal is processed using the second audio signal to reduce unwanted sound contributions, resulting in a processed audio signal. Further processing is then performed on the processed audio signal to determine a characteristic of the desired sound.