Spatial Audio Processing for Speech Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic devices with multiple microphones face challenges in accurately detecting and enhancing user speech signals due to interference from bystanders and environmental noise, which affects real-time applications like voice trigger phrase detection and automatic speech recognition.

Innovation Solution

The system employs a machine learning model and spatial probability model to determine if an audio signal corresponds to a user's speech, using spectro-temporal data and spatial information to enhance speech detection, and integrates non-audio signals from cameras and radar sensors to improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple microphones are used to capture audio signals, then the ability to detect user speech is improved, but interference from bystanders and environmental noise increases

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidbackground noise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent applies local quality by processing audio signals differently based on their spatial origin. The system identifies and enhances signals from the user's location while applying different processing to signals from other directions, effectively treating different spatial regions with different quality characteristics to suppress noise while preserving speech

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent transitions from one-dimensional audio signal processing to three-dimensional spatial processing by incorporating direction of arrival estimation and spatial filtering. This adds spatial dimensions to the audio processing, allowing the system to distinguish between speech and noise based on their spatial characteristics rather than just temporal or spectral features

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If machine learning models and spatial processing are used to enhance speech detection, then speech isolation accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvespeech isolation accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing spatial filters and processing parameters before actual speech detection is needed. The system prepares spatial processing configurations in advance based on anticipated user positions and environmental characteristics, reducing the computational burden during real-time operation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary spatial processing layer between the microphones and the final speech detection output. This intermediary layer performs direction of arrival estimation and spatial filtering, acting as a mediator that simplifies the subsequent speech detection task by pre-processing the audio signals with spatial information

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11514928B2Spatially informed audio signal processing for user speech
Publication Date: 2022.11.29 APPLE INC
  • US11514928B2 patent drawing
  • US11514928B2 patent drawing
  • US11514928B2 patent drawing

AI summary

A device implementing a system for processing speech in an audio signal includes at least one processor configured to receive an audio signal corresponding to at least one microphone of a device, and to determine, using a first model, a first probability that a speech source is present in the audio signal. The at least one processor is further configured to determine, using a second model, a second probability that an estimated location of a source of the audio signal corresponds to an expected position of a user of the device, and to determine a likelihood that the audio signal corresponds to the user of the device based on the first and second probabilities.