Spatio-Temporal Beamformer With Variable Frequency Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing beamforming techniques, particularly those using machine learning for voice activity detection, consume significant computing resources due to the need for high frequency resolution, which is not feasible for resource-constrained edge devices.

Innovation Solution

A spatio-temporal beamforming approach that transforms audio signals to a coarse frequency domain, applies a neural network for probability estimation, and then to a fine frequency domain, determining a minimum variance distortionless response (MVDR) beamforming filter based on uneven frequency resolution, reducing computing resource consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high frequency resolution is used for voice activity detection, then detection accuracy is improved, but computing resource consumption increases

Engineering Contradiction:
Improvevoice activity detection accuracyVSAvoidcomputing resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent divides the frequency spectrum into multiple bands and processes each band separately with different resolution levels. High-resolution processing is applied only to critical frequency bands where speech information is most important, while lower-resolution processing is used for less critical bands. This segmentation allows the system to maintain detection accuracy in critical regions while reducing overall computing resource consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different quality levels of processing to different parts of the frequency spectrum. Critical frequency bands receive high-resolution processing to ensure accurate voice activity detection, while non-critical bands receive lower-resolution processing. This local quality approach ensures that detection accuracy is maintained where it matters most while reducing computational burden in less critical regions.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If machine learning techniques are used for voice activity detection, then detection accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvevoice activity detection accuracyVSAvoidcomputing resource requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the frequency spectrum into multiple bands and applies simplified processing to each segment rather than using complex machine learning on the entire spectrum. This segmentation reduces the input dimensionality for any learning-based components while maintaining detection accuracy through the combined analysis of multiple frequency bands.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of frequency resolution from uniform high resolution across all bands to variable resolution where critical bands have high resolution and non-critical bands have lower resolution. This parameter change reduces the overall computational complexity while maintaining detection accuracy in critical frequency regions where speech information is concentrated.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250285636A1Spatio-temporal beamformer
Publication Date: 2025.09.11 SYNAPTICS INC
  • US20250285636A1 patent drawing
  • US20250285636A1 patent drawing
  • US20250285636A1 patent drawing

AI summary

This disclosure provides methods, devices, and systems for signal processing. The present implementations relate more specifically to a spatio-temporal beamformer. In some aspects, a beamforming system may receive an audio signal via a plurality of microphones, the audio signal including a number (B) of frames for each of the plurality of microphones, each of the B frames for each of the plurality of microphones including a number (N) of time-domain samples. For a first microphone, the beamforming system may transform the B*N time-domain samples into B*N/2 first frequency-domain samples; transform the B*N/2 first frequency-domain samples into B*N/2 second frequency-domain samples; and determine a probability of speech associated with the B*N/2 second frequency-domain samples based on a neural network model. The beamformer system may determine a minimum variance distortionless response (MVDR) beamforming filter based at least in part on the probability of speech for the first microphone.