Sub-Band Speech Onset Detection for Low-Power Noise Robustness

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-controlled devices struggle to accurately detect the onset of speech in noisy environments with low power consumption and complexity, leading to inefficient power usage and high false positive rates.

Innovation Solution

A low-power and low-complexity speech onset detector (SOD) using a fractional-band filter structure and spectral subtraction technique to derive sub-band energy profiles, which includes a cascade of low-pass filters and spectral subtraction to distinguish speech from noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the SOD remains in an active state to detect speech onset at any time, then the detection accuracy is improved, but the power consumption increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The SOD dynamically transitions between sleep mode and active state based on detection needs. The system activates the SOD only when speech onset detection is required, rather than maintaining a static active state, thus reducing power consumption while preserving detection capability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The SOD operates in periodic cycles of sleep and active states. By scheduling detection periods and using idle states between detections, the system reduces overall power consumption while maintaining the ability to detect speech onset when needed.

Inventive Principle:
Principle #19Periodic action

2Measurement precision

If the SOD uses complex algorithms to improve detection accuracy in noisy environments, then the detection precision is improved, but the device complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidalgorithm complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The detection process is segmented into multiple stages: a low-complexity energy-based detector first identifies potential speech onsets, followed by a more complex neural network classifier that processes only these candidate segments. This segmentation allows accurate detection while keeping the always-on portion simple.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An energy-based detection mechanism serves as an intermediary between the simple always-on sensor and the complex neural network classifier. This intermediary filters and pre-processes signals, reducing the computational burden on the complex algorithm while maintaining overall detection accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If the SOD activates the wake word detection algorithm frequently to ensure no speech is missed, then the detection reliability is improved, but the power consumption increases

Engineering Contradiction:
Improvedetection reliabilityVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system applies different quality levels of detection to different time periods and signal conditions. High-reliability neural network classification is applied only to candidate segments identified by the energy detector, while other periods use simpler or no detection, thus maintaining reliability where needed while reducing overall power consumption.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12537021B2Low complexity sub-band speech onset detection (SOD)
Publication Date: 2026.01.27 INFINEON TECHNOLOGIES AMERICAS CORP
  • US12537021B2 patent drawing
  • US12537021B2 patent drawing
  • US12537021B2 patent drawing

AI summary

Techniques are disclosed for a low-power and low-complexity speech onset detector (SOD) that uses a fractional-band filter structure and spectral subtraction technique to derive sub-band energy profiles to detect the onset of speech in the presence of noise. The SOD derives the sub-band energy profiles by filtering and down-sampling a full-band input audio signal using the fractional-bandwidth filter structure, which may be a low-pass filter with a cut-off frequency that is a fraction of the full bandwidth of the input signal. The SOD flexibly estimates the average noise energy across frames and the current frame speech energy in each sub-band to track noise and speech energy levels across the frames for each of the sub-bands to determine one or more band thresholds used to detect active speech. The sub-band energy profiles leverage any separation in frequency between noise and speech to detect the onset of speech in a target signal.