Audio Onset Detection via Multi-Band Spectrum Parameters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio onset detection methods fail to accurately detect note and syllable onsets in complex audio signals with weak rhythms, leading to frequent false and missing detections.

Innovation Solution

The method determines first and second speech spectrum parameters for frequency bands in a frequency domain signal, calculating mean values and differences to accurately identify onsets, reducing false and missing detections by referencing multiple frequency bands for more accurate second speech spectrum parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech spectrum parameter detection method is used, then detection simplicity is maintained, but detection accuracy deteriorates for complex audio signals with weak rhythms

Engineering Contradiction:
Improveonset detection accuracyVSAvoiddetection method complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio signal is divided into multiple frequency bands, and speech spectrum parameters are calculated separately for each band. This segmentation allows the system to capture local spectral characteristics that are sensitive to onsets, thereby improving detection accuracy for complex audio signals while maintaining manageable computational complexity through parallel processing of each band.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The method transitions from analyzing a single speech spectrum parameter to analyzing multiple parameters across different frequency bands. By adding the frequency band dimension, the system gains additional information about spectral changes that occur during onsets, improving accuracy without proportionally increasing overall system complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If single frequency band parameter is used, then computational complexity is reduced, but false detection and missing detection increase

Engineering Contradiction:
Improvedetection reliabilityVSAvoidparameter calculation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The method combines speech spectrum parameters from multiple frequency bands to make onset detection decisions. By merging information across bands, the system achieves more reliable detection with reduced false positives and missed detections, as the combined parameters provide a more robust indication of actual onsets versus noise or artifacts.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system calculates speech spectrum parameters for each frequency band and uses these multiple parameters as feedback to improve detection reliability. The comparative analysis of parameters across bands provides feedback that helps distinguish true onsets from false indicators, thereby reducing both false detection and missing detection.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12119023B2Audio onset detection method and apparatus
Publication Date: 2024.10.15 DOUYIN VISION CO LTD
  • US12119023B2 patent drawing
  • US12119023B2 patent drawing
  • US12119023B2 patent drawing

AI summary

An audio onset detection method and apparatus, an electronic device, and a computer readable storage medium. The audio onset detection method comprises: determining a first voice frequency spectrum parameter corresponding to each frequency band according to a frequency domain signal corresponding to an audio signal of an audio; for each frequency band, determining a second voice frequency spectrum parameter of a current frequency band according to the first voice frequency spectrum parameter of the current frequency band and the first voice frequency spectrum parameters of frequency bands positioned before the current frequency band according to a time sequence; and determining one or more onset positions of notes and syllables in the audio according to the second voice frequency spectrum parameters corresponding to the frequency bands.