Audio Transient Detection via Block Norm Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio signal processing techniques face challenges in accurately identifying transients within audio signals, leading to significant audible distortion due to poor time resolution when applying long transforms, and inadequate frequency resolution when using short transforms.
Innovation Solution
A multi-stage technique that compares maximum block norm values to determine the presence of transients by dividing audio signals into blocks, calculating norm values, and applying test criteria to identify significant signal strength changes across blocks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a long transform interval is applied to the entire frame, then good frequency resolution is achieved, but poor time resolution results in significant audible distortion during transients
Solution Approach 1:
The audio frame is divided into multiple subframes or blocks, allowing different transform lengths to be applied to different segments. This segmentation enables the system to use long transforms for quasi-stationary portions (good frequency resolution) and short transforms for transient portions (good time resolution), thereby resolving the contradiction between frequency and time resolution.
Solution Approach 2:
The transform interval length is made dynamic rather than fixed, adapting to the local characteristics of the audio signal. By detecting transient events and adjusting the transform window length accordingly, the system achieves optimal frequency resolution for stationary segments and optimal time resolution for transient segments, resolving the static contradiction between these two parameters.
2Loss of time
If a short transform interval is used to improve time resolution during transients, then time resolution is improved, but frequency resolution deteriorates
Solution Approach 1:
Different quality characteristics (transform lengths) are applied to different local regions of the audio signal based on their specific needs. Transient regions receive short transforms for high time resolution, while quasi-stationary regions receive long transforms for high frequency resolution. This local adaptation resolves the contradiction by allowing each region to have optimized parameters suited to its characteristics.
Solution Approach 2:
The transform interval is dynamically adjusted based on the detected signal characteristics in each region. The system transitions between long and short transforms depending on whether a transient is detected, making the frequency and time resolution properties dynamic rather than fixed, thereby resolving the contradiction adaptively.
3Extent of automation
If conventional transient detection methods are used, then transient identification is achieved, but accuracy is insufficient leading to inadequate processing decisions
Solution Approach 1:
The system uses feedback from multiple detection criteria and signal characteristics to improve transient detection accuracy. By evaluating multiple parameters (energy changes, spectral changes, temporal patterns) and using this feedback to make processing decisions, the system achieves more accurate transient identification than conventional single-criterion methods, resolving the contradiction between automation and precision.
Solution Approach 2:
The transient detection mechanism combines multiple detection criteria and signal analysis methods into a composite detection system. Rather than relying on a single simple criterion, the system integrates multiple indicators (amplitude changes, frequency changes, temporal patterns) to make more accurate transient detection decisions, thereby resolving the contradiction between automated detection and detection accuracy.
Data Source
AI summary
Provided are, among other things, systems, methods and techniques for detecting whether a transient exists within an audio signal. According to one representative embodiment, a segment of a digital audio signal is divided into blocks, and a norm value is calculated for each of a number of the blocks, resulting in a set of norm values for such blocks, each such norm value representing a measure of signal strength within a corresponding block. A maximum norm value is then identified across such blocks, and a test criterion is applied to the norm values. If the test criterion is not satisfied, a first signal indicating that the segment does not include any transient is output, and if the test criterion is satisfied, a second signal indicating that the segment includes a transient is output. According to this embodiment, the test criterion involves a comparison of the maximum norm value to a different second maximum norm value, subject to a specified constraint, within the segment.


