Audio Signal Classification via Frequency Sub-band Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio signal compression methods struggle to optimally classify and compress speech-like and music-like signals, as existing algorithms often require complex recognition methods and trade-offs between bit rate and quality, failing to efficiently handle mixed signal components.
Innovation Solution
The method divides audio signals into frequency sub-bands, analyzing energy level variations and relations between sub-bands to classify signals as speech-like or music-like, selecting appropriate excitation blocks for efficient compression, thereby improving sound quality without significantly affecting compression efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If complex recognition methods are used to classify speech-like and music-like signals, then classification accuracy is improved, but device complexity increases
Solution Approach 1:
The audio signal is divided into multiple frequency sub-bands, and classification is performed separately for each sub-band rather than treating the entire spectrum uniformly. This segmentation allows simpler local decisions to achieve better overall classification accuracy while reducing computational complexity compared to analyzing the full signal with complex algorithms.
Solution Approach 2:
Different excitation methods are selected for different frequency sub-bands based on local signal characteristics. Speech-like sub-bands use one excitation method while music-like sub-bands use another, allowing optimal local processing without requiring complex global classification of the entire signal.
2Productivity
If different compression algorithms are used for speech and music signals, then compression efficiency is improved, but excitation selection complexity increases
Solution Approach 1:
The signal is segmented into frequency sub-bands, and excitation selection is performed independently for each sub-band based on simple energy level comparisons rather than requiring complex global analysis. This enables efficient compression with reduced selection complexity.
Solution Approach 2:
The classification decision is based on changes in energy levels across frequency sub-bands rather than absolute energy values or complex spectral features. This parameter transformation simplifies the selection process while maintaining effective differentiation between speech-like and music-like segments.
3Measurement precision
If frequency sub-band analysis is performed to classify signals, then classification accuracy for mixed signals is improved, but processing power requirements increase
Solution Approach 1:
Dividing the frequency spectrum into sub-bands allows parallel processing of simpler tasks rather than one complex sequential analysis, improving classification accuracy for mixed signals while distributing computational load to reduce peak processing power requirements.
Solution Approach 2:
Each frequency sub-band is analyzed using simple energy level measurements rather than complex global signal analysis. This local approach achieves accurate classification of mixed speech and music components with significantly reduced processing power compared to comprehensive signal analysis methods.
Data Source
AI summary
An encoder comprising an input for inputting frames of an audio signal in a frequency band, at least a first excitation block for performing a first excitation for a speech like audio signal, and a second excitation block for performing a second excitation for a non-speech like audio signal. The encoder further comprises a filter for dividing the frequency band into a plurality of sub bands each having a narrower bandwidth than the frequency band. The encoder also comprises an excitation selection block for selecting one excitation block among the at least first excitation block and the second excitation block for performing the excitation for a frame of the audio signal on the basis of the properties of the audio signal at least at one of the sub bands. The invention also relates to a device, a system, a method and a storage medium for a computer program.


