Audio Signal Classification via Frequency Sub-band Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio signal compression methods struggle to optimally classify and compress speech-like and music-like signals, as existing algorithms often require complex recognition methods and trade-offs between bit rate and quality, failing to efficiently handle mixed signal components.

Innovation Solution

The method divides audio signals into frequency sub-bands, analyzing energy level variations and relations between sub-bands to classify signals as speech-like or music-like, selecting appropriate excitation blocks for efficient compression, thereby improving sound quality without significantly affecting compression efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If complex recognition methods are used to classify speech-like and music-like signals, then classification accuracy is improved, but device complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidalgorithm complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio signal is divided into multiple frequency sub-bands, and classification is performed separately for each sub-band rather than treating the entire spectrum uniformly. This segmentation allows simpler local decisions to achieve better overall classification accuracy while reducing computational complexity compared to analyzing the full signal with complex algorithms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different excitation methods are selected for different frequency sub-bands based on local signal characteristics. Speech-like sub-bands use one excitation method while music-like sub-bands use another, allowing optimal local processing without requiring complex global classification of the entire signal.

Inventive Principle:
Principle #3Local quality

2Productivity

If different compression algorithms are used for speech and music signals, then compression efficiency is improved, but excitation selection complexity increases

Engineering Contradiction:
Improvecompression efficiencyVSAvoidexcitation selection complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The signal is segmented into frequency sub-bands, and excitation selection is performed independently for each sub-band based on simple energy level comparisons rather than requiring complex global analysis. This enables efficient compression with reduced selection complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The classification decision is based on changes in energy levels across frequency sub-bands rather than absolute energy values or complex spectral features. This parameter transformation simplifies the selection process while maintaining effective differentiation between speech-like and music-like segments.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If frequency sub-band analysis is performed to classify signals, then classification accuracy for mixed signals is improved, but processing power requirements increase

Engineering Contradiction:
Improvesignal classification accuracyVSAvoidprocessing power requirements
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

Dividing the frequency spectrum into sub-bands allows parallel processing of simpler tasks rather than one complex sequential analysis, improving classification accuracy for mixed signals while distributing computational load to reduce peak processing power requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each frequency sub-band is analyzed using simple energy level measurements rather than complex global signal analysis. This local approach achieves accurate classification of mixed speech and music components with significantly reduced processing power compared to comprehensive signal analysis methods.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS8438019B2Classification of audio signals
Publication Date: 2013.05.07 NOKIA TECHNOLOGIES OY
  • US8438019B2 patent drawing
  • US8438019B2 patent drawing
  • US8438019B2 patent drawing

AI summary

An encoder comprising an input for inputting frames of an audio signal in a frequency band, at least a first excitation block for performing a first excitation for a speech like audio signal, and a second excitation block for performing a second excitation for a non-speech like audio signal. The encoder further comprises a filter for dividing the frequency band into a plurality of sub bands each having a narrower bandwidth than the frequency band. The encoder also comprises an excitation selection block for selecting one excitation block among the at least first excitation block and the second excitation block for performing the excitation for a frame of the audio signal on the basis of the properties of the audio signal at least at one of the sub bands. The invention also relates to a device, a system, a method and a storage medium for a computer program.