Tampering Audio Detection via Wavelet Transform and Mel Cepstrum Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for detecting tampering audio are limited in application scenarios and may not be effective for all audio formats, particularly those that have undergone secondary compression.

Innovation Solution

A method involving wavelet transforms, Mel cepstrum feature calculation, and deep learning models is employed to detect tampering audio by acquiring signals, performing wavelet and inverse wavelet transforms, calculating Mel cepstrum features, and concatenating them for detection using a trained deep learning model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing methods for detecting tampering audio are used, then detection can be performed on certain audio formats, but the application scenarios are limited and cannot be used in all scenarios

Engineering Contradiction:
Improveapplication scenario coverageVSAvoiddetection effectiveness
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent employs wavelet transform to decompose audio signals into multiple frequency sub-bands, enabling the detection system to process and analyze different audio formats and compression types uniformly. This multi-functional approach allows the same detection methodology to be applied across various audio scenarios including but not limited to secondarily compressed audio, thereby expanding application scope while maintaining reliability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transforms audio signals from the time domain to the frequency domain through wavelet transform, and further processes them through Mel cepstrum transformation. By changing the parameter representation of audio signals (from raw waveforms to frequency spectrum features), the system can effectively detect tampering across diverse audio formats and compression levels that were previously incompatible with fixed detection methods

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If wavelet transform and deep learning model are used for tampering audio detection, then applicability is expanded to broader scenarios, but the processing complexity increases

Engineering Contradiction:
Improvescenario applicabilityVSAvoidsignal processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent divides the audio signal processing into distinct stages: wavelet transform decomposition into frequency sub-bands, selective inverse wavelet transform on specific coefficients, Mel cepstrum feature extraction, and deep learning-based classification. This segmentation of the processing pipeline manages complexity by breaking down the overall task into manageable, specialized modules that can be implemented and optimized independently

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11636871B2Method and electronic apparatus for detecting tampering audio, and storage medium
Publication Date: 2023.04.25 INST OF AUTOMATION CHINESE ACAD OF SCI
  • US11636871B2 patent drawing
  • US11636871B2 patent drawing
  • US11636871B2 patent drawing

AI summary

Disclosed are a method, an electronic apparatus for detecting tampering audio and a storage medium. The method includes: acquiring a signal to be detected, and performing a wavelet transform of a first preset order on the signal to be detected so as to obtain a first low-frequency coefficient and a first high-frequency coefficient corresponding to the signal to be detected, the number of which is equal to that of the first preset order; performing an inverse wavelet transform on the first high-frequency coefficient having an order greater than or equal to a second preset order so as to obtain a first high-frequency component signal corresponding to the signal to be detected; calculating a first Mel cepstrum feature of the first high-frequency component signal in units of frame, and concatenating the first Mel cepstrum features of a current frame signal and a preset number of frame signals.