Reversible Audio Data Hiding via Variance-Based Partitioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing reversible audio data hiding techniques face limitations in embedding capacity and quality degradation, particularly in the spectral and compressed data domains, with insufficient space for payload and noticeable distortion in audio quality.
Innovation Solution
A method involving intelligent partitioning of 16-bit audio data into two portions based on variance calculation, using a generalized integer transform for data hiding and extraction, allowing for efficient embedding and restoration of audio while maintaining acceptable quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If data embedding is carried out directly on the audio waveform using waveform domain methods, then computation is reduced and simplicity is improved, but embedding capacity is limited
Solution Approach 1:
The audio signal is divided into multiple non-overlapping segments, and each segment is processed independently through integer transform and prediction error calculation. This segmentation enables parallel processing and increases overall embedding capacity while maintaining computational efficiency at the segment level.
Solution Approach 2:
The invention transitions from direct waveform domain embedding to spectral domain embedding by applying integer transform (e.g., intDCT) to convert time-domain audio signals into frequency-domain representations. This dimensional transformation opens up additional embedding spaces in the spectral coefficients while maintaining reversibility.
2Reliability
If integer transform and amplitude expansion are used in spectral domain for tamper detection, then tamper detection capability is improved, but most space is occupied by overhead including feature value and positional data
Solution Approach 1:
The invention extracts and separates the location map data (overhead) from the payload data. The location map is embedded first in the spectral domain using amplitude expansion, and then the payload is embedded in the remaining capacity. This separation optimizes the use of embedding space by placing critical metadata in the more robust spectral domain.
Solution Approach 2:
Different parts of the audio spectrum are treated differently based on their embedding suitability. High-frequency spectral coefficients with higher variance are selected for payload embedding, while lower-frequency components are used for location map and feature values. This local differentiation maximizes overall embedding capacity while maintaining audio quality.
3Quantity of substance
If compressed data domain methods with linear prediction model are used, then embedding space is provided by compressing unimportant parameters, but embedding capacity is limited
Solution Approach 1:
The invention changes the embedding domain from compressed time-domain parameters to transformed frequency-domain coefficients. By applying integer transform and using prediction error expansion in the spectral domain, the system achieves higher embedding capacity because spectral coefficients provide more variability and embedding opportunities compared to compressed time-domain representations.
Data Source
AI summary
The present invention provides a method of reversible audio data hiding. The method of data hiding and restoring comprises the steps of: protecting audio by embedding information into the audio according to variance calculation associated to the audio, wherein the quality of the protected audio is degraded after embedding the information into the audio; publishing the protected audio widely as a trial for listen version; and decoding the protected audio for a user who purchased the copyright of the audio by extracting the original audio from the protected audio.


