Robust audio watermarking method based on adaptive quantization strategy and feature classification

A robust audio watermarking method based on adaptive quantization strategy and feature classification solves the problem of adaptability of existing technologies in complex channel environments, and achieves highly robust and low-impact audio watermark embedding and extraction, adapting to nonlinear distortion attacks while maintaining audio quality.

CN121506155APending Publication Date: 2026-02-10XINYANG NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511706059.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing audio watermarking technologies have limited adaptability in complex channel environments, especially against nonlinear distortion attacks such as re-recording, and are difficult to effectively resist various signal processing attacks while ensuring audio quality.

Method used

A robust audio watermarking method employing adaptive quantization strategy and feature classification is proposed. A multi-dimensional frame feature vector is constructed using DWT-CLM features, zero-crossing rate, variance, and energy. The quantization step size is dynamically adjusted using a Sigmoid classifier, and the embedding strength is flexibly adjusted. Watermark embedding and extraction are achieved by combining discrete wavelet transform.

Benefits of technology

In complex channel environments, the watermark signal exhibits good signal-to-noise ratio and subjective distinguishability, with a low bit error rate. It can effectively resist attacks such as MP3 compression, resampling, low-pass filtering, and re-recording, ensuring accurate extraction of copyright information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506155A_ABST
    Figure CN121506155A_ABST
Patent Text Reader

Abstract

The invention discloses a robust audio watermarking method based on an adaptive quantization strategy and feature classification, and belongs to the technical field of digital audio copyright protection. The method comprises the following steps: framing an audio signal, extracting a logarithmic mean feature (DWT-CLM) of discrete wavelet transform, and combining a zero-crossing rate, a variance and energy to form a frame feature vector; a Sigmoid classifier is used to discriminate frame characteristics, and a fixed or variable quantization step size is adaptively selected to embed watermark bits into approximate components; and during extraction, the watermark is accurately extracted through the same feature analysis and classifier discrimination recovery quantization mode. Experiments show that the algorithm has better inaudible property and robustness, can effectively resist attacks such as MP3 compression, resampling, low-pass filtering and re-recording, and is suitable for digital audio copyright protection scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of digital audio copyright protection technology, and mainly relates to a robust audio watermarking method based on adaptive quantization strategy and feature classification. Background Technology

[0002] With the rapid development of network technology and digital signal processing technology, digital audio has become an important medium for internet information dissemination and content distribution, widely used in music platforms, short videos, podcasts, and online education. However, audio copyright infringement is becoming increasingly serious. Users can easily download, transcode, and re-upload audio resources, and even bypass copyright detection mechanisms, leading to frequent unauthorized copying and dissemination.

[0003] Digital watermarking technology has been widely researched and applied to achieve copyright protection and content authentication for digital audio. This technology embeds copyright information, identification marks, or traceability data into audio signals without affecting the perceived quality of the audio, thus enabling information hiding and traceability.

[0004] In existing technologies, the paper "A discrete wavelet transform-based audio watermarking technique for digital security" (V. Prabhu, L. Sundar, International Journal of Nonlinear Dynamics and Control, vol.2, no.2, pp.183–198, 2022) proposes an audio watermarking technique based on discrete wavelet transform, embedding watermark information into low-frequency subbands to enhance resistance to common attacks such as compression and filtering. It achieves high embedding capacity and low bit error rate while maintaining audio quality. However, it employs a fixed embedding strategy and fails to dynamically adjust the embedding strength according to the local characteristics of the audio frame, limiting its adaptability to complex channel environments (such as re-recording). The paper "A novel audio watermarking scheme based on fuzzy inference system in DCT domain" (M. Mosleh, S. Setayeshi, Multimedia Tools and The paper "Applications, vol.80, no.13, pp.20423–20447, 2021" proposes an audio watermarking method based on a fuzzy inference system. This method can resist common linear signal processing attacks such as compression and filtering, but its adaptability to nonlinear distortion attacks such as re-recording is limited. Summary of the Invention

[0005] To ensure that audio watermarking technology for copyright protection meets the needs of the new era and has better trainability and generalization ability, this invention provides a robust audio watermarking method based on adaptive quantization strategy and feature classification.

[0006] The technical solution adopted by this invention to solve its technical problem is as follows:

[0007] A robust audio watermarking method based on adaptive quantization strategy and feature classification includes a watermark embedding step S1 and a watermark extraction step S2.

[0008] The watermark embedding step S1 includes:

[0009] S11) Divide the carrier signal A into N points per frame to obtain the frame set. ,in The total number of frames; the i-th frame is denoted as . The length is N; Divided into two parts, denoted as follows: and The lengths are all ;

[0010] S12) respectively for and Perform a D-level discrete wavelet transform to obtain the D-level approximate component. and ; calculated from approximate components and The DWT-CLM features are denoted as and , , Where K represents the number of coefficients of the approximate component;

[0011] S13) Record For the watermark information to be embedded, where Take the difference feature Simultaneously calculate the zero-crossing rate ,variance and energy Constructing four-dimensional vector features ;

[0012] S14) Calculate the classification probability of the current frame based on the sigmoid classifier weights W and bias b. Based on probability values With threshold Compare and select the quantization step size; if Then step size Preset fixed step size Otherwise, step size ,in As an adaptive factor;

[0013] S15) Based on the current quantization results ,like Then adjust Make it satisfy Calculate the scaling factor Update the low-frequency approximation coefficients; otherwise, if If the embedding condition is met, no modification is needed.

[0014] S16) Perform inverse DWT on the modified approximate components and the original detail coefficients to complete the process. The watermark signal of the i-th frame is obtained by embedding the watermark. ;

[0015] S17) Repeat steps S11-S16 above to embed the watermark signal into other audio frames to obtain the watermarked audio signal.

[0016] The watermark extraction step S2 includes:

[0017] S21) The watermarked signal The frame is divided into frames of length N, and the i-th frame is denoted as . ,Will Divide into two equal parts, and denote them as follows: and ;

[0018] S22) Calculation and The DWT-CLM features are denoted as and And calculate the residuals ;

[0019] S23) Calculate the feature vector of the current frame. Step size is determined based on the sigmoid classifier. ,Depend on and Extract watermark , ;

[0020] S24) Repeat steps S21-S23 above to obtain complete watermark information.

[0021] Due to the adoption of the technical solution described above, the present invention has the following advantages:

[0022] 1) The adaptive quantization strategy enhances the balance between robustness and perceived quality. A sigmoid classifier is used to discriminate audio frame characteristics based on DWT-CLM features, zero-crossing rate, variance, and energy, dynamically selecting a fixed or variable quantization step size. For audio frames with different characteristics, the embedding strength is flexibly adjusted. In frames susceptible to interference, the step size is controlled by an adaptive factor, ensuring the stability of the watermark embedding while avoiding excessive modification that could affect audio perceived quality. This solves the problem of limited adaptability of existing fixed embedding strategies to complex channel environments (such as re-recording).

[0023] 2) Feature classification enhances the adaptability of the method. By combining the log-mean feature of discrete wavelet transform (DWT-CLM) with zero-crossing rate, variance, and energy to construct a multi-dimensional frame feature vector, a comprehensive characterization of the local properties of audio frames is achieved. The Sigmoid classifier achieves accurate discrimination based on these features, enabling the method to dynamically adjust the embedding strategy according to the specific characteristics of the frame. Compared with existing methods designed only for linear attacks, it has better adaptability to nonlinear distortion attacks such as re-recording.

[0024] 3) Experimental data show that the average signal-to-noise ratio (SNR) of the watermarked audio signal is 29.95 and the average subjective discrimination (SDG) is -0.61, indicating that the watermark embedding has a minimal impact on audio quality and the difference is difficult for the human ear to perceive, thus meeting the requirements for audio perception quality in practical applications.

[0025] 4) The bit error rate (BER) of watermark extraction is low in the face of common attacks such as MP3 compression (64kb / s, 128kb / s), resampling, low-pass filtering and re-recording. This shows that the method can effectively resist various signal processing and desynchronization attacks, and ensure that copyright information can still be accurately extracted in complex environments, providing reliable technical support for digital audio copyright protection. Attached Figure Description

[0026] Figure 1 A flowchart of the carrier signal framing and watermark embedding process; Figure 2 This is a flowchart of the audio watermark extraction process. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0028] A robust audio watermarking method based on adaptive quantization strategy and feature classification is characterized by including a watermark embedding step S1 and a watermark extraction step S2.

[0029] Watermark embedding step S1 includes:

[0030] S11) Carrier signal preprocessing:

[0031] (1) Framing and segmentation. The carrier signal is divided into frames and segments. Divide the data evenly into N points per frame to obtain a frame set. ,in Let i be the total number of frames. The i-th frame is denoted as... , with a length of N.

[0032] (2) Divided into two parts, denoted as follows: and The lengths are all .

[0033] S12) Feature extraction:

[0034] (1) For each and Performing a D-level Discrete Wavelet Transform (DWT) yields the D-level approximate components. and .

[0035] (2) Calculated from approximate components and The DWT-CLM features are denoted as and ,

[0036] ,

[0037] Where K represents the number of coefficients of the approximate component.

[0038] S13-16) Watermark Embedding

[0039] remember For the watermark information to be embedded, where .Will Embedded into In the middle. Take the difference feature. Simultaneously calculate the zero-crossing rate ,variance and energy Constructing four-dimensional vector features .

[0040] (1) Calculate the classification probability of the current frame based on the sigmoid classifier weights W and bias b. As shown in the following formula

[0041]

[0042] Based on probability values With threshold The system compares and selects the quantization step size. The system employs the following decision strategy:

[0043]

[0044] Where T is the threshold setting. To preset a fixed step size; It is an adaptive factor used to dynamically control the step size, ensuring that embedding quality and perception balance are maintained in susceptible frames.

[0045] (2) Current quantification results ,like Then adjust Make it satisfy Calculate the scaling factor Update the low-frequency approximation coefficients. Otherwise, if If the embedding condition is met, no modification is needed.

[0046] (3) Perform inverse DWT on the modified approximate components and the original detail coefficients to complete the process. The embedding is obtained to get the first Watermark signal of the frame .

[0047] S17) Repeat steps S11-S16 above to embed the watermark signal into other audio frames, thus obtaining the watermarked audio signal. The flowchart of the carrier signal framing and watermark embedding process is as follows: Figure 1 As shown.

[0048] Watermark extraction step S2 includes:

[0049] S21) The watermarked signal The frame is divided into frames of length N, and the i-th frame is denoted as . .Will Divide into two equal parts, and denote them as follows: and .

[0050] S22) Calculation and The DWT-CLM features are denoted as and And calculate the residuals .

[0051] S23) Calculate the feature vector of the current frame. Step size is determined based on the sigmoid classifier. .Depend on and Extract watermark As shown in the following formula:

[0052]

[0053] S24) Following steps S21-S23 above, extract the watermark from the remaining audio frames to obtain the complete watermark information. The audio watermark extraction process flowchart is as follows: Figure 2 As shown.

[0054] The effectiveness of the method of this invention can be verified through the following performance analysis:

[0055] Twenty-one audio signals of different types, each with 16-bit quantization and a sampling frequency of 44.1kHz, were randomly selected from a sample library. First, watermarks were embedded into the selected audio signals. Then, the watermarked signals were subjected to a re-recording attack using three different recording devices: a SONY PCM-D100 voice recorder, a Huawei P40, and an iPhone 14.

[0056] 1. Inaudibility

[0057] The inaudibility of the watermark was tested using Subjective Difference Grades (SDG) and Signal-to-Noise Ratio (SNR). Table 1 shows the maximum, minimum, and average SNR and SDG values ​​for 200 watermarked signal tests. The results in Table 1 demonstrate that this method achieves good inaudibility.

[0058] Table 1 Maximum value Minimum value mean SNR 32.40 27.94 29.95 SDG -0.87 -0.45 -0.61

[0059] Experimental data show that the average signal-to-noise ratio (SNR) of the watermarked audio signal is 29.95, and the average subjective discrimination (SDG) is -0.61, indicating that the watermark embedding has minimal impact on audio quality and the difference is difficult for the human ear to perceive, thus meeting the requirements for audio perception quality in practical applications.

[0060] 2. Robustness

[0061] The robustness of this method is tested using bit error rates (BER). A lower BER indicates a lower error rate in watermark extraction, and thus better robustness.

[0062] The BER values ​​of the watermarked signal after signal processing and re-recording attack are shown in Table 2. Signal processing included MP3 compression (64kb / s and 128kb / s), resampling (44.1→20.05→44.1 kHz), and low-pass filtering (16kHz). The BER values ​​of the watermarked signal after the attack are shown in Table 2.

[0063] Table 2 Attack type BER MP3 compression (64 kb / s) 2 MP3 compression (128 kb / s) 4 Resampling (44.1→20.05→44.1 kHz) 4 Low-pass filter (16 kHz) 0 Re-recording 5

[0064] As can be seen from the BER values ​​listed in Table 2, the BER value of the watermark extracted after the attack is relatively small, indicating that the method has good robustness against signal processing and re-recording attacks.

[0065] Faced with common attacks such as MP3 compression (64kb / s, 128kb / s), resampling, low-pass filtering, and re-recording, the watermark extraction has a low bit error rate, indicating that the method can effectively resist various signal processing and desynchronization attacks, ensuring that copyright information can still be accurately extracted in complex environments, and providing reliable technical support for digital audio copyright protection.

[0066] The parts not detailed above are existing technologies and therefore have not been described in detail.

Claims

1. A robust audio watermarking method based on adaptive quantization strategy and feature classification, characterized in that, This includes a watermark embedding step S1 and a watermark extraction step S2; The watermark embedding step S1 includes: S11) Divide the carrier signal A into N points per frame to obtain the frame set. ,in The total number of frames; the i-th frame is denoted as . The length is N; Divided into two parts, denoted as follows: and The lengths are all ; S12) respectively for and Perform a D-level discrete wavelet transform to obtain the D-level approximate component. and ; calculated from approximate components and The DWT-CLM features are denoted as and , , Where K represents the number of coefficients of the approximate component; S13) For the watermark information to be embedded, where Take the difference feature Simultaneously calculate the zero-crossing rate ,variance and energy Constructing four-dimensional vector features ; S14) Calculate the classification probability of the current frame based on the sigmoid classifier weights W and bias b. Based on probability values With threshold Compare and select the quantization step size; if Then step size Preset fixed step size Otherwise, step size ,in As an adaptive factor; S15) Based on the current quantization results ,like Then adjust Make it satisfy Calculate the scaling factor Update the low-frequency approximation coefficients; otherwise, if If the embedding condition is met, no modification is needed. S16) Perform inverse DWT on the modified approximate components and the original detail coefficients to complete the process. The watermark signal of the i-th frame is obtained by embedding the watermark. ; S17) Repeat steps S11-S16 above to embed the watermark signal into other audio frames to obtain the watermarked audio signal. The watermark extraction step S2 includes: S21) The watermarked signal The frame is divided into frames of length N, and the i-th frame is denoted as . ,Will Divide into two equal parts, and denote them as follows: and ; S22) Calculation and The DWT-CLM features are denoted as and And calculate the residuals ; S23) Calculate the feature vector of the current frame. Step size is determined based on the sigmoid classifier. ,Depend on and Extract watermark , ; S24) Repeat steps S21-S23 above to obtain complete watermark information.