A Blind Detection Method for Digital Audio Watermarking
By setting thresholds, framed FFT conversion and frequency adjustment, blind detection of digital audio watermarks is achieved, which solves the problem of difficulty in resisting synchronous attacks in the prior art, and significantly improves the anti-attack performance and detection efficiency.
Patent Information
- Application Number
- CN202211625698.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-01
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-12-01
AI Technical Summary
The existing digital audio watermarking technology is difficult to effectively resist synchronization attacks, resulting in watermark detection not being performed at a pre-set embedded position, and synchronization needs to be restored before detection.
A blind detection method for digital audio watermark is proposed, which detects and restores the watermark start frame and end frame information through setting threshold values, framed FFT transformation, frequency adjustment and linear correlation calculation.
This method significantly improves the anti-attack performance, especially when fighting against variable speed attacks, it can detect watermark information more accurately and efficiently, which is suitable for situations where the modulation factor is large.
Smart Images

Figure CN115985328B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of digital audio watermarking, and specifically relates to a method for blind detection of digital audio. Background Art
[0002] An audio digital watermark is a piece of identification generated or embedded based on audio digital content. This identification can be used to identify the copyright ownership of audio digital content and can also be used to protect the integrity of audio digital content. When an audio watermark is used to identify the copyright ownership of audio digital content, during the entire life cycle of audio digital content generation, management, distribution, and use, the audio watermark is added to the audio digital content in an imperceptible and non-removable manner. Once a copyright dispute occurs, only by extracting the embedded watermark information from the audio digital content can the copyright ownership of the audio digital content be proven. Similar to a product's trademark indicating the manufacturer of the product, the role of an audio watermark is to indicate the copyright owner of the audio digital content, but this is usually not the same concept as the owner. When an audio watermark is used to protect the integrity of audio digital content, once the audio digital content is tampered with and a certain part changes, the corresponding audio watermark information will also change, thereby enabling the detection and location of the tampering.
[0003] PSM (Pitch Scale Modification) attack refers to changing the pitch of a speaker without changing the audio playback speed. This attack method does not change the overall playback duration of the audio, but the pitch change will also make it difficult to extract the watermark. Most audio watermark algorithms, such as spread-spectrum watermarks, are location-based, that is, the watermark is embedded in a specific position and then detected from that position. However, the displacement caused by a synchronization attack will cause the watermark detection to not be performed at the embedding position, so synchronization needs to be restored before detection. Currently, many algorithms have been developed in the field of image watermarking to resist geometric distortion. In contrast, there are much fewer audio watermark algorithms that resist synchronization attacks. The synchronization mechanism is the primary problem that needs to be solved in audio watermark technology. Audio signals are one-dimensional signals based on time. Synchronization attacks will cause changes in the position and order of audio data, resulting in the watermark not being at the pre-set embedding position. Therefore, synchronization needs to be restored before detection. Currently, several methods for resisting synchronization attacks are: exhaustive search, explicit synchronization, embedding redundant watermarks, constant watermarks, implicit synchronization, cyclic correlation, etc. Each synchronization method has its own advantages and disadvantages and needs to be specifically designed according to the specific watermark embedding and extraction scheme as much as possible. Currently, the problem of resisting pitch change attacks is still a major difficult problem in the current field of audio watermark technology (the reason is that it has a very strong destructive force on the synchronization structure of audio, but has little impact on the auditory effect of audio).
[0004] The characteristics that a digital audio watermarking algorithm should possess mainly include: 1) The watermark must be embedded into the host audio data, rather than stored in the file header or a separate file. 2) The watermark should not cause audible distortion to the sound quality of the original audio, that is, it should be transparent. 3) The watermark must have a certain degree of robustness and be able to resist general signal processing operations such as compression, filtering, resampling, requantization, cropping, and adding noise on the host audio signal. 4) The watermark should be easy to embed, and the computational complexity of extraction and detection should be low to facilitate its integration into general electronic products. 5) The watermark algorithm must have a certain synchronization mechanism to counter synchronization attacks in the time domain. 6) In principle, the detection of the watermark should not require the original audio, that is, blind detection should be achieved, because it is very difficult to find the original audio. The watermark algorithm should be public, and the security preferably depends on the key rather than the secrecy of the algorithm. Summary of the Invention
[0005] The object of the present invention is to solve at least the above problems in the current digital audio watermark detection, and propose a digital audio watermark blind detection method, which includes the following steps: S101: Set a threshold T value ; S102: Perform frame-by-frame FFT transformation on the audio signal to be measured to obtain F t (n), where the frame length is set to N, 1 ≤ n ≤ N, and the watermark information of the signal to be measured consists of a synchronization start frame AA + watermark + synchronization end frame BB; S103: Set a frequency adjustment value, use a PSM attack range of -10% to 10%, set a starting adjustment value and a frequency adjustment interval, calculate the corresponding cycle length, and convert the corresponding frequency value into α; S104: Calculate an adjustment scale factor β, where β = 1 / α; S105: Coarsely adjust the frequency, use linear correlation to calculate the corresponding synchronization start frame information AA' and synchronization end frame information BB', and calculate the corresponding value CC' after frequency adjustment according to the AA' and BB'; S106: Determine whether the loop is completed, that is, whether the calculation within the entire PSM attack range has been completed. If the loop is not completed, increase the adjustment scale factor according to the frequency adjustment interval, and jump to step S104. If it is completed, continue to the next step; S107: Arrange CC' and find the minimum value of CC' and the corresponding scale factor β1 value; S108: First-level detection of the watermark, determine whether the minimum value of CC' is less than or equal to the set threshold. If the minimum value of CC' is less than or equal to the set threshold T value , then perform an inverse FFT transformation on the spectrum corresponding to the obtained β1 value, and use linear correlation to calculate the linear correlation information between the corresponding start frame information and end frame information, which is the embedded watermark; if the minimum value of CC' is greater than the set threshold T value, then proceed to the next step; S109: Recalculate to obtain the adjustment ratio factor β. The initial parameter of β is based on the value of β1, the sampling interval is 1 / 10 of the original, and new loop counts are determined according to this interval. The adjustment range is also centered around the value of β1 and takes 1 / 10 of the original; S110: Fine-tune the frequency, use linear correlation to calculate the corresponding synchronization start frame information AA” and synchronization end frame information BB”, and calculate the corresponding value CC” after frequency adjustment based on the AA” and BB”; S111: Determine whether the loop is completed, that is, whether the calculations within the entire PSM attack range have been completed. If the loop is not completed, increase the adjustment ratio factor according to the frequency adjustment interval and jump to step S109. If completed, continue to the next step; S112: Arrange CC”, and find the minimum value and the corresponding ratio factor β2 value; S113: Perform the second-level detection of the watermark, and determine whether the minimum value of CC” is less than or equal to the set threshold T value , if the minimum value of CC” is less than or equal to the set threshold T value , then directly perform an inverse FFT on the spectrum corresponding to the obtained β2 value, and use linear correlation to calculate the linear correlation information between the corresponding start frame information and end frame information, which is the embedded watermark; if the minimum value of CC” is greater than the set threshold T value , then adjust the threshold T value , and restart step S102 until the final CC’ or CC” is less than or equal to the set threshold T value .
[0006] Further, the adjustment ratio factor β calculated in step S104 and step S109 is specifically as follows: First, it is indexed by a vector i F , where i F = 1:N; Another index vector i X is calculated by rounding the result of i F / β to the nearest integer; The value of i X determines the vector D t after frequency adjustment; Then, judge the relationship between β and 1. If β > 1, calculate the corresponding i X . If the nth element of i X is a unique value m, then D t (m) = F t (n); When there are repeated values in the elements of i X (assuming there are x + 1 repeated elements), that is, the values of the nth, (n + 1)th until the (n + x)th elements are equal, then If β < 1, the values of i X are discontinuous; If the nth element of i X is n, then D t (n) = F t(n); if i X the value of the n-th element of is m, and m ≠ n, then D t (n) = F t (n), and the other elements of the vector can be obtained by interpolation method.
[0007] Further, in step S105, the specific method for calculating the corresponding value CC' after frequency adjustment according to the AA' and BB' is: perform an exclusive OR operation on each bit of the AA', BB' and the AA, BB of the signal to be measured, that is Then add up these values as the corresponding value CC' after this frequency adjustment.
[0008] Further, in step S110, the specific method for calculating the corresponding value CC'' after frequency adjustment according to the AA'' and BB'' is: perform an exclusive OR operation on each bit of the AA'', BB'' and the AA, BB of the signal to be measured, that is Then add up these values as the corresponding value CC'' after this frequency adjustment.
[0009] Further, in step S113, adjust the threshold T value Specifically: adjust the threshold T value to 1.5 - 2 times of the set threshold.
[0010] Further, F in step S102 t (n) is the positive frequency part, that is F t (1:N / 2).
[0011] The blind detection method for digital audio proposed by the present invention has the following beneficial effects:
[0012] (1) The anti-attack performance of the present invention is significantly improved. Especially in the face of speed change attacks, it can more accurately and efficiently detect the watermark information, especially applicable to the situation with a relatively large pitch shift factor.
[0013] (2) After introducing frequency adjustment, compared with the exhaustive search method, this method can relatively quickly find the corresponding frequency domain adjustment scale factor according to the synchronization start frame information and synchronization end frame information through the combination of coarse adjustment and fine adjustment, so as to accurately obtain the embedded watermark information. Description of the Drawings
[0014] Figure 1 is the flowchart of a blind detection method for digital audio watermark of the present invention; Detailed Embodiments
[0015] The present invention will be further described in detail below with reference to the accompanying drawings.
[0016] The main different part of the present invention lies in the watermark detection part of digital audio. For the watermark embedding part of digital audio, it is basically similar to the existing methods, except that a watermark start frame and a watermark end frame are added before and after the original watermark respectively. The specific method of embedding in the time domain is as follows: The watermark information bits are used to perform spread-spectrum modulation on the pseudo-random noise (PN) sequence, that is, a watermark information bit is XOR-multiplied with the entire binary PN sequence to form a watermark signal, and then it is embedded into the audio signal sampling value by addition operation. Before embedding, the masking effect of HAS is used to shape the watermark signal to ensure that it is not perceptible. The watermark information consists of a synchronization start frame plus the watermark information plus a synchronization end frame, that is, AA (synchronization start frame) + normal watermark + BB (synchronization end frame), where AA is the synchronization start frame information and BB is the synchronization end frame information.
[0017] The following mainly introduces the blind detection method of the digital audio watermark of the present invention.
[0018] Refer to Figure 1 , the embodiment of the present invention discloses a flowchart of a digital audio blind detection method, and the specific operation steps are as follows.
[0019] S101: Set the threshold T value ;
[0020] S102: Perform frame-by-frame FFT transformation on the audio signal to be measured to obtain F t (n), where the frame length is set to N, 1 ≤ n ≤ N, and the watermark information of the signal to be measured consists of the synchronization start frame AA + watermark + synchronization end frame BB. It should be noted that considering the conjugate symmetry of the spectrum, only the positive frequency part of F t (n) is processed here, that is, F t (1:N / 2).
[0021] S103: Set the frequency adjustment value, use the PSM attack range of -10% to 10%, set the starting adjustment value and the frequency adjustment interval, calculate the corresponding cycle length, and convert the corresponding frequency value to α.
[0022] S104: Calculate the adjustment scale factor β, where β = 1 / α.
[0023] S105: Coarsely adjust the frequency, use linear correlation to calculate the corresponding synchronization start frame information AA' and synchronization end frame information BB', and calculate the corresponding value CC' after frequency adjustment according to the AA' and BB'.
[0024] S106: Determine whether the loop is completed, that is, whether the calculations within the entire PSM attack range have been completed. If the loop is not completed, increase the adjustment ratio factor according to the frequency adjustment interval and jump to step S104. If completed, proceed to the next step.
[0025] S107: Arrange CC’ and find the minimum value of CC’ and the corresponding ratio factor β1 value.
[0026] S108: First-level detect the watermark. Determine whether the minimum value of CC’ is less than or equal to the set threshold. If the minimum value of CC’ is less than or equal to the set threshold T value , then perform an inverse FFT on the spectrum corresponding to the obtained β1 value, and use linear correlation to calculate the linear correlation information between the corresponding start frame information and end frame information, which is the embedded watermark; if the minimum value of CC’ is greater than the set threshold T value , then proceed to the next operation.
[0027] S109: Recalculate to obtain the adjustment ratio factor β. The initial parameter of β is based on the β1 value, the sampling interval is 1 / 10 of the original, determine the new number of loops according to this interval, and the adjustment range is also centered on the β1 value, taking 1 / 10 of the original.
[0028] S110: Fine-tune the frequency, use linear correlation to calculate the corresponding synchronous start frame information AA” and synchronous end frame information BB”, and calculate the corresponding value CC” after frequency adjustment according to the AA” and BB”.
[0029] S111: Determine whether the loop is completed, that is, whether the calculations within the entire PSM attack range have been completed. If the loop is not completed, increase the adjustment ratio factor according to the frequency adjustment interval and jump to step S109. If completed, proceed to the next step.
[0030] S112: Arrange CC”, find the minimum value and the corresponding ratio factor β2 value.
[0031] S113: Second-level detect the watermark. Determine whether the minimum value of CC” is less than or equal to the set threshold T value , if the minimum value of CC” is less than or equal to the set threshold T value , then directly perform an inverse FFT on the spectrum corresponding to the obtained β2 value, and use linear correlation to calculate the linear correlation information between the corresponding start frame information and end frame information, which is the embedded watermark; if the minimum value of CC” is greater than the set threshold T value , then adjust the threshold T value , restart step S102 until the final CC’ or CC” is less than or equal to the set threshold T valueNote that the threshold T is adjusted here value Specifically, the threshold T value is adjusted to 1.5 to 2 times the originally set threshold.
[0032] In this embodiment, the adjusted proportionality factor β calculated in steps S104 and S109 is specifically as follows: First, index with a vector i F where i F = 1:N; Another index vector i X is calculated by rounding the result of i F / β to the nearest integer; The value of i X determines the vector D t after frequency adjustment; Then, judge the relationship between β and 1. If β > 1, calculate the corresponding i X . If the nth element n of i X is a unique value m, then D t (m) = F t (n); When there are repeated values in the elements of i X (assuming there are x + 1 repeated elements), that is, the values of the nth, (n + 1)th until the (n + x)th elements are equal, then If β < 1, the values of i X are discontinuous; If the nth element value of i X is n, then D t (n) = F t (n); If the nth element value of i X is m and m ≠ n, then D t (n) = F t (n), and the other elements of the vector can be obtained by interpolation.
[0033] In this embodiment, in step S105, calculating the corresponding value CC' after frequency adjustment according to the AA' and BB' is specifically as follows: Exclusive-OR each bit of the AA', BB' and the AA, BB of the signal to be measured, that is Then add up these values as the corresponding value CC' after this frequency adjustment.
[0034] In this embodiment, in step S110, calculating the corresponding value CC'' after frequency adjustment according to the AA'' and BB'' is specifically as follows: Exclusive-OR each bit of the AA'', BB'' and the AA, BB of the signal to be measured, that is Then add up these values as the corresponding value CC'' after this frequency adjustment.
[0035] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A blind detection method for digital audio watermark, characterized in that, Including the following steps: S101: Set the threshold value T value ; S102: Perform frame-by-frame FFT transformation on the audio signal to be measured to obtain F t (n), where the frame length is set to N, 1 ≤ n ≤ N, and the watermark information of the signal to be measured consists of a synchronization start frame AA + a watermark + a synchronization end frame BB; S103: Set the frequency adjustment value, use the PSM attack range of -10% to 10%, set the starting adjustment value and the frequency adjustment interval, calculate the corresponding cycle length, and convert the corresponding frequency value to α; S104: Calculate the adjustment scale factor β, where β = 1 / α; S105: Coarse-tune the frequency, use linear correlation to calculate the corresponding synchronization start frame information AA' and synchronization end frame information BB', and calculate the corresponding value CC' after frequency adjustment according to the AA' and BB'; S106: Determine whether the cycle is completed, that is, whether the calculation within the entire PSM attack range has been completed. If the cycle is not completed, increase the adjustment scale factor according to the frequency adjustment interval, and jump to step S104. If it is completed, continue to the next step; S107: Arrange CC' and find the minimum value of CC', and the corresponding scale factor β1 value; S108: Detect the first-level watermark, and determine whether the minimum value of CC’ is less than or equal to the set threshold. If the minimum value of CC’ is less than or equal to the set threshold T value , then perform an inverse FFT on the spectrum corresponding to the obtained β1 value, and use linear correlation to calculate the linear correlation information between the corresponding start frame information and end frame information, which is the embedded watermark; if the minimum value of CC’ is greater than the set threshold T value , then proceed to the next step; S109: Recalculate the adjustment scale factor β. The initial parameter of β is based on the β1 value, the sampling interval is 1 / 10 of the original, determine the new number of cycles according to this interval, and the adjustment range is also centered on the β1 value and takes 1 / 10 of the original; S110: Fine-tune the frequency, use linear correlation to calculate the corresponding synchronization start frame information AA'' and synchronization end frame information BB'', and calculate the corresponding value CC'' after frequency adjustment according to the AA'' and BB''; S111: Determine whether the cycle is completed, that is, whether the calculation within the entire PSM attack range has been completed. If the cycle is not completed, increase the adjustment scale factor according to the frequency adjustment interval, and jump to step S109. If it is completed, continue to the next step; S112: Arrange CC'' and find the minimum value and the corresponding scale factor β2 value; S113: Detect the second-level watermark and determine whether the minimum value of CC” is less than or equal to the set threshold T value , if the minimum value of CC” is less than or equal to the set threshold T value , then directly perform an inverse FFT on the spectrum corresponding to the obtained β2 value, and use linear correlation to calculate the linear correlation information between the corresponding start frame information and end frame information, which is the embedded watermark; if the minimum value of CC” is greater than the set threshold T value , then adjust the threshold T value , restart step S102 until the final CC’ or CC” is less than or equal to the set threshold T value .
2. The blind detection method for digital audio watermark according to claim 1, characterized in that, The adjustment scale factor β calculated in step S104 and step S109 is specifically: First, use a vector i F to index, where i F = 1:N; Another index vector i X is calculated by rounding the result of i F / β to the nearest integer; The value of i X determines the vector D t after frequency adjustment; Next, determine the relationship between β and 1. If β > 1, calculate the corresponding i X , if the i X element n is a unique value m, then D t (m) = F t (n); when the i X values in the elements are repeated (assuming there are x + 1 repeated elements), that is, the values of the nth, (n + 1)th until the (n + x)th elements are equal, then If β < 1, then the values of i X are not continuous; if the value of the nth element of i X is n, then D t (n) = F t (n); if the value of the nth element of i X is m, and m ≠ n, then D t (n) = F t (n), and the other elements of the vector can be obtained by interpolation method.
3. The blind detection method for digital audio watermark according to claim 1, characterized in that, The value CC' corresponding to the frequency adjustment calculated according to the AA' and BB' in step S105 is specifically: Exclusive OR is performed on each bit of the AA’, BB’ and the AA, BB of the signal to be measured, that is Then add up these values as the corresponding value CC’ after frequency adjustment.
4. The blind detection method for digital audio watermark according to claim 1, characterized in that, The value CC'' corresponding to the frequency adjustment calculated according to the AA'' and BB'' in step S110 is specifically: Exclusive OR is performed on each bit of the AA”, BB” and the AA, BB of the signal to be measured, that is Then these values are added up to be used as the corresponding value CC” after frequency adjustment.
5. The blind detection method for digital audio watermark according to claim 1, characterized in that, Adjusting the threshold T in the step S113 value Specifically: Adjust the threshold T value to 1.5 to 2 times the set threshold value.
6. The blind detection method for digital audio watermark according to any one of claims 1-5, characterized in that, The F in step S102 t (n) is the positive frequency part, that is, F t (1:N / 2).
Citation Information
Patent Citations
Audio frequency watermark method and system based on embedding area selection
CN106409302A
Electronic watermark embedding method and extraction method for audio signal
JP2006330256A