Anti-rerecording audio watermark embedding method and device, and anti-rerecording audio watermark extracting method and device
Through frequency domain embedding and robust extraction audio watermarking technology, the challenges of audio watermarking in the prior art in sound quality protection and extraction accuracy are solved, and watermark extraction with low perceived distortion and high accuracy during the ripping process is achieved.
Patent Information
- Application Number
- CN202510829414.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing audio watermarking technology has challenges in sound quality protection, anti-distortion ability and extraction accuracy, especially in complex audio environments, where watermark position information is misaligned and difficult to automatically locate, and conventional methods are prone to causing audio artifacts and exposing watermarks.
The audio watermark technology of frequency domain embedding and robust extraction is adopted to enhance robustness by embedding watermark information into the frequency domain of the audio, and the watermark is extracted through a robust frequency domain analysis method to ensure high accuracy extraction under noise and distortion during the ripping process.
It achieves low perceptual distortion, high computing efficiency, low hearing interference and strong synchronization ability, and can extract watermark information stably and accurately to adapt to noise and distortion in audio ripping.
Smart Images

Figure CN120340508A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio watermarking, and particularly to an anti-duplication audio watermark embedding method, extraction method and device. Background Art
[0002] In the current environment where information dissemination is increasingly rapid, in applications such as video conferencing, distance education, and webcasting, digital audio content is extremely easy to be illegally recorded, tampered with or disseminated, resulting in problems such as commercial information leakage and copyright infringement. As an effective means of copyright protection, audio watermarking has become an important technology for content security by embedding imperceptible information in audio signals.
[0003] Traditional audio watermark embedding methods mostly rely on time-domain watermark embedding technology, but time-domain watermark embedding methods often cause obvious audio distortion and are prone to watermark information loss during duplication, compression or editing. To overcome these problems, frequency-domain watermark embedding technology has been proposed, which can effectively reduce the impact of watermarks on audio signals and enhance the robustness of watermark information.
[0004] At present, audio watermarking technology still faces certain challenges in terms of sound quality protection, anti-distortion ability and extraction accuracy. Especially in complex audio environments, duplication or playback processes, due to differences in audio device frequency responses, non-linear distortions or clock drifts, the watermark position information is misaligned and the watermark cannot be correctly extracted; the watermark signal persists, and if the energy is too large, it is easy to cause perceptible audio artifacts such as ringing or background noise; there is a lack of a robust synchronization mechanism, making it difficult to automatically locate watermarks from a large range of audio streams. In addition, conventional spread-spectrum or modulation-based methods mostly involve periodic signals with long durations, which are prone to form repeated peaks in the spectrum, affecting the listening experience and exposing the risk of watermark existence. Therefore, for the problem of audio duplication in real-time scenarios, a new audio watermark technology solution with high computational efficiency, low listening interference, strong synchronization ability and anti-distortion is urgently needed.
[0005] In view of this, the present invention is specifically proposed. Summary of the Invention
[0006] The technical problem solved by the present invention is: to overcome the deficiencies of the prior art and provide an anti-duplication audio watermark embedding method, extraction method and device. The anti-duplication audio watermark embedding and extraction methods are based on audio watermarking technology of frequency-domain embedding and robust extraction. By embedding watermark information into the frequency domain of audio and using frequency-domain features to enhance the robustness of the watermark, and at the same time extracting the watermark through a robust frequency-domain analysis method, it can effectively cope with the noise and distortion in audio duplication and ensure a high extraction accuracy of the watermark.
[0007] An embodiment of the present invention provides an anti-duplication audio watermark embedding method, including: S11. Divide the audio signal into frames to obtain multiple audio frames F k and the number of samples per frame is i. For the audio frame F k perform a fast Fourier transform to obtain the frequency-domain spectrum F1 k , where k ∈ N + , i ∈ N + ; S12. Map the bits 1 and 0 in the m-bit binary watermark sequence W to the preset frequency points f k in the frequency-domain spectrum F1 j to obtain the watermark frequency-domain signal W f (f j ), and expand the watermark frequency-domain signal W f (f j ) to a watermark frequency-domain signal W f (i) with a length of i, where ; where f0 is the reference low-frequency point, bit j is the value of the jth watermark bit, m ∈ N + , j ∈ N + and j ≤ m; S13. Enhance the amplitude spectrum of the frequency points with bit 1 in the frequency-domain spectrum F1 k proportionally while keeping the phase spectrum unchanged, and keep both the amplitude spectrum and the phase spectrum of the frequency points with bit 0 unchanged; S14. Perform an inverse fast Fourier transform on the watermark frequency-domain signal W f (i) to obtain the watermark time-domain signal W t (i); S15. Weightedly superimpose the watermark time-domain signal W t (i) and the audio frame F k (i) to obtain the watermark-containing audio frame , where s is the superimposition ratio parameter; S16. Combine the audio frames W k (i) in sequence to obtain the audio signal with the watermark embedded.
[0008] This application embodiment also provides an anti-duplication audio watermark extraction method, including: S21. Process the watermark-containing audio signal M1 after duplication or processing, including: S211. Decode the audio signal M1 into a two-channel 44100, floating-point format audio signal M2 through the audio decoding library of FFmpeg; S212. Divide the audio signal M2 into audio frames F2 of a fixed length k ; S22. Perform a fast Fourier transform on the audio frame F2 k to obtain a frequency-domain spectrum, select the amplitude spectra of N consecutive frames for averaging to obtain a robust average spectrogram; S23. Calculate the amplitude spectrum T1 of each watermark bit bit of the m-bit binary watermark sequence W j corresponding to the frequency point f j , the average amplitude spectrum T2 of the surrounding frequency points of the frequency point f j , and the average amplitude spectrum difference T3 of the surrounding frequency points of the frequency point f j , where m ∈ N + , j ∈ N + and j ≤ m, ; wherein, is the window width, is the neighborhood standard deviation, is the threshold multiple; S24. By comparing the amplitude spectrum T1 of each watermark bit bit of the watermark sequence W j corresponding to the frequency point f j with the average amplitude spectrum difference T3 of the surrounding frequency points of the frequency point f j , output the finally extracted m-bit watermark bit sequence, where .
[0009] The embodiment of the present application also provides an anti-duplication audio watermark embedding device, including: A frame division and frequency-domain conversion module D11 that divides an audio signal into multiple audio frames F k and the number of sampling points per frame is i, and performs a fast Fourier transform on the audio frame F k to obtain a frequency-domain spectrum F1 k , where k ∈ N + , i ∈ N + ; A watermark frequency-domain information acquisition module D12 that maps the bits 1 and 0 in the m-bit binary watermark sequence W to the frequency-domain spectrum F1 k at a preset frequency point f j to obtain a watermark frequency-domain signal W f (f j ), and expands the watermark frequency-domain signal W f (f j ) to a watermark frequency-domain signal W f (i) with a length of i, where ; wherein, f0 is the reference low-frequency frequency point, bit jis the value of the j-th watermark bit, m ∈ N + , j ∈ N + and j ≤ m; Frequency domain spectrum enhancement module D13 enhances the amplitude spectrum of the frequency points where the bit is 1 in the frequency domain spectrum F1 k in proportion, keeps the phase spectrum unchanged, and keeps both the amplitude spectrum and the phase spectrum of the frequency points where the bit is 0 unchanged; Watermark time domain information acquisition module D14 performs an inverse fast Fourier transform on the watermark frequency domain signal W f (i) to obtain the watermark time domain signal W t (i); Watermarked audio frame synthesis module D15 performs weighted superposition of the watermark time domain signal W t (i) and the audio frame F k (i) to obtain the watermarked audio frame , where s is the superposition ratio parameter; Watermarked audio synthesis module D16 combines the audio frames W k (i) in sequence to obtain the audio signal with the embedded watermark.
[0010] The embodiment of the present application also provides an anti-duplication audio watermark extraction device, including: Audio signal preprocessing module D21 processes the watermarked audio signal M1 after duplication or processing, including: decoding the audio signal M1 into a two-channel 44100, floating-point format audio signal M2 through the audio decoding library of FFmpeg; dividing the audio signal M2 into audio frames F2 of a fixed length k ; Average spectrogram acquisition module D22 performs a fast Fourier transform on the audio frame F2 k to obtain the frequency domain spectrum, selects the amplitude spectra of N consecutive frames for averaging to obtain a robust average spectrogram; Amplitude spectrum acquisition module D23 for each frequency point calculates the amplitude spectrum T1 of the frequency point f corresponding to each watermark bit bit of the m-bit binary watermark sequence W j , the average amplitude spectrum T2 of the surrounding frequency points of the frequency point f j , and the average amplitude spectrum difference T3 of the surrounding frequency points of the frequency point f j , where m ∈ N j , j ∈ N + , and j ≤ m, + ; ; Among them, is the window width, is the neighborhood standard deviation, is the threshold multiple; The watermark bit sequence extraction module D24 extracts each watermark bit bit of the watermark sequence W by comparing j the amplitude spectrum T1 corresponding to the frequency point f j with the average difference T3 of the amplitude spectra of the surrounding frequency points of the frequency point f j and outputs the finally extracted m-bit watermark bit sequence, where 。
[0011] An embodiment of the present application also provides a computer-readable storage medium storing computer-executable instructions for executing an anti-duplication audio watermark embedding method implemented as described in any one of the above.
[0012] An embodiment of the present application further provides a device for implementing an audio watermark, including a memory and a processor. The memory stores the following instructions executable by the processor: steps for executing an anti-duplication audio watermark embedding method as described in any one of the above.
[0013] Compared with the prior art, the beneficial effects of the present invention are: 1) Frequency-domain embedding, low perceptual distortion: The watermark embedding only affects selected low-frequency points, which can maximize the preservation of the original audio quality; 2) Frame spectrum averaging enhances robustness: By cross-frame spectrum averaging, the influence of temporary perturbations on watermark extraction is reduced; 3) Channel redundancy enhances fault tolerance: In a stereo structure, dual-channel redundancy extraction can be used to enhance the extraction success rate; 4) Synchronization mechanism ensures reliable decoding: The embedding position of the watermark is at a specific frequency index. It only needs to detect whether the amplitude of this frequency point is significantly higher than its neighborhood. The watermark extraction does not depend on the starting position. Even if there are distortions such as time-domain drift, delay, and head and tail cropping during the duplication process, this method can still stably and accurately extract the watermark information. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0015] Figure 1 is a flowchart of the anti-duplication audio watermark embedding method proposed by the present invention.
[0016] Figure 2 is a schematic structural diagram of the anti-duplication audio watermark embedding device proposed by the present invention.
[0017] Description of the reference numerals in the drawings: D11 is the frame division and frequency domain conversion module; D12 is the watermark frequency domain information acquisition module; D13 is the frequency domain spectrum enhancement module; D14 is the watermark time domain information acquisition module; D15 is the watermarked audio frame synthesis module; D16 is the watermarked audio synthesis module. Detailed implementation manners
[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0019] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plural" generally includes at least two.
[0020] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0021] Depending on the context, the words "if", "when" as used herein can be interpreted as "when" or "when...", or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0022] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such commodity or device. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the commodity or device including the said element.
[0023] The optional embodiments of the present invention will be described in detail below with reference to the drawings.
[0024] In the first embodiment of the present invention, a method for embedding anti-copying audio watermark is as follows Figure 1 shown, including: S11. Divide the audio signal into multiple audio frames F k and the number of sampling points per frame is i. Perform fast Fourier transform on the audio frame F k to obtain the frequency-domain spectrum F1 k , where k ∈ N + , i ∈ N + ; S12. Map the bits 1 and 0 in the m-bit binary watermark sequence W to the k preset frequency points f j in the frequency-domain spectrum F1 to obtain the watermark frequency-domain signal W f (f j ), and expand the watermark frequency-domain signal W f (f j ) to a watermark frequency-domain signal W f (i) with a length of i, where ; where f0 is the reference low-frequency point, bit j is the value of the jth watermark bit, m ∈ N + , j ∈ N + and j ≤ m; S13. Enhance the amplitude spectrum of the frequency points with bit 1 in the frequency-domain spectrum F1 k proportionally, keep the phase spectrum unchanged, and keep both the amplitude spectrum and the phase spectrum of the frequency points with bit 0 unchanged; S14. Perform inverse fast Fourier transform on the watermark frequency-domain signal W f (i) to obtain the watermark time-domain signal W t (i); S15. Weightedly superimpose the watermark time-domain signal W t (i) and the audio frame F k (i) to obtain the watermark-containing audio frame , where s is the superimposition ratio parameter; S16. Combine the audio frames W k (i) in time sequence to obtain the audio signal with the embedded watermark.
[0025] In this embodiment, in step S11, the fast Fourier transform performed on the audio frame F k obtains a complex frequency-domain spectrum F1 k , which includes an amplitude spectrum and a phase spectrum.
[0026] Since the number of sampling points per audio frame is i, in step S12, the watermark frequency-domain signal W in the frequency-domain data will be passed throughf (f j ) Insert a large number of zeros to make its total length reach i, and then perform an inverse Fourier transform on this extended frequency-domain data to obtain a high-sampling-rate time-domain signal. Then perform a Fourier transform on this time-domain signal to finally obtain the frequency-domain data watermark frequency-domain signal W f (i) to ensure alignment when the subsequent watermark signal is superimposed on the audio frame signal.
[0027] For ease of understanding, an example of how to determine the reference low-frequency frequency point and the watermark frequency-domain signal in this example is given.
[0028] When the 6-bit binary watermark sequence W is 101101 and it is determined that f0 = 10Hz is the reference low-frequency frequency point, 6 frequency points of 20Hz, 30Hz, 40Hz, 50Hz, 60Hz, and 70Hz can be determined at intervals of 10Hz, and the watermark sequence W is embedded in the frequency-domain spectrum F1 k where the watermark frequency-domain signal W f (f j ) is: .
[0029] Then, the watermark frequency-domain signal W f (f j ) is extended to a watermark frequency-domain signal W f (i) to ensure alignment when the subsequent watermark signal is superimposed on the audio frame signal.
[0030] In an exemplary example, the optimal number of samples per frame i is 2048 points.
[0031] In an exemplary example, the number of bits m of the watermark sequence W satisfies 1 ≤ m ≤ 6.
[0032] In an exemplary example, the reference low-frequency frequency point satisfies 10HZ ≤ f0 ≤ 50HZ.
[0033] In an exemplary example, the superimposition ratio parameter satisfies 0.5 ≤ s ≤ 1, and the optimal value is 0.8.
[0034] In an exemplary example, when the audio signal is stereo, the watermark embedding operations of S11 - S16 need to be performed separately on the audio of the left and right channels.
[0035] Advantages of the anti-duplication audio watermark embedding method of the present invention: In audio signal processing, converting a time-domain audio signal into a frequency-domain signal is a key step in extracting its frequency characteristics. In the process of embedding and extracting audio watermarks, the discrete Fourier transform is usually relied on. Since the discrete Fourier transform is a complex transform, if the amplitude and phase meet specific conditions, both the amplitude and phase coefficients can be used as the embedding positions of the watermark.
[0036] The fast Fourier transform is an efficient discrete Fourier transform. In the process of embedding and extracting audio watermarks in the present invention, the fast Fourier transform is adopted, which can greatly improve the operation efficiency. At the same time, in the present invention, the watermark information is only embedded in the amplitude part of the selected frequency points without adjusting the phase, which can ensure the stability of the sound quality.
[0037] The second embodiment of the present invention introduces an anti-duplication audio watermark extraction method, which is based on the first embodiment and includes: S21. Process the watermarked audio signal M1 after duplication or processing, including: S211. Decode the audio signal M1 into an audio signal M2 in double-channel 44100 and floating-point format through the audio decoding library of FFmpeg; S212. Divide the audio signal M2 into audio frames F2 with a fixed length k ; S22. Perform a fast Fourier transform on the audio frame F2 k to obtain a frequency-domain spectrum, and select the amplitude spectra of N consecutive frames for averaging to obtain a robust average spectrogram; S23. Calculate the amplitude spectrum T1 of the frequency point f corresponding to each watermark bit bit of the m-bit binary watermark sequence W j the average amplitude spectrum T2 of the surrounding frequency points of the frequency point f j and the average amplitude spectrum difference T3 of the amplitude spectra of the surrounding frequency points of the frequency point f j where m ∈ N j and j ∈ N + and j ≤ m, + ; ; where, is the window width, is the neighborhood standard deviation, is the threshold multiple; S24. By comparing the amplitude spectrum T1 of the frequency point f corresponding to each watermark bit bit of the watermark sequence W with the average amplitude spectrum difference T3 of the amplitude spectra of the surrounding frequency points of the frequency point f, output the finally extracted m-bit watermark bit sequence, where j the frequency point f j and the frequency point f j where .
[0038] In this embodiment, for the processing of the watermarked audio signal M1 after ripping or processing, the audio decoding library of FFmpeg mentioned in the present invention can be used, or other audio decoding libraries of the same type can also be used.
[0039] In step S22, before extracting the watermark information, the amplitude spectra of N consecutive frames are averaged first to construct a robust average spectrogram, which can effectively cope with the noise and distortion in audio ripping, ensure a high accuracy rate of watermark extraction, and the subsequent steps are all operated under the said robust average spectrogram.
[0040] In step S24, each watermark bit bit of the watermark sequence W j corresponds to the frequency point f j which is known when the watermark is embedded in the audio. Taking the synchronization frequency point as a reference, detecting the amplitude mutation thereof is used to determine the starting point of the watermark, that is, by comparing each watermark bit bit of the binary watermark sequence W j corresponding to the frequency point f j with the average difference T3 of the amplitude spectra of the surrounding frequency points of the said frequency point f j the starting point of the watermark can be confirmed. By the same method, by traversing each watermark bit bit j corresponding to the frequency point f j and measuring its amplitude, the final m-bit watermark bit sequence can be extracted. Aligning the time frames here is to ensure the correct order of watermark extraction.
[0041] In an exemplary instance, the optimal window width w = 4, the optimal threshold multiple , and the relevant experimental data are as follows: By testing 500 samples of watermarked audio samples, the watermark extraction accuracy rate is as follows:
[0042] Advantages of the anti-ripping audio watermark extraction method of the present technical invention: In audio signal processing, frame division is a basic operation for splitting a continuous audio stream into short-time segment signal units. Utilizing the multi-frame frequency domain results for data fusion is a key means to improve the robustness of analysis and the reliability of features. The multi-frame frequency domain results can be used for multi-frame frequency domain averaging (suppressing random noise), multi-frame frequency domain superposition (enhancing feature significance), and multi-frame frequency domain feature fusion (constructing a robust representation).
[0043] Traditional audio watermark extraction methods usually require the following synchronization mechanisms: pre-embedding a synchronization code and searching for the starting position through a sliding window during extraction; or detecting the peak change of a specific reference frequency for synchronization; however, such methods are extremely vulnerable to failure under ripping attacks, especially when there is frame cropping or delay.
[0044] This solution completely omits the synchronization step for the following reasons: The Fourier transform is invariant to time-domain displacement in the amplitude spectrum. The watermark embedding position is at a specific frequency index. It only needs to detect whether the amplitude spectrum at this frequency point is significantly higher than its neighborhood. Regardless of whether there is an initial offset in time for the audio, the structure of the spectral amplitude diagram remains unchanged. Therefore, watermark extraction does not depend on the starting position.
[0045] Therefore, even if there are distortions such as time-domain drift, delay, and head / tail clipping during the copying process, this method can still stably and accurately extract the watermark information. This design realizes the complete automatic synchronization and decoupling of the extraction process, greatly improving the practicability and robustness of the algorithm.
[0046] The embodiment of this application also provides an anti-copying audio watermark embedding device. Based on the first embodiment, it includes: A frame division and frequency-domain conversion module D11 that divides the audio signal into multiple audio frames F k and the number of samples per frame is i. For the audio frame F k perform a fast Fourier transform to obtain the frequency-domain spectrum F1 k , where k ∈ N + , i ∈ N + ; A watermark frequency-domain information acquisition module D12 that maps the bits 1 and 0 in the m-bit binary watermark sequence W to the frequency-domain spectrum F1 k at the preset frequency point f j to obtain the watermark frequency-domain signal W f (f j ), and expand the watermark frequency-domain signal W f (f j ) to a watermark frequency-domain signal W f (i) with a length of i, where ; where f0 is the reference low-frequency frequency point, bit j is the value of the jth watermark bit, m ∈ N + , j ∈ N + and j ≤ m; A frequency-domain spectrum enhancement module D13 that enhances the amplitude spectrum of the frequency points with bit 1 in the frequency-domain spectrum F1 k in proportion while keeping the phase spectrum unchanged, and keeping both the amplitude spectrum and phase spectrum of the frequency points with bit 0 unchanged; A watermark time-domain information acquisition module D14 that performs an inverse fast Fourier transform on the watermark frequency-domain signal W f (i) to obtain the watermark time-domain signal W t (i); A watermarked audio frame synthesis module D15 that combines the watermark time-domain signal Wt (i) weighted superposition with the audio frame F k (i) to obtain a watermarked audio frame through weighted superposition where s is the superposition ratio parameter; The watermarked audio synthesis module D16 combines the audio frames W k (i) in time sequence to obtain an audio signal with the watermark embedded.
[0047] The embodiment of the present application also provides an anti - ripping audio watermark extraction device. Based on the first embodiment, it includes: The audio signal pre - processing module D21 processes the ripped or processed watermarked audio signal M1, including: decoding the audio signal M1 into a two - channel 44100, floating - point format audio signal M2 through the audio decoding library of FFmpeg; dividing the audio signal M2 into audio frames F2 with a fixed length k ; The average spectrogram acquisition module D22 performs a fast Fourier transform on the audio frame F2 k to obtain the frequency - domain spectrum, selects the amplitude spectra of N consecutive frames for averaging to obtain a robust average spectrogram; The amplitude spectrum acquisition module D23 for each frequency point calculates the amplitude spectrum T1 of each watermark bit bit of the m - bit binary watermark sequence W j corresponding to the frequency point f j the average amplitude spectrum T2 of the surrounding frequency points of the frequency point f j and the average difference T3 of the amplitude spectra of the surrounding frequency points of the frequency point f j where m ∈ N + and j ∈ N + and j ≤ m, ; where, is the window width, is the neighborhood standard deviation, is the threshold multiple; The watermark bit sequence extraction module D24 outputs the finally extracted m - bit watermark bit sequence by comparing the amplitude spectrum T1 of each watermark bit bit of the watermark sequence W j corresponding to the frequency point f j with the average difference T3 of the amplitude spectra of the surrounding frequency points of the frequency point f j where .
[0048] The embodiment of the present application also provides a computer - readable storage medium storing computer - executable instructions for executing an anti - ripping audio watermark embedding method implemented as described in any one of the above.
[0049] An embodiment of the present application further provides a device for implementing audio watermarking, including a memory and a processor. Among them, the memory stores the following instructions executable by the processor: for performing the steps of any one of the above-mentioned anti-duplication audio watermark embedding methods.
[0050] Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device is made to read and execute the program code stored in the storage medium.
[0051] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0052] Embodiments of the storage medium for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.
[0053] In addition, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by the computer, but also by making the operating system and the like operating on the computer complete part or all of the actual operations based on the instructions of the program code.
[0054] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then the CPU and the like installed on the expansion board or the expansion unit are made to execute part and all of the actual operations based on the instructions of the program code, so as to realize the functions of any one of the above embodiments.
[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An anti-copying audio watermark embedding method, characterized in that, including: S11. Divide the audio signal into frames to obtain multiple audio frames F k And the number of samples per frame is i. For the audio frame F k Perform a fast Fourier transform to obtain the frequency-domain spectrum F1 k , where k ∈ N + , i ∈ N + ; S1 2. Map the bits 1 and 0 in the m-bit binary watermark sequence W to the frequency-domain spectrum F1 k Preset frequency point f j Obtain the watermark frequency-domain signal W f (f j ), and expand the watermark frequency-domain signal W f (f j ) to a watermark frequency-domain signal W f (i) with a length of i, where ; where f0 is the reference low-frequency frequency point, bit j is the value of the j-th watermark bit, m ∈ N + , j ∈ N + and j ≤ m; S13. For the frequency-domain spectrum F1 k enhance the amplitude spectrum of the frequency points where the bit is 1 in proportion, keep the phase spectrum unchanged, and keep both the amplitude spectrum and the phase spectrum of the frequency points where the bit is 0 unchanged; S14. For the watermark frequency-domain signal W f (i) Perform an inverse fast Fourier transform to obtain the watermark time-domain signal W t (i); S15. Weight the watermark time-domain signal W t (i) and the audio frame F k (i) and perform weighted superposition to obtain a watermark-containing audio frame , where s is the superposition ratio parameter; S16. Merge the audio frame W k (i) in chronological order to obtain the audio signal embedded with the watermark.
2. The anti-copying audio watermark embedding method according to claim 1, characterized in that, The optimal number of sampling points i per frame is 2048 points.
3. The anti-copying audio watermark embedding method according to claim 1, characterized in that The number of bits m of the watermark sequence W satisfies 1 ≤ m ≤ 6.
4. The anti-copying audio watermark embedding method according to claim 1, characterized in that The reference low-frequency frequency point satisfies 10HZ ≤ f0 ≤ 50HZ.
5. The anti-copying audio watermark embedding method according to claim 1, characterized in that The superposition ratio parameter satisfies 0.5 ≤ s ≤ 1, and the optimal value is 0.
8.
6. The anti-copying audio watermark embedding method according to claim 1, characterized in that When the audio signal is a stereo signal, the watermark embedding operations of S11-S16 described in Claim 1 need to be separately performed on the audio of the left channel and the right channel.
7. An anti-copying audio watermark extraction method, which adopts the anti-copying audio watermark embedding method described in any one of claims 1-6, is characterized in that, including: S21. Process the watermarked audio signal M1 after dubbing or processing, including: S211. Decode the audio signal M1 into a stereo 44100, floating-point format audio signal M2 through the audio decoding library of FFmpeg; S212. Divide the audio signal M2 into audio frames F2 of a fixed length k ; S22. Perform a fast Fourier transform on the audio frame F2 k to obtain a frequency-domain spectrum, select the amplitude spectra of N consecutive frames for averaging, and obtain a robust average spectrogram; S23. Calculate each watermark bit bit of the m-bit binary watermark sequence W j corresponding to the frequency point f j of the amplitude spectrum T1, the frequency point f j average amplitude spectrum T2 of the surrounding frequency points, the frequency point f j average difference T3 of the amplitude spectra of the surrounding frequency points, where m ∈ N + , j ∈ N + and j ≤ m ; Among them, is the window width, is the neighborhood standard deviation, is the threshold multiple; S24. By comparing each watermark bit bit of the watermark sequence W j corresponding to the frequency point f j with the average difference T3 of the amplitude spectra of the surrounding frequency points of its amplitude spectrum T1 at the frequency point f j output the finally extracted m-bit watermark bit sequence, where 。 8. The anti-copying audio watermark extraction method according to claim 7, characterized in that, The optimal window width w = 4, the optimal threshold multiple .
9. An anti - piracy audio watermark embedding device, which adopts the anti - piracy audio watermark embedding method according to any one of claims 1 - 6, is characterized in that including: Frame division and frequency domain conversion module D11 divides the audio signal into multiple audio frames F k and the number of samples per frame is i. For the audio frame F k a fast Fourier transform is performed to obtain the frequency domain spectrum F1 k , where k ∈ N + , i ∈ N + ; The watermark frequency-domain information acquisition module D12 maps the bits 1 and 0 in the m-bit binary watermark sequence W to the frequency-domain spectrum F1 k The preset frequency point f j To obtain the watermark frequency-domain signal W f (f j ), and the watermark frequency-domain signal W f (f j ) is extended to a watermark frequency-domain signal W of length i f (i), where ; where f0 is the reference low-frequency frequency point, bit j is the value of the j-th watermark bit, m ∈ N + , j ∈ N + and j ≤ m; Frequency domain spectrum enhancement module D13 enhances the amplitude spectrum of the frequency points with bit 1 in the frequency domain spectrum F1 k in proportion while keeping the phase spectrum unchanged, and keeps both the amplitude spectrum and the phase spectrum of the frequency points with bit 0 unchanged; Watermark time-domain information acquisition module D14, for the watermark frequency-domain signal W f (i) performs an inverse fast Fourier transform to obtain the watermark time-domain signal W t (i); The watermark-containing audio frame synthesis module D15 combines the watermark time-domain signal W t (i) with the audio frame F k (i) through weighted superposition to obtain a watermark-containing audio frame , where s is the superposition ratio parameter; The watermark-containing audio synthesis module D16 combines the audio frames W k (i) in sequence to obtain the audio signal with the embedded watermark.
10. An anti-duplication audio watermark extraction device, which adopts the anti-duplication audio watermark extraction method as described in claim 7, is characterized in that, including: The audio signal preprocessing module D21 processes the watermarked audio signal M1 after dubbing or processing, including: decoding the audio signal M1 into a stereo 44100, floating-point format audio signal M2 through the audio decoding library of FFmpeg; dividing the audio signal M2 into audio frames F2 of a fixed length k ; Average spectrum diagram acquisition module D22, for the audio frame F2 k Perform fast Fourier transform to obtain the frequency-domain spectrum, select the amplitude spectra of N consecutive frames for averaging, and obtain a robust average spectrum diagram; The amplitude spectrum acquisition module D23 for each frequency point calculates the amplitude spectrum T1 of the corresponding frequency point f of each watermark bit bit of the m-bit binary watermark sequence W, the average amplitude spectrum T2 of the surrounding frequency points of the frequency point f, and the average amplitude difference T3 of the amplitude spectra of the surrounding frequency points of the frequency point f, where m ∈ N, j ∈ N and j ≤ m j corresponding frequency point f j and the average amplitude spectrum T2 of the surrounding frequency points of the frequency point f j average amplitude spectrum T2 of the surrounding frequency points of the frequency point f j average amplitude difference T3 of the amplitude spectra of the surrounding frequency points of the frequency point f, where m ∈ N + , j ∈ N + and j ≤ m ; Among them, is the window width, is the standard deviation of the neighborhood, is the threshold multiple; The watermark bit sequence extraction module D24 outputs the finally extracted m-bit watermark bit sequence by comparing the amplitude spectrum T1 of each watermark bit bit of the watermark sequence W with the average difference T3 of the amplitude spectra of the surrounding frequency points corresponding to the frequency point f j at the corresponding frequency point f j and the average difference T3 of the amplitude spectra of the surrounding frequency points at the frequency point f j where 。
Citation Information
Patent Citations
Digital audio watermarking method capable of resisting re-recording attack
CN103208289A
Robust digital audio watermarking algorithm in view of time-frequency analysis
CN106898358A
Method and device for realizing audio watermarking
CN114743555A
Auditory sense non-inductive simulation watermark embedding method in voice generation process
CN116524940A
Watermark incorporation
CN1969487A