Anti-copying audio watermark embedding method, extraction method and device
The audio watermark technology of frequency domain embedding and robust extraction solves the problem of unstable extraction of audio watermarks during the dubbing process in the existing technology, and achieves an audio watermark extraction effect with low distortion, high efficiency and strong synchronization capability.
Patent Information
- Application Number
- CN202510829414.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Existing audio watermarking technology faces challenges in sound quality protection, distortion resistance and extraction accuracy. Especially in complex audio environments and during the dubbing process, it is difficult to correctly extract watermark information, and conventional methods are prone to cause audio artifacts and expose watermark risks.
The audio watermarking technology that uses frequency domain embedding and robust extraction embeds the watermark information into the frequency domain of the audio, uses frequency domain features to enhance robustness, and extracts the watermark through a robust frequency domain analysis method. It can effectively deal with noise and distortion in audio ripping.
It achieves low perceptual distortion, high computational efficiency, small auditory interference, and strong synchronization capability. It can extract watermark information stably and accurately during the dubbing process, enhancing the robustness and extraction success rate of audio watermarks.
Smart Images

Figure CN120340508B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio watermarking, and in particular to an anti-copying audio watermark embedding method, an extraction method and a device. Background Art
[0002] In today's rapidly evolving information landscape, digital audio content in applications such as video conferencing, distance education, and live streaming can be easily illegally recorded, tampered with, or disseminated, leading to commercial information leaks and copyright infringement. As an effective copyright protection method, audio watermarking, which embeds imperceptible information in audio signals, has become a crucial technology for content security.
[0003] Traditional audio watermarking methods rely on time-domain watermarking, but this often results in noticeable audio distortion and can easily lead to watermark loss during dubbing, compression, or editing. To overcome these issues, frequency-domain watermarking has been proposed, which can effectively reduce the impact of watermarking on audio signals and enhance the robustness of watermark information.
[0004] Currently, audio watermarking technology still faces certain challenges in terms of sound quality protection, distortion resistance, and extraction accuracy. This is especially true in complex audio environments, during dubbing or playback. Due to differences in the frequency response of audio devices, nonlinear distortion, or clock drift, the watermark position information is misplaced and cannot be correctly extracted. The watermark signal persists, and if the energy is too high, it can easily cause perceptible audio artifacts such as ringing or background noise. The lack of a robust synchronization mechanism makes it difficult to automatically locate the watermark in a large audio stream. In addition, conventional spread spectrum or modulation-based methods often involve periodic signals of longer duration, which are prone to forming repeated peaks in the spectrum, affecting the listening experience and exposing the watermark to risks. Therefore, to address the problem of audio dubbing in real-time scenarios, there is an urgent need for a new audio watermarking technology solution with high computational efficiency, low auditory interference, strong synchronization capabilities, and resistance to distortion.
[0005] In view of this, the present invention is proposed. Summary of the Invention
[0006] The present invention addresses the technical problem of overcoming the shortcomings of existing technologies by providing a method, method, and device for embedding and extracting audio watermarks that are resistant to dubbing. The method, based on audio watermarking technology using frequency domain embedding and robust extraction, embeds watermark information into the frequency domain of the audio, utilizing frequency domain characteristics to enhance the robustness of the watermark. Furthermore, the method extracts the watermark using a robust frequency domain analysis method, effectively addressing noise and distortion in audio dubbing and ensuring high accuracy in watermark extraction.
[0007] An embodiment of the present invention provides a method for embedding an anti-copying audio watermark, comprising:
[0008] S11. Divide the audio signal into frames to obtain multiple audio frames F k The number of sampling points per frame is i. k Perform fast Fourier transform to obtain the frequency domain spectrum F1 k , where k∈N + , i∈N + ;
[0009] S12. Map the bits 1 and 0 in the m-bit binary watermark sequence W to the frequency domain spectrum F1 k Preset frequency point f j Get the watermark frequency domain signal W f (f j ), and the watermark frequency domain signal W f (f j ) is expanded to a watermark frequency domain signal W of length i f (i), where
[0010] ;
[0011] Among them, f0 is the reference low frequency point, bit j is the value of the jth watermark bit, m∈N + , j∈N + and j≤m;
[0012] S13. The frequency domain spectrum F1 k The amplitude spectrum of the frequency point where the bit is 1 is proportionally enhanced, and the phase spectrum remains unchanged; the amplitude spectrum and phase spectrum of the frequency point where the bit is 0 remain unchanged;
[0013] S14. The watermark frequency domain signal W f (i) Perform inverse fast Fourier transform to obtain the watermark time domain signal W t (i);
[0014] S15. The watermark time domain signal W t (i) with the audio frame F k (i) Perform weighted superposition to obtain the audio frame containing the watermark , where s is the superposition scale parameter;
[0015] S16. The audio frame W k (i) Merge in time sequence to obtain the audio signal with embedded watermark.
[0016] The present application also provides an anti-copying audio watermark extraction method, comprising:
[0017] S21. Processing the ripped or processed watermarked audio signal M1, including:
[0018] S211. The audio signal M1 is decoded into a dual-channel 44100 floating-point format audio signal M2 by the FFmpeg audio decoding library;
[0019] S212. Divide the audio signal M2 into audio frames F2 of fixed length k ;
[0020] S22. The audio frame F2 k Perform fast Fourier transform to obtain the frequency domain spectrum, select the amplitude spectrum of N consecutive frames and average them to obtain a robust average spectrum graph;
[0021] S23. Calculate each watermark bit of the m-bit binary watermark sequence W j Corresponding frequency point f j The amplitude spectrum T1, the frequency point f j The average amplitude spectrum T2 of the surrounding frequency points, the frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points T3, where m∈N + , j∈N + and j≤m,
[0022] ;
[0023] in, is the window width, is the neighborhood standard deviation, is the threshold multiple;
[0024] S24. By comparing the watermark sequence W each watermark bit bit j Corresponding frequency point f j The amplitude spectrum T1 and its frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points is T3, and the final extracted m-bit watermark bit sequence is output, where
[0025] .
[0026] The present application also provides an anti-copying audio watermark embedding device, comprising:
[0027] The frame division and frequency domain conversion module D11 divides the audio signal into frames to obtain multiple audio frames F k The number of sampling points per frame is i. k Perform fast Fourier transform to obtain the frequency domain spectrum F1 k , where k∈N + , i∈N + ;
[0028] The watermark frequency domain information acquisition module D12 maps the bits 1 and 0 in the m-bit binary watermark sequence W to the frequency domain spectrum F1 k Preset frequency point f j Get the watermark frequency domain signal W f (f j ), and the watermark frequency domain signal W f (f j ) is expanded to a watermark frequency domain signal W of length i f (i), where
[0029] ;
[0030] Among them, f0 is the reference low frequency point, bit j is the value of the jth watermark bit, m∈N + , j∈N + and j≤m;
[0031] The frequency domain spectrum enhancement module D13 performs the frequency domain spectrum F1 k The amplitude spectrum of the frequency point where the bit is 1 is proportionally enhanced, and the phase spectrum remains unchanged; the amplitude spectrum and phase spectrum of the frequency point where the bit is 0 remain unchanged;
[0032] The watermark time domain information acquisition module D14 obtains the watermark frequency domain signal W f (i) Perform inverse fast Fourier transform to obtain the watermark time domain signal W t (i);
[0033] The watermarked audio frame synthesis module D15 converts the watermarked time domain signal W t (i) with the audio frame F k (i) Perform weighted superposition to obtain the audio frame containing the watermark , where s is the superposition scale parameter;
[0034] The watermarked audio synthesis module D16 converts the audio frame W k (i) Merge in time sequence to obtain the audio signal with embedded watermark.
[0035] The embodiment of the present application also provides an anti-copying audio watermark extraction device, comprising:
[0036] The audio signal preprocessing module D21 processes the watermarked audio signal M1 after ripping or processing, including: decoding the audio signal M1 into a dual-channel 44100 floating-point audio signal M2 through the FFmpeg audio decoding library; dividing the audio signal M2 into fixed-length audio frames F2 k ;
[0037] The average spectrum acquisition module D22 obtains the average spectrum of the audio frame F2 k Perform fast Fourier transform to obtain the frequency domain spectrum, select the amplitude spectrum of N consecutive frames and average them to obtain a robust average spectrum graph;
[0038] The frequency spectrum acquisition module D23 calculates the m-bit binary watermark sequence W for each watermark bit. j Corresponding frequency point f j The amplitude spectrum T1, the frequency point f j The average amplitude spectrum T2 of the surrounding frequency points, the frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points T3, where m∈N + , j∈N + and j≤m,
[0039] ;
[0040] in, is the window width, is the neighborhood standard deviation, is the threshold multiple;
[0041] The watermark bit sequence extraction module D24 compares each watermark bit of the watermark sequence W j Corresponding frequency point f j The amplitude spectrum T1 and its frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points is T3, and the final extracted m-bit watermark bit sequence is output, where
[0042] .
[0043] An embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute an anti-copying audio watermark embedding method implemented as described in any one of the above.
[0044] An embodiment of the present application further provides a device for implementing audio watermarking, including a memory and a processor, wherein the memory stores the following instructions that can be executed by the processor: used to execute the steps of any of the above-mentioned anti-copying audio watermark embedding methods.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] 1) Frequency domain embedding, low perceptual distortion: Watermark embedding only affects selected low-frequency points, which can preserve the original audio quality to the greatest extent;
[0047] 2) Frame spectrum averaging to enhance robustness: By averaging the spectrum across frames, the impact of temporary disturbances on watermark extraction is reduced;
[0048] 3) Channel redundancy enhancement and fault tolerance: In a stereo structure, dual-channel redundancy extraction can be used to enhance the extraction success rate;
[0049] 4) Synchronization mechanism ensures reliable decoding: The watermark is embedded at a specific frequency index. It is only necessary to detect whether the amplitude of this frequency point is significantly higher than that of its neighborhood. Watermark extraction does not depend on the starting position. Even if there are distortions such as time domain drift, delay, and head and tail cropping during the dubbing process, this method can still extract the watermark information stably and accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present invention, and together with the description, serve to explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and it is clear that those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0051] Figure 1 This is a flow chart of the anti-copying audio watermark embedding method proposed by the present invention.
[0052] Figure 2 It is a structural diagram of the anti-copying audio watermark embedding device proposed by the present invention.
[0053] Explanation of the accompanying symbols: D11 framing and frequency domain conversion module; D12 watermark frequency domain information acquisition module; D13 frequency domain spectrum enhancement module; D14 watermark time domain information acquisition module; D15 watermarked audio frame synthesis module; D16 watermarked audio synthesis module. DETAILED DESCRIPTION
[0054] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only some, not all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.
[0055] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a," "an," "the," and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, and unless the context clearly indicates otherwise, "a plurality" generally includes at least two.
[0056] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0057] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0058] It should also be noted that the terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such product or device. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the product or device comprising the element.
[0059] The optional embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0060] The first embodiment of the present invention is a method for embedding an anti-copying audio watermark. Figure 1 Shown, including:
[0061] S11. Divide the audio signal into frames to obtain multiple audio frames F k The number of sampling points per frame is i. k Perform fast Fourier transform to obtain the frequency domain spectrum F1 k , where k∈N + , i∈N + ;
[0062] S12. Map the bits 1 and 0 in the m-bit binary watermark sequence W to the frequency domain spectrum F1 k Preset frequency point f j Get the watermark frequency domain signal W f (f j ), and the watermark frequency domain signal W f (f j ) is expanded to a watermark frequency domain signal W of length i f (i), where
[0063] ;
[0064] Among them, f0 is the reference low frequency point, bit j is the value of the jth watermark bit, m∈N + , j∈N + and j≤m;
[0065] S13. The frequency domain spectrum F1 k The amplitude spectrum of the frequency point where the bit is 1 is proportionally enhanced, and the phase spectrum remains unchanged; the amplitude spectrum and phase spectrum of the frequency point where the bit is 0 remain unchanged;
[0066] S14. The watermark frequency domain signal W f (i) Perform inverse fast Fourier transform to obtain the watermark time domain signal W t (i);
[0067] S15. The watermark time domain signal W t (i) with the audio frame F k (i) Perform weighted superposition to obtain the audio frame containing the watermark , where s is the superposition scale parameter;
[0068] S16. The audio frame W k (i) Merge in time sequence to obtain the audio signal with embedded watermark.
[0069] In this embodiment, in step S11, the audio frame F k The fast Fourier transform obtains the complex frequency domain spectrum F1 k , which includes amplitude spectrum and phase spectrum.
[0070] Since the number of sampling points per audio frame is i, the frequency domain signal W is watermarked in the frequency domain data in step S12. f (f j ) to make its total length reach i, and then perform inverse Fourier transform on the expanded frequency domain data to obtain a time domain signal with a high sampling rate, and then perform Fourier transform on this time domain signal to finally obtain the frequency domain data watermark frequency domain signal W with the number of sampling points i f (i) To ensure that the subsequent watermark signal is aligned with the audio frame signal when superimposed.
[0071] To facilitate understanding, an example is given to illustrate how to determine the reference low-frequency point and the watermark frequency domain signal in this embodiment.
[0072] When the 6-bit binary watermark sequence W is 101101 and f0 = 10 Hz is determined to be the reference low-frequency point, the 6 frequency points of 20 Hz, 30 Hz, 40 Hz, 50 Hz, 60 Hz, and 70 Hz can be determined at every 10 Hz interval to embed the watermark sequence W into the frequency domain spectrum F1. k Where the watermark frequency domain signal W f (f j )for:
[0073] .
[0074] Then the watermark frequency domain signal W f (f j ) is expanded to a watermark frequency domain signal W of length i f (i) To ensure that the subsequent watermark signal is aligned with the audio frame signal when superimposed.
[0075] In an exemplary embodiment, the optimal number of sampling points i per frame is 2048 points.
[0076] In an exemplary embodiment, the number of bits of the watermark sequence W is 1≤m≤6.
[0077] In an exemplary embodiment, the reference low frequency point is 10 Hz ≤ f0 ≤ 50 Hz.
[0078] In an exemplary embodiment, the stacking ratio parameter is 0.5≤s≤1, and the optimal value is 0.8.
[0079] In an exemplary embodiment, when the audio signal is dual-channel, the watermark embedding operations of S11 to S16 need to be performed separately on the audio of the left channel and the audio of the right channel.
[0080] Advantages of the anti-copying audio watermark embedding method of this technology invention:
[0081] In audio signal processing, converting time-domain audio signals into frequency-domain signals is a key step in extracting their frequency characteristics. The embedding and extraction of audio watermarks typically relies on the discrete Fourier transform (DFT). This is because the DFT is a complex transform. If the amplitude and phase meet certain conditions, both the amplitude and phase coefficients can be used as the embedding location for the watermark.
[0082] Fast Fourier transform is an efficient discrete Fourier transform. The present invention adopts fast Fourier transform in both embedding and extraction of audio watermarks, which can greatly improve the computing efficiency. At the same time, the present invention embeds the watermark information only in the amplitude part of the selected frequency point without adjusting the phase, which can ensure stable sound quality.
[0083] The second embodiment of the present invention introduces a method for extracting an anti-copying audio watermark, based on the first embodiment, comprising:
[0084] S21. Processing the ripped or processed watermarked audio signal M1, including:
[0085] S211. The audio signal M1 is decoded into a dual-channel 44100 floating-point format audio signal M2 by the FFmpeg audio decoding library;
[0086] S212. Divide the audio signal M2 into audio frames F2 of fixed length k ;
[0087] S22. The audio frame F2 k Perform fast Fourier transform to obtain the frequency domain spectrum, select the amplitude spectrum of N consecutive frames and average them to obtain a robust average spectrum graph;
[0088] S23. Calculate each watermark bit of the m-bit binary watermark sequence W j Corresponding frequency point f j The amplitude spectrum T1, the frequency point f j The average amplitude spectrum T2 of the surrounding frequency points, the frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points T3, where m∈N + , j∈N + and j≤m,
[0089] ;
[0090] in, is the window width, is the neighborhood standard deviation, is the threshold multiple;
[0091] S24. By comparing the watermark sequence W each watermark bit bit j Corresponding frequency point f j The amplitude spectrum T1 and its frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points is T3, and the final extracted m-bit watermark bit sequence is output, where
[0092] .
[0093] In this embodiment, step S21 processes the ripped or processed watermarked audio signal M1, which may be processed by using the FFmpeg audio decoding library mentioned in the present invention, or other similar audio decoding libraries.
[0094] In step S22, the amplitude spectra of N consecutive frames are averaged before watermark information extraction to construct a robust average spectrum diagram, which can effectively deal with noise and distortion in audio dubbing and ensure high accuracy of watermark extraction. Subsequent steps are all performed under the robust average spectrum diagram.
[0095] In step S24, each watermark bit of the watermark sequence W is j Corresponding frequency point f j When the watermark is embedded in the audio, it is known that the amplitude mutation is detected with the synchronous frequency point as a reference to determine the starting point of the watermark, that is, by comparing the binary watermark sequence W with each watermark bit bit j Corresponding frequency point f j The frequency point f j The average difference T3 of the amplitude spectrum of the surrounding frequency points can be used to confirm the starting point of the watermark. In the same way, by traversing each watermark bit j Corresponding frequency point f j , measure its amplitude, and then extract the final m-bit watermark bit sequence. The time frame alignment here is to ensure the correct order of watermark extraction.
[0096] In an exemplary embodiment, the optimal window width w=4, the optimal threshold multiple , the relevant experimental data are as follows:
[0097] By testing 500 samples of watermarked audio samples, the watermark extraction accuracy is as follows:
[0098]
[0099] Advantages of the anti-copying audio watermark extraction method of this technology invention:
[0100] In audio signal processing, frame division is the basic operation of dividing a continuous audio stream into short-term signal units. Using multi-frame frequency domain results for data fusion is a key means to improve analysis robustness and feature reliability. Multi-frame frequency domain results can be used for multi-frame frequency domain averaging (suppressing random noise), multi-frame frequency domain superposition (enhancing feature significance), and multi-frame frequency domain feature fusion (building robust representation).
[0101] Traditional audio watermark extraction methods usually require the following synchronization mechanisms: pre-embedded synchronization codes, searching for the starting position through a sliding window during extraction; or detecting the peak changes of a specific reference frequency for synchronization; but such methods are extremely susceptible to failure under ripping attacks, especially when there is frame cropping or delay.
[0102] This solution completely omits the synchronization step for the following reasons: the Fourier transform is invariant to the time domain displacement in the amplitude spectrum. The watermark is embedded at a specific frequency index, and it is only necessary to detect whether the amplitude spectrum of this frequency point is significantly higher than its neighboring frequency. Regardless of whether the audio has a starting offset in time, the structure of the spectrum amplitude graph remains unchanged, so watermark extraction does not depend on the starting position.
[0103] Therefore, even if there are distortions such as temporal drift, delay, and head-end cropping during the dubbing process, this method can still extract the watermark information stably and accurately. This design achieves fully automatic synchronization and decoupling of the extraction process, greatly improving the practicality and robustness of the algorithm.
[0104] The present application also provides an anti-copying audio watermark embedding device, based on the first embodiment, comprising:
[0105] The frame division and frequency domain conversion module D11 divides the audio signal into frames to obtain multiple audio frames F k The number of sampling points per frame is i. k Perform fast Fourier transform to obtain the frequency domain spectrum F1 k , where k∈N + , i∈N + ;
[0106] The watermark frequency domain information acquisition module D12 maps the bits 1 and 0 in the m-bit binary watermark sequence W to the frequency domain spectrum F1 k Preset frequency point f j Get the watermark frequency domain signal W f (f j ), and the watermark frequency domain signal W f (f j ) is expanded to a watermark frequency domain signal W of length i f (i), where
[0107] ;
[0108] Among them, f0 is the reference low frequency point, bit j is the value of the jth watermark bit, m∈N + , j∈N + and j≤m;
[0109] The frequency domain spectrum enhancement module D13 performs the frequency domain spectrum F1 k The amplitude spectrum of the frequency point where the bit is 1 is proportionally enhanced, and the phase spectrum remains unchanged; the amplitude spectrum and phase spectrum of the frequency point where the bit is 0 remain unchanged;
[0110] The watermark time domain information acquisition module D14 obtains the watermark frequency domain signal W f(i) Perform inverse fast Fourier transform to obtain the watermark time domain signal W t (i);
[0111] The watermarked audio frame synthesis module D15 converts the watermarked time domain signal W t (i) with the audio frame F k (i) Perform weighted superposition to obtain the audio frame containing the watermark , where s is the superposition scale parameter;
[0112] The watermarked audio synthesis module D16 converts the audio frame W k (i) Merge in time sequence to obtain the audio signal with embedded watermark.
[0113] The present application also provides an anti-copying audio watermark extraction device based on the first embodiment, including:
[0114] The audio signal preprocessing module D21 processes the watermarked audio signal M1 after ripping or processing, including: decoding the audio signal M1 into a dual-channel 44100 floating-point audio signal M2 through the FFmpeg audio decoding library; dividing the audio signal M2 into fixed-length audio frames F2 k ;
[0115] The average spectrum acquisition module D22 obtains the average spectrum of the audio frame F2 k Perform fast Fourier transform to obtain the frequency domain spectrum, select the amplitude spectrum of N consecutive frames and average them to obtain a robust average spectrum graph;
[0116] The frequency spectrum acquisition module D23 calculates the m-bit binary watermark sequence W for each watermark bit. j Corresponding frequency point f j The amplitude spectrum T1, the frequency point f j The average amplitude spectrum T2 of the surrounding frequency points, the frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points T3, where m∈N + , j∈N + and j≤m,
[0117] ;
[0118] in, is the window width, is the neighborhood standard deviation, is the threshold multiple;
[0119] The watermark bit sequence extraction module D24 compares each watermark bit of the watermark sequence W j Corresponding frequency point f jThe amplitude spectrum T1 and its frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points is T3, and the final extracted m-bit watermark bit sequence is output, where
[0120] .
[0121] An embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute an anti-copying audio watermark embedding method implemented as described in any one of the above.
[0122] An embodiment of the present application further provides a device for implementing audio watermarking, including a memory and a processor, wherein the memory stores the following instructions that can be executed by the processor: used to execute the steps of any of the above-mentioned anti-copying audio watermark embedding methods.
[0123] Specifically, a system or device equipped with a storage medium can be provided, on which software program codes that implement the functions of any of the above-mentioned embodiments are stored, and a computer (or CPU or MPU) of the system or device can be enabled to read and execute the program codes stored in the storage medium.
[0124] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute part of the present invention.
[0125] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (e.g., CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RYMs, DVD-RWs, DVD+RWs), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer via a communications network.
[0126] In addition, it should be clear that the functions of any of the above embodiments can be achieved not only by executing the program code read by the computer, but also by enabling the operating system operating on the computer to complete part or all of the actual operations based on the instructions of the program code.
[0127] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU installed on the expansion board or expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above embodiments.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for embedding and extracting anti-copying audio watermarks, characterized in that: include: S11. Divide the audio signal into frames to obtain multiple audio frames F k The number of sampling points per frame is i. k Perform fast Fourier transform to obtain the frequency domain spectrum F1 k , where k∈N + , i∈N + ; S1 2. Map the bits 1 and 0 in the m-bit binary watermark sequence W to the frequency domain spectrum F1 k Preset frequency point f j Get the watermark frequency domain signal W f (f j ), and the watermark frequency domain signal W f (f j ) is expanded to a watermark frequency domain signal W of length i f (i), where Among them, f0 is the reference low frequency point and 10HZ≤f0≤50HZ, bit j is the value of the jth watermark bit, m∈N + , j∈N + and j≤m; S13. The frequency domain spectrum F1 k The amplitude spectrum of the frequency point where the bit is 1 is proportionally enhanced, and the phase spectrum remains unchanged; the amplitude spectrum and phase spectrum of the frequency point where the bit is 0 remain unchanged; S14. The watermark frequency domain signal W f (i) Perform inverse fast Fourier transform to obtain the watermark time domain signal W t (i); S15. The watermark time domain signal W t (i) with the audio frame F k Perform weighted superposition to obtain the watermarked audio frame W k (i) = F k +s*W t (i), where s is the superposition scale parameter; S16. The audio frame W k (i) Merge in time sequence to obtain the audio signal after embedding the watermark; S21. Processing the ripped or processed watermarked audio signal M1, including: S211. The audio signal M1 is decoded into a dual-channel 44100 floating-point format audio signal M2 by the FFmpeg audio decoding library; S212. Divide the audio signal M2 into audio frames F2 of fixed length k ; S22. The audio frame F2 k Perform fast Fourier transform to obtain the frequency domain spectrum, select the amplitude spectrum of N consecutive frames and average them to obtain a robust average spectrum graph; S23. Calculate each watermark bit of the m-bit binary watermark sequence W j Corresponding frequency point f j The amplitude spectrum T1, the frequency point f j The average amplitude spectrum T2 of the surrounding frequency points, the frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points T3, where m∈N + , j∈N + and j≤m, T3(f j )=T2(f j )+σ*γ; Among them, w is the window width, σ is the neighborhood standard deviation, and γ is the threshold multiple; S24. By comparing the watermark sequence W each watermark bit bit j Corresponding frequency point f j The amplitude spectrum T1 and its frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points T3 is output, and the final extracted m-bit watermark sequence W is output, where 2. The method for embedding and extracting an anti-copying audio watermark according to claim 1, characterized in that: The number of sampling points i per frame is 2048 points.
3. The method for embedding and extracting an anti-copying audio watermark according to claim 1, characterized in that: The number of bits of the watermark sequence W is 1≤m≤6.
4. The method for embedding and extracting an anti-copying audio watermark according to claim 1, characterized in that: The stacking ratio parameter is 0.5≤s≤1.
5. The method for embedding and extracting an anti-copying audio watermark according to claim 1, characterized in that: When the audio signal in S11 is dual-channel, the watermark embedding operations of S11 to S16 as claimed in claim 1 need to be performed separately on the left channel and the right channel audio.
6. The method for embedding and extracting an anti-copying audio watermark according to claim 1, characterized in that: The window width w=4, and the threshold multiple γ=4.
7. An anti-copying audio watermark embedding and extraction device, using the anti-copying audio watermark embedding and extraction method according to any one of claims 1 to 6, characterized in that: include: The frame division and frequency domain conversion module D11 divides the audio signal into frames to obtain multiple audio frames F k The number of sampling points per frame is i. k Perform fast Fourier transform to obtain the frequency domain spectrum F1 k , where k∈N + , i∈N + ; The watermark frequency domain information acquisition module D12 maps the bits 1 and 0 in the m-bit binary watermark sequence W to the frequency domain spectrum F1 k Preset frequency point f j Get the watermark frequency domain signal W f (f j ), and the watermark frequency domain signal W f (f j ) is expanded to a watermark frequency domain signal W of length i f (i), where Among them, f0 is the reference low frequency point and 10HZ≤f0≤50HZ, bit j is the value of the jth watermark bit, m∈N + , j∈N + and j≤m; The frequency domain spectrum enhancement module D13 performs the frequency domain spectrum F1 k The amplitude spectrum of the frequency point where the bit is 1 is proportionally enhanced, and the phase spectrum remains unchanged; the amplitude spectrum and phase spectrum of the frequency point where the bit is 0 remain unchanged; The watermark time domain information acquisition module D14 obtains the watermark frequency domain signal W f (i) Perform inverse fast Fourier transform to obtain the watermark time domain signal W t (i); The watermarked audio frame synthesis module D15 converts the watermarked time domain signal W t (i) with the audio frame F k (i) Perform weighted superposition to obtain the watermarked audio frame W k (i) = F k +s*W t (i), where s is the superposition scale parameter; The watermarked audio synthesis module D16 converts the audio frame W k (i) Merge in time sequence to obtain the audio signal after embedding the watermark; The audio signal preprocessing module D21 processes the watermarked audio signal M1 after ripping or processing, including: decoding the audio signal M1 into a dual-channel 44100 floating-point audio signal M2 through the FFmpeg audio decoding library; dividing the audio signal M2 into fixed-length audio frames F2 k ; The average spectrum acquisition module D22 obtains the average spectrum of the audio frame F2 k Perform fast Fourier transform to obtain the frequency domain spectrum, select the amplitude spectrum of N consecutive frames and average them to obtain a robust average spectrum graph; The frequency spectrum acquisition module D23 calculates the m-bit binary watermark sequence W for each watermark bit. j Corresponding frequency point f j The amplitude spectrum T1, the frequency point f j The average amplitude spectrum T2 of the surrounding frequency points, the frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points T3, where m∈N + , j∈N + and j≤m, T3(f j )=T2(f j )+σ*γ; Among them, w is the window width, σ is the neighborhood standard deviation, and γ is the threshold multiple; The watermark bit sequence extraction module D24 compares each watermark bit of the watermark sequence W j Corresponding frequency point f j The amplitude spectrum T1 and its frequency point f j The average difference of the amplitude spectrum of the surrounding frequency points is T3, and the final extracted m-bit watermark bit sequence is output, where
Citation Information
Patent Citations
Auditory sense non-inductive simulation watermark embedding method in voice generation process
CN116524940A