Audio histogram shape watermarking method against synchronization attacks
By performing segmented DWT transformation and frequency domain feature extension on the audio histogram, the problem of balancing embedding capacity and robustness in audio watermarking algorithms is solved, realizing a high-capacity and highly robust audio watermarking technology suitable for various signal processing scenarios.
Patent Information
- Application Number
- CN202511203594.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing audio histogram watermarking algorithms struggle to balance embedding capacity and robustness, especially when facing synchronization attacks and MP3 compression. Furthermore, the lack of fine-grained histogram shape design leads to wasted embedding space.
By reasonably modifying the shape of the low-frequency domain statistical histogram of the audio carrier, and using piecewise DWT transformation, the time-domain features are extended to the frequency domain. Combined with bin shape adjustment, watermark embedding and extraction are achieved, thereby improving embedding capacity and enhancing robustness.
While ensuring audio quality, it significantly improves the embedding capacity and is robust against various attacks such as MP3 compression, time-domain stretching, filtering, and Gaussian noise. It is suitable for different audio formats and can effectively extract watermark information under different signal processing conditions.
Smart Images

Figure CN120748415B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of digital watermarking, in particular to an audio histogram shape watermarking method against synchronization attacks. BACKGROUND
[0002] In recent years, novel digital watermarking algorithms have been proposed to gradually fill the gaps in the field of digital watermarking, but synchronization attacks are still a difficult problem to solve in traditional audio watermarking. Through research on audio features, many scholars have found that the histogram shape and mean value in the time domain have invariance to synchronization attacks. By using this characteristic, a technology is proposed to embed watermark by modifying the histogram shape. However, in the existing algorithms that use histogram shape modification to achieve watermark embedding, the design of the histogram shape lacks detailed consideration, resulting in a large amount of embedding space being wasted, which ultimately limits the improvement of embedding capacity.
[0003] Robustness to synchronization attacks and MP3 compression is an important aspect of the performance of audio robust watermarking algorithms. Most current audio watermarking algorithms can only address one of these attacks.
[0004] An earlier audio histogram robust watermarking algorithm is the synchronization attack-resistant watermarking algorithm based on time-domain statistical features proposed by Xiang et al. This algorithm mainly uses the theoretical invariance of the proportion relationship between the number of audio samples of different amplitudes before and after linear stretching. By modifying the shape proportion between three statistical histogram sample intervals (hereinafter referred to as "bins"), watermark embedding is achieved. Later, Xiang et al. proposed a method of segmenting DWT to extend the time-domain invariance feature to the frequency domain, solving the problem of poor robustness to Gaussian and compression attacks in the previous algorithm. Hu et al. proposed a robust watermarking high-capacity algorithm based on the DCT transform domain. The idea is to embed the watermark in the DCT low-frequency band of the audio frame. Liang et al. improved the histogram-based audio histogram by introducing a high-order difference statistical model. The histogram can be regarded as a robust feature and moved to embed the watermark sequence. By using the inverse operation of histogram shifting, the original audio file can be restored losslessly. The research on histogram-based robust watermarking technology is still ongoing. In the above-mentioned existing histogram-related methods, they cannot achieve a good balance between high capacity and high robustness.
[0005] In the current research progress, traditional audio robust watermarking technology still has strong development and application occasions, and with the progress of communication technology, the requirements for information capacity and security are also gradually increasing. SUMMARY
[0006] The application aims at providing a watermarking technology with strong robustness to common attacks and synchronization attacks and high embedding capacity based on a statistical histogram. The embedding capacity is greatly improved while ensuring the audio hearing quality and robustness by reasonably modifying the statistical histogram shape of the low frequency domain of the audio carrier, and the problem that the traditional histogram watermarking method is difficult to embed when the sample number is insufficient or the watermark data volume is too large is solved, so the idea embodied by the application has a great enlightening effect on the watermarking technology in the scene where capacity and robustness are simultaneously required. The application has a great effect on the integrity authentication of digital media, and the effect comes from the fact that the application can effectively resist various attacks such as MP3 compression, time domain stretching, filtering, Gaussian noise and volume change.
[0007] To achieve the above object, the application provides the following scheme.
[0008] The audio histogram shape watermarking method against synchronization attacks comprises the following steps.
[0009] According to the number of pre-allocated bins, the pre-selected audio samples are divided to obtain a first time domain histogram;
[0010] The first time domain histogram is subjected to a segmented DWT transformation and an inverse DWT transformation to obtain an embedded watermark audio;
[0011] A second time domain histogram corresponding to the embedded watermark audio is obtained, and the second time domain histogram is subjected to the segmented DWT transformation to obtain a low frequency domain second time domain histogram;
[0012] Information is extracted from the low frequency domain second time domain histogram to obtain an embedded watermark.
[0013] Optionally, obtaining the first time domain histogram comprises the following steps.
[0014] The initial audio is pre-processed, and the absolute mean value of the audio is calculated;
[0015] Based on the audio absolute mean value and a preset constant value λ, an embedding range, i.e. the audio samples, is selected;
[0016] The number of bins is allocated according to the number of watermark bits to be embedded; wherein 2 bits are embedded using 3 bins;
[0017] The samples are divided according to the number of bins to obtain the time domain histogram.
[0018] Optionally, obtaining the embedded watermark audio comprises the following steps.
[0019] The first time domain histogram is subjected to a segmented DWT transformation, and the breakpoints between each discontinuous bin are recorded to obtain a low frequency domain first time domain histogram;
[0020] performing information embedding on the first time-domain histogram of the low-frequency domain, modifying the shape of the bin, and obtaining a modified first time-domain histogram;
[0021] performing inverse two-level DWT transformation on the modified first time-domain histogram combined with the recorded breakpoint to obtain the watermark-embedded audio.
[0022] Optionally, the segmented DWT transformation on the first time-domain histogram comprises:
[0023] the audio sample [-λA, λA] is regarded as having been allocated to a bin with equal width; the samples in the first bin, i.e., the range [-λA, -λA + M), are subjected to two-level DWT transformation, where M represents the width of each bin, the samples in the second bin, i.e., the range [-λA + M, -λA + 2M), are subjected to two-level DWT transformation, and each subsequent bin interval is increased by M on the basis of the range of the previous bin, and so on to obtain the interval of each bin; each M range is subjected to DWT transformation individually to obtain the first time-domain histogram of the low-frequency domain.
[0024] Optionally, the method for modifying the shape of the bin comprises:
[0025] the bin interval is grouped into groups of every three, and each group is used to embed two bits of information; for the four different two-bit information units: 00, 01, 10, and 11, present in the watermark sequence, different embedding strategies are used respectively, wherein:
[0026] when the current embedded information is 00, the shape of the bin is modified to low-middle-high;
[0027] when the current embedded information is 01, the shape of the middle bin is modified to the lowest;
[0028] when the current embedded information is 10, the shape of the middle bin is modified to the highest;
[0029] when the current embedded information is 11, the shape is modified to high-middle-low;
[0030] low, middle, and high represent the size relationship of three consecutive bin intervals in terms of sample quantity.
[0031] Optionally, obtaining the second time-domain histogram corresponding to the watermark-embedded audio comprises:
[0032] performing preprocessing on the watermark-embedded audio, and calculating the absolute mean value of the watermark-embedded audio;
[0033] based on the absolute mean value of the watermark-embedded audio and a preset constant value λ, selecting an extraction range;
[0034] The binning operation is performed on the samples in the extraction range to obtain the second time domain histogram.
[0035] Optionally, the information extraction on the second time domain histogram in the low frequency domain comprises:
[0036] The watermark information is extracted in 3 bins as a group, and the sample number of each group of bins is a, b and c:
[0037] if The current bin group watermark information is 11.
[0038] if The current bin group watermark information is 00.
[0039] if The current bin group watermark information is 10.
[0040] if The current bin group watermark information is 01.
[0041] Optionally, the audio absolute mean value is:
[0042] ;
[0043] Wherein, A is the audio absolute mean value, N is the length of the original audio, is the i-th sample in each bin.
[0044] The present application has the following advantages:
[0045] The present application is based on the histogram shape, and the time domain features of the histogram are transferred to the frequency domain by the segmented DWT method, which greatly improves the robustness of the algorithm. Compared with the previous algorithm, the present application utilizes the diversity of the histogram shape to realize higher capacity embedding on the same number of bins, thereby solving the problem of difficulty in embedding watermark in some speech files due to the small number of samples. This has a great enlightening effect on the research of robust watermarking in modifying the histogram.
[0046] The present application utilizes the robustness of the low frequency domain coefficient after DWT transformation to compression, Gaussian noise and other attacks, and realizes the extension of the time domain stretching invariant feature to the frequency domain by the segmented DWT method under the premise of preserving the histogram shape, which improves the resistance to conventional attacks and retains the robustness to stretching attacks.
[0047] The present application can effectively extract watermark information under different signal processing, such as Gaussian noise, MP3 compressed sound attack, and meets the requirements of daily digital forensics and digital authentication; meanwhile, the algorithm can be applied to different audios, and good results are obtained on different audios. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0049] Figure 1 A flowchart of the robust watermark embedding algorithm of the present embodiment using audio histogram shape to resist synchronous attack;
[0050] Figure 2 A flowchart of the present embodiment for extracting watermark;
[0051] Figure 3 A histogram shape diagram of the present embodiment for different embedded information modification;
[0052] Figure 4 A time-domain histogram segmentation DWT diagram of the present embodiment. DETAILED DESCRIPTION
[0053] The technical solutions in the embodiments of the present application will be described clearly and completely in the embodiments of the present application combined with the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0054] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail in combination with the drawings and specific embodiments.
[0055] The present embodiment proposes an audio histogram shape watermark method resisting synchronous attack, including: robust watermark embedding and robust watermark extraction;
[0056] The robust watermark embedding part: according to the pre-allocated number of bins, the pre-selected audio samples are divided to obtain a first time-domain histogram;
[0057] The first time-domain histogram is subjected to segmented DWT transformation and inverse DWT transformation to obtain an embedded watermark audio;
[0058] robust watermark extraction part: obtaining a second time domain histogram corresponding to the embedded watermark audio, and performing the segmented DWT transformation on the second time domain histogram to obtain a low-frequency domain second time domain histogram;
[0059] performing information extraction on the low-frequency domain second time domain histogram to obtain the embedded watermark.
[0060] Further, obtaining the first time domain histogram comprises:
[0061] performing preprocessing on the initial audio to calculate an audio absolute mean value;
[0062] selecting an embedding range, i.e., the audio samples, based on the audio absolute mean value and λ, wherein λ is a fixed constant value, and is usually taken as a value between [2, 4]. The role of λ is to artificially control the audio range to be processed according to [-λA, λA]. The larger λ is, the larger the selected range is. However, if λ is too large, the range may exceed the audio itself, and if λ is too small, the number of samples in the range may be insufficient. Therefore, according to the experimental results, a better range [2.4, 4] is obtained. The main role of selecting this range is to ensure that the number of samples can complete embedding while removing the lower and higher amplitude values in the audio signal. These relatively extreme samples are less and are easy to affect the embedding quality.
[0063] allocating the number of bins according to the number of watermark bits to be embedded; wherein 3 bins are used to embed 2 bits;
[0064] dividing the samples according to the number of bins to obtain the time domain histogram.
[0065] Further, obtaining the embedded watermark audio comprises:
[0066] performing the segmented DWT transformation on the first time domain histogram to record the breakpoints between each discontinuous bin to obtain a low-frequency domain first time domain histogram;
[0067] modifying the shape of the bin in the low-frequency domain first time domain histogram to obtain a modified first time domain histogram;
[0068] performing inverse two-level DWT transformation on the modified first time domain histogram combined with the recorded breakpoints to obtain the embedded watermark audio.
[0069] Further, performing the segmented DWT transformation on the first time domain histogram comprises:
[0070] The audio sample [-λA, λA] is regarded as being allocated into bins with equal width; the samples in the first bin, i.e. in the range [-λA, -λA + 2λA / a), are subjected to a two-level DWT transformation, the samples in the second bin, i.e. in the range [-λA + 2λA / a, -λA + 4λA / a), are subjected to a two-level DWT transformation, and so on, so that a DWT transformation is performed on each piecewise function in the range M separately, to obtain the first time-domain histogram in the low-frequency domain; wherein M represents the width of each bin, and a is the number of bins.
[0071] Further, the method for modifying the shape of the bin comprises:
[0072] To realize the embedding of the watermark information, the bin intervals are grouped by every 3 as a group, and each group is used to embed 2 bits of information. For the four different 2-bit information units that may exist in the watermark sequence: 00, 01, 10, 11, different embedding strategies are used respectively, wherein,
[0073] When the current embedded information is "00", the shape of the bin is modified to low-middle-high;
[0074] When the current embedded information is "01", the shape is modified to b lowest; middle-low-high, high-low-middle, all of which are acceptable;
[0075] When the current embedded information is "10", the shape is modified to b highest; middle-high-low, low-high-middle, all of which are acceptable;
[0076] When the current embedded information is "11", the shape is modified to high-middle-low;
[0077] Wherein, "low, middle, high" represent the size relationship of the three consecutive bin intervals in the number of samples. The adjustment strategy aims to: under the premise of maintaining the quality of the audio, the shape of the number of samples is adjusted to realize the robust embedding of the watermark information.
[0078] The embodiment is based on a robust watermark algorithm proposed by an audio histogram, mainly uses two features which are not changed after linear scaling modification in time domain, i.e., histogram shape and average value, and expands the time domain features to the frequency domain through the method of segmented DWT, has stronger robustness in resisting synchronous attacks than existing audio robust watermark algorithms, guarantees strong robustness to regular attacks, and reasonably designs the distribution between histogram shapes to further improve the embedded watermark information capacity. The embodiment solves the problem between traditional audio robustness and capacity, and the idea adopted in the embodiment has great enlightenment effect on the implementation of the digital watermark algorithm with insufficient capacity in reality. Under the background of future information technology development, the integrity authentication and copyright problem of digital products will inevitably become the focus of research, the histogram shape embedding method proposed in the embodiment greatly improves the embedded capacity under the condition of ensuring the invariance of audio quality and robustness, and is a new development direction of digital watermark application technology.
[0079] Specifically, in the embodiment, the audio histogram shape watermark method resisting synchronous attacks, the watermark embedding process is as shown in the figure, and the specific steps are as follows: Figure 1
[0080] 1. Preprocess the initial audio , calculate the audio absolute average value, and define that the normalized sum of the value of the initial audio signal in the duration is always close to zero before and after the attack, so the audio absolute average value can be applied to one of the two invariant features proposed, so as to better estimate the statistical properties of the audio signal. Considering that the embodiment is based on the audio histogram Figure 3 bin is 1 group of embedded information, and if the number of samples in a certain group is too small, the embedding quality will be seriously affected, so the embedding range [-λA, λA] is selected according to the original audio average value A and λ, wherein the range of λ is generally [2.4, 3], and 2.4 is taken in the embodiment. The larger λ represents the larger selection range, the more samples, and the stronger robustness but the lower signal to noise ratio (SNR), so a suitable value needs to be selected to balance the SNR and the robustness; specifically:
[0081] 101, assuming a binary watermark sequence , the original audio is represented as , wherein is the length of the embedded information, N represents the length of the original audio, indicates the specific i-th watermark in the W sequence. The absolute average value A is calculated as the sum of the absolute values of the audio samples in their duration:
[0082] ;
[0083] 2. Distribute the number of bins according to the number of watermark bits embedded. This embodiment uses 3 bins to embed 2 bits, and the number of bins required under normal circumstances = (number of watermark bits / 2) 3. Divide the samples according to the number of bins to obtain the time-domain histogram H; specifically:
[0084] 201. Embed the watermark as Lw bits, and the algorithm requires (Lw / 2) 3 bins to complete the embedding. According to the mean A calculated in step 1, select the embedding range [-λA, λA], and divide it into (Lw / 2) 3+3 bins, each with a width of Assign the samples in the range to the corresponding bins according to their amplitude, and obtain the time-domain histogram , L = (Lw / 2) 3+3; where, represents the specific i-th bin in the bin set H.
[0085] 202. The extra 3 bins are due to the small number of samples in the bins near the middle region of the histogram, so they are excluded. In addition, within the same audio range, the larger the width of each bin, the more samples it contains, and the more stable the histogram shape, the stronger the robustness, but the SNR will also be reduced. Therefore, a trade-off must be made between robustness, capacity, and SNR. After experiments, the embedded watermark information in this embodiment is 80 bits, and the SNR is generally 42db~45db.
[0086] 3. Perform statistics on the time-domain histogram , and separately perform secondary DWT transformation on the samples within the same bin range. Record the breakpoints between each discontinuous bin to ensure that the subsequent inverse transformation can proceed normally, and obtain the low-frequency domain histogram ; specifically:
[0087] 301. In step 2, the division of the sample region has been completed, and we need to perform a segmented DWT transformation on the histogram. The purpose of this is that, on the one hand, the frequency domain has stronger resistance to Gaussian and compression attacks, and on the other hand, it is to preserve the shape characteristics of the histogram, so that the time linear scaling invariance of the histogram shape is extended to the DWT domain.
[0088] 302、Assuming the target range [-λA, λA] samples have been assigned to bins with equal width, the first bin, i.e. the range [-λA, -λA + M), is subjected to a two-level DWT transform, followed by the second bin, i.e. the range [-λA + M, -λA + 2M), and so on, with each bin interval increasing M over the previous bin range, such as the third bin is [-λA + 2M, -λA + 3M) and the fourth bin is [-λA + 3M, -λA + 4M). As shown in the following figure, Figure 4 belongs to [-1.5M, -0.5M), , , belongs to [-0.5M, 0.5M), belongs to [0.5M, 1.5M), and a DWT transform is needed for each M range of the piecewise function, resulting in , L = (Lw / 2) 3+3. At the same time, because the samples are usually randomly distributed in the audio, the samples in the range are not necessarily continuous in the time domain, and the breakpoint coordinates between the discontinuous samples need to be recorded during the transform for subsequent inverse transform; wherein, represents the sample at the start of the function, represents the sample with an amplitude falling at -0.5M boundary, , , , represents the sample with an amplitude falling at 0.5M boundary, represents the sample with an amplitude falling at 0, represents the i-th bin in the frequency domain function set obtained after DWT transform of L M range bins.
[0089] 4. According to Figure 3 , every 2 bits of the watermark modifies the shape of 3 bins to meet the expected target, resulting in the modified histogram ; Specifically:
[0090] 401、After completing the DWT transform, the frequency domain histogram shape is obtained, and each group of shapes is modified according to the corresponding watermark information with 3 bins as a group, and the specific modification shape is as follows Figure 3 As shown. Assuming the number of samples in each bin is a, b, and c, and the algorithm has thresholds T=[2, 3], T1=2.5, T2=3, the rule changes are as follows:
[0091] Assuming the current embedded information w(i) is "00", modify the shape to low-medium-high:
[0092] ;
[0093] This represents the i-th sample in the first bin before modification. This represents the i-th sample in the second bin before modification. For the i-th sample in each bin, The corresponding modified value, M, represents the width of the bin in the time-domain histogram. However, due to the two-stage DWT transformation, based on the characteristic of the signal transitioning from the time domain to the frequency domain, the bin width in the frequency domain will double, becoming 2M. In the formula, a total of I1 + I2 samples in the bin are modified; samples I1 and I2 are transferred to bin2 and bin3 respectively. Sample migration is performed in the order from a to c, ensuring... If the adjustment is not met Transfer I1 samples from a to b to obtain the adjusted... and ;ensure If the adjustment is not met Transfer I2 samples from a to c to obtain the adjusted... and ;in, This represents the threshold set in advance, T1=2.5.
[0094] 402. Assuming the current embedded information w(i) is "01", modify the shape to the minimum value of b:
[0095] ;
[0096] Sample transfer is performed from b to a, c to ensure If the adjustment is not met , in b Each sample is transferred to a. Each sample is transferred to c to obtain the adjusted sample. , and ;in, This represents the threshold set in advance, T2=3.
[0097] 403. Assuming the current embedded information w(i) is "10", modifying the shape to b yields the highest result:
[0098] ;
[0099] Sample migration from a, c to b, ensuring If the adjustment , transfer I1 samples in a to b, and transfer I2 samples in c to b, to obtain the adjusted , and .
[0100] 404, assuming that the current embedded information w(i) is "00", the shape is modified to high-middle-low:
[0101] ;
[0102] wherein, represents the i-th sample in the third bin before modification;
[0103] Sample migration from c to a in order, ensuring If the adjustment , transfer I1 samples in c to b, to obtain the adjusted and ; ensuring If the adjustment , transfer I2 samples in to a, to obtain the adjusted and .
[0104] 5. The modified histogram is combined with the breakpoints recorded in step 3 to perform inverse two-level DWT to obtain the embedded watermark audio ; specifically:
[0105] 501, after modifying the shape of each bin group of the histogram, the frequency domain histogram Hw of the embedded information is obtained, and the modified frequency domain histogram Hw is combined with the breakpoints between each discontinuous sample recorded in step 3 and the high and low frequency coefficients to perform inverse two-level DWT to obtain the embedded watermark audio Iw.
[0106] Further, obtaining the second time domain histogram corresponding to the embedded watermark audio includes:
[0107] Pretreating the embedded watermark audio to calculate the absolute mean value of the embedded watermark audio;
[0108] Based on the absolute mean value of the embedded watermark audio and λ, the extraction range is selected;
[0109] The samples in the extraction range are subjected to binning operation to obtain the second time-domain histogram.
[0110] Specifically, in the embodiment, the audio histogram shape is used to resist synchronous attack, and the watermark extraction process is as shown in the following steps. Figure 2 The specific steps are as follows:
[0111] 6. The initial audio I is preprocessed to calculate the absolute mean value of the audio, and the extraction range [-λA, -λA] is determined according to the mean value and λ.
[0112] 7. The samples in the range [-λA, -λA] are subjected to binning operation to obtain the time-domain statistical histogram .
[0113] 8. The samples belonging to the same bin are subjected to secondary DWT transformation to obtain the histogram in the low-frequency domain .
[0114] 9. The information is extracted according to the shape of the low-frequency histogram, every 3 bins as a group Figure 3 . Specifically:
[0115] 901. After the DWT transformation, the frequency domain histogram shape is obtained, and the watermark information is extracted every 3 bins as a group. Assuming that the number of samples in each bin group is a, b, and c.
[0116] 902. If , the watermark information w'(i) of the current bin group is "11".
[0117] 903. If , the watermark information w'(i) of the current bin group is "00".
[0118] 904. If , the watermark information w'(i) of the current bin group is "10".
[0119] 905. If , the watermark information w'(i) of the current bin group is "01".
[0120] In the example of the present application, three different audios of 90s popular music track1, 121s rock music track2 and 150s conversation in noisy environment track3 are used as experimental objects. The three groups of audios have different characteristics, and various audios in daily life have these characteristics, so the three groups of audios are used as experimental objects, so that the experimental results have generalizability. The sampling rates of the selected audios are all 44.1 kHz, and the differences between different audios are not large, so the present application can be extended to various audios.
[0121] The application example provides an audio robust watermarking algorithm against synchronization attacks and conventional signal processing, mainly including a watermark embedding method, a watermark extraction method and a histogram shape modification method.
[0122] The watermark embedding method provided by the application example comprises the following steps:
[0123] The original audio is preprocessed;
[0124] The binary watermark sequence , the original audio is expressed as , wherein is the length of embedded information, and N represents the length of the audio. The absolute average value A is calculated as the sum of absolute values of audio samples within their duration time:
[0125] ;
[0126] b) a time-domain histogram is obtained by counting;
[0127] Supposing that the embedded watermark is Lw bits, the algorithm requires (Lw / 2) 3 bins to complete watermark embedding. The embedding range [-λA, λA] selected in step 1 is divided into (Lw / 2) 3+3 parts, and each bin has a width of , and the samples in the range are distributed into corresponding bins according to their amplitude, so as to obtain a time-domain statistical histogram , L=(Lw / 2) 3+3.
[0128] c) segmenting the time-domain histogram by DWT;
[0129] After the samples in the target range [-λA, λA] are distributed into bins with equal width, the samples in each bin (i) in the same i M range are transformed by two-level DWT one by one, and the breakpoint coordinates between discontinuous samples are recorded during the transformation, so as to be used during subsequent inverse transformation, so as to obtain a frequency-domain statistical histogram , L=(Lw / 2) 3+3.
[0130] d) modifying the frequency-domain histogram shape to embed watermark information;
[0131] After the DWT transformation is completed, the frequency-domain histogram shape is obtained, and each group of shapes is modified according to corresponding watermark information by taking 3 bins as a group. Supposing that the sample numbers of each group of bins are a, b and c, the change rule is expressed as follows:
[0132] When the embedded information w(i) is "00", the modified shape is low-middle-high:
[0133] ;
[0134] Sample migration is performed in the order from a to c, ensuring that If the adjustment condition is not met, adjust I1 samples in a are transferred to b to obtain the adjusted and ; ensure that If the adjustment condition is not met, adjust I2 samples in a are transferred to c to obtain the adjusted and .
[0135] When the embedded information w(i) is "01", the modified shape is b lowest:
[0136] ;
[0137] Sample migration is performed from b to a, c, ensuring that If the adjustment condition is not met, adjust I1 samples in b are transferred to a, and I2 samples are transferred to c to obtain the adjusted , and . When the embedded information w(i) is "10", the modified shape is b highest:
[0138]
[0139] ;
[0140] Sample migration is performed from a, c to b, ensuring that If the adjustment condition is not met, adjust I1 samples in a are transferred to b, and I2 samples in c are transferred to b to obtain the adjusted , and . When the embedded information w(i) is "00", the modified shape is high-middle-low:
[0141]
[0142] ;
[0143] Sample migration is performed in the order from c to a, ensuring that If the adjustment condition is not met, adjust , the I1 samples in c are transferred to b, and the adjusted and ; ensure , if the adjustment , the I2 samples in d are transferred to a, and the adjusted and . .
[0144] e) extract watermark information:
[0145] After the frequency domain histogram Hw with watermark information is obtained by preprocessing the watermark audio Iw, the watermark information is extracted in groups of 3 bins. Assume that the number of samples in each group of bins is a, b, and c.
[0146] If a , the watermark information w'(i) of the current bin group is "11".
[0147] If b , the watermark information w'(i) of the current bin group is "00".
[0148] If c , the watermark information w'(i) of the current bin group is "10".
[0149] If a , the watermark information w'(i) of the current bin group is "01".
[0150] f) convert the modified frequency domain histogram back to the watermark audio:
[0151] After modifying the shape of each bin group of the histogram, the frequency domain histogram Hw with embedded information is obtained, and the modified frequency domain histogram Hw is inversely transformed by the second-level DWT to obtain the audio Iw with embedded watermark, combined with the breakpoints between each discontinuous sample and the high and low frequency coefficients recorded in step 3.
[0152] The experimental results are shown in the following table.
[0153] For audio robust watermarking algorithm, the error rate of audio with watermark information after being attacked is below 20%, which can be considered as good robustness.
[0154] The experimental results based on audio track1 are shown in Tables 1-5, which show that the algorithm can resist MP3 compression with a quality factor of 48Kbps, stretching attack, amplitude transformation attack, and high-pass filter attack with a Gaussian noise of 30db and a low-pass filter of 4k or more.
[0155] The results of the experiments based on audio track 2 are shown in Tables 6-10, which show that the algorithm can resist MP3 compression attacks with a quality factor of 48 Kbps, stretching attacks, amplitude transformation attacks, Gaussian noise attacks with 30 db, and low-pass filtering attacks with 4k or more.
[0156] The results of the experiments based on audio track 3 are shown in Tables 11-15, which show that the algorithm can resist MP3 compression attacks with a quality factor of 48 Kbps, stretching attacks, amplitude transformation attacks, Gaussian noise attacks with 30 db, and low-pass filtering attacks with 4k or more.
[0157] Table 1 Error rate of audio track 1 under Gaussian noise attack (robust watermark embedded is 80 bits)
[0158] Gaussian noise (dB) 30 35 40 45 Bit error rate 19% 3% 0% 0%
[0159] Table 2 Error rate of audio track 1 under compression factor attack
[0160] MP3 compression (compression factor) 48 56 64 80 96 128 Bit error rate 16% 11.3% 6% 2% 2% 3%
[0161] Table 3 Error rate of audio track 1 under amplitude transformation attack
[0162] Amplitude transform (%) -40 -30 -20 -10 +10 +20 +30 +40 Bit error rate 5% 0% 0% 0% 0% 0% 0% 1.67%
[0163] Table 4 Error rate of audio track 1 under stretching factor attack
[0164] Stretch (stretch factor) % -30 -20 -15 -10 -5 +5 +10 +15 +20 +30 Bit error rate 0% 0% 0% 0% 0% 0% 0% 0% 0% 5%
[0165] Table 5 Error rate of audio track 1 under low-pass filtering attack
[0166] Low pass filtering 4k 5k 6k 7k 8k Bit error rate 7.4% 0% 0% 0% 0%
[0167] Table 6 Error rate of audio track 2 under Gaussian noise attack (robust watermark embedded is 80 bits)
[0168] Gaussian noise (dB) 30 35 40 45 Bit error rate 13.7% 0% 0% 0%
[0169] Table 7 Error rate of audio track 2 under compression factor attack
[0170] MP3 compression (compression factor) 48 56 64 80 96 128 Bit error rate 20% 15% 12.7% 8% 5% 0%
[0171] Table 8 Error rate of audio track 2 under amplitude transformation attack
[0172] Amplitude transform (%) -40 -30 -20 -10 +10 +20 +30 +40 Bit error rate 10% 7% 0% 0% 0% 0% 0% 5%
[0173] Table 9: BER of audio track 2 under attack of stretch factor
[0174] Stretch (stretch factor) % -30 -20 -15 -10 -5 +5 +10 +15 +20 +30 Bit error rate 0% 0% 0% 0% 0% 0% 0% 0% 0% 0%
[0175] Table 10: BER of audio track 2 under attack of low-pass filter
[0176] Low pass filtering 4k 5k 6k 7k 8k Bit error rate 13.7% 4% 2% 2% 2%
[0177] Table 11: BER of audio track 3 under attack of Gaussian noise (robust watermarking of 80 bits)
[0178] Gaussian noise (dB) 30 35 40 45 Bit error rate 15% 8% 0% 0%
[0179] Table 12: BER of audio track 3 under attack of compression factor
[0180] MP3 compression (compression factor) 48 56 64 80 96 128 Bit error rate 9% 4% 0% 0% 5% 0%
[0181] Table 13: BER of audio track 3 under attack of amplitude transformation
[0182] Amplitude transform (%) -40 -30 -20 -10 +10 +20 +30 +40 Bit error rate 0% 0% 0% 0% 0% 0% 0% 0%
[0183] Table 14: BER of audio track 3 under attack of stretch factor
[0184] Stretch (stretch factor) % -30 -20 -15 -10 -5 +5 +10 +15 +20 +30 Bit error rate 0% 0% 0% 0% 0% 0% 0% 0% 0% 0%
[0185] Table 15: BER of audio track 3 under attack of low-pass filter
[0186] Low pass filtering 4k 5k 6k 7k 8k Bit error rate 15% 14% 12% 12% 12%
[0187] The robust watermarking algorithm in the present application greatly improves the embedding capacity by reasonable design of the histogram embedding shape, and expands the time domain characteristics of the histogram to the frequency domain using the segmented DWT method to enhance the robustness.
[0188] Robustness to synchronization attack and MP3 compression is an important aspect of the performance of an audio robust watermarking algorithm, and most of the previous audio watermarking algorithms can only resist one of the attacks. In the present embodiment, good robustness to both synchronization attack and compression attack is achieved, and more diverse histogram shapes are designed to improve the embedding capacity while ensuring robustness. Under the same number of histogram sampling samples, the algorithm has a higher embedding capacity than similar algorithms. This characteristic shows that even for some audio carriers with fewer sampling points (such as a sampling rate of 8000hz), the algorithm still has the ability to embed, thereby improving the practical application value of the algorithm.
[0189] Compared with the prior art, the embodiment can resist various stretching attacks, has stronger robustness after being subjected to synchronous attacks such as MP3 compression, resampling and stretching or conventional signal operation, and the watermark can be effectively extracted, in particular:
[0190] The embodiment is based on histogram shape, and by converting the time domain features of the histogram into the frequency domain through the segmented DWT method, the robustness of the algorithm is greatly improved. Compared with the previous algorithm, the application utilizes the diversity of the histogram shape to achieve higher capacity embedding on the same number of bins, thereby solving the problem that it is difficult to embed a watermark in some speech files due to the small number of samples, which has a great enlightening effect on the robust watermark research in the aspect of modifying the histogram.
[0191] The embodiment utilizes the robustness of the low-frequency domain coefficients after DWT transformation to compression, Gaussian noise and other attacks, and by the segmented DWT method, the time domain stretching invariant characteristic is extended to the frequency domain under the premise of preserving the histogram shape, thereby improving the resistance to conventional attacks and retaining the robustness to stretching attacks.
[0192] The embodiment can effectively extract the watermark information under different signal processing such as Gaussian noise, MP3 compression and sound attacks, which meets the requirements of daily digital forensics and digital authentication; at the same time, the embodiment can be applied to different audios, and good results are achieved on different audios.
[0193] The above-described embodiments only describe the preferred modes of the application, and do not limit the scope of the application. Without departing from the design spirit of the application, various modifications and improvements to the technical solutions of the application made by those skilled in the art shall fall within the protection scope determined by the claims of the application.
Claims
1. A method of audio histogram shape watermarking against synchronization attacks, characterized in that, The method comprises the following steps: According to the number of pre-allocated bins, the pre-selected audio samples are divided to obtain a first time-domain histogram; wherein, the bin is a statistical histogram sample interval; The first time-domain histogram is subjected to a segmented DWT transformation and an inverse DWT transformation to obtain a watermark-embedded audio; The watermark-embedded audio is obtained by: The first time-domain histogram is subjected to a segmented DWT transformation, and the breakpoints between each discontinuous bin are recorded to obtain a low-frequency-domain first time-domain histogram; The low-frequency-domain first time-domain histogram is subjected to information embedding to modify the shape of the bin to obtain a modified first time-domain histogram; The method for modifying the shape of the bin comprises: The bin interval is grouped by every 3 as a group, and each group is used to embed 2 bits of information; for the four different 2-bit information units: 00, 01, 10, 11, in the watermark sequence, different embedding strategies are used respectively, wherein: When the current embedded information is 00, the shape of the bin is modified to low-middle-high; When the current embedded information is 01, the shape of the middle bin is modified to the lowest; When the current embedded information is 10, the shape of the middle bin is modified to the highest; When the current embedded information is 11, the shape is modified to high-middle-low; Low, middle and high represent the size relationship of three consecutive bin intervals in the sample quantity; The modified first time-domain histogram is subjected to an inverse two-level DWT transformation combined with the recorded breakpoints to obtain the watermark-embedded audio; The second time-domain histogram corresponding to the watermark-embedded audio is obtained, and the segmented DWT transformation is performed on the second time-domain histogram to obtain a low-frequency-domain second time-domain histogram; Information extraction is performed on the low-frequency-domain second time-domain histogram to obtain the watermark-embedded audio.
2. The synchronization attack resistant audio histogram shape watermarking method of claim 1, wherein, The first time-domain histogram is obtained by: The initial audio is preprocessed to calculate the audio absolute mean value; Based on the audio absolute mean value and a preset constant value λ, the embedding range, i.e., the audio sample, is selected; The number of bins is allocated according to the number of watermark bits to be embedded; wherein, 3 bins are used to embed 2 bits; The samples are divided according to the number of bins to obtain the first time-domain histogram.
3. The synchronization attack resistant audio histogram shape watermarking method of claim 1, wherein, The segmented DWT transformation of the first time-domain histogram comprises: For the audio samples [-λA, λA] that have been allocated to the bins with equal width, wherein A is the audio absolute mean value, and λ is a preset constant value used to control the range of audio to be processed; the samples in the first bin, i.e., the range [-λA, -λA + M), are subjected to a two-level DWT transformation, wherein M represents the width of each bin, the samples in the second bin, i.e., the range [-λA + M, -λA + 2M), are subjected to a two-level DWT transformation, and each subsequent bin interval is increased by M on the range of the previous bin, and so on to obtain the interval of each bin; the segmented function in each M range is subjected to a DWT transformation to obtain the low-frequency-domain first time-domain histogram.
4. The synchronization attack resistant audio histogram shape watermarking method of claim 1, wherein, The second time-domain histogram corresponding to the watermark-embedded audio is obtained by: Preprocessing the embedded watermark audio, calculating an absolute mean value of the embedded watermark audio; Selecting an extraction range based on the absolute mean value of the embedded watermark audio and a preset constant value λ; Performing a bin operation on samples in the extraction range to obtain a second time-domain histogram.
5. The synchronization attack resistant audio histogram shape watermarking method of claim 1, wherein, The information extraction on the second time-domain histogram in the low-frequency domain includes: Extracting watermark information in 3 bins as a group, and setting the sample numbers of each group of bins as a, b, and c: if then the current bin group watermark information is 11; if then the current bin group watermark information is 00; if then the current bin group watermark information is 10; if then the current bin group watermark information is 01.
6. The synchronization attack resistant audio histogram shape watermarking method of claim 2, wherein, The absolute mean value of the audio is: ; where A is the audio absolute mean value, N is the length of the original audio, is the i-th sample in each bin.
Citation Information
Patent Citations
Digital audio watermarking method based on invariant characteristic of histogram
CN102074237A
High-capacity digital audio reversible watermark processing method
CN103050120A