Real-time robust speech watermarking method based on logarithmic polar coordinates with adjustable frequency domain
Through the logarithmic bottom-tunable frequency domain logarithmic polar coordinate method, large-capacity watermarks are embedded in the voice watermark in real time, and normalized frequency correlation value fusion extraction is used to solve the problems of real-time embedding difficulties and poor robustness in the prior art, and the extraction accuracy and inability to perceive watermarks are improved.
Patent Information
- Application Number
- CN202310129081.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-15
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-02-15
AI Technical Summary
The existing voice watermarking technology cannot realize the real-time embedding of large-capacity watermarks, and the embedding is poor in perceptuality and robustness.
Real-time and robust voice watermarking method based on logarithmic bottom-tunable frequency domain logarithmic polar coordinates is adopted, and watermark embedding is performed in real-time audio clips at the frame level according to the characteristics of amplitude on the logarithmic coordinates, and extracted by fusion of normalized frequency correlation values.
While real-time embedding of large-capacity watermarks, it improves the accuracy and imperceptibility of watermark extraction.
Smart Images

Figure CN116110408B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimedia signal processing, and more particularly to a real-time robust voice watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates. Background Art
[0002] Speech is the most important form of human communication, carrying valuable information about who, what, and how the speaker speaks. Currently, there are three main reasons for applying speech signals to computer science: (1) speech is easy to generate, capture, and transmit; (2) speech signals can be acquired over long distances; and (3) speech also carries other types of information, such as emotion, age, and gender.
[0003] Digital watermarking is a technology that exploits the redundant information and human perception inherent in digital media (images, video, audio, text, etc.) to embed additional information (watermarks) without compromising the quality of the original media. The original purpose and primary use of digital watermarking was to protect the copyright of digital works. With years of research and development, its applications have expanded to include access control, digital fingerprinting, content authentication, and implicit annotation. Digital voice watermarking is a type of digital watermark. With the increasing application of advanced communication technologies such as mobile wireless and internet-based telephony in our daily lives, digital voice watermarking has become crucial for ensuring the security of voice signals.
[0004] In recent years, scholars have conducted research on digital speech watermarking technology. Due to the unique characteristics of audio signals, desynchronization attacks can cause the framing of the watermarked audio after the attack to become out of sync with that before the attack, making it impossible to correctly extract the watermark. To combat desynchronization attacks, a common approach is to embed identification information within each audio frame. During watermark detection, the identification information of each frame is first detected to locate the content of the corresponding frame. However, current research often only embeds recognizable identification information within the audio file, but cannot achieve real-time watermark embedding. Real-time watermark embedding requires that the audio that can be embedded at one time is extremely short, only a few milliseconds, making watermark embedding extremely difficult.
[0005] Prior art discloses a robust audio watermarking method based on Fourier discrete logarithmic coordinate transform. This method embeds a watermark in the discrete Fourier amplitude coefficients of the audio, determining the embedded Fourier amplitude coefficients via their discrete logarithmic coordinates. This method ensures robustness and accurate watermark extraction, but when applied to human speech, it suffers from poor perceptibility, is unable to achieve real-time frame-level embedding, and has a limited embedded information capacity. Summary of the Invention
[0006] In order to solve the problems that large-capacity watermarks cannot be embedded in real time in current speech, and the watermarks are imperceptible and have poor robustness after embedding, the present invention proposes a real-time robust speech watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates, which expands the capacity of embedded information, realizes real-time frame-level embedding, and improves the accuracy and imperceptibility of watermark extraction.
[0007] In order to achieve the above technical effects, the technical solutions of the present invention are as follows:
[0008] A real-time robust speech watermarking method based on logarithmic polar coordinates with adjustable logarithmic base in the frequency domain, comprising a real-time robust speech watermark embedding process at the frame level with a length of 0.1 to 0.2 seconds and a watermark extraction process. The real-time watermark embedding process embeds multi-bit information into the host audio according to the amplitude characteristics of the logarithmic coordinates with adjustable logarithmic base in the real-time audio segment at the frame level. The watermark extraction process, based on the real-time watermark embedding process, performs a watermark extraction operation on the host audio that has been embedded with the watermark and has been attacked by fusing normalized frequency-related values, thereby extracting the watermark information embedded in the host audio.
[0009] Real-time embedding means embedding the watermark while the person is speaking through the device, and playing the watermarked audio. The delay of audio and video data on the device side can reach 30-200ms, so the watermark embedding in this technical solution is embedded in audio clips within 200ms.
[0010] The real-time robust speech watermarking method based on logarithmic polar coordinates with logarithmic base adjustable frequency domain proposed in this technical solution addresses the problems that existing speech cannot embed large-capacity watermarks in real time, and the watermarks are imperceptible and robust after embedding. By utilizing the mutually coordinated real-time watermark embedding process and watermark extraction process, the watermark is embedded in the real-time audio clip at the frame level according to the characteristics of the amplitude on the logarithmic coordinates, and the watermark is extracted by fusing normalized frequency-related values, thereby improving the accuracy and imperceptibility of watermark extraction.
[0011] Preferably, the real-time watermark embedding process at least includes:
[0012] S1. Determine the lengths of the pseudo-random modulation sequence and template sequence based on the number of bits in the watermark bit information sequence to be embedded. Combined with the key, randomly generate a bipolar pseudo-random modulation sequence and template sequence. Perform spread spectrum calculation on each bit of the watermark bit information sequence based on the bipolar pseudo-random modulation sequence, and combine with the template sequence to obtain the watermark sequence to be embedded.
[0013] S2. Determine the host audio output from the device in real time, perform pre-emphasis on the host voice audio at a very short frame level, obtain the determined host audio, and expand the determined host audio;
[0014] S3. According to the extended host audio frame length, the logarithmic base value of the logarithmic base adjustable frequency domain logarithmic coordinate transform is adaptively determined between 1 and 2 to ensure real-time watermark embedding;
[0015] S4. Perform a one-dimensional discrete Fourier transform on the expanded host audio segment, move the transformed DC component to the center of the Fourier magnitude spectrum, use the right half of the Fourier magnitude spectrum as the embedding region, map the coordinates of points within the embedding region to log-polar coordinates, and embed the watermark sequence based on the amplitude characteristics.
[0016] S5. Spread each bit of the watermark bit information sequence. Calculate the Fourier amplitude spectrum averages of the regions corresponding to the sequence obtained after the spread, at the positions corresponding to the embedded +1, -1, and the sequence obtained after the spread. Based on these three amplitude averages, embed the watermark bit information according to the situation.
[0017] S6. Based on the central symmetry of the Fourier amplitude spectrum, the symmetrical coefficients of the right half are copied to the left half, and then an inverse Fourier transform is performed. The watermark embedding process is completed, and the entire audio is de-emphasized to obtain the watermarked audio. The de-emphasis operation is the opposite of the pre-emphasis operation process.
[0018] Preferably, the process of step S1 includes:
[0019] S101. Set the length of the bipolar pseudo-random modulation sequence and the template sequence, denoted as L ps and L TS ; Let the key be key, and use the key to generate the bipolar pseudo-random modulation sequence ps and template sequence TS:
[0020] ps={ps i ; 1≤i≤L ps ,ps i ∈{-1,1}}
[0021] TS={TS j ; 1≤j≤L TS ,TS j ∈{-1,1}}
[0022] Among them, ps i TS j They represent the i-th element of the bipolar pseudo-random modulation sequence and the j-th element of the template sequence respectively;
[0023] S102. Let the length be L ms The watermark bit information sequence is ms:
[0024] ms={ms i ; 1≤i≤L ms,ms i ∈{-1,1}}
[0025] Among them, ms i Represents the i-th element of the watermark bit information sequence;
[0026] Use bipolar pseudo-random modulation sequence ps to modulate each bit ms i Perform spread spectrum modulation: If ms i =1, then it is spread spectrum modulated into a ps in-phase sequence, that is, W i =+1×ps; if ms i =-1, then it is spread spectrum modulated into the inverse sequence of ps, that is, W i = -1×ps, and finally get the meaningful watermark information array W = {W i ; 1≤i≤L ms ,W i ∈{-ps,ps}};
[0027] S103. The obtained meaningful watermark information array W = {W i ; 1≤i≤L ms ,W i ∈{-ps,ps}} and the template sequence TS are arranged in order to form a sequence of length M=L ms ×L ps +L TS The sequence WT={WT i ; 1≤i≤M}, as the watermark sequence to be embedded, where the element WT in WT i It is composed of "1" and "-1", and the template sequence TS is stored in the last L of the watermark sequence WT in order. TS locations.
[0028] Preferably, the process of step S2 includes:
[0029] S201. Assume that the host audio voice signal is S={S t ; 1≤t≤L}, where S t represents the t-th sample point, L represents the length of the signal S, and a pre-emphasis operation is performed on it. The expression of the pre-emphasis operation is:
[0030] in, is the pre-emphasized speech signal, α=0.97;
[0031] S202. If the host audio voice duration is x milliseconds, the pre-emphasized voice signal duration is also x milliseconds; before watermark embedding, the pre-emphasized voice signal x milliseconds Expand to y millisecond segments: Initialize an all-zero matrix of length y milliseconds, store the x millisecond audio segment in the last x millisecond matrix area of the all-zero matrix, and use the expanded y millisecond matrix as the segment to be embedded to obtain matrix J.
[0032] Preferably, in step S4, a one-dimensional discrete Fourier transform of length d is performed on the y millisecond matrix J, the transformed DC component is moved to the center of the Fourier amplitude spectrum, and the center of the Fourier amplitude spectrum is used as the origin of the rectangular coordinate system. The watermark is embedded in the right half of the Fourier amplitude spectrum, and the embedding area is located at the normalized frequency value f of the Fourier coefficient amplitude spectrum. n nearby;
[0033] The process of mapping the coordinates of the points in the embedded region to logarithmic polar coordinates is as follows: the rectangular coordinates r of the Fourier coefficients of the embedded region are transformed into discrete logarithmic polar coordinates lρ. The transformation formula is:
[0034]
[0035]
[0036] Among them, a is a constant greater than 1 but close to 1, f n ×d is the origin of the logarithmic coordinate, offset is an offset constant that ensures the discrete logarithmic coordinate is not less than zero, and the floor() function represents the floor function to balance the robustness and invisibility of the watermark.
[0037] Preferably, the process of step S5 includes:
[0038] S501. For each information bit ms i , spread spectrum modulation is the in-phase sequence or reverse sequence W of ps i , before watermark embedding, calculate W i The average amplitude value amp in the corresponding rectangular coordinate area avg , and then calculate W i The average amplitude of the Fourier coefficients at the positions of +1 and -1 embedded in the corresponding rectangular coordinate area are recorded as and
[0039] S502. When When W i The embedding formula used in the corresponding rectangular coordinate area is as follows:
[0040]
[0041] when When W i The embedding formula used in the corresponding rectangular coordinate area is as follows:
[0042]
[0043] Among them, amp0 k is the Fourier coefficient amplitude of the original audio, ampw k is the amplitude of the audio Fourier coefficient after embedding the watermark, β=0.00001, δ is the watermark embedding strength, w k Indicates that the watermark bit to be embedded when the rectangular coordinate k is mapped to the logarithmic coordinate is "1" or "-1".
[0044] Preferably, the watermark extraction process at least includes:
[0045] SA determines the audio to be tested, performs pre-emphasis operation on the entire audio segment to be tested, and intercepts the pre-emphasized audio to obtain a collection of audio segments;
[0046] SB. Perform a one-dimensional discrete Fourier transform on the pre-emphasized audio and the set of further truncated audio segments, moving the transformed DC component to the center of the Fourier amplitude spectrum. Using the center of the amplitude spectrum as the origin of the rectangular coordinate system, extract the watermark from the right half of the Fourier coefficient amplitude spectrum. Map the rectangular coordinates of the amplitude coefficients within the extraction range to log-polar coordinates. Sum the amplitude coefficients with the same log-polar coordinates after mapping, and use this sum as an element of the Fourier amplitude coefficient sequence.
[0047] SC. Based on the phase correlation principle, the original template sequence and the Fourier amplitude coefficient sequence are quickly matched and calculated to preliminarily determine the synchronization position of the embedded watermark. The amplitude matrix of the synchronized amplitude coefficient sequence is intercepted from the center position, and the final synchronization position is further determined using the neighborhood search method.
[0048] SD. uses a pseudo-random modulation sequence to despread the amplitude matrix and integrate the sub-segment correlation values to extract the watermark information.
[0049] Preferably, step SA comprises the following steps:
[0050] SA01. Assume that the input watermarked speech signal is SW:
[0051] SW={SW t ; 1≤t≤L ′}
[0052] Among them, SW t represents the t-th sample point, L′ represents the length of the signal SW; t Perform pre-emphasis operation to obtain the pre-emphasized watermarked speech signal SW * .
[0053] SA02. To SW * Cut with a step size of z2 milliseconds and a window length of y milliseconds until the sliding window reaches the end and record the audio segment as k represents the kth segment during sliding capture.
[0054] Preferably, step SB comprises the following steps:
[0055] SB01. To SW * and Perform one-dimensional discrete Fourier transform respectively, and move the DC component to the center of the Fourier amplitude spectrum; take the center of the amplitude spectrum as the origin of the rectangular coordinate system, and normalize the right half plane of the Fourier coefficient amplitude spectrum to the frequency f n The rectangular coordinates r′ of the Fourier coefficients near are transformed into discrete logarithmic coordinates lρ′. The transformation formula is:
[0056]
[0057]
[0058] Wherein, M′=λM, λ≥1, offset′ is an offset constant that ensures that the discrete logarithmic coordinate is not less than zero, and the floor() function represents the floor rounding function;
[0059] SB02. Initialize an all-zero matrix amp of length M″=λ×μ×M, where μ is a positive integer not less than 1; map the rectangular coordinates to the logarithmic polar coordinate system, and sum the Fourier coefficient amplitudes with the same discrete logarithmic coordinate lρ′ as the Fourier coefficient amplitude sequence amp lρ′ An element of , thus obtaining a Fourier coefficient amplitude sequence amp.
[0060] Preferably, step SC comprises the following steps:
[0061] SC01. Use the same key as the embedding algorithm to generate a key with a length of L TS The template sequence TS={TS i ; 1≤i≤L TS ,TS i ∈{-1,1}}, each TS i Expanded into μ TS i , get TS1; initialize the full zero matrix TS with length M″ m , store TS1 in TS m The second half of the synchronization template TS is obtained m ;
[0062] SC02. Obtain the translation phase correlation value through phase correlation fast matching calculation. The formula for calculating the translation phase correlation value is:
[0063]
[0064] Among them, X(k) is the correlation value sequence, φ amp (u) is the phase angle of amp(u), G * (u)=(DFT(TS m (i))) * Is the synchronization template TS m The complex conjugate of the one-dimensional Fourier transform coefficient, "*" indicates the complex conjugate;
[0065] According to the maximum value of the correlation value sequence X(k), the position of the embedded watermark WT in the amplitude sequence amp is preliminarily determined, which is recorded as k max ;
[0066] SC03. max The positions of the left and right neighbors, and k = 1 as an alternative sequence of synchronization positions col = {1, k max -1,k max ,k max +1}, take the synchronization position col i , the first col in the amplitude sequence amp i -1 position is moved to the end of the amplitude sequence amp to obtain the synchronized amplitude sequence, and the synchronized amplitude sequence is cut out from the center position to obtain an amplitude matrix of length μ×M, which is recorded as
[0067] SC04. Calculation The mean of the non-zero amplitudes at each μ position in is stored in the amplitude matrix amp2 of length M in sequence;
[0068] SC05. Calculate each synchronization position col in col i The correlation value Y(i) between the corresponding amplitude matrix amp2 and the template sequence TS is calculated as follows:
[0069]
[0070] The final synchronization position is determined according to the maximum value of the correlation value sequence Y(i), which is denoted as col f .
[0071] Preferably, step SD comprises the following steps:
[0072] SD01. Despread spectrum modulate the amplitude matrix amp2 using the original pseudo-random modulation sequence ps and calculate the normalized correlation value corresponding to each information bit. The process is as follows:
[0073] Take L from amp2 in order msSegments do not overlap and are of length L ps The sequence W′ i :
[0074]
[0075] Calculate the normalized correlation value Q between the amplitude corresponding to the meaningful information sequence and the original pseudo-random modulation sequence ps. The calculation formula is:
[0076]
[0077] Q={Q i ; 1≤i≤L ms}
[0078] Calculate the normalized correlation value H between the amplitude corresponding to the synchronization template sequence and the template sequence TS. The calculation formula is:
[0079]
[0080] SD02. According to step SD01, calculate SW * and Normalized correlation values Q1, Q2 k , and the normalized correlation values H, H2 of the template sequence TS k , where k = 0,…,K-1;
[0081] SD03. Filter out audio clips The corresponding template sequence normalized correlation value H is greater than the threshold fragment, and a new fragment set is obtained Among them, 1≤c≤C, C is the number of fragments that meet the conditions; The correlation value with the total fragment SW * The correlation values of are integrated to calculate the integrated correlation value of each information bit. The integrated correlation value calculation formula is as follows:
[0082]
[0083] Among them, q i is the correlation value of the i-th information bit, i = 1, ..., L, N1 is the set of audio segments The number of segments with normalized correlation values greater than 0 in the synchronization sequence;
[0084] If q i >0, the embedded information bit is judged to be '1', otherwise, the embedded information bit is judged to be '-1' and the watermark extraction process ends.
[0085] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0086] The present invention proposes a real-time robust speech watermarking method based on logarithmic polar coordinates with adjustable logarithmic base frequency domain. To address the problems that existing speech cannot embed large-capacity watermarks in real time and the watermarks are imperceptible and have poor robustness after embedding, the present invention utilizes a coordinated real-time watermark embedding process and a watermark extraction process to complete watermark embedding in real-time audio clips at the frame level according to the characteristics of the amplitude on the logarithmic coordinates, and completes watermark extraction by fusing normalized frequency-related values, thereby improving the accuracy and imperceptibility of watermark extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 A schematic diagram showing the overall process of the real-time robust voice watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates proposed in Example 1 of the present invention;
[0088] Figure 2 A schematic diagram showing a real-time watermark embedding process of a real-time robust speech watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates proposed in Example 1 of the present invention;
[0089] Figure 3 A schematic diagram showing a watermark extraction process of a real-time robust speech watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates proposed in Example 1 of the present invention;
[0090] Figure 4 A schematic diagram showing a real-time watermark embedding framework of a real-time robust speech watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates proposed in Example 2 of the present invention;
[0091] Figure 5 A schematic diagram showing the arrangement of the watermark sequence WT generated by the base proposed in Example 2 of the present invention;
[0092] Figure 6 A schematic diagram showing a watermark extraction framework of a real-time robust speech watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates proposed in Example 3 of the present invention; DETAILED DESCRIPTION
[0093] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;
[0094] In order to better illustrate this embodiment, some parts of the drawings may be omitted, enlarged, or reduced, and do not represent the actual size;
[0095] It is understandable to those skilled in the art that descriptions of certain well-known contents may be omitted in the drawings.
[0096] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0097] The positional relationships described in the drawings are for illustrative purposes only and should not be construed as limiting this patent;
[0098] Example 1
[0099] like Figure 1 As shown, this embodiment proposes a real-time robust speech watermarking method based on logarithmic polar coordinates with adjustable logarithmic base in the frequency domain. The method includes a real-time robust speech watermark embedding process at the frame level with a length of 0.1 to 0.2 seconds and a watermark extraction process. The real-time watermark embedding process embeds multi-bit information into the host audio according to the amplitude characteristics of the logarithmic coordinates with adjustable logarithmic base in the real-time audio segment at the frame level; the watermark extraction process is based on the real-time watermark embedding process, and performs a watermark extraction operation on the host audio that has been embedded with the watermark and attacked by fusing normalized frequency-related values to extract the watermark information embedded in the host audio.
[0100] Specifically, see Figure 2 , the watermark embedding process at least includes:
[0101] S1. Determine the lengths of the pseudo-random modulation sequence and template sequence based on the number of bits in the watermark bit information sequence to be embedded. Combined with the key, randomly generate a bipolar pseudo-random modulation sequence and template sequence. Perform spread spectrum calculation on each bit of the watermark bit information sequence based on the bipolar pseudo-random modulation sequence, and combine with the template sequence to obtain the watermark sequence to be embedded.
[0102] S2. Determine the host audio output from the device in real time, perform a pre-emphasis operation on the very short frame-level host voice audio to obtain the determined host audio, and then expand the determined host audio. In actual implementation, the real-time output host audio is generally very short, making watermark embedding more difficult. The expansion operation facilitates embedding the watermark in very short host audio.
[0103] S3. According to the extended host audio frame length, the logarithmic base value of the logarithmic base adjustable frequency domain logarithmic coordinate transform is adaptively determined between 1 and 2 to ensure real-time watermark embedding;
[0104] S4. Perform a one-dimensional discrete Fourier transform on the expanded host audio segment, move the transformed DC component to the center of the Fourier amplitude spectrum, map the coordinates of points within the embedding region in the right half of the Fourier coefficient amplitude spectrum to log-polar coordinates, and embed the watermark sequence based on the amplitude characteristics. In the prior art, the logarithmic base of the frequency domain logarithmic coordinate transform is generally fixed to 2. However, in this method, the logarithmic base is adaptively selected between 1 and 2, such as 1.3, based on the length of a very short audio embedding frame, such as 0.1 to 0.2 seconds.
[0105] S5. Spread each bit of the watermark bit information sequence. Calculate the Fourier amplitude spectrum averages of the regions corresponding to the sequence obtained after the spread, at the positions corresponding to the embedded +1, -1, and the sequence obtained after the spread. Based on these three amplitude averages, embed the watermark bit information according to the situation.
[0106] S6. Based on the central symmetry of the Fourier amplitude spectrum, the symmetrical coefficients of the right half are copied to the left half, and then an inverse Fourier transform is performed. The watermark embedding process is completed, and the entire audio is de-emphasized to obtain the watermarked audio. The de-emphasis operation is the opposite of the pre-emphasis operation process.
[0107] See also Figure 3 , the watermark extraction process at least includes:
[0108] SA determines the audio to be tested, performs pre-emphasis operation on the entire audio segment to be tested, and intercepts the pre-emphasized audio to obtain a collection of audio segments;
[0109] SB. Perform a one-dimensional discrete Fourier transform on the pre-emphasized audio and the set of further truncated audio segments, moving the transformed DC component to the center of the Fourier amplitude spectrum. Using the center of the amplitude spectrum as the origin of the rectangular coordinate system, extract the watermark from the right half of the Fourier coefficient amplitude spectrum. Map the rectangular coordinates of the amplitude coefficients within the extraction range to log-polar coordinates. Sum the amplitude coefficients with the same log-polar coordinates after mapping, and use this sum as an element of the Fourier amplitude coefficient sequence.
[0110] SC. Based on the phase correlation principle, the original template sequence and the Fourier amplitude coefficient sequence are quickly matched and calculated to preliminarily determine the synchronization position of the embedded watermark. The amplitude matrix of the synchronized amplitude coefficient sequence is intercepted from the center position, and the final synchronization position is further determined using the neighborhood search method.
[0111] SD. uses a pseudo-random modulation sequence to despread the amplitude matrix and integrate the sub-segment correlation values to extract the watermark information.
[0112] Example 2
[0113] Based on Example 1, Figure 4 In this embodiment, the real-time watermark embedding process includes:
[0114] S101. Set the length of the bipolar pseudo-random modulation sequence and the template sequence. In this embodiment, according to the number of bits of the multi-bit meaningful information to be embedded, 128, the lengths of the pseudo-random modulation sequence and the template sequence are determined to be 2 and 128, respectively, and are denoted as L ps and L TS ; Let the key be key, use the key to generate a length of Lps =2 bipolar pseudo-random modulation sequence ps and length L TS =128 template sequence TS:
[0115] ps={ps i ; 1≤i≤L ps ,ps i ∈{-1,1}}
[0116] TS={TS j ; 1≤j≤L TS ,TS j ∈{-1,1}}
[0117] Among them, ps i TS j They represent the i-th element of the bipolar pseudo-random modulation sequence and the j-th element of the template sequence respectively;
[0118] S102. Assume length L ms =128 watermark bit information sequence is ms:
[0119] ms={ms i ; 1≤i≤L ms ,ms i ∈{-1,1}}
[0120] Among them, ms i Represents the i-th element of the watermark bit information sequence;
[0121] Use bipolar pseudo-random modulation sequence ps to modulate each bit ms i Perform spread spectrum modulation: If ms i =1, then it is spread spectrum modulated into a ps in-phase sequence, that is, W i =+1×ps; if ms i =-1, then it is spread spectrum modulated into the inverse sequence of ps, that is, W i = -1×ps, and finally get the meaningful watermark information array W = {W i ; 1≤i≤L ms ,W i ∈{-ps,ps}};
[0122] S103. The obtained meaningful watermark information array W = {W i ; 1≤i≤L ms ,W i ∈{-ps, ps}} and the template sequence TS are arranged in order to form a sequence of length M=L ms ×L ps +L TS The sequence WT={WT i; 1≤i≤M}, as the watermark sequence to be embedded, where the element WT in WT i It is composed of "1" and "-1", and the template sequence TS is stored in the last L of the watermark sequence WT in order. TS Locations, see Figure 5 .
[0123] S201. Determine the host audio output from the device in real time for 128ms. Suppose the host audio voice signal is S = {S t ; 1≤t≤L}, where S t represents the t-th sample point, L represents the length of the signal S, and a pre-emphasis operation is performed on it. The expression of the pre-emphasis operation is:
[0124] in, is the pre-emphasized speech signal, α=0.97;
[0125] S202. In this embodiment, the host audio has a speech duration of 128 milliseconds. The host audio is very short and has a high requirement for the real-time performance of watermark embedding. The speech signal after pre-emphasis is also 128 milliseconds long. Before watermark embedding, the 128 milliseconds of the pre-emphasized speech signal are Expand to 256 millisecond segments: Initialize an all-zero matrix with a length of 256 milliseconds, store the 128 millisecond audio segment in the last 128 milliseconds of the all-zero matrix, and use the expanded 256 millisecond matrix as the segment to be embedded to obtain matrix J.
[0126] S3. Adaptively determine the logarithmic base value of the logarithmic base adjustable frequency domain logarithmic coordinate transformation between 1 and 2 according to the frame length of the expanded host audio. In this embodiment, the logarithmic base value is 1.55.
[0127] S4. Perform a one-dimensional discrete Fourier transform of length d on the 256-millisecond matrix J. Move the transformed DC component to the center of the Fourier magnitude spectrum. Use the right half of the Fourier magnitude spectrum as the embedding region. Map the coordinates of the points within the embedding region to log-polar coordinates. The embedding region is located at the normalized frequency value f of the Fourier coefficient magnitude spectrum. n nearby;
[0128] The process of mapping the coordinates of the points in the embedded region to logarithmic polar coordinates is as follows: the rectangular coordinates r of the Fourier coefficients of the embedded region are transformed into discrete logarithmic polar coordinates lρ. The transformation formula is:
[0129]
[0130]
[0131] Where a is a constant greater than 1 but close to 1 In this embodiment, a=1.55, f n ×d is the origin of the logarithmic coordinate, offset is an offset constant that ensures the discrete logarithmic coordinate is not less than zero, and the floor() function represents the floor function to balance the robustness and invisibility of the watermark.
[0132] S501. For each information bit ms i , spread spectrum modulation is the in-phase sequence or reverse sequence W of ps i , before watermark embedding, calculate W i The average amplitude value amp in the corresponding rectangular coordinate area avh , and then calculate W i The average amplitude of the Fourier coefficients at the positions of +1 and -1 embedded in the corresponding rectangular coordinate area are recorded as and
[0133] S502. When When W i The embedding formula used in the corresponding rectangular coordinate area is as follows:
[0134]
[0135] when When W i The embedding formula used in the corresponding rectangular coordinate area is as follows:
[0136]
[0137] Among them, amp0 k is the Fourier coefficient amplitude of the original audio, ampw k is the amplitude of the audio Fourier coefficient after embedding the watermark, β=0.00001, δ is the watermark embedding strength, w k Indicates that the watermark bit to be embedded when the rectangular coordinate k is mapped to the logarithmic coordinate is "1" or "-1".
[0138] S6. Based on the central symmetry of the Fourier amplitude spectrum, the symmetrical coefficients of the right half are copied to the left half, and then an inverse Fourier transform is performed. The watermark embedding process is completed, and the entire audio is de-emphasized to obtain the watermarked audio.
[0139] Example 3
[0140] Based on Example 2, see Figure 6 In this embodiment, the watermark extraction process includes:
[0141] SA01. Assume that the input watermarked speech signal is SW:
[0142] SW={SW t ; 1≤t≤L′}
[0143] Among them, SW t represents the t-th sample point, L′ represents the length of the signal SW; t Perform pre-emphasis operation to obtain the pre-emphasized watermarked speech signal SW * .
[0144] SA02. To SW * The audio segment is intercepted with a step size of 128 milliseconds and a window length of 256 milliseconds until the sliding window reaches the end and the audio segment is recorded as k represents the kth segment during sliding capture.
[0145] SB01. To SW * and Perform one-dimensional discrete Fourier transform respectively, and move the DC component to the center of the Fourier amplitude spectrum; take the center of the amplitude spectrum as the origin of the rectangular coordinate system, and normalize the right half plane of the Fourier coefficient amplitude spectrum to the frequency f n The rectangular coordinates r′ of the Fourier coefficients near are transformed into discrete logarithmic coordinates lρ′. The transformation formula is:
[0146]
[0147]
[0148] Wherein, M′=λM, λ≥1, offset′ is an offset constant that ensures that the discrete logarithmic coordinate is not less than zero, and the floor() function represents the floor rounding function;
[0149] SB02. Initialize an all-zero matrix amp of length M″=λ×μ×M, where μ is a positive integer not less than 1; map the rectangular coordinates to the logarithmic polar coordinate system, and sum the Fourier coefficient amplitudes with the same discrete logarithmic coordinate lρ′ as the Fourier coefficient amplitude sequence amp lρ′ An element of , thus obtaining a Fourier coefficient amplitude sequence amp.
[0150] SC01. Use the same key as the embedding algorithm to generate a key with a length of L TS The template sequence TS={TS i ; 1≤i≤L TS ,TS i ∈{-1,1}}, each TS i Expanded into μ TS i , get TS1; initialize the full zero matrix TS with length M″ m, store TS1 in TS m The second half of the synchronization template TS is obtained m ;
[0151] SC02. Obtain the translation phase correlation value through phase correlation fast matching calculation. The formula for calculating the translation phase correlation value is:
[0152]
[0153] Among them, X(k) is the correlation value sequence, φ amp (u) is the phase angle of amp(u), G * (u)=(DFT(TS m (i))) * Is the synchronization template TS m The complex conjugate of the one-dimensional Fourier transform coefficient, "*" indicates the complex conjugate;
[0154] According to the maximum value of the correlation value sequence X(k), the position of the embedded watermark WT in the amplitude sequence amp is preliminarily determined, which is recorded as k max ;
[0155] SC03. max The positions of the left and right neighbors, and k = 1 as an alternative sequence of synchronization positions col = {1, k max -1,k max ,k max +1}, take the synchronization position col i , the first col in the amplitude sequence amp i -1 position is moved to the end of the amplitude sequence amp to obtain the synchronized amplitude sequence, and the synchronized amplitude sequence is cut out from the center position to obtain an amplitude matrix of length μ×M, which is recorded as
[0156] SC04. Calculation The mean of the non-zero amplitudes at each μ position in is stored in the amplitude matrix amp2 of length M in sequence;
[0157] SC05. Calculate each synchronization position col in col i The correlation value Y(i) between the corresponding amplitude matrix amp2 and the template sequence TS is calculated as follows:
[0158]
[0159] The final synchronization position is determined according to the maximum value of the correlation value sequence Y(i), which is denoted as col f .
[0160] SD01. Despread spectrum modulate the amplitude matrix amp2 using the original pseudo-random modulation sequence ps and calculate the normalized correlation value corresponding to each information bit. The process is as follows:
[0161] Take L from amp2 in order ms Segments do not overlap and are of length L ps The sequence W′ i :
[0162]
[0163] Calculate the normalized correlation value Q between the amplitude corresponding to the meaningful information sequence and the original pseudo-random modulation sequence ps. The calculation formula is:
[0164]
[0165] Q={Q i ; 1≤i≤L ms}
[0166] Calculate the normalized correlation value H between the amplitude corresponding to the synchronization template sequence and the template sequence TS. The calculation formula is:
[0167]
[0168] SD02. According to step SD01, calculate SW * and Normalized correlation values Q1, Q2 k , and the normalized correlation values H, H2 of the template sequence TS k , where k = 0,…,K-1;
[0169] SD03. Filter out audio clips The corresponding template sequence normalized correlation value H is greater than the threshold fragment, and a new fragment set is obtained Among them, 1≤c≤C, C is the number of fragments that meet the conditions; The correlation value with the total fragment SW * The correlation values of are integrated to calculate the integrated correlation value of each information bit. The integrated correlation value calculation formula is as follows:
[0170]
[0171] Among them, q i is the correlation value of the i-th information bit, i = 1, ..., L, N1 is the set of audio segments The number of segments with normalized correlation values greater than 0 in the synchronization sequence;
[0172] If q i>0, the embedded information bit is judged to be '1', otherwise, the embedded information bit is judged to be '-1' and the watermark extraction process ends.
[0173] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.
Claims
1. A real-time robust speech watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates, characterized in that: The method includes a real-time robust speech watermark embedding process and a watermark extraction process at a frame level of 0.1 to 0.2 seconds in length, wherein the real-time watermark embedding process at least includes: S1. Determine the lengths of the pseudo-random modulation sequence and template sequence based on the number of bits in the watermark bit information sequence to be embedded. Combined with the key, randomly generate a bipolar pseudo-random modulation sequence and template sequence. Perform spread spectrum calculation on each bit of the watermark bit information sequence based on the bipolar pseudo-random modulation sequence, and combine with the template sequence to obtain the watermark sequence to be embedded. S2. Determine the host audio output from the device in real time, perform pre-emphasis on the host voice audio at a very short frame level, obtain the determined host audio, and expand the determined host audio; S3. According to the extended host audio frame length, the logarithmic base value of the logarithmic base adjustable frequency domain logarithmic coordinate transform is adaptively determined between 1 and 2 to ensure real-time watermark embedding; S4. Perform a one-dimensional discrete Fourier transform on the expanded host audio, move the transformed DC component to the center of the Fourier magnitude spectrum, use the right half of the Fourier magnitude spectrum as the embedding region, map the coordinates of points within the embedding region to log-polar coordinates, and embed the watermark sequence based on the amplitude characteristics. S5. Spread each bit of the watermark bit information sequence. Calculate the Fourier amplitude spectrum averages of the regions corresponding to the sequence obtained after the spread, at the positions corresponding to the embedded +1, -1, and the sequence obtained after the spread. Based on these three amplitude averages, embed the watermark bit information according to the situation. S6. Based on the central symmetry of the Fourier amplitude spectrum, the symmetrical coefficients of the right half are copied to the left half, and then an inverse Fourier transform is performed. The watermark embedding process is completed, and the entire audio segment is de-emphasized to obtain the watermarked audio; The real-time watermark embedding process embeds multi-bit information into the host audio according to the characteristics of the amplitude on the logarithmic coordinate with an adjustable logarithmic base in the real-time audio clip at the frame level; the watermark extraction process is based on the real-time watermark embedding process, and performs a watermark extraction operation on the host audio that has been embedded with the watermark and attacked by fusing normalized frequency-related values to extract the watermark information embedded in the host audio.
2. The real-time robust voice watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates according to claim 1 is characterized in that: The process of step S1 includes: S101. Set the length of the bipolar pseudo-random modulation sequence and the template sequence, denoted as L ps and L TS ; Let the key be key, and use the key to generate the bipolar pseudo-random modulation sequence ps and template sequence TS: ps={ps i ;1≤i≤L ps ,ps i ∈{-1,1}} TS={TS j ;1≤j≤L TS ,TS j ∈{-1,1}} Among them, ps i TS j They represent the i-th element of the bipolar pseudo-random modulation sequence and the j-th element of the template sequence respectively; S102. Let the length be L ms The watermark bit information sequence is ms: ms={ms i ;1≤i≤L ms ,ms i ∈{-1,1}} Among them, ms i Represents the i-th element of the watermark bit information sequence; Use bipolar pseudo-random modulation sequence ps to modulate each bit ms i Perform spread spectrum modulation: If ms i =1, then it is spread spectrum modulated into a ps in-phase sequence, that is, W i =+1×ps; if ms i =-1, then it is spread spectrum modulated into the inverse sequence of ps, that is, W i = -1×ps, and finally get the meaningful watermark information array W = {W i ; 1≤i≤L ms ,W i ∈{-ps,ps}}; S103. The obtained meaningful watermark information array W = {W i ; 1≤i≤L ms ,W i ∈{-ps,ps}} and the template sequence TS are arranged in order to form a sequence of length M=L ms ×L ps +L TS The sequence WT={W T i; 1≤i≤M}, as the watermark sequence to be embedded, where the element WT in WT i It is composed of "1" and "-1", and the template sequence TS is stored in the last L of the watermark sequence WT in order. TS locations.
3. The real-time robust voice watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates according to claim 2 is characterized in that: The process of step S2 includes: S201. Assume that the host audio voice signal is S={S t ; 1≤t≤L}, where S t represents the t-th sample point, L represents the length of the signal S, and a pre-emphasis operation is performed on it. The expression of the pre-emphasis operation is: in, is the pre-emphasized speech signal, α=0.97; S202. If the host audio voice duration is x milliseconds, the pre-emphasized voice signal duration is also x milliseconds; before watermark embedding, the pre-emphasized voice signal x milliseconds Expand to y millisecond segments: Initialize an all-zero matrix of length y milliseconds, store the x millisecond audio segment in the last x millisecond matrix area of the all-zero matrix, and use the expanded y millisecond matrix as the segment to be embedded to obtain matrix J.
4. The real-time robust voice watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates according to claim 3 is characterized in that: In step S4, a one-dimensional discrete Fourier transform of length d is performed on the y millisecond matrix J, and the transformed DC component is moved to the center of the Fourier amplitude spectrum. The center of the Fourier amplitude spectrum is used as the origin of the rectangular coordinate system, and a watermark is embedded in the right half of the Fourier amplitude spectrum. The embedding area is located at the normalized frequency value f of the Fourier coefficient amplitude spectrum. n nearby; The process of mapping the coordinates of the points in the embedded region to logarithmic polar coordinates is as follows: the rectangular coordinates r of the Fourier coefficients of the embedded region are transformed into discrete logarithmic polar coordinates lρ. The transformation formula is: Among them, a is a constant greater than 1 but close to 1, f n ×d is the origin of the logarithmic coordinates, offset is an offset constant that ensures that the discrete logarithmic coordinates are not less than zero, and the floor() function represents the floor function.
5. The real-time robust voice watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates according to claim 4 is characterized in that: The process of step S5 includes: S501. For each information bit ms i , spread spectrum modulation is the in-phase sequence or reverse sequence W of ps i , before watermark embedding, calculate W i The average amplitude value amp in the corresponding rectangular coordinate area avg , and then calculate W i The average amplitude of the Fourier coefficients at the positions of +1 and -1 embedded in the corresponding rectangular coordinate area are recorded as and S502. When When W i The embedding formula used in the corresponding rectangular coordinate area is as follows: when When W i The embedding formula used in the corresponding rectangular coordinate area is as follows: Among them, amp0 k is the Fourier coefficient amplitude of the original audio, ampw k is the amplitude of the audio Fourier coefficient after embedding the watermark, β=0.00001, δ is the watermark embedding strength, w k Indicates that the watermark bit to be embedded when the rectangular coordinate k is mapped to the logarithmic coordinate is "1" or "-1".
6. The real-time robust voice watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates according to claim 1 is characterized in that: The watermark extraction process at least includes: SA determines the audio to be tested, performs pre-emphasis operation on the entire audio segment to be tested, and intercepts the pre-emphasized audio to obtain a collection of audio segments; SB. Perform a one-dimensional discrete Fourier transform on the pre-emphasized audio and the set of further truncated audio segments, moving the transformed DC component to the center of the Fourier amplitude spectrum. Using the center of the amplitude spectrum as the origin of the rectangular coordinate system, extract the watermark from the right half of the Fourier coefficient amplitude spectrum. Map the rectangular coordinates of the amplitude coefficients within the extraction range to log-polar coordinates. Sum the amplitude coefficients with the same log-polar coordinates after mapping, and use this sum as an element of the Fourier amplitude coefficient sequence. SC. Based on the phase correlation principle, the original template sequence and the Fourier amplitude coefficient sequence are quickly matched and calculated to preliminarily determine the synchronization position of the embedded watermark. The amplitude matrix of the synchronized amplitude coefficient sequence is intercepted from the center position, and the final synchronization position is further determined using the neighborhood search method. SD. uses a pseudo-random modulation sequence to despread the amplitude matrix and integrate the sub-segment correlation values to extract the watermark information.
7. The real-time robust voice watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates according to claim 6 is characterized in that: Step SA includes the following steps: SA01. Assume that the input watermarked speech signal is SW: SW={SW t ;1≤t≤L'} Among them, SW t represents the t-th sample point, L' represents the length of the signal SW; t Perform pre-emphasis operation to obtain the pre-emphasized watermarked speech signal SW * ; SA02. To SW * Cut with a step size of z2 milliseconds and a window length of y milliseconds until the sliding window reaches the end and record the audio segment as k represents the kth segment during sliding capture.
8. The real-time robust voice watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates according to claim 7 is characterized in that: Step SB includes the following steps: SB01. To SW * and Perform one-dimensional discrete Fourier transform respectively, and move the DC component to the center of the Fourier amplitude spectrum; take the center of the amplitude spectrum as the origin of the rectangular coordinate system, and normalize the right half plane of the Fourier coefficient amplitude spectrum to the frequency f n The rectangular coordinates r' of the nearby Fourier coefficients are transformed into discrete logarithmic coordinates lρ'. The transformation formula is: Where, M'=λM, λ≥1, offset' is an offset constant that ensures that the discrete logarithmic coordinate is not less than zero, and the floor() function represents the floor function; SB02. Initialize an all-zero matrix amp of length M″ = λ×μ×M, where μ is a positive integer not less than 1; map the rectangular coordinates to the logarithmic polar coordinate system, and sum the Fourier coefficient amplitudes with the same discrete logarithmic coordinate lρ' as the Fourier coefficient amplitude sequence amp lρ' An element of , thus obtaining a Fourier coefficient amplitude sequence amp; Step SC comprises the following steps: SC01. Use the same key as the embedding algorithm to generate a key with a length of L TS The template sequence TS={TS i ; 1≤i≤L TS ,TS i ∈{-1,1}}, each TS i Expanded into μ TS i , get TS1; initialize the full zero matrix TS with length M" m , store TS1 in TS m The second half of the synchronization template TS is obtained m ; SC02. Obtain the translation phase correlation value through phase correlation fast matching calculation. The formula for calculating the translation phase correlation value is: Among them, X(k) is the correlation value sequence, φ amp (u) is the phase angle of amp(u), G * (u)=(DFT(TS m (i))) * Is the synchronization template TS m The complex conjugate of the one-dimensional Fourier transform coefficient, "*" represents the complex conjugate; According to the maximum value of the correlation value sequence X(k), the position of the embedded watermark WT in the amplitude sequence amp is preliminarily determined, which is recorded as k max ; SC03. max The positions of the left and right neighbors, and k = 1 as an alternative sequence of synchronization positions col = {1, k max -1,k max ,k max +1}, take the synchronization position col i , the first col in the amplitude sequence amp i -1 position is moved to the end of the amplitude sequence amp to obtain the synchronized amplitude sequence, and the synchronized amplitude sequence is cut out from the center position to obtain an amplitude matrix of length μ×M, which is recorded as SC04. Calculation The mean of the non-zero amplitudes at each μ position in is stored in the amplitude matrix amp2 of length M in sequence; SC05. Calculate each synchronization position col in col i The correlation value Y(i) between the corresponding amplitude matrix amp2 and the template sequence TS is calculated as follows: The final synchronization position is determined according to the maximum value of the correlation value sequence Y(i), which is denoted as col f .
9. The real-time robust voice watermarking method based on logarithmic base adjustable frequency domain logarithmic polar coordinates according to claim 8, characterized in that: Step SD includes the following steps: SD01. Despread spectrum modulate the amplitude matrix amp2 using the original pseudo-random modulation sequence ps and calculate the normalized correlation value corresponding to each information bit. The process is as follows: Take L from amp2 in order ms Segments do not overlap and are of length L ps The sequence W' i : Calculate the normalized correlation value Q between the amplitude corresponding to the meaningful information sequence and the original pseudo-random modulation sequence ps. The calculation formula is: Q={Q i ;1≤i≤L ms } Calculate the normalized correlation value H between the amplitude corresponding to the synchronization template sequence and the template sequence TS. The calculation formula is: SD02. According to step SD01, calculate SW * and Normalized correlation values Q1, Q2 k , and the normalized correlation values H1 and H2 of the template sequence TS k , where k = 0,…,K-1; SD03. Filter out audio clips The corresponding template sequence normalized correlation value H is greater than the threshold fragment, and a new fragment set is obtained Among them, 1≤c≤C, C is the number of fragments that meet the conditions; The correlation value with the total fragment SW * The correlation values of are integrated to calculate the integrated correlation value of each information bit. The integrated correlation value calculation formula is as follows: Among them, q i is the correlation value of the i-th information bit, i = 1, ..., L, N1 is the set of audio segments The number of segments with normalized correlation values greater than 0 in the synchronization sequence; If q i >0, the embedded information bit is judged to be "1", otherwise, the embedded information bit is judged to be "-1", and the watermark extraction process ends.
Citation Information
Patent Citations
Method for compressing watermark using voice based on bone conduction
CN101350198A
Voice and video watermark for exfiltration prevention
US20150373032A1