A digital audio watermarking method, device, electronic device and storage medium
The digital audio watermarking method addresses synchronization attack vulnerabilities by segmenting and processing audio signals to determine correct watermark positions and lengths, ensuring efficient and accurate extraction without synchronization codes.
Patent Information
- Application Number
- CN202510518016.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing anti-synchronous audio watermarking schemes usually require synchronization codes, resulting in reduced payload and complex extraction process, making it difficult to quickly and accurately determine the correct position and length of audio information embedded in the watermark.
By performing segmentation processing and transforming domain processing on the audio information, combining digital audio watermark embedding algorithm and feature value statistics, the method with its own verification function eliminates synchronization code, and quickly and accurately determines the candidate position and length to extract watermark information.
It realizes the rapid and accurate determination of the correct position and length of the audio information embedded in the watermark, simplifies the extraction process, and improves the extraction efficiency and accuracy of the watermark information.
Smart Images

Figure CN120048271B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of digital audio watermark processing, and in particular, to a digital audio watermark method, device, electronic device, and storage medium. Background Technique
[0002] Digital audio watermark is a type of digital watermark, mainly used to protect the copyright of audio works. Its principle is to utilize the redundancy of audio signals and the masking effect of the human ear to embed watermark information into the host audio without affecting the perceptual quality of the audio signal, so as to achieve purposes such as copyright protection, reliability, and integrity identification of the audio. Like other digital watermark technologies, one problem that must be considered in the practical application of audio watermark technology is various attacks that the watermark may be subjected to. For digital audio watermarks, synchronization attack is a powerful attack method. This attack disrupts the normal timing relationship of the audio, resulting in the inability to locate and retrieve the corresponding embedding interval when the watermark is embedded, thus causing the failure of watermark extraction. Current anti-synchronization audio watermark schemes often rely on synchronization codes to achieve. One is that the effective payload is reduced because a part of the embedding capacity needs to be set aside for the synchronization code. The other is that the extraction process is too complex, and the sliding window matching of the synchronization code needs to be considered outside the double loop. Therefore, how to perform digital audio watermarking has become a technical problem that cannot be underestimated. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a digital audio watermark method, device, electronic device, and storage medium. The digital audio watermark method provided by this application has a built-in verification function, can eliminate the synchronization code in the anti-synchronization algorithm, and can quickly and accurately determine the correct candidate positions and candidate lengths of the audio information after embedding the watermark, so as to accurately extract the watermark information.
[0004] An embodiment of this application provides a digital audio watermark method, and the digital audio watermark method includes:
[0005] Perform segmented processing on the first audio information to generate multiple audio segments, perform transform domain processing on each audio segment, and determine the watermark data embedding domain in each audio segment;
[0006] Based on the digital audio watermark embedding algorithm, perform digital watermark information embedding processing on the audio data in each watermark data embedding domain, determine multiple audio segments after embedding the watermark, and combine the multiple audio segments after embedding the watermark to determine the second audio information;
[0007] Based on a preset plurality of candidate positions and candidate segment lengths, perform segmented processing and eigenvalue statistical processing on the second audio information, and determine the statistical eigenvalue of each candidate segment;
[0008] Based on the statistical feature values of each candidate segment, the first preset threshold, and the second preset threshold, it is determined whether the current candidate position and candidate length are correct. If so, based on the correct candidate position and correct candidate length, the watermark bit values of each candidate segment are decoded in sequence to determine the watermark information of the second audio information.
[0009] In a possible implementation manner, the digital watermark information embedding process is performed on the audio data in each watermark data embedding domain based on the digital audio watermark embedding algorithm to determine a plurality of audio segments embedded with watermarks, including:
[0010] If the watermark bit value is 1, a constant value is added to the audio data of the first sub-audio segment among two adjacent sub-audio segments corresponding to the watermark data embedding domain, and a constant value is subtracted from the audio data in the second sub-audio segment to complete the digital watermark information embedding process;
[0011] If the watermark bit value is 0, a constant value is added to the audio data of the first sub-audio segment among two adjacent sub-audio segments corresponding to the watermark data embedding domain, and a constant value is subtracted from the audio data in the second sub-audio segment to complete the digital watermark information embedding process.
[0012] In a possible implementation manner, for each candidate position and candidate segment length, the second audio information is segmented and the statistical feature values are calculated based on a preset plurality of candidate positions and candidate segment lengths to determine the statistical feature values of each candidate segment, including:
[0013] The second audio information is segmented based on the candidate position and candidate segment length to determine a plurality of candidate segments of the second audio information;
[0014] For each candidate segment, an equal-share first audio data set and a second audio data set are divided based on the audio data in the candidate segment, and the average values are calculated for the first audio data set and the second audio data set respectively to determine the first statistical feature value of the first audio data set and the second statistical feature value of the second audio data set.
[0015] In a possible implementation manner, determining whether the current candidate position and candidate length are correct based on the statistical feature values of each candidate segment, the first preset threshold, and the second preset threshold includes:
[0016] Based on the statistical feature value of the candidate segment and the first preset threshold, the watermark state of the current candidate segment is determined; wherein, the watermark state includes an invalid watermark state and a valid watermark state, the watermark bit value of the valid watermark state is 0 or 1, and the watermark bit value of the invalid watermark state is 2;
[0017] Check whether the number of the invalid watermark states of each detected candidate segment exceeds a second preset threshold;
[0018] If so, the current candidate position and candidate length are incorrect; if not, the current candidate position and candidate length are correct.
[0019] In a possible implementation manner, determining the watermark state of the current candidate segment based on the statistical feature value of the candidate segment and a first preset threshold includes:
[0020] Check whether the absolute value of the difference between the first statistical feature value and the second statistical feature value is greater than the first preset threshold;
[0021] If so, the watermark state of the current candidate segment is a valid watermark state; wherein, if the first statistical feature value is greater than the second statistical feature value, the decoded watermark bit value is 0, and if the first statistical feature value is less than the second statistical feature value, the decoded watermark bit is 1;
[0022] If not, the watermark state of the current candidate segment is an invalid watermark state.
[0023] In a possible implementation manner, after segmenting the first audio information to generate a plurality of audio segments, performing transform domain processing on each of the audio segments, and determining a watermark data embedding domain in each of the audio segments, the digital audio watermark method further includes:
[0024] Perform a summation process on the audio data in the watermark data embedding domain to determine a quantized statistical feature value;
[0025] Determine a quantization step based on the statistical characteristics of the first audio information and an auditory threshold, and multiply the quantization step by a preset fraction to determine a target value;
[0026] If the watermark bit value is 0, perform a subtraction process on the quantized statistical feature value and the target value;
[0027] If the watermark bit value is 1, perform an addition process on the quantized statistical feature value and the target value, so as to complete the digital watermark information embedding process.
[0028] In a possible implementation manner, the digital audio watermark method further includes:
[0029] Segment the second audio information based on a plurality of preset candidate positions and candidate segment lengths to determine a plurality of candidate segments, and count the statistical feature value of each candidate segment;
[0030] Quantize the statistical feature values of each candidate segment based on the quantization step size to determine the quantization residual value of each candidate segment;
[0031] If the interval where the quantization residual value is located is the first preset interval or the second preset interval, the watermark state of the candidate segment is a valid watermark state;
[0032] If the interval where the quantization residual value is located is the third preset interval, the watermark state of the candidate segment is an invalid watermark state; wherein, the ranges of the first preset interval, the third preset interval, and the second preset interval increase in sequence;
[0033] Based on the number of invalid watermark states of each candidate segment counted and a preset second threshold, determine whether the current candidate position and candidate length are correct.
[0034] An embodiment of the present application further provides a digital audio watermarking device, and the digital audio watermarking device includes:
[0035] A first processing module, configured to segment the first audio information to generate a plurality of audio segments, perform transform domain processing on each of the audio segments to determine the watermark data embedding domain in each of the audio segments;
[0036] A first watermark information embedding module, configured to perform digital watermark information embedding processing on the audio data in each of the watermark data embedding domains based on a digital audio watermark embedding algorithm to determine a plurality of audio segments with embedded watermarks, and combine the plurality of audio segments with embedded watermarks to determine the second audio information;
[0037] A second processing module, configured to perform segmentation processing and eigenvalue statistical processing on the second audio information based on a plurality of preset candidate positions and candidate segment lengths to determine the statistical eigenvalue of each candidate segment;
[0038] A first watermark information extraction module, configured to determine whether the current candidate position and candidate length are correct based on the statistical eigenvalue of each candidate segment, a first preset threshold, and a second preset threshold. If so, decode the watermark bit values of each candidate segment in sequence based on the correct candidate position and correct candidate length to determine the watermark information of the second audio information.
[0039] An embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the digital audio watermarking method as described above are executed.
[0040] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the digital audio watermarking method as described above.
[0041] A digital audio watermarking method, device, electronic device and storage medium provided by an embodiment of the present application. The digital audio watermarking method includes: segmenting first audio information to generate a plurality of audio segments, performing transform domain processing on each of the audio segments to determine a watermark data embedding domain in each of the audio segments; performing digital watermark information embedding processing on the audio data in each of the watermark data embedding domains based on a digital audio watermark embedding algorithm to determine a plurality of audio segments after embedding watermarks, and combining the plurality of audio segments after embedding watermarks to determine second audio information; performing segmenting processing and eigenvalue statistical processing on the second audio information based on a plurality of preset candidate positions and candidate segment lengths to determine a statistical eigenvalue of each candidate segment; determining whether the current candidate position and candidate length are correct based on the statistical eigenvalue of each candidate segment, a first preset threshold, and a second preset threshold. If so, decoding the watermark bit values of each candidate segment in sequence based on the correct candidate position and correct candidate length to determine the watermark information of the second audio information. The digital audio watermarking method provided by the present application has a built-in verification function, can eliminate the synchronization code in the anti-synchronization algorithm, and can quickly and accurately determine the correct candidate position and candidate length of the audio information after embedding the watermark, so as to accurately extract the watermark information.
[0042] To make the above objects, features and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Description of the Drawings
[0043] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can also be obtained based on these drawings without creative efforts.
[0044] Figure 1 It is a flowchart of a digital audio watermarking method provided by an embodiment of the present application;
[0045] Figure 2 It is a schematic structural diagram one of a digital audio watermarking device provided by an embodiment of the present application;
[0046] Figure 3 It is a schematic structural diagram two of a digital audio watermarking device provided by an embodiment of the present application;
[0047] Figure 4 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Specific implementation manners
[0048] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. Components of the embodiments of the present application generally described and illustrated in the accompanying drawings herein may be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0049] First, an applicable application scenario of the present application will be introduced. The present application can be applied to the technical field of digital audio watermark processing.
[0050] After research, it is found that digital audio watermark is a kind of digital watermark, mainly used to protect the copyright of audio works. Its principle is to utilize the redundancy of audio signals and the masking effect of the human ear to embed watermark information into the host audio without affecting the perceptual quality of the audio signal, so as to achieve the purposes of copyright protection, reliability and integrity identification of the audio. Like other digital watermark technologies, one problem that must be considered in the actual application of audio watermark technology is various attacks that the watermark may be subjected to. For digital audio watermarks, synchronization attack is a powerful attack method. This attack destroys the normal timing relationship of the audio, resulting in the inability to locate and retrieve the corresponding embedding interval when the watermark is embedded, thus causing the failure of watermark extraction. Current anti-synchronization audio watermarking schemes often rely on synchronization codes to achieve. One is that the effective payload is reduced because a part of the embedding capacity needs to be set aside for the synchronization code. The other is that the extraction process is too complex, and the sliding window matching of the synchronization code needs to be considered outside the double loop. Therefore, how to perform digital audio watermarking has become a technical problem that cannot be underestimated.
[0051] Based on this, the embodiments of the present application provide a digital audio watermarking method. The digital audio watermarking method provided by the present application has a built-in verification function, can eliminate the synchronization code in the anti-synchronization algorithm, and can quickly and accurately determine the correct candidate position and candidate length of the audio information after the watermark is embedded, so as to accurately extract the watermark information.
[0052] Refer to Figure 1 , Figure 1The flowchart of a digital audio watermarking method provided by an embodiment of this application. As Figure 1 shown in the figure, the digital audio watermarking method provided by the embodiment of this application includes:
[0053] S101: Segment the first audio information to generate multiple audio segments, perform transform domain processing on each of the audio segments, and determine the watermark data embedding domain in each of the audio segments.
[0054] In this step, segment the first audio information to generate multiple audio segments, perform transform domain processing on each audio segment, and determine the watermark data embedding domain in each audio segment.
[0055] Among them, transform domain processing is to convert the audio signal from the time domain (i.e., the original sampling points on the time axis) to other mathematical domains (such as the frequency domain, wavelet domain, etc.) to more efficiently analyze or modify certain characteristics of the signal. Common transform domain processing methods include discrete Fourier transform, discrete cosine transform, discrete wavelet transform, and short-time Fourier transform, etc.
[0056] Here, the first audio information is the original audio information.
[0057] Among them, the purpose of segmenting the first audio information is to embed a watermark bit payload in each audio segment.
[0058] S102: Perform digital watermark information embedding processing on the audio data in each of the watermark data embedding domains based on the digital audio watermark embedding algorithm, determine multiple audio segments with embedded watermarks, and combine the multiple audio segments with embedded watermarks to determine the second audio information.
[0059] In this step, perform digital watermark information embedding processing on the audio data in each watermark data embedding domain according to the digital audio watermark embedding algorithm, determine multiple audio segments with embedded watermarks, and combine the multiple audio segments with embedded watermarks to determine the second audio information.
[0060] Here, modify the audio data in each watermark data embedding domain according to the digital audio watermark embedding algorithm, and make different modifications to the audio data according to whether the watermark load bit is 0 or 1, so that the data carries watermark information.
[0061] Among them, the digital audio watermark embedding algorithm can be a patchwork - type algorithm.
[0062] In a possible implementation manner, the performing digital watermark information embedding processing on the audio data in each of the watermark data embedding domains based on the digital audio watermark embedding algorithm to determine multiple audio segments with embedded watermarks includes:
[0063] If the watermark bit value is 1, a constant value is added to the audio data of the first sub-audio segment among two adjacent sub-audio segments corresponding to the watermark data embedding domain, and the constant value is subtracted from the audio data in the second sub-audio segment, so as to complete the digital watermark information embedding process; if the watermark bit value is 0, a constant value is added to the audio data of the first sub-audio segment among two adjacent sub-audio segments corresponding to the watermark data embedding domain, and the constant value is subtracted from the audio data in the second sub-audio segment, so as to complete the digital watermark information embedding process.
[0064] Among them, the constant value added and the constant value subtracted should be the same constant.
[0065] Here, the two adjacent sub-audio segments corresponding to the watermark data embedding domain are obtained by dividing the audio segmentation into two equal parts.
[0066] S103: Based on a plurality of preset candidate positions and candidate segment lengths, perform segmentation processing and eigenvalue statistical processing on the second audio information to determine the statistical eigenvalue of each candidate segment.
[0067] In this step, according to a plurality of preset candidate positions and candidate segment lengths, perform segmentation processing and eigenvalue statistical processing on the second audio information to determine the statistical eigenvalue of each candidate segment.
[0068] In a possible implementation manner, for each candidate position and candidate segment length, the performing segmentation processing and eigenvalue statistical processing on the second audio information based on a plurality of preset candidate positions and candidate segment lengths to determine the statistical eigenvalue of each candidate segment includes:
[0069] A: Perform segmentation processing on the second audio information based on the candidate position and the candidate segment length to determine a plurality of candidate segments of the second audio information.
[0070] Here, perform segmentation processing on the second audio information according to the candidate position and the candidate segment length to determine a plurality of candidate segments of the second audio information.
[0071] Among them, the role of the candidate position is to use multiple positions as the starting points of the candidate segments.
[0072] B: For each candidate segment, divide the audio data in the candidate segment into equal first audio data set and second audio data set, calculate the mean values of the first audio data set and the second audio data set respectively, and determine the first statistical eigenvalue of the first audio data set and the second statistical eigenvalue of the second audio data set.
[0073] Here, for each candidate segment, an equal - sized first audio data set and second audio data set are divided according to the audio data in the candidate segment. The mean values of the first audio data set and the second audio data set are calculated respectively, and a first statistical feature value of the first audio data set and a second statistical feature value of the second audio data set are determined.
[0074] Here, the statistical feature value can also be determined according to energy, variance, singular value, etc., and this part is not specifically limited.
[0075] S104: Based on the statistical feature value of each candidate segment, a first preset threshold, and a second preset threshold, determine whether the current candidate position and candidate length are correct. If so, decode the watermark bit values of each candidate segment in sequence based on the correct candidate position and correct candidate length, and determine the watermark information of the second audio information.
[0076] In this step, according to the statistical feature value of each candidate segment and the first preset threshold, determine whether the current candidate position and candidate length are correct. If so, decode the watermark bit values of each candidate segment in sequence based on the correct candidate position and correct candidate length, and determine the watermark information of the second audio information.
[0077] Here, even if the audio segment has been embedded with a watermark, if the segment position or segment length used during decoding is incorrect, the decoding process will still regard the signal as random data. The reasons are as follows: If segmentation starts from the wrong position, the rules for embedding the watermark (such as the mean difference of the Patchwork - like algorithm or the quantization value of the QIM - like algorithm) will be violated, resulting in the decoding result being unable to match the valid watermark state (0 or 1). If the wrong segment length is used, the data with the embedded watermark will be divided into incorrect small segments, thus violating the embedding rules. Therefore, it is necessary to determine the correct candidate position and candidate length.
[0078] In a possible implementation manner, the determining whether the current candidate position and candidate length are correct based on the statistical feature value of each candidate segment, a first preset threshold, and a second preset threshold includes:
[0079] a: Based on the statistical feature value of the candidate segment and the first preset threshold, determine the watermark state of the current candidate segment; wherein, the watermark state includes an invalid watermark state and a valid watermark state, the watermark bit value of the valid watermark state is 0 or 1, and the watermark bit value of the invalid watermark state is 2.
[0080] Here, according to the statistical feature value of the candidate segment and the first preset threshold, determine the watermark state of the current candidate segment.
[0081] Among them, the innovative watermark states in this application include the invalid watermark state and the valid watermark state. The watermark bit values in all valid watermark states are 0 or 1, and the watermark bit value in the invalid watermark state is 2.
[0082] In a possible implementation manner, determining the watermark state of the current candidate segment based on the statistical feature value of the candidate segment and the first preset threshold includes:
[0083] (1): Detect whether the absolute value of the difference between the first statistical feature value and the second statistical feature value is greater than the first preset threshold.
[0084] Here, detect whether the absolute value of the difference between the first statistical feature value and the second statistical feature value is greater than the first preset threshold.
[0085] Among them, the first preset threshold is determined according to expert experience.
[0086] (2): If so, the watermark state of the current candidate segment is the valid watermark state; among them, if the first statistical feature value is greater than the second statistical feature value, the decoded watermark bit value is 0, and if the first statistical feature value is less than the second statistical feature value, the decoded watermark bit is 1; if not, the watermark state of the current candidate segment is the invalid watermark state.
[0087] Here, if so, the watermark state of the current candidate segment is the valid watermark state; if not, the watermark state of the current candidate segment is the invalid watermark state.
[0088] Among them, the three-state watermark algorithm proposed in this application transforms the patchwork method into 3 categories, as follows:
[0089]
[0090] Among them, when the difference between the two calculated values and is large (greater than the first preset threshold ), the watermark is valid. According to the and relative size relationship, the corresponding watermark bit is 0 or 1; when the absolute value of the difference between and is less than the threshold , it means that the watermark is invalid, that is, no watermark is embedded in the current data segment.
[0091] b: Detect whether the number of the invalid watermark states of each candidate segment counted exceeds the second preset threshold; if so, the current candidate position and candidate length are incorrect, and if not, the current candidate position and candidate length are correct.
[0092] Here, it is detected whether the number of invalid watermark states of each candidate segment statistically detected exceeds a second preset threshold; if so, the current candidate position and candidate length are incorrect, and the candidate position and candidate length are continuously determined; if not, the current candidate position and candidate length are correct.
[0093] Among them, the second preset threshold is determined according to expert experience.
[0094] In this application, the new three-state watermark cannot simplify the dual search, and multiple positions still need to be considered, with multiple segment lengths considered for each position. However, the new algorithm can simplify the subsequent synchronization code matching. In fact, since the new algorithm has a built-in verification function, the synchronization code can be completely removed, and all embedded data are payloads, without the need to set aside a part of the embedding capacity for the synchronization code. For the new algorithm, at each candidate position, multiple segments are continuously decoded using the candidate segment lengths. Then, it is only necessary to count how many '2's appear in the decoding results. According to the previous algorithm description, for an audio segment without a watermark, or an audio segment with a watermark but incorrect segment parameters, the proportion of '2's after decoding is relatively large. If the number of watermark bit values '2' after decoding is less than or equal to the second preset threshold, it indicates that the correct segment position and segment length have been found; otherwise, it indicates that the parameters of the segment length or segment position are incorrect.
[0095] In a possible implementation manner, after segmenting the first audio information to generate multiple audio segments, performing transform domain processing on each audio segment, and determining the watermark data embedding domain in each audio segment, the digital audio watermark method further includes:
[0096] I: Performing a summation process on the audio data in the watermark data embedding domain to determine a quantized statistical eigenvalue.
[0097] Here, a summation process is performed on the audio data in the watermark data embedding domain to determine a quantized statistical eigenvalue.
[0098] II: Determining a quantization step based on the statistical characteristics and auditory threshold of the first audio information, and multiplying the quantization step by a preset fraction to determine a target value.
[0099] Here, a quantization step Δ is determined according to the statistical characteristics and auditory threshold of the first audio information, and the quantization step is multiplied by a preset fraction to determine a target value Δ / 6.
[0100] III: If the watermark bit value is 0, subtracting the target value from the quantized statistical eigenvalue; if the watermark bit value is 1, adding the target value to the quantized statistical eigenvalue, so as to complete the digital watermark information embedding process.
[0101] Here, if the watermark bit value is 0, the quantized statistical feature value is subtracted from the target value; if the watermark bit value is 1, the quantized statistical feature value is added to the target value, so as to complete the digital watermark information embedding process.
[0102] In this application, the three-state watermark algorithm proposed based on the original QIM algorithm transforms the above method into outputting 3 categories. When embedding the digital watermark, according to whether the watermark bit is 0 or 1 after quantization, the quantization result is added with -Δ / 6 or Δ / 6.
[0103] In a possible implementation manner, the digital audio watermark method further includes:
[0104] i: Segmenting the second audio information based on a plurality of preset candidate positions and candidate segment lengths to determine a plurality of candidate segments, and calculating the statistical feature value of each candidate segment.
[0105] Here, segmenting the second audio information based on a plurality of preset candidate positions and candidate segment lengths to determine a plurality of candidate segments, and calculating the statistical feature value of each candidate segment.
[0106] ii: Quantizing the statistical feature value of each candidate segment based on the quantization step size to determine the quantization residual value of each candidate segment.
[0107] Here, quantizing the statistical feature value of each candidate segment based on the quantization step size to determine the quantization residual value of each candidate segment.
[0108] iii: If the interval where the quantization residual value is located is the first preset interval or the second preset interval, the watermark state of the candidate segment is the valid watermark state; if the interval where the quantization residual value is located is the third preset interval, the watermark state of the candidate segment is the invalid watermark state; wherein, the ranges of the first preset interval, the third preset interval and the second preset interval increase in sequence.
[0109] Here, if the interval where the quantization residual value is located is the first preset interval or the second preset interval, the watermark state of the candidate segment is the valid watermark state; if the interval where the quantization residual value is located is the third preset interval, the watermark state of the candidate segment is the invalid watermark state.
[0110] Among them, the first preset interval is (0, 2 / 6], the second preset interval is (2 / 6, 4 / 6], and the third preset interval is (4 / 6, 1].
[0111] iv: Based on the number of invalid watermark states of each candidate segment calculated and the second preset threshold, determining whether the current candidate position and candidate length are correct.
[0112] Here, according to the number of invalid watermark states of each candidate segment counted and a preset second threshold, it is determined whether the current candidate position and candidate length are correct.
[0113] A digital audio watermarking method provided by an embodiment of the present application, the digital audio watermarking method includes: segmenting first audio information to generate a plurality of audio segments, performing transform domain processing on each of the audio segments to determine a watermark data embedding domain in each of the audio segments; performing digital watermark information embedding processing on the audio data in each of the watermark data embedding domains based on a digital audio watermark embedding algorithm to determine a plurality of audio segments after embedding watermarks, and combining the plurality of audio segments after embedding watermarks to determine second audio information; performing segmenting processing and eigenvalue statistics processing on the second audio information based on a plurality of preset candidate positions and candidate segment lengths to determine a statistical eigenvalue of each candidate segment; determining whether the current candidate position and candidate length are correct based on the statistical eigenvalue of each candidate segment, a first preset threshold, and a second preset threshold. If so, decoding the watermark bit values of each candidate segment in sequence based on the correct candidate position and the correct candidate length to determine the watermark information of the second audio information. The digital audio watermarking method provided by the present application has a built-in verification function, can eliminate the synchronization code in the anti-synchronization algorithm, and can quickly and accurately determine the correct candidate position and candidate length of the audio information after embedding the watermark, so as to accurately extract the watermark information.
[0114] Please refer to Figure 2 、 Figure 3 , Figure 2 which is one of the structural schematic diagrams of a digital audio watermarking device provided by an embodiment of the present application; Figure 3 which is another structural schematic diagram of a digital audio watermarking device provided by an embodiment of the present application. As Figure 2 shown in
[0115] A first processing module 210, configured to segment first audio information to generate a plurality of audio segments, perform transform domain processing on each of the audio segments, and determine a watermark data embedding domain in each of the audio segments;
[0116] A first watermark information embedding module 220, configured to perform digital watermark information embedding processing on the audio data in each of the watermark data embedding domains based on a digital audio watermark embedding algorithm, determine a plurality of audio segments after embedding watermarks, and combine the plurality of audio segments after embedding watermarks to determine second audio information;
[0117] The second processing module 230 is configured to perform segmentation processing and eigenvalue statistical processing on the second audio information based on a plurality of preset candidate positions and candidate segment lengths, and determine the statistical eigenvalues of each candidate segment.
[0118] The first watermark information extraction module 240 is configured to determine whether the current candidate position and candidate length are correct based on the statistical eigenvalues of each candidate segment, a first preset threshold, and a second preset threshold. If so, it decodes the watermark bit values of each candidate segment in sequence based on the correct candidate position and correct candidate length, and determines the watermark information of the second audio information.
[0119] Furthermore, when the first watermark information embedding module 220 is used to perform digital watermark information embedding processing on the audio data in each watermark data embedding domain based on the digital audio watermark embedding algorithm and determine a plurality of audio segments with embedded watermarks, the first watermark information embedding module 220 is specifically configured to:
[0120] If the watermark bit value is 1, add a constant value to the audio data of the first sub-audio segment in two adjacent sub-audio segments corresponding to the watermark data embedding domain, and subtract a constant value from the audio data in the second sub-audio segment, so as to complete the digital watermark information embedding processing.
[0121] If the watermark bit value is 0, add a constant value to the audio data of the first sub-audio segment in two adjacent sub-audio segments corresponding to the watermark data embedding domain, and subtract a constant value from the audio data in the second sub-audio segment, so as to complete the digital watermark information embedding processing.
[0122] Furthermore, when the second processing module 230 is used to perform segmentation processing and eigenvalue statistical processing on the second audio information based on a plurality of preset candidate positions and candidate segment lengths for each candidate position and candidate segment length, and determine the statistical eigenvalues of each candidate segment, the second processing module 230 is specifically configured to:
[0123] Perform segmentation processing on the second audio information based on the candidate position and candidate segment length, and determine a plurality of candidate segments of the second audio information.
[0124] For each candidate segment, divide the audio data in the candidate segment into equal first audio data sets and second audio data sets, calculate the mean values of the first audio data set and the second audio data set respectively, and determine the first statistical eigenvalue of the first audio data set and the second statistical eigenvalue of the second audio data set.
[0125] Further, when the first watermark information extraction module 240 is used to determine whether the current candidate position and candidate length are correct based on the statistical feature values of each candidate segment, the first preset threshold, and the second preset threshold, the first watermark information extraction module 240 specifically is configured to:
[0126] Based on the statistical feature values of the candidate segment and the first preset threshold, determine the watermark status of the current candidate segment; wherein, the watermark status includes an invalid watermark status and a valid watermark status, the watermark bit value of the valid watermark status is 0 or 1, and the watermark bit value of the invalid watermark status is 2;
[0127] Detect whether the number of the invalid watermark status of each candidate segment statistically obtained exceeds the second preset threshold;
[0128] If so, the current candidate position and candidate length are incorrect; if not, the current candidate position and candidate length are correct.
[0129] Further, when the first watermark information extraction module 240 is used to determine the watermark status of the current candidate segment based on the statistical feature values of the candidate segment and the first preset threshold, the first watermark information extraction module 240 specifically is configured to:
[0130] Detect whether the absolute value of the difference between the first statistical feature value and the second statistical feature value is greater than the first preset threshold;
[0131] If so, the watermark status of the current candidate segment is a valid watermark status; wherein, if the first statistical feature value is greater than the second statistical feature value, the decoded watermark bit value is 0, and if the first statistical feature value is less than the second statistical feature value, the decoded watermark bit is 1;
[0132] If not, the watermark status of the current candidate segment is an invalid watermark status.
[0133] Further, as Figure 3 shown, the digital audio watermark device 300 further includes a second watermark information embedding module 250, and the second watermark information embedding module 250 is configured to:
[0134] Perform a summation process on the audio data in the watermark data embedding domain to determine a quantized statistical feature value;
[0135] Based on the statistical characteristics of the first audio information and the auditory threshold, determine a quantization step size, and multiply the quantization step size by a preset fraction to determine a target value;
[0136] If the watermark bit value is 0, perform a subtraction process on the quantized statistical feature value and the target value;
[0137] If the watermark bit value is 1, the quantized statistical feature value and the target value are added to complete the embedding process of the digital watermark information.
[0138] Further, as Figure 3 shown, the digital audio watermark device 300 further includes a second watermark information extraction module 260, and the second watermark information extraction module 260 is configured to:
[0139] Segment the second audio information based on a plurality of preset candidate positions and candidate segment lengths to determine a plurality of candidate segments, and calculate the statistical feature value of each candidate segment;
[0140] Quantize the statistical feature value of each candidate segment based on a quantization step to determine the quantization residual value of each candidate segment;
[0141] If the interval where the quantization residual value is located is the first preset interval or the second preset interval, the watermark state of the candidate segment is a valid watermark state;
[0142] If the interval where the quantization residual value is located is the third preset interval, the watermark state of the candidate segment is an invalid watermark state; wherein, the ranges of the first preset interval, the third preset interval, and the second preset interval increase in sequence;
[0143] Based on the number of invalid watermark states of each candidate segment and a preset second threshold, determine whether the current candidate position and candidate length are correct.
[0144] A digital audio watermarking device provided by an embodiment of the present application, the digital audio watermarking device includes: a first processing module, configured to perform segmentation processing on first audio information to generate a plurality of audio segments, perform transform domain processing on each of the audio segments, and determine a watermark data embedding domain in each of the audio segments; a first watermark information embedding module, configured to perform digital watermark information embedding processing on the audio data in each of the watermark data embedding domains based on a digital audio watermark embedding algorithm, determine a plurality of audio segments after embedding watermarks, and combine the plurality of audio segments after embedding watermarks to determine second audio information; a second processing module, configured to perform segmentation processing and eigenvalue statistical processing on the second audio information based on a plurality of preset candidate positions and candidate segment lengths, and determine a statistical eigenvalue of each candidate segment; a first watermark information extraction module, configured to determine whether the current candidate position and candidate length are correct based on the statistical eigenvalue of each candidate segment, a first preset threshold, and a second preset threshold. If so, decode the watermark bit values of each candidate segment in sequence based on the correct candidate position and correct candidate length, and determine the watermark information of the second audio information. The digital audio watermarking method provided by the present application has a built-in verification function, can eliminate the synchronization code in the anti-synchronization algorithm, and can quickly and accurately determine the correct candidate position and candidate length of the audio information after embedding the watermark, so as to accurately extract the watermark information.
[0145] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 4 shown in
[0146] the electronic device 400 includes a processor 410, a memory 420, and a bus 430. Figure 1 The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 runs, the processor 410 communicates with the memory 420 through the bus 430. When the machine-readable instructions are executed by the processor 410, the steps of the digital audio watermarking method in the method embodiment shown above can be executed. The specific implementation manner can refer to the method embodiment and will not be elaborated here.
[0147] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the digital audio watermarking method in the method embodiment shown above can be executed. The specific implementation manner can refer to the method embodiment and will not be elaborated here. Figure 1
[0148] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0149] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some communication interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0150] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0151] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0152] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0153] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the technical field of the present application can still modify the technical solutions recorded in the foregoing embodiments or can easily conceive of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A digital audio watermarking method, characterized in that The digital audio watermarking method includes: Segment the first audio information to generate multiple audio segments, perform transform domain processing on each of the audio segments, and determine the watermark data embedding domain in each of the audio segments; Based on the digital audio watermark embedding algorithm, perform digital watermark information embedding processing on the audio data in each of the watermark data embedding domains, determine multiple audio segments with embedded watermarks, and combine the multiple audio segments with embedded watermarks to determine the second audio information; Based on a plurality of preset candidate positions and candidate segment lengths, perform segmentation processing and eigenvalue statistics processing on the second audio information, and determine the statistical eigenvalues of each candidate segment; Based on the statistical eigenvalues of each candidate segment, a first preset threshold, and a second preset threshold, determine whether the current candidate position and candidate length are correct. If so, decode the watermark bit values of each candidate segment in sequence based on the correct candidate position and correct candidate length, and determine the watermark information of the second audio information; After segmenting the first audio information to generate multiple audio segments, performing transform domain processing on each of the audio segments, and determining the watermark data embedding domain in each of the audio segments, the digital audio watermarking method further includes: Perform summation processing on the audio data in the watermark data embedding domain to determine the quantized statistical eigenvalue; Determine the quantization step size based on the statistical characteristics of the first audio information and the auditory threshold, and multiply the quantization step size by a preset fraction to determine the target value; If the watermark bit value is 0, perform subtraction processing on the quantized statistical eigenvalue and the target value; If the watermark bit value is 1, perform addition processing on the quantized statistical eigenvalue and the target value to complete the digital watermark information embedding processing; The digital audio watermarking method further includes: Based on a plurality of preset candidate positions and candidate segment lengths, perform segmentation processing on the second audio information to determine multiple candidate segments, and count the statistical eigenvalues of each candidate segment; Quantize the statistical eigenvalues of each candidate segment based on the quantization step size to determine the quantization residual value of each candidate segment; If the interval where the quantization residual value is located is the first preset interval or the second preset interval, the watermark state of the candidate segment is the valid watermark state; If the interval where the quantization residual value is located is the third preset interval, the watermark state of the candidate segment is the invalid watermark state; wherein, the ranges of the first preset interval, the third preset interval, and the second preset interval increase in sequence; Based on the number of invalid watermark states of each candidate segment counted and a preset second threshold, determine whether the current candidate position and candidate length are correct.
2. The digital audio watermarking method according to claim 1, characterized in that, The performing digital watermark information embedding processing on the audio data in each of the watermark data embedding domains based on the digital audio watermark embedding algorithm to determine multiple audio segments with embedded watermarks includes: If the watermark bit value is 1, add a constant value to the audio data of the first sub-audio segment among two adjacent sub-audio segments corresponding to the watermark data embedding domain, and subtract the constant value from the audio data in the second sub-audio segment, so as to complete the digital watermark information embedding process; If the watermark bit value is 0, add a constant value to the audio data of the first sub-audio segment among two adjacent sub-audio segments corresponding to the watermark data embedding domain, and subtract the constant value from the audio data in the second sub-audio segment, so as to complete the digital watermark information embedding process.
3. The digital audio watermarking method according to claim 1, characterized in that For each candidate position and candidate segment length, the second audio information is segmented and the eigenvalue statistics process is performed based on a plurality of preset candidate positions and candidate segment lengths to determine the statistical eigenvalue of each candidate segment, including: Segment the second audio information based on the candidate position and candidate segment length to determine a plurality of candidate segments of the second audio information; For each of the candidate segments, divide the audio data in the candidate segment into equal first audio data set and second audio data set, calculate the mean value of the first audio data set and the second audio data set respectively, and determine the first statistical eigenvalue of the first audio data set and the second statistical eigenvalue of the second audio data set.
4. The digital audio watermarking method according to claim 3, wherein Determining whether the current candidate position and candidate length are correct based on the statistical eigenvalue of each candidate segment, the first preset threshold, and the second preset threshold includes: Based on the statistical eigenvalue of the candidate segment and the first preset threshold, determine the watermark state of the current candidate segment; wherein, the watermark state includes an invalid watermark state and a valid watermark state, the watermark bit value of the valid watermark state is 0 or 1, and the watermark bit value of the invalid watermark state is 2; Detect whether the number of the invalid watermark states of each candidate segment counted exceeds the second preset threshold; If so, the current candidate position and candidate length are incorrect, if not, the current candidate position and candidate length are correct.
5. The digital audio watermarking method according to claim 4, wherein Determining the watermark state of the current candidate segment based on the statistical eigenvalue of the candidate segment and the first preset threshold includes: Detect whether the absolute value of the difference between the first statistical eigenvalue and the second statistical eigenvalue is greater than the first preset threshold; If so, the watermark state of the current candidate segment is a valid watermark state; wherein, if the first statistical eigenvalue is greater than the second statistical eigenvalue, the decoded watermark bit value is 0, and if the first statistical eigenvalue is less than the second statistical eigenvalue, the decoded watermark bit is 1; If not, the watermark state of the current candidate segment is an invalid watermark state.
6. A digital audio watermarking device, characterized in that, The digital audio watermark device includes: A first processing module, configured to segment the first audio information to generate a plurality of audio segments, perform a transform domain process on each of the audio segments, and determine a watermark data embedding domain in each of the audio segments; The first watermark information embedding module is used to perform digital watermark information embedding processing on the audio data in each of the watermark data embedding domains based on a digital audio watermark embedding algorithm, determine multiple audio segments with embedded watermarks, and combine the multiple audio segments with embedded watermarks to determine the second audio information; The second processing module is used to perform segmentation processing and eigenvalue statistical processing on the second audio information based on a plurality of preset candidate positions and candidate segment lengths, and determine the statistical eigenvalues of each candidate segment; The first watermark information extraction module is used to determine whether the current candidate position and candidate length are correct based on the statistical eigenvalues of each candidate segment, a first preset threshold, and a second preset threshold. If so, decode the watermark bit values of each candidate segment in sequence based on the correct candidate position and correct candidate length, and determine the watermark information of the second audio information; The digital audio watermark device further includes a second watermark information embedding module, and the second watermark information embedding module is used for: Perform a summation process on the audio data in the watermark data embedding domain to determine the quantized statistical eigenvalue; Determine a quantization step size based on the statistical characteristics and auditory threshold of the first audio information, multiply the quantization step size by a preset fraction to determine a target value; If the watermark bit value is 0, perform a subtraction process on the quantized statistical eigenvalue and the target value; If the watermark bit value is 1, perform an addition process on the quantized statistical eigenvalue and the target value to complete the digital watermark information embedding process; The digital audio watermark device further includes a second watermark information extraction module, and the second watermark information extraction module is used for: Perform segmentation processing on the second audio information based on a plurality of preset candidate positions and candidate segment lengths to determine multiple candidate segments, and count the statistical eigenvalues of each candidate segment; Perform quantization processing on the statistical eigenvalues of each candidate segment based on the quantization step size to determine the quantization residual value of each candidate segment; If the interval where the quantization residual value is located is the first preset interval or the second preset interval, the watermark state of the candidate segment is a valid watermark state; If the interval where the quantization residual value is located is the third preset interval, the watermark state of the candidate segment is an invalid watermark state; wherein, the ranges of the first preset interval, the third preset interval, and the second preset interval increase in sequence; Based on the number of invalid watermark states of each candidate segment counted and a preset second threshold, determine whether the current candidate position and candidate length are correct.
7. An electronic device, characterized in that, Comprising: A processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are run by the processor, the steps of the digital audio watermark method according to any one of claims 1 to 5 are executed.
8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of the digital audio watermark method according to any one of claims 1 to 5 are executed.