Video conference real-time audio tamper-proofing watermark detection method and device
By performing frame segmentation and DCT processing on the audio signal, modifying the singular value relationship of the intermediate frequency coefficients, and combining weighted voting decision-making and synchronous search mechanisms, the problem of synchronization difficulties and insufficient robustness of audio watermarks in video conferencing is solved, achieving high-accuracy audio tampering detection.
Patent Information
- Application Number
- CN202511326471.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-07-08
- Filing Date
- 2025-09-17
- Publication Date
- 2025-11-21
AI Technical Summary
Existing audio watermarking technologies suffer from synchronization difficulties and insufficient robustness in video conferencing, especially under low bitrate compression and network latency, making it difficult to achieve high-accuracy audio tampering detection.
The audio signal is divided into multiple frames and DCT is performed. Watermark bits are embedded by modifying the singular value relationship of the intermediate frequency coefficients of the subframes. Combined with weighted voting decision and synchronous search mechanism, real-time detection and tampering alarm of audio signal are realized.
It achieves a detection accuracy of up to 99.99% at audio transmission rates of 256kbps and above, and is robust against compression, packet loss and resampling to ensure audio quality.
Smart Images

Figure CN120998211A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital audio watermarking, in particular to a video conference real-time audio tamper-proof watermark detection method and device. BACKGROUND
[0002] With the popularization of video conferencing, the security of conference content is increasingly important. Attackers may insert themselves into the communication link between the conference terminal and the server through ARP spoofing, DNS hijacking or routing hijacking, realize two-way data eavesdropping, and then impersonate the participants to speak or tamper with the key audio content of the conference to commit fraud. Therefore, there is an urgent need for an audio watermarking technology that can detect in real time and alarm in time when the audio is tampered with.
[0003] Existing audio watermarking technologies mainly include time domain watermarking and frequency domain watermarking: 1. Time domain watermarking: watermark is embedded by modifying the amplitude of the sampling points, which has fast calculation speed but poor robustness and is easily affected by compression and noise; 2. Frequency domain watermarking: usually embedded in the discrete cosine transform (DCT) or wavelet transform domain, which has good robustness, but is mainly used for embedding and extracting audio files, and does not consider the watermark distortion caused by compression, resampling and packet loss in real-time audio communication scenarios.
[0004] The existing audio watermarking technology has the following problems: 1. Synchronization difficulty: watermark detection requires accurate starting position, and most existing methods are sensitive to audio cropping and displacement; 2. Insufficient robustness: under low bit rate compression (such as 256kbps), the watermark extraction accuracy significantly decreases; packet loss caused by network delay and audio-video synchronization also causes false detection, making it difficult to apply in video conference tamper-proof scenarios with extremely high accuracy requirements.
[0005] Therefore, the present application is proposed. SUMMARY
[0006] The technical problem solved by the present application is to overcome the shortcomings of the prior art and provide a video conference real-time audio tamper-proof watermark detection method and device. The video conference real-time audio tamper-proof watermark detection method comprises: dividing the audio signal into multiple audio frames, each frame being divided into two sub-frames and performing DCT; selecting a middle frequency band and embedding watermark bits by modifying the singular value relationship of the middle frequency coefficients in the sub-frames; determining the starting position of the watermark information by analyzing the coefficient ratio of the corresponding frequency band and weighting the voting decision, and combining the synchronization search mechanism. The video conference real-time audio tamper-proof watermark detection method can process audio streams in real time, ensure high sound quality, and has robustness against compression, packet loss and resampling, and the detection accuracy is more than 99.99% under 256kbps and above audio transmission rate.
[0007] The first embodiment of the present application introduces a real-time audio anti-forgery watermark detection method for video conference, comprising: S11. Sampling the real-time audio stream at a sampling rate fs to obtain an audio signal F with a length of m sampling points, where m∈N + ; taking the peak absolute value of the audio signal F, and presetting a threshold T1; when the peak absolute value < the threshold T1, skipping watermark embedding; when the peak absolute value ≥ the threshold T1, equally dividing the audio signal F into j audio frames f t , where j is the information length of a binary watermark sequence W, t, j∈N + , t≤j; S12. Equally dividing the audio frame f t into two sub-frames, and performing DCT on the two sub-frames to obtain frequency domain coefficients; S13. According to the sampling rate, adaptively calculating the start point b of the medium frequency range of the DCT frequency domain coefficients of the two sub-frames, and taking n continuous medium frequency coefficients l 1i 、 l 2i , where b, i, n∈N + , 3≤n≤7, i≤n; S14. Adjusting the medium frequency coefficients by modifying the singular value relationship of the medium frequency coefficients of the two sub-frames to obtain two new sub-frames, comprising: a. When bit t =0, if , no modification is needed; if , the medium frequency coefficients l 1i and l 2i are modified to make ; b. When bit t =1, if , no modification is needed; if , the medium frequency coefficients l 1i and l 2i are modified to make ; where bit t is the bit value of the watermark sequence W, is a watermark embedding strength factor; S15. Performing IDCT on the two new sub-frames to obtain a watermark-containing audio frame, and sequentially combining the watermark-containing audio frames to obtain a watermark-containing audio signal; S16. The real-time audio stream is sampled at a sampling rate fs to obtain an audio signal F1 with a length of 2m sampling points; S17. Take the absolute value of the peak value of the audio signal F1. When the absolute value of the peak value is less than the threshold T1, skip the watermark detection; when the absolute value of the peak value is greater than or equal to the threshold T1, perform synchronous detection on the audio signal F1 to determine the starting sampling point of the candidate watermark sequence W1, including: a. Let the sampling point sequence of the audio signal F1 be |0|1|2|...|k|k+1|...|k+1023|...|2m|, and select the offset k as the starting sampling point of the candidate watermark sequence W1; b. Calculate the audio signal F for each audio frame f. t Number of sampling points M t Watermark extraction score q t_i Consistency score q t Specifically: ; in T1 is the weighting factor, and T2 is the preset detection threshold; c. Calculate the j audio frames f t The average consistency score Q is as follows: ; d. Traverse the offset position k, and denote the information columns of the top r ranked by the average consistency score Q as {Qr, k}. r , W1r}, where k r W1r represents the corresponding offset and candidate watermark sequence. The candidate watermark sequence W1r is cyclically shifted right by s positions to obtain the watermark sequence W2r. The Hamming distance between the watermark sequence W2r and the watermark sequence W is calculated, and the offset k corresponding to the minimum Hamming distance is selected. r As a synchronous sampling point, the starting sampling point ; A preset threshold T3 is set. If the minimum Hamming distance is less than the threshold T3, then synchronization is successful and the starting sampling point k is recorded, and S18-S19 continue; if the minimum Hamming distance is greater than or equal to the threshold T3, then synchronization fails and returns to S16 to start again. S18. Sample the real-time audio stream at a sampling rate fs to obtain an audio signal F2 with a length of 2m sampling points. Using the starting sampling point k as the starting point, sample the audio signal F2 at a sampling rate fs to obtain an audio signal F3 with a length of m sampling points. Extract the watermark from the audio signal F3 to obtain a watermark sequence W', including: a. Divide the audio signal F3 into j equal audio frames f t1 and for the audio frame f t1performing DCT on the two sub-frames respectively to obtain frequency domain coefficients; b. dividing the audio frame f t1 into two sub-frames, and performing DCT on the two sub-frames respectively to obtain frequency domain coefficients; c. extracting medium frequency coefficients of the two sub-frames respectively, wherein a medium frequency range of the medium frequency coefficients is the same as a medium frequency range when the watermark sequence W is embedded; d. calculating a ratio of the medium frequency coefficients of the two sub-frames, and determining a watermark bit bit t1 using a weighted voting manner to obtain a watermark sequence W', wherein t1∈N + , t1≤j, and specifically: calculating a consistency score q t1 of each audio frame f t1 of the audio signal F3, when q , bit t1 =1, and when q , bit t1 =0; S19. calculating a Hamming distance between the watermark sequence W' and the watermark sequence W, when the Hamming distance < the threshold value T3, watermark verification is passed, that is, the audio signal F1 is not tampered with; when the Hamming distance ≥ the threshold value T3 for x consecutive times, watermark verification is not passed, that is, the audio signal F1 is tampered with, and an alarm prompt is given.
[0008] The second embodiment of the application introduces a video conference real-time audio tamper-proofing watermark detection device, based on the first embodiment, comprising: D11 framing module, sampling a real-time audio stream at a sampling rate fs to obtain an audio signal F with a length of m sampling points, wherein m∈N + ; taking a peak absolute value of the audio signal F, and presetting a threshold value T1, when the peak absolute value < the threshold value T1, skipping watermark embedding; when the peak absolute value ≥ the threshold value T1, dividing the audio signal F into j audio frames f t , wherein j is an information length of a binary watermark sequence W, t, j∈N + , t≤j; D12 DCT module, dividing the audio frame f t into two sub-frames, and performing DCT on the two sub-frames respectively to obtain frequency domain coefficients; D13 coefficient selection module, adaptively calculating a medium frequency range starting point b of DCT frequency domain coefficients of the two sub-frames according to the sampling rate, and taking n continuous medium frequency coefficients of the two sub-frames respectively l 1i , l 2i , wherein b, i, n∈N +, 3≤n≤7, i≤n; The D14 coefficient modification module adjusts the intermediate frequency coefficients in the two sub-frames to obtain two new sub-frames by modifying singular value relations of the intermediate frequency coefficients, comprising: a. When bit t =0, if , no modification is needed; if , the intermediate frequency coefficients l 1i and l 2i are modified to ; b. When bit t =1, if , no modification is needed; if , the intermediate frequency coefficients l 1i and l 2i are modified to ; wherein bit t is a bit value of the watermark sequence W, and is a watermark embedding strength factor; The D15 audio output module combines the two new sub-frames by IDCT to obtain a watermark-containing audio frame, and combines the watermark-containing audio frames in time sequence to obtain a watermark-containing audio signal. The D16 preprocessing module samples the real-time audio stream at a sampling rate fs to obtain an audio signal F1 with a length of 2m sampling points. The D17 synchronization detection module takes the peak absolute value of the audio signal F1, and when the peak absolute value is < the threshold T1, skips watermark detection; when the peak absolute value is ≥ the threshold T1, performs synchronization detection on the audio signal F1 to determine the starting sampling point of the candidate watermark sequence W1, comprising: a. Taking the sampling point sequence diagram of the audio signal F1 as |0|1|2|...|k|k+1|...|k+1023|...|2m|, the offset bit k is selected as the starting sampling point of the candidate watermark sequence W1; b. The sampling point number M t , the watermark extraction score q t , and the consistency score q t_i of each audio frame f t of the audio signal F are calculated, specifically: ; wherein is a weight factor, and T2 is a preset detection threshold; c. The j audio frames f tThe average consistency score Q is as follows: ; d. Traverse the offset position k, and denote the information columns of the top r ranked by the average consistency score Q as {Qr, k}. r , W1r}, where k r W1r represents the corresponding offset and candidate watermark sequence. The candidate watermark sequence W1r is cyclically shifted right by s positions to obtain the watermark sequence W2r. The Hamming distance between the watermark sequence W2r and the watermark sequence W is calculated, and the offset k corresponding to the minimum Hamming distance is selected. r As a synchronous sampling point, the starting sampling point ; If the minimum Hamming distance is less than the threshold T3, the synchronization is successful and the starting sampling point k is recorded. Then, proceed to steps D18 (watermark extraction module) and D19 (verification alarm module). If the minimum Hamming distance is greater than or equal to the threshold T3, the synchronization fails and returns to step D16 (preprocessing module) to start over. The D18 watermark extraction module samples the real-time audio stream at a sampling rate fs to obtain an audio signal F2 with a length of 2m sampling points. Starting from the initial sampling point k, it samples the audio signal F2 at a sampling rate fs to obtain an audio signal F3 with a length of m sampling points. The module then extracts the watermark from the audio signal F3 to obtain a watermark sequence W', including: a. Divide the audio signal F3 into j equal audio frames f t1 and for the audio frame f t1 Perform DCT; b. Transfer the audio frame f t1 The image is divided into two subframes, and DCT is performed on each of the two subframes to obtain the frequency domain coefficients. c. Extract the intermediate frequency coefficients of the two subframes respectively, wherein the intermediate frequency range of the intermediate frequency coefficients is the same as the intermediate frequency range when the watermark sequence W is embedded; d. Calculate the ratio of the intermediate frequency coefficients of the two subframes, and determine the watermark bit using a weighted voting method. t1 This yields the watermark sequence W', where t1∈N + t1≤j, specifically: Calculate the audio frame f of the audio signal F3 per frame. t1 Consistency score q t1 ,when At that time, bit t1 =1, when At that time, bit t1 =0; D19 verification alarm module, the Hamming distance of the watermark sequence W' and the watermark sequence W is calculated, when the Hamming distance < the threshold value T3, the watermark verification passes, that is, the audio signal F1 is not tampered with;When the Hamming distance is greater than or equal to the threshold value T3 for x consecutive times, the watermark verification fails, that is, the audio signal F1 has been tampered with and an alarm prompt is issued.
[0009] The embodiment of the application further provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are used for executing the video conference real-time audio tamper-proof watermark detection method.
[0010] The embodiment of the application further provides an electronic device, which comprises a memory and a processor, wherein the memory stores instructions executable by the processor, and the instructions are used for executing the video conference real-time audio tamper-proof watermark detection method.
[0011] Compared with the prior art, the application has the following beneficial effects: 1. Multi-band joint embedding and synchronization mechanism: the same watermark bit is embedded in multiple intermediate frequency coefficients at the same time, and the consistency of the multi-band detection result is used as the synchronization basis, so that the watermark synchronization problem of real-time audio stream is solved; 2. Adaptive frequency band selection: the intermediate frequency range is dynamically adjusted according to the adaptive sampling rate, so that the best frequency band can be selected under different sampling rates, and the anti-resampling capability is improved; 3. Weighted voting detection algorithm: the multiple DCT intermediate frequency coefficients are given decreasing weights from low to high frequency, and the low frequency coefficient has a higher weight, so that the anti-compression capability is enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the application and, together with the specification, serve to explain the principles of the application. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained from these drawings without creative labor for those skilled in the art.
[0013] Figure 1 The video conference real-time audio tamper-proof watermark detection method flowchart is provided by the application.
[0014] Figure 2 The structure diagram of the video conference real-time audio tamper-proof watermark detection device provided by the application.
[0015] D11 frame dividing module; D12 DCT module; D13 coefficient selecting module; D14 coefficient modifying module; D15 audio output module; D16 pre-processing module; D17 synchronization detecting module; D18 watermark extracting module; D19 verification alarming module. DETAILED DESCRIPTION
[0016] In order to make the objects, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0017] The terms used in the embodiments of the present application are only for the purpose of describing the specific embodiments, and are not intended to limit the present application. The singular forms "a", "an" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plural" generally includes at least two.
[0018] It should be understood that the term "and / or" used herein only describes an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects.
[0019] Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (a stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (a stated condition or event)" or "in response to detecting (a stated condition or event)".
[0020] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that the products or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such products or devices. Without more limitations, the element defined by the sentence "including a" does not exclude the presence of another identical element in the product or device including the element.
[0021] The optional embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0022] AsFigure 1 As shown, the first embodiment of the present application introduces a real-time audio anti-forgery watermark detection method for video conference, comprising: S11. Sampling the real-time audio stream at a sampling rate fs to obtain an audio signal F with a length of m sampling points, where m∈N + ; taking the peak absolute value of the audio signal F, and presetting a threshold T1, when the peak absolute value < the threshold T1, skipping watermark embedding; when the peak absolute value ≥ the threshold T1, equally dividing the audio signal F into j audio frames f t , where j is the information length of the binary watermark sequence W, t, j∈N + , t≤j; S12. Equally dividing the audio frame f t into two sub-frames, and performing DCT on the two sub-frames to obtain frequency domain coefficients; S13. According to the sampling rate, adaptively calculating the start point b of the medium frequency range of the DCT frequency domain coefficients of the two sub-frames, and taking n continuous medium frequency coefficients of the two sub-frames respectively l 1i 、 l 2i , where b, i, n∈N + , 3≤n≤7, i≤n; S14. Adjusting the medium frequency coefficients by modifying the singular value relationship of the medium frequency coefficients of the two sub-frames to obtain two new sub-frames, comprising: a. When bit t =0, if , no modification is needed; if , modify the medium frequency coefficients l 1i and l 2i to make ; b. When bit t =1, if , no modification is needed; if , modify the medium frequency coefficients l 1i and l 2i to make ; where bit t is the bit value of the watermark sequence W, is a watermark embedding strength factor; S15. Performing IDCT on the two new sub-frames to obtain a watermark-containing audio frame, and merging the watermark-containing audio frames in time sequence to obtain a watermark-containing audio signal; S16. The real-time audio stream is sampled at a sampling rate fs to obtain an audio signal F1 with a length of 2m sampling points; S17. Take the absolute value of the peak value of the audio signal F1. When the absolute value of the peak value is less than the threshold T1, skip the watermark detection; when the absolute value of the peak value is greater than or equal to the threshold T1, perform synchronous detection on the audio signal F1 to determine the starting sampling point of the candidate watermark sequence W1, including: a. Let the sampling point sequence of the audio signal F1 be |0|1|2|...|k|k+1|...|k+1023|...|2m|, and select the offset k as the starting sampling point of the candidate watermark sequence W1; b. Calculate the audio signal F for each audio frame f. t Number of sampling points M t Watermark extraction score q t_i Consistency score q t Specifically: ; in T1 is the weighting factor, and T2 is the preset detection threshold; c. Calculate the j audio frames f t The average consistency score Q is as follows: ; d. Traverse the offset position k, and denote the information columns of the top r ranked by the average consistency score Q as {Qr, k}. r , W1r}, where k r W1r represents the corresponding offset and candidate watermark sequence. The candidate watermark sequence W1r is cyclically shifted right by s positions to obtain the watermark sequence W2r. The Hamming distance between the watermark sequence W2r and the watermark sequence W is calculated, and the offset k corresponding to the minimum Hamming distance is selected. r As a synchronous sampling point, the starting sampling point ; A preset threshold T3 is set. If the minimum Hamming distance is less than the threshold T3, then synchronization is successful and the starting sampling point k is recorded, and S18-S19 continue; if the minimum Hamming distance is greater than or equal to the threshold T3, then synchronization fails and returns to S16 to start again. S18. Sample the real-time audio stream at a sampling rate fs to obtain an audio signal F2 with a length of 2m sampling points. Using the starting sampling point k as the starting point, sample the audio signal F2 at a sampling rate fs to obtain an audio signal F3 with a length of m sampling points. Extract the watermark from the audio signal F3 to obtain a watermark sequence W', including: a. Divide the audio signal F3 into j equal audio frames f t1 and for the audio frame f t1performing DCT; b. dividing the audio frame f t1 into two sub-frames and performing DCT on the two sub-frames respectively to obtain frequency domain coefficients; c. extracting medium frequency coefficients of the two sub-frames respectively, wherein the medium frequency range of the medium frequency coefficients is the same as the medium frequency range when the watermark sequence W is embedded; d. calculating the ratio of the medium frequency coefficients of the two sub-frames, and determining the watermark bit bit t1 using a weighted voting method to obtain the watermark sequence W', wherein t1∈N + , t1≤j, specifically: calculating the consistency score q t1 of each audio frame f t1 of the audio signal F3, when bit t1 =1, and when bit t1 =0; S19. calculating the Hamming distance between the watermark sequence W' and the watermark sequence W, when the Hamming distance < the threshold value T3, the watermark verification passes, that is, the audio signal F1 is not tampered with; when the Hamming distance ≥ the threshold value T3 for x consecutive times, the watermark verification fails, that is, the audio signal F1 has been tampered with and an alarm prompt is issued.
[0023] In step S11, the audio signal F usually has 1024 sampling points, about 23 milliseconds of audio. When the audio signal F is divided into j audio frames f t , if the total length of the audio signal F cannot be evenly divided by the number of frames, usually truncation or zero padding processing needs to be selected.
[0024] In step S13, when the sampling rate is 44.1 kHz, 1 kHz~3 kHz can be selected as the medium frequency range.
[0025] The right circular shift in step S17 is a method of operating the bits of a binary number. Its core rule is: shift each bit of the binary number to the right by a specified number of bits, and the bits removed from the right (the lowest bit) are not discarded, but are refilled to the empty bits on the left (the highest bit).
[0026] For further illustration, the original number 10001 is right circularly shifted by 1 bit to obtain 11000; the original number 00110 is right circularly shifted by 1 bit to obtain 00011.
[0027] In an exemplary example, the parameter n=7 of n consecutive medium frequency coefficients of the two sub-frames in step S13 is the optimal solution.
[0028] In an exemplary instance, when the audio signal F1 sample point number is 2048, the synchronization sample point k r = 32 is the optimal solution.
[0029] In an exemplary instance, the watermark embedding strength factor in step S14 is .
[0030] In an exemplary instance, the preset detection threshold T2 in step S17 is 0.5≤T2≤1.5, wherein T2=0.5 is the optimal solution.
[0031] In an exemplary instance, the weight factor in step S17 is , which is the optimal solution.
[0032] In an exemplary instance, the preset threshold T3 in step S19 is T3=3, which is the optimal solution.
[0033] In an exemplary instance, the continuous x=5 in step S19 when the Hamming distance is greater than the threshold T3 is the optimal solution.
[0034] In an exemplary instance, when the audio signal F1 is a dual-channel, the channel with higher volume is selected to detect the audio signal F1.
[0035] The video conference real-time audio tamper-proof watermark detection method has the following advantages: The present application can process audio stream in real time, has robustness such as anti-compression, anti-packet loss and anti-resampling while ensuring high sound quality, and the detection accuracy is more than 99.99% under the audio transmission rate of 256 kbps and above.
[0036] As Figure 2 shown, the second embodiment of the present application introduces a video conference real-time audio tamper-proof watermark detection device, based on the first embodiment, comprising: D11 frame module, sampling the real-time audio stream at a sampling rate fs to obtain an audio signal F with a length of m sample points, wherein m∈N + ; taking the peak absolute value of the audio signal F, and presetting a threshold T1, when the peak absolute value is less than the threshold T1, skipping watermark embedding; when the peak absolute value is greater than or equal to the threshold T1, equally dividing the audio signal F into j audio frames f t , wherein j is the information length of the binary watermark sequence W, t, j∈N + , t≤j; D12 DCT module, equally dividing the audio frame f t into two sub-frames, and performing DCT on the two sub-frames to obtain frequency domain coefficients; D13 coefficient selection module, according to the sampling rate, adaptively calculates the middle frequency range starting point b of the DCT frequency domain coefficient of the two subframes, and takes n continuous middle frequency coefficients of the two subframes respectively l 1i 、 l 2i , wherein b, i, n ∈ N + , 3 ≤ n ≤ 7, i ≤ n D14 coefficient modification module, by modifying the singular value relationship of the middle frequency coefficients of the two subframes, adjusting the middle frequency coefficients to obtain two new subframes, comprising: a. When bit t =0, if , no modification is needed; if , the middle frequency coefficients l 1i and l 2i are modified ; b. When bit t =1, if , no modification is needed; if , the middle frequency coefficients l 1i and l 2i are modified ; wherein bit t is the bit value of the watermark sequence W, is a watermark embedding strength factor; D15 audio output module, IDCT merging the two new subframes to obtain a watermark-containing audio frame, and merging the watermark-containing audio frame in time sequence to obtain a watermark-containing audio signal; D16 preprocessing module, sampling the real-time audio stream at a sampling rate fs to obtain an audio signal F1 with a length of 2m sampling points; D17 synchronization detection module, taking the peak absolute value of the audio signal F1, when the peak absolute value < the threshold T1, skipping watermark detection; when the peak absolute value ≥ the threshold T1, performing synchronization detection on the audio signal F1 to determine the starting sampling point of the candidate watermark sequence W1, comprising: a. recording the sampling point sequence diagram of the audio signal F1 as |0|1|2|...|k|k+1|...|k+1023|...|2m|, selecting the offset bit k as the starting sampling point of the candidate watermark sequence W1; b. calculating the sampling point number M t of each audio frame f t of the audio signal F, and the watermark extraction score q t_iConsistency score q t , specifically: ; wherein is a weight factor, and T2 is a preset detection threshold; c. Calculate the average value Q of the consistency scores q of the j audio frames f t , specifically: ; d. Traverse the offset bit k, and record the top r information in the average value Q of the consistency scores as {Qr, k r , W1r}, wherein k r , W1r is the corresponding offset bit and candidate watermark sequence, the watermark sequence W2r is obtained by cyclically right shifting the candidate watermark sequence W1r by s bits, the Hamming distance between the watermark sequence W2r and the watermark sequence W is calculated, and the offset bit k r corresponding to the minimum Hamming distance is selected as the synchronization sampling point, and the starting sampling point ; is the preset threshold T3, if the minimum Hamming distance < the threshold T3, the synchronization is successful and the starting sampling point k is recorded, and the steps D18 watermark extraction module and D19 verification alarm module are continued; if the minimum Hamming distance ≥ the threshold T3, the synchronization fails and returns to the preprocessing module D16 to start again; D18 watermark extraction module, the real-time audio stream is sampled at a sampling rate fs to obtain an audio signal F2 with a length of 2m sampling points, the audio signal F2 is sampled at a sampling rate fs starting from the starting sampling point k to obtain an audio signal F3 with a length of m sampling points, and the watermark sequence W' is extracted from the audio signal F3, including: a. The audio signal F3 is equally divided into j audio frames f t1 , and the DCT is performed on the audio frames f t1 ; b. The audio frames f t1 are equally divided into two sub-frames, and the DCT is performed on the two sub-frames respectively to obtain frequency domain coefficients; c. The intermediate frequency coefficients of the two sub-frames are extracted respectively, wherein the intermediate frequency range of the intermediate frequency coefficients is the same as the intermediate frequency range when the watermark sequence W is embedded; d. The ratio of the intermediate frequency coefficients of the two sub-frames is calculated, and the weighted voting method is used to determine the watermark bit bit t1 , to obtain the watermark sequence W', wherein t1 ∈ N + , t1 ≤ j, specifically: Calculate the consistency score q t1 of each audio frame f t1 of the audio signal F3, when bit = 1 when t1 = 1, and bit = 0 when t1 = 0. D19 verification alarm module, calculate the Hamming distance between the watermark sequence W' and the watermark sequence W, when the Hamming distance < the threshold T3, the watermark verification is passed, that is, the audio signal F1 is not tampered with; when the Hamming distance is greater than or equal to the threshold T3 for x consecutive times, the watermark verification fails, that is, the audio signal F1 has been tampered with and an alarm prompt is issued.
[0037] The embodiment of the present application also provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are used for executing the video conference real-time audio tamper-proofing watermark detection method.
[0038] The embodiment of the present application further provides an electronic device, which comprises a memory and a processor, wherein the memory stores instructions executable by the processor, and the instructions are used for executing the video conference real-time audio tamper-proofing watermark detection method.
[0039] Specifically, a system or device provided with a storage medium can be provided, and the storage medium stores software program codes for realizing the functions of any one of the above embodiments, and the computer (or CPU or MPU) of the system or device reads and executes the program codes stored in the storage medium.
[0040] In this case, the program codes read from the storage medium can realize the functions of any one of the above embodiments, and thus the program codes and the storage medium storing the program codes constitute a part of the present application.
[0041] The storage medium for providing the program codes includes floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, nonvolatile memory cards, and ROMs. Alternatively, the program codes can be downloaded from a server computer through a communication network.
[0042] In addition, it should be clear that not only the program codes read by the computer can be executed, but also part or all of the actual operations can be completed by the operating system and the like operating on the computer based on the instructions of the program codes, so as to realize the functions of any one of the above embodiments.
[0043] Further, it is understood that the program code read by the storage medium can be written into the memory provided in the expansion board inserted into the computer or the memory provided in the expansion unit connected to the computer, and then the CPU or the like mounted on the expansion board or the expansion unit is caused to perform part or all of the actual operation based on the instruction of the program code, thereby realizing the function of any of the above-described embodiments.
[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not limited thereto; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or part or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting a real-time audio tamper-proof watermarking of a video conference, characterized in that, Comprising: S11. Sampling a real-time audio stream at a sampling rate fs to obtain an audio signal F with a length of m samples, where m ∈ N + ; taking a peak absolute value of the audio signal F, and presetting a threshold T1, when the peak absolute value < the threshold T1, skipping watermark embedding; when the peak absolute value ≥ the threshold T1, equally dividing the audio signal F into j audio frames f t , where j is the information length of a binary watermark sequence W, t, j ∈ N + , t ≤ j; S12. divide the audio frame f t into two sub-frames, and perform DCT on the two sub-frames respectively to obtain frequency domain coefficients; S13. Calculate the middle frequency range start point b of DCT frequency domain coefficient of the two subframes according to the sampling rate, and take n continuous middle frequency coefficients for the two subframes respectively l 1i 、 l 2i wherein b, i, n ∈ N + , 3 ≤ n ≤ 7, i ≤ n; S14. Adjusting the intermediate frequency coefficients of the two subframes by modifying the singular value relationship of the intermediate frequency coefficients to obtain two new subframes, comprising: a. When bit t = 0, if then no modification is needed; if then the intermediate coefficient l 1i and l 2i is made ; b. When bit t = 1, if then no modification is needed; if then the intermediate coefficient l 1i and l 2i is made ; where bit t is the bit value of the watermark sequence W, is a watermark embedding strength factor; S15. Merging the two new subframes by IDCT to obtain a watermark-containing audio frame, and merging the watermark-containing audio frames in time sequence to obtain a watermark-containing audio signal; S16. Sampling the real-time audio stream at a sampling rate fs to obtain an audio signal F1 with a length of 2m sampling points; S17. Taking the peak absolute value of the audio signal F1, when the peak absolute value < the threshold value T1, skip watermark detection; when the peak absolute value ≥ the threshold value T1, perform synchronous detection on the audio signal F1 to determine the starting sampling point of the candidate watermark sequence W1, comprising: a. Record the sampling point sequence diagram of the audio signal F1 as |0|1|2|...|k|k+1|...|k+1023|...|2m|, select the offset bit k as the starting sampling point of the candidate watermark sequence W1; b. computing a number M of sample points of the audio signal F per audio frame f t of the audio signal F t , a watermark extraction score q t_i , a consistency score q t , in particular: ; wherein is a weight factor, T2 is a preset detection threshold value; c. computing an average value Q of the consistency scores of the j audio frames f t , in particular: ; d. Traverse the offset bit k, record the information column with the top r consistency score average value Q as {Qr, k r , W1r}, wherein k r , W1r are the corresponding offset bit and candidate watermark sequence, the candidate watermark sequence W1r is cyclically right shifted by s bits to obtain a watermark sequence W2r, the Hamming distance between the watermark sequence W2r and the watermark sequence W is calculated, and the offset bit k r corresponding to the minimum Hamming distance is selected as the synchronization sampling point, and the starting sampling point ; Pre-set threshold value T3, if the minimum Hamming distance < the threshold value T3, synchronization is successful and the starting sampling point k is recorded, continue S18-S19; if the minimum Hamming distance ≥ the threshold value T3, synchronization fails and restarts from S16; S18. Sampling the real-time audio stream at a sampling rate fs to obtain an audio signal F2 with a length of 2m sampling points, and sampling the audio signal F2 at a sampling rate fs to obtain an audio signal F3 with a length of m sampling points starting from the starting sampling point k, extracting a watermark from the audio signal F3 to obtain a watermark sequence W', comprising: a. dividing the audio signal F3 into j audio frames f t1 and performing a DCT on the audio frames f t1 ; b. dividing the audio frame f t1 into two sub-frames and performing DCT on the two sub-frames respectively to obtain frequency domain coefficients; c. Extract the intermediate frequency coefficients of the two subframes respectively, wherein the intermediate frequency range of the intermediate frequency coefficients is the same as that when the watermark sequence W is embedded; d. Calculate the ratio of the intermediate frequency coefficients of the two subframes, and determine the watermark bit bit using weighted voting method t1 , to obtain the watermark sequence W', where t1∈N + , t1≤j, specifically: Calculate the audio frame f of the audio signal F3 per frame. t1 Consistency score q t1 ,when At that time, bit t1 =1, when At that time, bit t1 =0; S19. Calculate the Hamming distance between the watermark sequence W' and the watermark sequence W, when the Hamming distance < the threshold value T3, the watermark verification is passed, that is, the audio signal F1 has not been tampered with; when the Hamming distance ≥ the threshold value T3 for x consecutive times, the watermark verification fails, that is, the audio signal F1 has been tampered with and an alarm prompt is issued.
2. The video conference real-time audio tamper-proof watermark detection method according to claim 1, wherein the parameter n=7 of n consecutive intermediate frequency coefficients is the optimal solution in S13.
3. The method of claim 1, wherein the watermark embedding strength factor in S14 is in the range of .
4. The method of claim 1, wherein, When the audio signal F1 sample number is 2048, the synchronization sample point k r = 32 is the optimal solution.
5. The method of claim 1, wherein, The preset detection threshold value 0.5≤T2≤1.5 in S17, wherein T2=0.5 is the optimal solution.
6. The method of claim 1, wherein, The weight factor in S17 is the optimal solution.
7. The method of claim 1, wherein, The preset threshold value T3=3 in S19 is the optimal solution.
8. The method of claim 1, wherein, The consecutive x=5 in S19 when the Hamming distance ≥ the threshold value T3 is the optimal solution.
9. The method of claim 1, wherein, When the audio signal F1 is a double-channel, the channel with higher volume is selected to detect the audio signal F1.
10. A video conference real-time audio tamper-proof watermark detection apparatus employing the video conference real-time audio tamper-proof watermark detection method according to any one of claims 1, characterized by, Comprising: D11 framing module, sampling the real-time audio stream at a sampling rate fs to obtain an audio signal F with a length of m samples, where m ∈ N + ; taking the peak absolute value of the audio signal F, presetting a threshold T1, when the peak absolute value < the threshold T1, skipping watermark embedding; when the peak absolute value ≥ the threshold T1, equally dividing the audio signal F into j audio frames f t , where j is the information length of the binary watermark sequence W, t, j ∈ N + , t ≤ j; D12 a DCT module, to transform the audio frame f t is equally divided into two subframes, and the two subframes are respectively transformed by DCT to obtain frequency domain coefficients; D13 coefficient selection module, according to the sampling rate, adaptively calculates the middle frequency range starting point b of the DCT frequency domain coefficient of the two subframes, and takes n continuous middle frequency coefficients for the two subframes respectively l 1i 、 l 2i , wherein b, i, n ∈ N + , 3 ≤ n ≤ 7, i ≤ n; D14 coefficient modification module, adjusting the intermediate frequency coefficients of the two subframes by modifying the singular value relationship of the intermediate frequency coefficients to obtain two new subframes, comprising: a. When bit t = 0, if then no modification is needed; if then the intermediate coefficient l 1i and l 2i is made ; b. When bit t = 1, if then no modification is needed; if then the intermediate coefficient l 1i and l 2i is made ; where bit t is the bit value of the watermark sequence W, is a watermark embedding strength factor; D15 audio output module, merging the two new subframes by IDCT to obtain a watermark-containing audio frame, and merging the watermark-containing audio frames in time sequence to obtain a watermark-containing audio signal; D16 pre-processing module, sampling the real-time audio stream at a sampling rate fs to obtain an audio signal F1 with a length of 2m sampling points; D17 synchronization detection module, taking the peak absolute value of the audio signal F1, when the peak absolute value < the threshold T1, skipping watermark detection; when the peak absolute value ≥ the threshold T1, performing synchronization detection on the audio signal F1 to determine the starting sampling point of the candidate watermark sequence W1, including: a. recording the sampling point sequence diagram of the audio signal F1 as |0|1|2|...|k|k+1|...|k+1023|...|2m|, selecting the offset bit k as the starting sampling point of the candidate watermark sequence W1; b. calculating the number M of sample points of the audio signal F per audio frame f t b. calculating the number M of sample points of the audio signal F per audio frame f t b. calculating the number M of sample points of the audio signal F per audio frame f t_i b. calculating the number M of sample points of the audio signal F per audio frame f t b. calculating the number M of sample points of the audio signal F per audio frame f ; wherein is a weight factor, and T2 is a preset detection threshold value. c. computing an average value Q of the consistency scores of the j audio frames f t , in particular: ; d. Traverse the offset bit k, record the information column with the top r consistency score average value Q as {Qr, k r , W1r}, wherein k r , W1r are the corresponding offset bit and candidate watermark sequence, the candidate watermark sequence W1r is cyclically right shifted by s bits to obtain a watermark sequence W2r, the Hamming distance between the watermark sequence W2r and the watermark sequence W is calculated, and the offset bit k r corresponding to the minimum Hamming distance is selected as the synchronization sampling point, and the starting sampling point ; preset threshold T3, if the minimum Hamming distance < the threshold T3, synchronization is successful and the starting sampling point k is recorded, and the watermark extraction module D18 and the verification alarm module D19 are continued; if the minimum Hamming distance ≥ the threshold T3, synchronization fails and returns to the pre-processing module D16 to start again; D18 watermark extraction module, sampling the real-time audio stream at a sampling rate fs to obtain an audio signal F2 with a length of 2m sampling points, taking the starting sampling point k as the starting point to sample the audio signal F2 at a sampling rate fs to obtain an audio signal F3 with a length of m sampling points, extracting the watermark from the audio signal F3 to obtain a watermark sequence W', including: a. dividing the audio signal F3 into j audio frames f t1 and performing a DCT on the audio frames f t1 ; b. dividing the audio frame f t1 into two sub-frames and performing DCT on the two sub-frames respectively to obtain frequency domain coefficients; c. extracting the intermediate frequency coefficients of the two subframes respectively, wherein the intermediate frequency range of the intermediate frequency coefficients is the same as the intermediate frequency range when the watermark sequence W is embedded; d. Calculate the ratio of the intermediate frequency coefficients of the two subframes, and determine the watermark bit bit using weighted voting method t1 , to obtain the watermark sequence W', where t1∈N + , t1≤j, specifically: computing a consistency score q of the audio signal F3 per audio frame f t1 t1 bit = 1 when bit t1 = 1, and bit = 0 when bit t1 = 0; D19 verification alarm module, calculating the Hamming distance between the watermark sequence W' and the watermark sequence W, when the Hamming distance < the threshold T3, the watermark verification passes, that is, the audio signal F1 has not been tampered with; when the Hamming distance ≥ the threshold T3 for x consecutive times, the watermark verification fails, that is, the audio signal F1 has been tampered with and an alarm prompt is issued.