Video watermark embedding processing method, video watermark extraction processing method, video watermark embedding processing device, video watermark extraction processing device and video watermark equipment
By performing color space conversion, time domain segmentation and aliquoting of video data, the minimum perceptible difference threshold for the airspace is dynamically calculated, and watermark embedding is performed based on this, which solves the problem of poor robustness of existing video watermark algorithms in screen shooting scenarios, achieving higher stability and robustness.
Patent Information
- Application Number
- CN202510518813.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In the screen shooting scenario, when facing complex screen shooting attacks, there is a big contradiction between visual masking and embedding intensity, resulting in poor robustness.
By performing color space conversion, time domain segmentation and aliquoting processing on the video data, the minimum perceptible difference threshold in the airspace is dynamically calculated, and watermark embedding processing is performed based on this. This method combines the active prediction mechanism of the visual cortex and the analysis of heterogeneous visual feature to achieve the invisibility and robustness of watermarks.
Effectively resist geometric deformation and composite noise interference caused by the screen-camera channel, enhancing the stability and robustness of screen-camera attacks.
Smart Images

Figure CN120075545A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer information processing, and particularly to a video watermark embedding processing method, an extraction processing method, a device and a device. Background Art
[0002] As a basic paradigm in the field of multimedia security, the core goal of information hiding technology is to embed encrypted or encoded sensitive information into a carrier in an imperceptible manner by analyzing the redundant features of the host data, so as to achieve the functions of covert communication and copyright marking while maintaining the perceptual quality of the original content. Under this technical framework, digital watermarking has become a key research direction due to its wide applicability. It hides a unique copyright identifier or metadata into digital media such as images and audio through a specific algorithm, constructing a verifiable and robust marking system, thus providing technical support for content traceability, tampering detection and copyright declaration. Video watermarking technology, with various formats of video files as the research carrier, shows an extremely urgent demand in the industrial field. A video watermarking system usually uses an imperceptible information embedding mechanism to achieve the covert transmission of copyright information while maintaining visual quality. This technical framework needs to construct an embedding model that fuses multi-dimensional features. For example, combined with the quantization parameters of video coding standards, an adaptive embedding strategy is designed in the time domain - spatial domain - transform domain.
[0003] According to the differences in the embedding domain characteristics and performance indicators, the existing video watermarking algorithms can be divided into two mainstream methods: the frequency domain and the spatial domain. Their technical paths and applicable scenarios show significant differences. The video watermarking technology based on frequency domain characteristics embeds watermark information covertly in the transform domain coefficients by analyzing the spectral distribution characteristics of video signals. Typical methods use the discrete cosine transform to construct a block transform matrix, or combine the multi-resolution analysis framework of the discrete wavelet transform, and preferentially select low-frequency components for quantization modulation. Benefiting from the frequency domain energy concentration characteristics, this type of algorithm has natural compatibility with compression coding such as H.264 / HEVC, and shows strong robustness when resisting attacks such as re-encoding and noise interference in traditional channels. However, when facing the composite distortion accumulated during the processes of video decoding, screen display, optical acquisition, re-encoding, etc. in the screen capture attack chain, the robustness is poor.
[0004] Spatial video watermarking technology generally directly operates on the spatial pixel values of video frames. Usually, some strategies are adopted to improve the robustness and invisibility of watermarks. For example, by utilizing the characteristics of the Human Visual System (HVS), watermarks are embedded in the pixels of areas with lower human eye sensitivity. A common approach is to embed copyright information by changing components such as the brightness and color of pixels. Such algorithms avoid complex frequency-domain transformation calculations and have significant advantages in real-time streaming media transmission and low-power devices. However, in the screen capture scenario, there is a large contradiction between visual masking and embedding strength in existing spatial video watermarking algorithms, and instability occurs when facing complex screen capture attacks. Summary of the Invention
[0005] The present invention provides a video watermark embedding processing method, an extraction processing method, a device, and a device, which can effectively resist the geometric deformation and composite noise interference problems brought by the screen-camera channel and enhance the stability during screen capture attacks.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] An embodiment of the present invention provides a video watermark embedding processing method, including:
[0008] Obtain video data to be processed;
[0009] Perform color space conversion processing on the video data to be processed to obtain original video data;
[0010] Perform segmentation processing on the original video data according to a preset duration to obtain a plurality of time-domain groups, where each time-domain group includes consecutive video frames;
[0011] Perform equal division processing on the consecutive video frames to obtain a first video frame and a second video frame, where the first video frame and the second video frame both include a plurality of channel data, and the plurality of channel data includes luminance channel data and chrominance channel data;
[0012] Perform threshold dynamic calculation processing on the plurality of channel data to obtain the spatial just noticeable difference thresholds of multiple channels;
[0013] Perform watermark embedding processing on the plurality of channel data according to the spatial just noticeable difference thresholds of multiple channels to obtain video data with embedded watermarks.
[0014] Optionally, performing threshold dynamic calculation processing on the plurality of channel data to obtain the spatial just noticeable difference thresholds of multiple channels includes:
[0015] Determine residual masking effect data according to the plurality of channel data;
[0016] Perform a first preprocessing on the luminance channel data to obtain luminance adaptive masking data;
[0017] Determine composite masking effect data according to the residual masking effect data and the luminance adaptive masking data;
[0018] Determine the self-information of color components according to the chrominance channel data;
[0019] Perform a second preprocessing on the luminance channel data in combination with edge features to obtain luminance-edge coupled saliency;
[0020] Determine perceptual weight data according to the self-information of the color components and the luminance-edge coupled saliency;
[0021] Determine the spatial just noticeable difference (JND) thresholds of multiple channels according to the composite masking effect data and the perceptual weight data.
[0022] Optionally, determining the spatial JND thresholds of multiple channels according to the composite masking effect data and the perceptual weight data includes:
[0023] According to:
[0024] ,
[0025] Obtain the spatial JND thresholds of multiple channels;
[0026] where JND θ,k is the spatial JND threshold of the θ channel, θ ∈ {Y, U, V}, Y is luminance, U and V are chrominance, and DM k,θ is the composite masking effect data, and PW is the perceptual weight data.
[0027] Optionally, performing watermark embedding processing on the multiple channel data according to the spatial JND thresholds of the multiple channels to obtain video data with an embedded watermark includes:
[0028] Obtain watermark data;
[0029] Encode the watermark data to obtain watermark bit data;
[0030] Perform watermark embedding processing on the multiple channel data according to the spatial JND thresholds of the multiple channels and the watermark bit data to obtain video data with an embedded watermark.
[0031] An embodiment of the present invention further provides a method for video watermark extraction processing, including:
[0032] Obtain video data with a watermark;
[0033] Perform color space conversion processing on the watermarked video data to obtain the video data after color space conversion;
[0034] Perform segmentation processing on the video data after color space conversion according to a preset duration to obtain a plurality of time-domain groups;
[0035] Perform equal division processing on the video frames in the plurality of time-domain groups to obtain a first video frame and a second video frame;
[0036] Perform difference statistics processing on the first video frame and the second video frame to obtain a comprehensive difference value;
[0037] Determine the watermark bit position according to the comprehensive difference value;
[0038] Process the watermark bit position to obtain watermark data.
[0039] Optionally, performing difference statistics processing on the first video frame and the second video frame to obtain a comprehensive difference value includes:
[0040] According to:
[0041] ,
[0042] Obtain a comprehensive difference value;
[0043] Where D j is the comprehensive difference value, is the mean value of the j-th frame of the first part of the video in channel θ, is the mean value of the j-th frame of the second part of the video in channel θ, Y is the luminance, and U and V are the chrominance.
[0044] Optionally, determining the watermark bit position according to the comprehensive difference value includes:
[0045] According to:
[0046] ,
[0047] Obtain the watermark bit position;
[0048] Where, is the watermark bit position.
[0049] An embodiment of the present invention further provides a video watermark embedding processing device, including:
[0050] A first acquisition module, configured to acquire video data to be processed;
[0051] A first processing module, configured to perform color space conversion processing on the video data to be processed to obtain original video data; perform segmentation processing on the original video data according to a preset duration to obtain a plurality of time-domain groups, where each time-domain group includes consecutive video frames; perform equal division processing on the consecutive video frames to obtain a first video frame and a second video frame, where the first video frame and the second video frame both include a plurality of channel data, and the plurality of channel data includes luminance channel data and chrominance channel data; perform threshold dynamic calculation processing on the plurality of channel data to obtain spatial minimum perceptible difference thresholds for multiple channels; and perform watermark embedding processing on the plurality of channel data according to the spatial minimum perceptible difference thresholds for multiple channels to obtain video data with an embedded watermark.
[0052] An embodiment of the present invention further provides a video watermark extraction processing device, including:
[0053] A second acquisition module, configured to acquire video data with a watermark;
[0054] A second processing module, configured to perform color space conversion processing on the video data with a watermark to obtain video data after color space conversion; perform segmentation processing on the video data after color space conversion according to a preset duration to obtain a plurality of time-domain groups; perform equal division processing on the video frames in the plurality of time-domain groups to obtain a first video frame and a second video frame; perform difference statistics processing on the first video frame and the second video frame to obtain a comprehensive difference value; determine a watermark bit position according to the comprehensive difference value; and process the watermark bit position to obtain watermark data.
[0055] An embodiment of the present invention further provides a computing device, including: a processor and a memory storing a computer program, where when the computer program is run by the processor, the above method is executed.
[0056] The technical solution of the present invention at least includes the following effects:
[0057] The above solution of the present invention performs color space conversion processing on the video data to be processed to obtain original video data; performs segmentation processing on the original video data according to a preset duration to obtain a plurality of time-domain groups, where each time-domain group includes consecutive video frames; performs equal division processing on the consecutive video frames to obtain a first video frame and a second video frame, where both the first video frame and the second video frame include a plurality of channel data, and the plurality of channel data includes luminance channel data and chrominance channel data; performs threshold dynamic calculation processing on the plurality of channel data to obtain the spatial just noticeable difference thresholds of the plurality of channels; performs watermark embedding processing on the plurality of channel data according to the spatial just noticeable difference thresholds of the plurality of channels to obtain video data with an embedded watermark; relies on the collaborative modeling of the double masking effect driven by the active prediction mechanism of the visual cortex for the process of obtaining the spatial just noticeable distortion threshold, combines heterogeneous visual feature analysis and dynamic perception weight adaptive regulation mechanism, and finally realizes the quantization of the just noticeable difference (JND) threshold. The obtained JND threshold will be applied to the embedding link of adding alternating perturbation signals to the image frames processed by the grouping strategy. In this process, by constructing a periodically cumulative differential gradient field in the time domain of the carrier video, this feature shows considerable stability after being attacked, so as to show good robustness when facing screen capture attacks. In addition, the redundant information after BCH coding enhances the fault tolerance of the watermark sequence during watermark extraction, and forms an implicit identifier in cooperation with the grouping mechanism. Even if the attack causes frame rate deviation, the watermark timing information can still be restored by time window alignment during extraction, avoiding the problem of phase mismatch. Description of the Drawings
[0058] Figure 1 is a flowchart of the video watermark embedding processing method provided by an embodiment of the present invention;
[0059] Figure 2 is a flowchart of the video watermark extraction processing method provided by an embodiment of the present invention;
[0060] Figure 3 is a schematic diagram of the time-domain grouping strategy of the video watermark embedding processing method provided by an embodiment of the present invention;
[0061] Figure 4 is a schematic diagram of the process of the video watermark extraction processing method provided by an embodiment of the present invention;
[0062] Figure 5 is a structural diagram of the video watermark embedding processing device provided by an embodiment of the present invention;
[0063] Figure 6 is a structural diagram of the video watermark extraction processing device provided by an embodiment of the present invention;
[0064] Figure 7 It is a schematic structural diagram of a computing device provided by an embodiment of the present invention. Detailed implementation manners
[0065] Hereinafter, exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.
[0066] As Figure 1 shown, an embodiment of the present invention provides a video watermark embedding processing method, including:
[0067] Step 11, obtaining video data to be processed;
[0068] Step 12, performing color space conversion processing on the video data to be processed to obtain original video data;
[0069] Step 13, dividing the original video data according to a preset duration to obtain a plurality of time-domain groups, where each time-domain group includes consecutive video frames;
[0070] Step 14, equally dividing the consecutive video frames to obtain a first video frame and a second video frame, where the first video frame and the second video frame both include a plurality of channel data, and the plurality of channel data includes luminance channel data and chrominance channel data;
[0071] Step 15, performing threshold dynamic calculation processing on the plurality of channel data to obtain the spatial just noticeable difference thresholds of the plurality of channels;
[0072] Step 16, performing watermark embedding processing on the plurality of channel data according to the spatial just noticeable difference thresholds of the plurality of channels to obtain video data with an embedded watermark.
[0073] In this embodiment, first, original video data to be watermarked is obtained from sources such as locally stored video files, network video streams, and video capture devices. The obtained video data is segmented according to a preset duration. For example, the video can be segmented into a time-domain group every 10 s, and each time-domain group contains consecutive video frames within that time period.
[0074] Within each time-domain group, the consecutive video frames are equally divided. The equally divided video frames provide smaller processing units for subsequent watermark embedding. By embedding watermarks in different small blocks, the distribution range of the watermark can be increased, and the robustness of the watermark can be improved.
[0075] Convert the evenly processed video frames from one color space, such as the RGB color space, to another color space, such as the YUV color space. In the YUV color space, the video frames are decomposed into a luminance channel (Y channel) and chrominance channels (U channel and V channel);
[0076] Among them, the luminance channel is more sensitive to the human eye, while the chrominance channels are relatively less sensitive to the human eye. By converting the video frames from RGB to the YUV color space, different channels can be processed separately, and appropriate watermark embedding positions and intensities can be selected according to the visual characteristics of the human eye, thereby improving the invisibility and robustness of the watermark.
[0077] For the data of each channel, a dynamic calculation method is used to determine the spatial minimum perceptible difference threshold (JND threshold). The JND threshold represents the minimum luminance or chrominance change that the human eye can perceive in this channel. The dynamic calculation method considers factors such as the content, luminance, and contrast of the video frames to ensure that the calculated JND threshold can accurately reflect the visual perception characteristics of the human eye.
[0078] According to the JND threshold of each channel, watermark embedding processing is performed on the channel data. Specifically, a specific watermark signal can be added to the channel data, and the intensity of the watermark signal does not exceed the JND threshold of this channel. For example, the watermark information can be embedded by adjusting the pixel values of the luminance channel or chrominance channels. By reasonably using the JND thresholds of different channels, effective embedding and invisibility of the watermark can be achieved. The video data after embedding the watermark can be used in various application scenarios. When it is necessary to verify the video copyright, the watermark information can be extracted through the corresponding watermark extraction algorithm, thereby proving the copyright ownership of the video.
[0079] This technical solution dynamically quantifies the distortion threshold by constructing a spatial JND model based on the visual system, and fuses multi-dimensional heterogeneous features such as color, edges, and luminance, effectively balancing the relationship between the embedding strength and the visual masking effect, ensuring the invisibility of the watermark; at the same time, through the time-domain grouping strategy and the modulation of the YUV color space, the reshaping of the statistical characteristics of each channel is realized, and in cooperation with the BCH coding (Bose Chaudhuri Hocquenghem), it can effectively resist the problems of frame rate offset and timing mismatch, and significantly improve the robustness against screen capture.
[0080] In an optional embodiment of the present invention, in step 12, performing color space conversion processing on the video data to be processed to obtain the original video data may include:
[0081] According to the significant difference in sensitivity of rods and cones in the human retina to brightness and chromaticity, in order to facilitate the subsequent HVS brightness-chromaticity perception decoupling characteristics, channel-independent feature analysis is performed. The formula for converting the video image frame from RGB to YUV color space is:
[0082]
[0083] In this embodiment, different color spaces have different characteristics and application scenarios. RGB (red, green, blue) color space is a color space commonly used in computer graphics and digital image processing, which directly corresponds to the color output of the display device. YUV (brightness, chrominance) color space is a color space that separates brightness (Y) and chrominance (U, V) components, which is more in line with the human visual system (HVS) perception characteristics of brightness and chrominance. The color space conversion processing is performed on the video data to be processed, aiming to convert the video image frame from the RGB color space to the YUV color space. The purpose is to facilitate the subsequent use of the brightness-chrominance perception decoupling characteristics of the human visual system (HVS) for channel-independent feature analysis.
[0084] Since the rods in the human retina are mainly responsible for perceiving brightness, while the cones are responsible for perceiving chromaticity. There are significant differences in the sensitivity of rods and cones to brightness and chromaticity, which enables the human visual system to process brightness and chromaticity information more effectively. The YUV color space separates the brightness (Y) and chromaticity (U, V) components, which is more in line with the perceptual characteristics of the human visual system. In the YUV color space, the brightness component (Y) can be processed independently, and the chromaticity components (U, V) can also be processed specifically according to their characteristics. The formula for converting video image frames from RGB to YUV color space is usually based on certain mathematical transformation formulas. These formulas are designed to convert the red, green, and blue components in the RGB color space into the brightness and chromaticity components in the YUV color space.
[0085] Conversion process: For each frame of the video data to be processed, first obtain its RGB components; then, use the above conversion formula to convert the RGB components into YUV components; the converted YUV components will be used as the original video data for subsequent processing and analysis.
[0086] In an optional embodiment of the present invention, in step 13, the original video data is segmented according to a preset time length to obtain a plurality of time domain groups, wherein each time domain group includes continuous video frames, which may include:
[0087] Assume the original video sequence is ,in Represents the kth frame, and the video is fixed in length By splitting, we can get a time-domain grouping, denoted as , .
[0088] In the field of video processing and analysis, long videos are segmented for more detailed analysis; time-domain grouping is a method of dividing video data into multiple consecutive segments in chronological order, and each segment contains a certain number of consecutive video frames.
[0089] In this embodiment, long videos are segmented for more detailed analysis; time-domain grouping is a method of dividing video data into multiple consecutive segments in chronological order, and each segment contains a certain number of consecutive video frames. The specific process includes:
[0090] Representation of the original video sequence: Let the original video sequence be , where represents the k-th frame, T is the total number of frames of the video, indicating the length of the video sequence.
[0091] Setting the segmentation duration: The preset duration ΔT is the basis for segmenting the video, and it determines the number of video frames included in each time-domain grouping. ΔT can be set according to specific requirements, such as several seconds, several minutes, or other time units.
[0092] Generating time-domain groupings: By dividing the video at a fixed duration ΔT, M time-domain groupings can be obtained, denoted as G i , where i = 1, 2, ⋯, M. Each time-domain grouping G i contains consecutive video frames, which are continuous in time and the total duration does not exceed ΔT.
[0093] In an alternative embodiment of the present invention, in step 14, the consecutive video frames are equally divided to obtain a first video frame and a second video frame, where the first video frame and the second video frame both include a plurality of channel data, and the plurality of channel data includes luminance channel data and chrominance channel data, and may include:
[0094] Through a video segmentation method with a symmetric time-domain structure, that is, the sequence image frames of each group are further equally divided into two parts, front and back, denoted as part A and part B respectively.
[0095] Since screen capture attacks are often accompanied by frame rate offsets, traditional fixed-frame number grouping will cause temporal breaks. Therefore, this strategy aligns the watermark signal period with the physical time axis strictly through time-base segmentation, avoiding phase misalignment caused by resampling. And part A and part B of each group are complementary distributed on the time axis, laying the foundation for adopting differential operations when embedding watermarks later.
[0096] It should be noted that each frame group will correspond to 1 bit of information encoded by BCH, and is one embedding cycle.
[0097] In an optional embodiment of the present invention, step 15 may include:
[0098] Step 151, determining residual masking effect data according to the multiple channel data;
[0099] Step 152, performing first preprocessing on the luminance channel data to obtain luminance adaptive masking data;
[0100] Step 153, determining composite masking effect data according to the residual masking effect data and the luminance adaptive masking data;
[0101] Step 154, determining the self-information of color components according to the chrominance channel data;
[0102] Step 155, performing second preprocessing on the luminance channel data in combination with edge features to obtain luminance-edge coupled saliency;
[0103] Step 156, determining perceptual weight data according to the self-information of the color components and the luminance-edge coupled saliency;
[0104] Step 157, determining the spatial minimum perceptible difference threshold of multiple channels according to the composite masking effect data and the perceptual weight data.
[0105] In this embodiment, the calculation of the spatial just-noticeable distortion threshold is the core link to achieve the balance between concealment and robustness. Its necessity stems from the non-linear characteristic of the HVS's sensitivity to distortion, and it is necessary to accurately quantify the distortion perception boundary through physiological-psychophysical modeling. The design of its JND model is mainly composed of the following two main links, namely the composite masking model and the perceptual weight adjustment dominated by heterogeneous visual features.
[0106] To simulate the active prediction mechanism of the human brain's visual cortex for image content, a neural prediction model that combines residual masking (RM) and luminance adaptation masking (LAM) is specifically implemented as follows:
[0107] Quantification of residual masking effect: The human brain's visual cortex has the ability to actively predict image content, and an autoregressive prediction model is constructed to generate a predicted image; since the YUV three channels are independently modeled and analyzed, which is beneficial to accurately quantify the masking effect of the HVS on different color space components, so according to the calculation of the pixel values of a certain channel of the input image frame and the absolute value of the residual ψ of the predicted imagek,θ , which is used to characterize the prediction uncertainty; and θ is used to distinguish color channels, where θ ∈ {Y, U, V}. The residual masking effect data RM k,θ By adjusting the coefficient η θ weights the residuals to quantify the masking intensity, ψ k,θ and RM k,θ is calculated as follows:
[0108]
[0109]
[0110] where is the weight coefficient of the i-th surrounding pixel, is the pixel value of the surrounding pixel position z in the k-th frame in channel θ i , where i and k are natural numbers, is the error compensation term, , 1.0, 1.0;
[0111] Adaptive luminance masking dynamic modeling: The human eye's perception sensitivity to noise in local areas with different brightness levels shows significant differences. It is necessary to dynamically adjust the masking threshold to match the non-linear characteristics of brightness perception. The piecewise function design conforms to the physiological response law of the HVS to ensure stricter distortion control in low-brightness areas.
[0112]
[0113] where is the pixel value of the Y channel, i.e., the luminance channel, of this coordinate in the YUV color space; the luminance masking threshold of the k-th frame maintains a piecewise characteristic, with the threshold decreasing in the low-brightness area and increasing linearly in the high-brightness area; is the adaptive luminance masking data. By performing the first preprocessing process on the luminance channel data as described above, the adaptive luminance masking data can be obtained ;
[0114] Non-linear superposition of double masking effects: The residual masking RM k,θ and the adaptive luminance masking are fused through a non-linear superposition formula, specifically:
[0115]
[0116] where C θ (taking 0.3, 0.25, and 0.2 for the YUV channels respectively) is used to eliminate the overlap of the masking effect and avoid overestimation of the threshold; DM k,θ is the composite masking effect data.
[0117] Visual attention is driven by the competition of heterogeneous features such as color and edges. At the same time, regions with high saliency (such as bright objects in the center, prominent edges) require a lower distortion threshold to match the attention priority of the HVS. To comprehensively consider various influencing factors, the analysis is carried out from three dimensions: color heterogeneity, spatial structure sensitivity, and brightness difference. The entire process is based on the clustering analysis of the color space distribution of the Gaussian Mixture Model (GMM), the extraction of the edge gradient field features using the Canny multi-directional gradient filter, and the quantification of the brightness region contrast based on the Kullback-Leibler (KL) divergence, and then the perceptual weight is obtained to regulate the results of the previous process.
[0118] Color perception saliency modeling: Several color categories obtained by pre-clustering with GMM are used to comprehensively represent the color regions with similar perceptual characteristics in the image, denoting the th color component. Then, through the quantification of color contrast, spatial distribution dispersion, and central preference, etc., the color perception saliency is finally characterized by self-information , and the formula is:
[0119]
[0120] where, is the Gaussian weight of the color component in GMM, reflecting the statistical proportion or importance of this color in the overall image. The larger this weight, the more extensive or prominent the distribution of this color; is the contrast intensity of the color component , which comprehensively considers color contrast, spatial distribution dispersion (quantified by color variance), the average distance from the color region to the center of the image, etc.; p(c i ) is the probability of the color component c i ; I(c i ) is the self-information of the color component c i . The larger its value, the brighter, more concentrated, and closer to the center of the image this color component is, and the stronger the visual saliency.
[0121] Brightness-edge coupling saliency modeling: Given that flat regions (low contrast) can tolerate more distortion due to the brightness masking effect and have lower saliency, while edge regions require strict distortion limitation due to high sensitivity. This link conducts a joint analysis of brightness and edges. First, the Canny edge detection is used to simulate the processing of edge information by the human eye through retinal ganglion cells, and the edge weight is extracted. Further, through the design of an exponential function, the edge saliency is dynamically enhanced, and the edge saliency adjustment factor , finally calculate the luminance-edge coupling saliency , and the specific formula is:
[0122]
[0123]
[0124] Wherein, is the Gaussian weight of the luminance component after clustering the luminance values through GMM; is the contrast intensity of this luminance component ; p(δ i ) is the probability of the luminance component δ i ; I(δ i ) is the luminance-edge coupling saliency. Through the above luminance channel data and combining with edge features for the second preprocessing process, the luminance-edge coupling saliency can be obtained.
[0125] The JND threshold formula for dynamically perceiving weight regulation is:
[0126]
[0127]
[0128] Wherein, JND θ,k is the spatial minimum perceptible difference threshold of the θ channel, θ ∈ {Y, U, V}, Y is luminance, U and V are chrominance, DM k,θ is the composite masking effect data, and PW is the perception weight data. Unifying the above features affecting perception into a single weight to avoid threshold conflicts caused by independent adjustment, which conforms to the characteristics of multi-feature collaborative perception of HVS.
[0129] In an optional embodiment of the present invention, step 16 may include:
[0130] Step 161, obtaining watermark data;
[0131] Step 162, encoding the watermark data to obtain watermark bit data;
[0132] Step 163, according to:
[0133] ,
[0134] obtaining the video data embedded with the watermark;
[0135] Wherein, F' k,θ (x, y) is the pixel value at coordinates (x, y) in the θ channel of the kth frame, and F k,θ (x, y) is the video frame F kThe pixel value of the θ channel, A is the first part of the video frame, B is the second part of the video frame, and b is the watermark bit data.
[0136] In this embodiment, under the framework of the time-domain grouping strategy, the video sequence is divided into several groups of frames with a fixed duration of ∆T, and each group corresponds to a watermark bit after BCH error correction coding. Each group of frames is further divided into two subsequences A and B along the time axis. Through the multi-pixel collaborative adjustment mechanism, controlled chromaticity and luminance offsets are introduced in the YUV three channels, and finally stable inter-group statistical feature differences are formed to achieve watermark coding. The specific process is as follows:
[0137] ,
[0138] Among them, F' k,θ (x, y) is the pixel value of the k-th frame at the coordinate (x, y) in the θ channel, and F k,θ (x, y) is the θ-channel pixel value of the video frame F k , A is the first part of the video frame, B is the second part of the video frame, and b is the BCH-encoded bit corresponding to the current frame group. Through the symbol alternation mechanism, the A and B subsequences show positive and negative perturbation trends on the YUV channels respectively, ensuring the stability of the statistical feature differences.
[0139] As Figure 2 shown, an embodiment of the present invention proposes a video watermark extraction processing method, including:
[0140] Step 21, obtaining video data with a watermark;
[0141] Step 22, performing color space conversion processing on the video data with the watermark to obtain the video data after color space conversion;
[0142] Step 23, dividing the video data after color space conversion according to a preset duration to obtain a plurality of time-domain groups;
[0143] Step 24, equally dividing the video frames in the plurality of time-domain groups to obtain a first video frame and a second video frame;
[0144] Step 25, performing difference statistical processing on the first video frame and the second video frame to obtain a comprehensive difference value;
[0145] Step 26, determining the watermark bit according to the comprehensive difference value;
[0146] Step 27, processing the watermark bit to obtain the watermark data.
[0147] In this embodiment, first, video data with watermarks is obtained from a network video platform, locally stored video files, etc.; through color space conversion, the video data can be converted from the RGB color space to the YUV color space. In the YUV color space, the Y component represents luminance information, and the U and V components represent chrominance information. Since watermark information is usually less associated with luminance information and has a certain correlation with chrominance information, converting to the YUV color space can facilitate the extraction of watermarks more conveniently. For the conversion from RGB to YUV, the following formula can be used:
[0148] The video data has continuity in the time dimension. By dividing the video data according to a preset duration, the video can be divided into multiple time-domain groups. Each time-domain group contains video frames within a certain time range, facilitating subsequent segmented processing and analysis of the watermark information. The preset duration can be adjusted according to the actual situation. For example, the appropriate preset duration can be determined based on factors such as the embedding period of the watermark and the content characteristics of the video. If the watermark is embedded periodically, the preset duration can be set as an integer multiple of the watermark embedding period. According to the preset duration, starting from the first frame of the video data, video segments of the corresponding duration are intercepted in sequence to form multiple time-domain groups.
[0149] By equally dividing the video frames in each time-domain group, the video frames can be divided into two parts, facilitating subsequent differential statistical processing of these two parts of video frames. Multiple equal division methods are adopted. For example, the video frames in each time-domain group are equally divided according to the middle position of the number of frames, obtaining the first video frames in the first half and the second video frames in the second half.
[0150] Watermark information usually manifests as some subtle differences in video frames. By performing differential statistical processing on the first video frames and the second video frames, these differences can be quantified, thereby extracting features related to the watermark. Multiple differential statistical methods can be adopted, such as calculating the grayscale value difference and color difference of corresponding pixel points between two frames. Then, the differences of all pixel points are statistically analyzed to obtain a comprehensive difference value. The comprehensive difference value can reflect the overall difference degree between the first video frame and the second video frame.
[0151] Since watermark information manifests as a specific difference pattern in video frames, there is a certain corresponding relationship between the comprehensive difference value and the watermark bit. By analyzing the characteristics such as the magnitude and distribution of the comprehensive difference value, the corresponding watermark bit can be determined. Some thresholds or rules can be preset in advance, and according to the comparison result of the comprehensive difference value with these thresholds or rules, the watermark bit is determined. For example, if the comprehensive difference value is greater than a certain threshold, the watermark bit is determined to be 1; otherwise, the watermark bit is determined to be 0.
[0152] Finally, according to the encoding method and structure of the watermark, operations such as arranging, combining, and decoding the watermark bit positions are performed. For example, if a specific encoding algorithm is used for the watermark, the corresponding decoding algorithm needs to be used to decode the watermark bit positions to obtain the final watermark data.
[0153] In an alternative embodiment of the present invention, step 25 may include:
[0154] According to:
[0155] ,
[0156] obtain the comprehensive difference value;
[0157] where D j is the comprehensive difference value, is the mean value of the j-th frame of the first part of the video in channel θ, is the mean value of the j-th frame of the second part of the video in channel θ, Y is the luminance, and U and V are the chrominance.
[0158] In this embodiment, for a certain frame group with a fixed duration ∆T, it can be based on:
[0159] ,
[0160] obtain the comprehensive difference value; where j is the relative index of the image frame within the frame group, and D j is the difference in the total mean value of the combined YUV three channels, which is the statistical feature of the two subsets (parts A and B) on each channel, directly reflecting the distribution shift caused by the embedding perturbation, that is, reflecting the overall difference degree between the j-th frame of the first video frame and the j-th frame of the second video frame; the magnitude of this value is closely related to the differences between the two frames on each color channel.
[0161] In an alternative embodiment of the present invention, step 26 may include:
[0162] According to:
[0163] ,
[0164] obtain the watermark bit positions;
[0165] where, is the watermark bit position.
[0166] In this embodiment, during the watermark extraction process, the comprehensive difference value D j is obtained through the previous steps, and this value reflects the overall difference degree between the first video frame and the second video frame at the j-th frame. And this step aims to determine the corresponding watermark bit positions j based on the comprehensive difference value D , by setting simple judgment rules, the comprehensive difference value is converted into binary watermark bit information. For each calculated comprehensive difference value D j , make a judgment according to the set rules:
[0167] When D j ≥0, the watermark bit is determined to be 1. This means that when the comprehensive difference between the first video frame and the second video frame on the j-th frame is greater than or equal to zero, the watermark bit corresponding to this position is considered to be 1.
[0168] When D j <0, the watermark bit is determined to be 0. That is, when the comprehensive difference is less than zero, the watermark bit corresponding to this position is considered to be 0.
[0169] The above steps can convert the comprehensive difference value into specific watermark bit information. Through threshold judgment, an effective conversion from the difference value to binary bits is achieved; this method of judging the watermark bit based on the comprehensive difference value has a certain generality and can adapt to different watermark embedding methods. As long as the watermark is embedded by introducing specific difference patterns on certain channels of the video frame, the watermark information can be extracted by reasonably setting the judgment rules and thresholds using this method.
[0170] A specific embodiment of the video watermark processing method provided by the embodiment of the present invention is:
[0171] This anti-screen capture video watermark processing method includes two processes: embedding a watermark into video data and extracting watermark data from the video data;
[0172] (1) When embedding the watermark, first perform preprocessing according to the time-domain grouping strategy, then independently model each channel to calculate the visually optimal JND threshold, and then construct a periodically cumulative differential gradient field on the time domain of the carrier video according to the embedding strategy. Thus, through the synergistic effect of the time-domain segmentation strategy and the spatial domain perception constraint model, the robustness and invisibility of watermark embedding are optimized. Specifically, it includes:
[0173] Step 31, time-domain grouping;
[0174] Let the original video sequence be , where represents the k-th frame. The video is segmented according to a fixed duration , and time-domain groups can be obtained, denoted as , , and then through the video segmentation method of symmetric time-domain structure, that is, the sequence image frames of each group are equally divided into two parts before and after according to time, such as Figure 3As described above, they are denoted as part A and part B respectively. Since screen capture attacks are often accompanied by frame rate offsets, traditional fixed-frame grouping will lead to temporal breaks. Therefore, this strategy uses time base segmentation to strictly align the watermark signal period with the physical time axis, avoiding phase misalignment caused by resampling. Moreover, part A and part B of each group are complementary distributed on the time axis, laying the foundation for subsequent differential operations when embedding watermarks. It should be noted that each frame group will correspond to 1 bit of information encoded by BCH, and is one round of embedding cycle.
[0175] Step 32, color space conversion;
[0176] There are significant differences in the sensitivity of rod cells and cone cells in the human retina to brightness and chrominance. To facilitate the subsequent luminance-chrominance perception decoupling characteristics of the HVS and perform channel-independent feature analysis, the arithmetic formula for converting the video image frame from the RGB color space to the YUV color space is:
[0177]
[0178] Step 33, calculation of JND threshold in the spatial domain;
[0179] In the video watermark embedding framework, the calculation of the just-noticeable distortion threshold in the spatial domain is the core link to achieve the balance between invisibility and robustness. Its necessity stems from the non-linear characteristics of the HVS's sensitivity to distortion, and it is necessary to accurately quantify the distortion perception boundary through physiological-psychological physical modeling. The design of its JND model mainly consists of the following two main links, namely the composite masking model and the perception weight adjustment dominated by heterogeneous visual features.
[0180] To simulate the active prediction mechanism of the human brain's visual cortex for image content, a neural prediction model that combines Residual Masking (RM) and Luminance Adaptation Masking (LAM) is used. The specific implementation is as follows.
[0181] Quantification of the residual masking effect: The human brain's visual cortex has the ability to actively predict image content. In this method, an autoregressive prediction model is constructed to generate a predicted image. Also, since the YUV three channels are independently modeled and analyzed, it is beneficial to accurately quantify the masking effect of the HVS on different color space components. Therefore, according to the calculation of the absolute value of the residual between a certain channel of the input image frame and the predicted image, which is used to represent the prediction uncertainty; while is used to distinguish color channels, . The residual masking effect weights the residual through the adjustment coefficient to quantify the masking intensity; the specific arithmetic formula is:
[0182]
[0183]
[0184] Among them, is the weight coefficient of the i-th surrounding pixel, is the position z of the surrounding pixel in the k-th frame in channel θ i of the pixel value, where i and k are natural numbers, is the error compensation term, , 1.0, 1.0;
[0185] Brightness Adaptive Mask Dynamic Modeling: The human eye's perception sensitivity to noise in local areas with different brightness levels shows significant differences. It is necessary to dynamically adjust the masking threshold to match the non-linear characteristics of brightness perception. The piecewise function design conforms to the physiological response law of HVS, ensuring stricter distortion control in low-brightness areas.
[0186]
[0187] Among them, is the pixel value of the Y channel, that is, the brightness channel, of this coordinate in the YUV color space; the brightness masking threshold of the k-th frame maintains the piecewise characteristic, with the threshold decreasing in the low-brightness area and increasing linearly in the high-brightness area; is the brightness adaptive mask.
[0188] Non-linear Superposition of Double Masking Effects: The residual mask RM k,θ and the brightness adaptive mask are fused through a non-linear superposition formula, specifically:
[0189]
[0190] Among them, C θ (taking 0.3, 0.25, 0.2 for the YUV channels respectively) is used to eliminate the overlap of masking effects and avoid overestimation of the threshold; DM k,θ is the data of the composite masking effect.
[0191] Visual attention is driven by the competition of heterogeneous features such as color and edges. Regions with high saliency (such as bright objects in the center and prominent edges) require a lower distortion threshold to match the attention priority of the HVS. To comprehensively consider various influencing factors, analysis is carried out from three dimensions: color heterogeneity, spatial structure sensitivity, and brightness difference. The entire process is based on the clustering analysis of the color space distribution of the Gaussian Mixture Model (GMM), the extraction of the edge gradient field features using the Canny multi-directional gradient filter, and the quantification of the brightness region contrast based on the Kullback-Leibler (KL) divergence, and then the perceptual weight is obtained to regulate the results of the previous process.
[0192] Color perception saliency modeling: Several color categories obtained by pre-clustering with GMM are used to comprehensively represent color regions with similar perceptual characteristics in the image, denoted as the th color component. Then, through the quantification of color contrast, spatial distribution dispersion, and central preference, etc., the color perception saliency is finally characterized by self-information, and the formula is:
[0193]
[0194] where is the Gaussian weight of the color component in GMM, reflecting the statistical proportion or importance of this color in the overall image. The larger this weight, the more extensive or prominent the distribution of this color. is the contrast intensity of the color component , which combines color contrast, spatial distribution dispersion (quantified by color variance), the average distance from the color region to the center of the image, etc. The larger its value, the brighter, more concentrated, and closer to the center of the image this color component is, and the stronger the visual saliency.
[0195] Brightness-edge coupling saliency modeling: Given that flat regions (low contrast) can tolerate more distortion due to the brightness masking effect and have low saliency, while edge regions need to strictly limit distortion due to high sensitivity. This step conducts a joint analysis of brightness and edge. First, the Canny edge detection is used to simulate the processing of edge information by the human eye through retinal ganglion cells, and the edge weight is extracted. Further, through the design of an exponential function, the edge saliency is dynamically enhanced to obtain the edge saliency adjustment factor that conforms to the neural response characteristics. Finally, is calculated, and the specific formula is:
[0196]
[0197]
[0198] Among them, is the Gaussian weight of the luminance component after clustering the luminance values through GMM ; is the contrast intensity of this luminance component ; p(δ i ) is the probability of the luminance component δ i ; through the above luminance channel data, combined with the edge features for the second preprocessing process, I(δ i ) can be obtained.
[0199] The JND threshold formula for dynamic perception weight regulation is:
[0200]
[0201]
[0202] Among them, JND θ,k is the spatial minimum perceptible difference threshold of the θ channel, θ ∈ {Y, U, V}, Y is luminance, U and V are chrominance, DM k,θ is the composite masking effect data, and PW is the perception weight data. Unifying the above features affecting perception into a single weight can avoid threshold conflicts caused by independent adjustment and conform to the characteristics of multi-feature collaborative perception of HVS.
[0203] Step 34, alternating perturbation injection;
[0204] Under the framework of the time-domain grouping strategy, the video sequence is divided into several groups of frames with a fixed duration ∆T, and each group corresponds to a watermark bit after BCH error correction coding. Each group of frames is further divided into two sub-sequences A and B along the time axis. Through the multi-pixel collaborative adjustment mechanism, controlled chrominance and luminance offsets are introduced in the YUV three channels, and finally stable inter-group statistical feature differences are formed to achieve watermark coding. The specific process is as follows:
[0205] ,
[0206] Among them, F' k,θ (x, y) is the pixel value at coordinates (x, y) in the θ channel of the k-th frame, F k,θ (x, y) is the pixel value of the θ channel of the video frame F k , A is the first part of the video frame, B is the second part of the video frame, b is the BCH coding bit corresponding to the current frame group. Through the symbol alternating mechanism, the A and B sub-sequences show positive and negative perturbation trends on the YUV channels respectively, ensuring the stability of the statistical feature differences.
[0207] (2)When extracting the watermark, first convert the captured video to be detected into the YUV color space, and perform time-domain grouping according to the fixed duration ∆T set in the embedding stage. Each group is further symmetrically divided into two parts, A and B, in terms of time, ensuring that the grouping strategy is consistent with the embedding process to resist the timing phase mismatch caused by frame rate offset. Considering that the grouping starting point may shift after a screen capture attack, this method first takes the existing starting point as the starting point, and then uses BCH error correction and cyclic step size to adjust the time window to gradually align with the original embedding period. Specifically, the steps after grouping the captured video are as follows:
[0208] Step 41, calculation and fusion of multi-channel differences;
[0209] Different from the traditional single-channel (such as the U channel) dependence, this method needs to fuse the mean differences of the YUV three channels during extraction. For a certain frame group with a fixed duration ∆T, the calculation formula is:
[0210] ,
[0211] where j is the relative index of the image frame within the frame group, and are the pixel means of the A part or the B part in the θ channel of this frame group respectively; the calculation formula for the corresponding bit information of this frame group is:
[0212] ,
[0213] Step 42, calculation and fusion of multi-channel differences;
[0214] To cope with the frame rate offset or group break caused by screen capture attacks, this link restores the group timing through a dynamic time window alignment mechanism to ensure that the grouping starting point is aligned with the embedding period. In this process, the assistance of BCH error correction coding is required. As Figure 4 shown, restructure the binary watermark sequence extracted in the previous link according to the unit embedding period into a symbolized data stream, and map it to the BCH codeword space for iterative decoding. If the decoder output is an invalid identifier, trigger the adaptive re-capture mechanism: based on the preset step size ∆t, dynamically adjust the video capture start time by increasing n∙∆t at each segmentation starting point (n is the current iteration number), and re-initialize the group extraction process.
[0215] This process continues to iterate until the termination condition, that is, the remaining video duration that can be captured is less than a complete embedding period, or the payload is successfully parsed. This closed-loop feedback architecture significantly improves the robustness against the time-varying damage of the screen capture channel through double error correction in the time and space dimensions.
[0216] To verify the generalization ability and robustness of the algorithm in complex screen capture scenarios, the present invention designs a multi-dimensional experimental verification framework. Five groups of heterogeneous video samples (v1.mp4 to v5.mp4) are used in the experiment. The sample parameters are shown in Table 1. Their resolutions and frame rates are diversified, and the scene complexity includes typical types such as static text, dynamic motion, and high-texture areas, which are used to simulate the diversity characteristics of actual application scenarios.
[0217] Table 1 Basic Information of Video Samples
[0218] sample resolution frame rate duration / s v1 450×260 25.00 60 v2 684×380 25.00 60 v3 936×526 25.00 80 v4 1006×566 25.00 120 v5 1356×680 24.67 120
[0219] In terms of attack simulation, a screen capture attack chain is constructed through a multi-angle optical shooting device, and attacked videos are collected by combining different lighting conditions and shooting angles (0°, 5°, 10°) to evaluate the anti-interference ability of the algorithm under physical channel distortion. The experimental results show that through the collaborative design of spatial JND modeling and temporal periodic gradient fields, the present invention realizes visually imperceptible watermark embedding in the YUV three channels, and at the same time significantly improves the robustness against screen capture. The video frames after watermark embedding do not introduce visible distortion compared with the original frames. Although there are optical noises and geometric distortions in the captured frames after screen capture attacks, the watermark features still maintain a stable distribution.
[0220] Quantitative analysis shows that under various screen capture angle conditions before BCH correction, the correct extraction rate of watermark sampling in a single embedding period is higher than 89%. Thanks to the collaborative effectiveness of the grouping and error correction alignment framework and the information redundant embedding of multiple embedding periods, the watermarked videos in the following experiments can all be successfully extracted. The experimental results of this experiment confirm that the present invention achieves a good balance in terms of invisibility, robustness, and adaptability.
[0221] The anti-screen capture video watermark processing method proposed by the present invention dynamically quantifies the distortion threshold by constructing a spatial JND model based on HVS, and fuses multi-dimensional heterogeneous features such as color, edge, and brightness, effectively balancing the relationship between the embedding strength and the visual masking effect, and ensuring invisibility; at the same time, through the temporal grouping strategy and the modulation of the YUV color space, the statistical features of each channel are reshaped, and in cooperation with BCH coding, it can effectively resist the problems of frame rate offset and timing mismatch, and significantly improve the robustness against screen capture.
[0222] As Figure 5 shown, the embodiment of the present invention also provides a video watermark embedding processing device 50, including:
[0223] A first acquisition module 51, configured to acquire video data to be processed;
[0224] The first processing module 52 is configured to perform color space conversion processing on the video data to be processed to obtain original video data; perform segmentation processing on the original video data according to a preset duration to obtain a plurality of time-domain groups, where each time-domain group includes consecutive video frames; perform equal division processing on the consecutive video frames to obtain a first video frame and a second video frame, where the first video frame and the second video frame both include a plurality of channel data, and the plurality of channel data includes luminance channel data and chrominance channel data; perform threshold dynamic calculation processing on the plurality of channel data to obtain spatial minimum perceptible difference thresholds for the plurality of channels; and perform watermark embedding processing on the plurality of channel data according to the spatial minimum perceptible difference thresholds for the plurality of channels to obtain video data with an embedded watermark.
[0225] Optionally, the first processing module 52 is specifically configured to:
[0226] Determine residual masking effect data according to the plurality of channel data;
[0227] Perform first preprocessing on the luminance channel data to obtain luminance adaptive masking data;
[0228] Determine composite masking effect data according to the residual masking effect data and the luminance adaptive masking data;
[0229] Determine the self-information of the color components according to the chrominance channel data;
[0230] Perform second preprocessing on the luminance channel data in combination with edge features to obtain luminance-edge coupled saliency;
[0231] Determine perceptual weight data according to the self-information of the color components and the luminance-edge coupled saliency;
[0232] Determine spatial minimum perceptible difference thresholds for the plurality of channels according to the composite masking effect data and the perceptual weight data.
[0233] Optionally, the first processing module 52 is further specifically configured to:
[0234] According to:
[0235] ,
[0236] Obtain spatial minimum perceptible difference thresholds for the plurality of channels;
[0237] Where JND θ,k Is the spatial minimum perceptible difference threshold of the θ channel, θ ∈ {Y, U, V}, Y is luminance, U and V are chrominance, DM k,θ Is the composite masking effect data, and PW is the perceptual weight data.
[0238] Optionally, the first processing module 52 is further specifically configured to:
[0239] Obtain watermark data;
[0240] Encode the watermark data to obtain watermark bit data;
[0241] According to the spatial just noticeable difference thresholds of the multiple channels and the watermark bit data, perform watermark embedding processing on the multiple channel data to obtain video data with an embedded watermark.
[0242] It should be noted that this device corresponds to the above watermark embedding method. All implementation manners in the above watermark embedding method embodiment are applicable to this embodiment and can achieve the same technical effects.
[0243] As Figure 6 shown, an embodiment of the present invention further provides a video watermark extraction processing device 60, including:
[0244] A second acquisition module 61, configured to acquire video data with a watermark;
[0245] A second processing module 62, configured to perform color space conversion processing on the video data with a watermark to obtain video data after color space conversion; perform segmentation processing on the video data after color space conversion according to a preset duration to obtain a plurality of time domain groups; perform equal division processing on the video frames in the plurality of time domain groups to obtain a first video frame and a second video frame; perform difference statistics processing on the first video frame and the second video frame to obtain a comprehensive difference value; determine watermark bits according to the comprehensive difference value; and process the watermark bits to obtain watermark data.
[0246] Optionally, the second processing module 62 is specifically configured to:
[0247] According to:
[0248] ,
[0249] Obtain a comprehensive difference value;
[0250] where D j is the comprehensive difference value, is the mean value of the j-th frame of the first part of the video in channel θ, is the mean value of the j-th frame of the second part of the video in channel θ, Y is the luminance, and U and V are the chrominance.
[0251] Optionally, the second processing module 62 is further specifically configured to:
[0252] According to:
[0253] ,
[0254] Obtain the watermark bit;
[0255] Wherein, is the watermark bit.
[0256] It should be noted that this device corresponds to the above watermark extraction method. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0257] For example Figure 7 As shown, an embodiment of the present invention further provides a computing device 70, including a processor 71, a memory 72, a program or instruction stored on the memory 72 and executable on the processor 71. When the program or instruction is executed by the processor 71, it implements each process of the above video watermark embedding processing method and extraction processing method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. It should be noted that the computing device in the embodiment of the present invention includes the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0258] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0259] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.
[0260] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.
[0261] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0262] In addition, in each embodiment of the present invention, each functional unit may be integrated in a processing unit, may exist physically separately for each unit, or two or more units may be integrated in one unit.
[0263] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0264] In addition, it should be noted that in the devices and methods of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. And, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel or independently of each other. For those of ordinary skill in the art, it can be understood that all or any steps or components of the methods and devices of the present invention can be implemented in any computing device (including processors, storage media, etc.) or in a network of computing devices in the form of hardware, firmware, software, or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present invention.
[0265] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing device. The computing device can be a well-known general-purpose device. Therefore, the object of the present invention can also be achieved only by providing a program product containing program code for implementing the method or device. That is to say, such a program product also constitutes the present invention, and a storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future. It should also be noted that in the device and method of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. And, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel or independently of each other.
[0266] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A video watermark embedding method, characterized in that: include: Obtain video data to be processed; Performing color space conversion processing on the video data to be processed to obtain original video data; The original video data is segmented according to a preset time length to obtain a plurality of time domain groups, wherein each time domain group includes continuous video frames; The continuous video frames are equally divided to obtain a first video frame and a second video frame, wherein the first video frame and the second video frame each include a plurality of channel data, and the plurality of channel data include luminance channel data and chrominance channel data; Performing dynamic threshold calculation processing on the multiple channel data to obtain spatial domain just noticeable difference thresholds of the multiple channels; According to the spatial domain just noticeable difference thresholds of the multiple channels, watermark embedding processing is performed on the multiple channel data to obtain video data embedded with watermarks.
2. The video watermark embedding method according to claim 1, characterized in that: Performing dynamic threshold calculation processing on the multiple channel data to obtain spatial domain just noticeable difference thresholds of the multiple channels includes: Determining residual masking effect data according to the plurality of channel data; Performing a first preprocessing on the brightness channel data to obtain brightness adaptive masking data; Determining composite masking effect data according to the residual masking effect data and the brightness adaptive masking data; Determining self-information of a color component according to the chromaticity channel data; Performing a second preprocessing on the brightness channel data in combination with edge features to obtain brightness-edge coupling significance; Determining perceptual weight data based on the self-information of the color component and the brightness-edge coupling significance; Determine spatial domain just noticeable difference thresholds of multiple channels according to the composite masking effect data and the perception weight data.
3. The video watermark embedding method according to claim 2, characterized in that: Determining spatial domain just noticeable difference thresholds of multiple channels according to the composite masking effect data and the perception weight data, including: according to: , Obtaining the spatial least noticeable difference thresholds of multiple channels; Among them, JND θ,k is the spatial minimum perceptible difference threshold of the θ channel, θ∈{Y,U,V}, Y is brightness, U and V are chromaticity, DM k,θ is the composite masking effect data, and PW is the perceptual weight data.
4. The video watermark embedding method according to claim 3, characterized in that: According to the spatial least noticeable difference thresholds of the multiple channels, watermark embedding processing is performed on the multiple channel data to obtain video data embedded with watermarks, including: Get watermark data; Encoding the watermark data to obtain watermark bit data; According to the spatial domain least noticeable difference thresholds and watermark bit data of the multiple channels, watermark embedding processing is performed on the multiple channel data to obtain video data embedded with watermarks.
5. A video watermark extraction and processing method, characterized in that: include: Get video data with watermark; Performing color space conversion processing on the video data with the watermark to obtain video data after color space conversion; Segmenting the video data after the color space conversion according to a preset time length to obtain a plurality of time domain groups; Equally divide the video frames in the multiple time domain groups to obtain a first video frame and a second video frame; Performing statistical processing on the difference between the first video frame and the second video frame to obtain a comprehensive difference value; Determining a watermark bit position according to the comprehensive difference value; The watermark bits are processed to obtain watermark data.
6. The video watermark extraction and processing method according to claim 5 is characterized in that: Performing statistical processing on the difference between the first video frame and the second video frame to obtain a comprehensive difference value includes: according to: , Get the comprehensive difference value; Among them, D j is the comprehensive difference value, is the mean value of the jth frame in the first part of the video in channel θ, is the mean value of the jth frame in channel θ of the second part of the video, Y is the brightness, and U and V are the chromaticity.
7. The video watermark extraction and processing method according to claim 6 is characterized in that: Determining a watermark bit according to the comprehensive difference value includes: according to: , Get the watermark bit; in, The watermark bit.
8. A video watermark embedding processing device, characterized in that: include: A first acquisition module, used to acquire video data to be processed; The first processing module is used to perform color space conversion processing on the video data to be processed to obtain original video data; segment the original video data according to a preset time length to obtain multiple time domain groups, wherein each time domain group includes continuous video frames; divide the continuous video frames into equal parts to obtain a first video frame and a second video frame, wherein the first video frame and the second video frame both include multiple channel data, and the multiple channel data include brightness channel data and chrominance channel data; perform threshold dynamic calculation processing on the multiple channel data to obtain spatial minimum noticeable difference thresholds of the multiple channels; perform watermark embedding processing on the multiple channel data according to the spatial minimum noticeable difference thresholds of the multiple channels to obtain video data embedded with watermarks.
9. A video watermark extraction and processing device, characterized in that: include: The second acquisition module is used to acquire the video data with the watermark; A second processing module is used to perform color space conversion processing on the video data with the watermark to obtain video data after color space conversion; The video data after the color space conversion is segmented according to a preset time length to obtain multiple time domain groups; the video frames in the multiple time domain groups are equally divided to obtain a first video frame and a second video frame; the first video frame and the second video frame are statistically processed to obtain a comprehensive difference value; according to the comprehensive difference value, a watermark bit position is determined; the watermark bit position is processed to obtain watermark data.
10. A computing device, characterized in that include: A processor and a memory storing a computer program, wherein when the computer program is executed by the processor, the method according to any one of claims 1 to 4 or the method according to any one of claims 5 to 7 is executed.
Citation Information
Patent Citations
Self-adaptive perception time-space domain quantization method for video coding
CN112825557A
Data processing method and device, electronic equipment and storage equipment
CN113395475A
DCT domain just noticeable distortion model construction method based on entropy masking
CN115086682A
Video watermark information processing method, device and equipment
CN117376664A
Video watermarking embedding and detection apparatus and method using temporal modulation and error-correcting code
KR1020120068084A
Cited By
Watermark extraction method and device, storage medium and computer readable storage medium
CN120893024A