Method, device, equipment and medium for extracting video watermark based on quick response code

By combining quick response code and discrete cosine transform, the embedding size of video watermark is dynamically adjusted, which solves the problems of large computational complexity and insufficient security in the original domain video watermark algorithm and realizes efficient and secure watermark information extraction.

CN120302055BActive Publication Date: 2025-09-16TIANJIN XITONG ELECTRONICS EQUIP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510780761.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-16
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing raw domain video watermarking algorithms use fixed-size embedded watermarks in smooth areas of video frames, which leads to information redundancy, increased computational complexity, low extraction efficiency and insufficient security.

Method used

A video watermark extraction method based on quick response code is proposed. The watermark information is embedded and extracted by dynamically adjusting the size of the embedded watermark and combining discrete cosine transform and encryption processing.

Benefits of technology

The computational complexity is reduced, the efficiency of watermark extraction is improved, the security of the watermark is enhanced, the encryption strategy for different video contents is adapted, and the robustness and security of the watermark are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302055B_ABST
    Figure CN120302055B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device, and medium for extracting video watermarks based on a quick response code, relating to the field of image processing technology. The method comprises: obtaining a first scene change frame and a quick response code matrix of a first video, processing the above contents separately to obtain an encrypted quick response code matrix and multiple image blocks of different sizes; performing a discrete cosine transform on the luminance component of the first scene change frame, combining the encrypted quick response code matrix to obtain DC coefficients and AC coefficients; then performing an inverse discrete cosine transform on the multiple image blocks of different sizes to embed watermark information and obtain a first video with a watermark; processing the first video with a watermark to obtain a target estimated value of the encrypted quick response code matrix; decrypting and decoding the target estimated value in sequence to obtain target watermark information, thereby extracting the watermark information. The method can dynamically adjust the size of the embedded watermark, improve extraction efficiency, and enhance the security of the watermark.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, device, equipment and medium for extracting video watermarks based on a quick response code. Background Art

[0002] With the development of the Internet, the dissemination of digital media content has become increasingly convenient, and the rights of authors have also been seriously damaged. However, as an algorithm for embedding specific copyright information into digital media content, digital watermarking algorithms can very effectively protect the copyright of digital media content. Therefore, the research on digital watermarking algorithms is of great significance, so as to better protect the rights of authors. In existing technologies, video watermarking algorithms can be divided into compressed domain video watermarking algorithms and original domain video watermarking algorithms according to the different embedding locations. Compressed domain video watermarking algorithms refer to algorithms that embed and extract watermarks in compressed video streams or partially decoded compressed videos, while original domain video watermarking algorithms refer to algorithms that embed and extract watermarks directly in uncompressed video data.

[0003] Compared with the compressed domain video watermarking algorithm, the advantage of the original domain video watermarking algorithm is that it is not affected by the video format. Generally, a fixed size is used to embed the watermark. However, in the smooth area of ​​the video frame, since the pixel value changes little, using a fixed size to embed the watermark will cause many blocks to contain a large amount of similar information, resulting in information redundancy, increased computational complexity, and low watermark extraction efficiency. Summary of the Invention

[0004] The present application provides a method, apparatus, device and medium for extracting video watermarks based on quick response codes, which can dynamically adjust the size of the embedded watermark, eliminate the need to process video frames according to a fixed size, reduce the amount of calculation, improve the extraction efficiency, and add encryption processing to improve the security of the watermark.

[0005] To achieve the above objectives, this application adopts the following technical solutions:

[0006] In a first aspect, the present application provides a method for extracting a video watermark based on a quick response code, the method comprising:

[0007] Obtaining a first video and a quick response code matrix, and selecting a first scene change frame in the first video;

[0008] Processing the quick response code matrix to obtain an encrypted quick response code matrix;

[0009] performing size-block processing on a first scene change frame in the first video to obtain a plurality of size image blocks;

[0010] Performing discrete cosine transform on a luminance component of a first scene change frame in the first video to obtain a first DC coefficient and a first low-frequency AC coefficient;

[0011] Processing the first DC coefficient, the first low-frequency AC coefficient, and the encrypted quick response code matrix to obtain a second DC coefficient and a second low-frequency AC coefficient;

[0012] Performing inverse discrete cosine transform on the image blocks of multiple sizes according to the second DC coefficient and the second low-frequency AC coefficient to embed watermark information and obtain a first video with a watermark;

[0013] Processing the first video with the watermark according to the second DC coefficient and the second low-frequency AC coefficient to obtain a target estimated value of the encrypted quick response code matrix;

[0014] The target estimated value of the encrypted quick response code matrix is ​​decrypted to obtain a target quick response code matrix, and the target quick response code matrix is ​​decoded to obtain target watermark information, thereby extracting the watermark information.

[0015] In some possible implementations, the performing size-blocking processing on the first scene change frame in the first video to obtain multiple size image blocks includes:

[0016] A visual saliency value of a pixel point of a first scene change frame in the first video is calculated based on the first scene change frame in the first video, and the first scene change frame in the first video is segmented into blocks of different sizes based on the visual saliency value to obtain a plurality of image blocks of different sizes.

[0017] In some possible implementations, processing the watermarked first video according to the second DC coefficient and the second low-frequency AC coefficient to obtain a target estimated value of an encrypted quick response code matrix includes:

[0018] The second DC coefficient and the second low-frequency AC coefficient are restored to obtain a first partial estimated value and a second partial estimated value of the encrypted quick response code matrix, and the first partial estimated value and the second partial estimated value of the encrypted quick response code matrix are processed to obtain a target estimated value of the encrypted quick response code matrix.

[0019] In some possible implementations, processing the quick response code matrix to obtain an encrypted quick response code matrix includes:

[0020] A dynamic noise matrix is ​​generated for a quick response code matrix using Bernoulli distribution, and the quick response code matrix, the dynamic noise matrix and a time attenuation factor are processed to obtain an encrypted quick response code matrix.

[0021] In some possible implementations, the method further includes:

[0022] Performing encryption processing on the dynamic noise matrix according to the dynamic noise matrix and the time attenuation factor to obtain encrypted noise parameters;

[0023] During the encryption process, a random private key is generated, the security of the random private key is verified, a hash value is obtained based on the random private key, the hash value is compared with a preset secure hash value range to obtain a first judgment result, and if the first judgment result indicates that the hash value is within the preset secure hash value range, it is confirmed that the random private key is safe and usable, the encrypted noise parameters are decrypted based on the random private key, and the watermark information is extracted.

[0024] In some possible implementations, the method further includes:

[0025] If the first judgment result indicates that the hash value is not within the preset secure hash value range, a new random private key is regenerated and the judgment is repeated until the new random private key is safely available.

[0026] In a second aspect, the present application provides a device for extracting a video watermark based on a quick response code, the device comprising:

[0027] An acquisition module is configured to acquire a first video and a quick response code matrix, select a first scene change frame in the first video, and process the quick response code matrix to obtain an encrypted quick response code matrix.

[0028] an embedding module configured to perform size-blocking processing on a first scene change frame in the first video to obtain a plurality of size image blocks; perform discrete cosine transform on a luminance component of the first scene change frame in the first video to obtain a first DC coefficient and a first low-frequency AC coefficient; process the first DC coefficient, the first low-frequency AC coefficient, and the encrypted quick response code matrix to obtain a second DC coefficient and a second low-frequency AC coefficient; and perform inverse discrete cosine transform on the plurality of size image blocks based on the second DC coefficient and the second low-frequency AC coefficient to embed watermark information, thereby obtaining a first video with a watermark;

[0029] The extraction module is configured to process the first video with the watermark according to the second DC coefficient and the second low-frequency AC coefficient to obtain a target estimated value of the encrypted quick response code matrix; decrypt the target estimated value of the encrypted quick response code matrix to obtain a target quick response code matrix; and decode the target quick response code matrix to obtain target watermark information, thereby extracting the watermark information.

[0030] In a third aspect, the present application provides a computing device, including a memory and a processor;

[0031] One or more computer programs are stored in the memory, and the one or more computer programs include instructions; when the instructions are executed by the processor, the computing device executes the method as described in any one of the first aspects.

[0032] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program for executing the method as described in any one of the first aspects.

[0033] In a fifth aspect, the present application provides a computer program product, which includes one or more computer instructions. When the computer instructions are executed by a computer, the computer executes the method as described in any one of the first aspects.

[0034] It can be seen from the above technical solution that this application has at least the following beneficial effects:

[0035] In the present application, a first video and a quick response code matrix are obtained, and the first scene change frame in the first video is selected; the quick response code matrix is ​​processed to obtain an encrypted quick response code matrix; the first scene change frame in the first video is subjected to size block processing to obtain multiple size image blocks, and the size can be dynamically adjusted according to the first scene change frame of the first video, and then the block processing is performed to reduce the amount of calculation caused by the fixed size and improve the accuracy of subsequent watermark information extraction; then a discrete cosine transform is performed according to the brightness component of the first scene change frame in the first video to obtain a first DC coefficient and a first low-frequency AC coefficient; the first DC coefficient , the first low-frequency AC coefficient and the encrypted quick response code matrix are processed to obtain the second DC coefficient and the second low-frequency AC coefficient; according to the second DC coefficient and the second low-frequency AC coefficient, a plurality of image blocks of different sizes are subjected to inverse discrete cosine transform to embed the watermark information and obtain the first video with the watermark; according to the second DC coefficient and the second low-frequency AC coefficient, the first video with the watermark is processed to obtain the target estimated value of the encrypted quick response code matrix; the target estimated value of the encrypted quick response code matrix is ​​decrypted to obtain the target quick response code matrix, the target quick response code matrix is ​​decoded to obtain the target watermark information, and the watermark information is extracted.

[0036] In traditional schemes, video watermark algorithms can be divided into compressed domain video watermark algorithms and original domain video watermark algorithms. Among them, the compressed domain video watermark algorithm has a faster extraction efficiency, but due to the need to compress the video, it is affected by the format and can only use a fixed format to process the watermark in the video, while the original domain video watermark algorithm is not affected by the format. The original domain video watermark algorithm generally uses a fixed size to embed the watermark, but in the smooth area of ​​the video frame, due to the small change in pixel value, using a fixed size to embed the watermark will cause many blocks to contain a large amount of similar information, resulting in information redundancy, increasing the amount of calculation, and resulting in very low watermark extraction efficiency. It can be seen that the present application provides a method for extracting video watermarks based on a quick response code, which can dynamically adjust the size of the embedded watermark, does not need to process all video frames according to a fixed size, reduces the amount of calculation, improves the extraction efficiency, and also adds encryption processing to improve the security of the watermark.

[0037] It should be understood that the description of technical features, technical solutions, beneficial effects or similar language in this application does not imply that all features and advantages can be realized in any single embodiment. On the contrary, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution or beneficial effect is included in at least one embodiment. Therefore, the description of a technical feature, technical solution or beneficial effect in this specification does not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions and beneficial effects described in the present embodiment can also be combined in any appropriate manner. Those skilled in the art will understand that the embodiment can be implemented without one or more specific technical features, technical solutions or beneficial effects of a specific embodiment. In other embodiments, additional technical features and beneficial effects can also be identified in specific embodiments that do not embody all embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 A flowchart of a method for extracting video watermarks based on a quick response code provided in an embodiment of the present application;

[0039] Figure 2 A schematic diagram of a video watermark extraction device based on a quick response code provided in an embodiment of the present application;

[0040] Figure 3 A schematic diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] The terms "first", "second" and "third" in this application specification and the accompanying drawings are used to distinguish different objects rather than to limit a specific order.

[0042] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0043] To make the description of the following embodiments clear and concise, a brief introduction to the related technologies is first given:

[0044] The Quick Response Code (QR code) is a black and white pattern composed of dark and light modules distributed in a regular pattern in a two-dimensional plane with specific geometric shapes. These modules are arranged and combined to record data symbol information and can store various types of data. The Quick Response Code can be divided into 40 versions based on size. Version 1 has a size of 21 21, version 40 has a size of 177 177. Each higher version has 4 additional modules on each side. According to the error correction capability, QRCode is divided into L level, M level, Q level and H level. The approximate correction amount of each level is shown in Table 1:

[0045] Table 1:

[0046]

[0047] The Discrete Cosine Transform (DCT) is a mathematical transformation method widely used in digital signal processing and image processing. It is a tool for converting a set of discrete data points from the spatial domain to the frequency domain. In the imaging field, an image can be viewed as spatially distributed information consisting of a large number of pixels, each with a specific brightness or color value. The DCT converts the spatial information represented by these pixel values ​​into a combination of different frequency components. It can also decompose an image into a superposition of cosine waves of different frequencies. The low-frequency component represents basic information such as the image's main outline, general shape, and brightness, similar to the basic framework and primary color tone of a painting. The high-frequency component corresponds to image details, texture, and edge information, such as the veins of leaves or the fine edges of objects.

[0048] At present, video watermarking algorithms based on discrete cosine transform can be divided into compressed domain video watermarking algorithms and original domain video watermarking algorithms. Among them, the compressed domain video watermarking algorithm has a faster extraction efficiency, but due to the need to compress the video, it is affected by the format and can only use a fixed format to process the watermark in the video, while the original domain video watermarking algorithm is not affected by the format. The original domain video watermarking algorithm generally uses a fixed size to embed the watermark, but in the smooth area of ​​the video frame, due to the small change in pixel value, using a fixed size to embed the watermark will make many blocks contain a large amount of similar information, resulting in information redundancy, increased computational complexity, and low watermark extraction efficiency.

[0049] In view of this, an embodiment of the present application provides a method for extracting video watermarks based on a quick response code, in which a first video and a quick response code matrix are obtained, and a first scene change frame in the first video is selected; the quick response code matrix is ​​processed to obtain an encrypted quick response code matrix; the first scene change frame in the first video is subjected to size block processing to obtain multiple size image blocks, and the size can be dynamically adjusted according to the first scene change frame of the first video, and then the block processing is performed to reduce the amount of calculation caused by the fixed size and improve the accuracy of subsequent watermark information extraction; then a discrete cosine transform is performed according to the brightness component of the first scene change frame in the first video to obtain a first DC coefficient and a first low-frequency AC coefficient; processing the first DC coefficient, the first low-frequency AC coefficient and the encrypted quick response code matrix to obtain a second DC coefficient and a second low-frequency AC coefficient; performing an inverse discrete cosine transform on multiple sized image blocks according to the second DC coefficient and the second low-frequency AC coefficient to embed watermark information and obtain a first video with a watermark; processing the first video with a watermark according to the second DC coefficient and the second low-frequency AC coefficient to obtain a target estimated value of the encrypted quick response code matrix; decrypting the target estimated value of the encrypted quick response code matrix to obtain a target quick response code matrix; decoding the target quick response code matrix to obtain target watermark information and extracting the watermark information.

[0050] It can be seen that this application provides a method for extracting video watermarks based on quick response codes, which can dynamically adjust the size of the embedded watermark. There is no need to process the video frames according to a fixed size, which reduces the amount of calculation, improves the extraction efficiency, and adds encryption processing to improve the security of the watermark.

[0051] In order to make the technical solution of this application clearer and easier to understand, the following describes a method for extracting video watermarks based on a quick response code provided by an embodiment of this application in conjunction with the accompanying drawings. Figure 1 As shown in the figure, this figure is a flow chart of a method for extracting video watermarks based on a quick response code provided by an embodiment of the present application. The method for extracting video watermarks based on a quick response code includes:

[0052] S101: Obtain a first video and a quick response code matrix, and select a first scene change frame in the first video.

[0053] This application selects a sample video in the sample library and sets it as the first video for embedding and extracting watermarks. Before embedding the watermark in the first video, it is necessary to select the first scene change frame of the first video. Since the video is composed of a series of continuous frames, scene change detection is crucial for accurately embedding the watermark. Here, frames at different moments in the first video are selected, and the correlation coefficient between the histogram of the Y component of the frame at the mth moment and the histogram of the Y component of the frame at the m-1th moment is calculated to determine whether the scene has changed. In the YUV color space, the Y component represents brightness information, which has a greater impact on human visual perception. Y component: represents brightness, which describes the brightness of the image. The Y component values ​​of different areas in the image reflect whether the area is bright or dark; U component and V component: represent chromaticity, which are used to describe the color information of the image. The U component and the V component jointly determine the color attributes of the image, such as hue and saturation.

[0054] Using the YUV color space offers numerous advantages. First, the human eye is much more sensitive to brightness than to color. Separating brightness allows for more refined processing of brightness information during image processing and video encoding, while appropriately simplifying the color information. This helps improve encoding efficiency and saves storage space and transmission bandwidth. Second, in many video analysis tasks, such as scene change detection and object recognition, analyzing the Y component first can quickly capture the image's basic brightness and shading characteristics, laying the foundation for subsequent, more complex analysis.

[0055] The correlation coefficient between the Y component histogram of the frame at time m and the Y component histogram of the frame at time m-1 is calculated by combining the covariance of the Y component histogram of the frame at time m with the Y component histogram of the frame at time m-1 and the variance of the Y component histogram of the frame at time m with the Y component histogram of the frame at time m-1. The covariance of the Y component histogram of the frame at time m and the Y component histogram of the frame at time m-1 measures the similarity between the Y component histograms of the two adjacent frames. The variance of the Y component histogram of the frame at time m and the Y component histogram of the frame at time m-1 and the correlation coefficient between the Y component histogram of the frame at time m and the Y component histogram of the frame at time m-1 respectively represent the degree of dispersion of the Y component histograms of the two frames. The value calculated in this way can reflect the difference between adjacent frames.

[0056] In order to make the content described in this application clearer, the following formula for calculating the correlation coefficient between the histogram of the Y component of the frame at the mth moment and the histogram of the Y component of the frame at the m-1th moment is provided, as shown in the following formula:

[0057]

[0058] in, is the histogram of the Y component of the frame at the m-1th moment, is the histogram of the Y component of the frame at the mth moment, yes and The covariance of yes The variance of yes The variance of is the correlation coefficient between the histogram of the Y component of the frame at the mth moment and the histogram of the Y component of the frame at the m-1th moment.

[0059] Set the threshold T and judge Is it less than the threshold T? If it is less than the threshold T, The first scene change frame of the first video is identified, and T is a constant. In order to ensure that more frames that may contain watermarks are found during the extraction process, the T value in the extraction process is usually set to be higher than the T value in the embedding process.

[0060] In the embodiment of the present application, the threshold value T is set to 0.9. <0.9, it is determined as a scene change frame. After a large number of experiments, it is verified that when When the value is less than 0.9, it indicates that the content difference between adjacent frames is large, and it is likely that a scene change has occurred. For example, in a video, the previous frame is an indoor scene, and the next frame switches to an outdoor scene. It will be significantly smaller than 0.9, and thus be judged as a scene change frame.

[0061] S102: Process the quick response code matrix to obtain an encrypted quick response code matrix.

[0062] The size of the quick response code matrix is ​​r×r. In this application, r=21 is set. This size can ensure that it contains sufficient watermark information while also being well adapted to subsequent processing operations. Here, the Bernoulli distribution is used to generate a dynamic noise matrix for the quick response code matrix. The quick response code matrix, the dynamic noise matrix, and the time decay factor are processed to obtain the encrypted quick response code matrix. The calculation formula is as follows:

[0063]

[0064] in, is the dynamic noise matrix, is the element at row i and column j in the dynamic noise matrix, i is the row index of the element in the dynamic noise matrix, j is the column index of the element in the dynamic noise matrix, r represents the number of rows and columns of the dynamic noise matrix, , For a Bernoulli distribution with probability p=0.3, each element in the dynamic noise matrix By probability p =0.3 takes the value as 1, with probability The value is 0. In this way, noise elements are randomly generated within the range of the r×r quick response code matrix to prepare for the subsequent noise superposition; is the element at row i and column j in the encrypted quick response code matrix, is the time decay factor, t is the timestamp of the frame at the mth moment, is the element at row i and column j in the quick response code matrix, Represents the exclusive OR operation.

[0065] Then, according to the dynamic noise matrix and the time attenuation factor, the dynamic noise matrix is ​​encrypted to obtain the encrypted noise parameters:

[0066]

[0067] Among them, k is a random private key, G is the elliptic curve base point, is the dynamic noise matrix, is the time decay factor, is the encrypted noise parameter, The public key is then calculated using the base point of the elliptic curve.

[0068] During the encryption process, a random private key is generated, and the security of the random private key is verified. A hash value is obtained based on the random private key, and the hash value is judged with the preset secure hash value range to obtain a first judgment result. If the first judgment result indicates that the hash value is within the preset secure hash value range, the random private key is confirmed to be safe and available, and the encrypted noise parameters are decrypted according to the random private key to extract the watermark information; if the first judgment result indicates that the hash value is not within the preset secure hash value range, a new random private key is regenerated and rejudged until the new random private key is safe and available. Only those who have the corresponding private key can obtain the watermark information. k Only the legitimate receiver can decrypt the encrypted noise parameters and correctly extract the watermark, which greatly improves the security of the watermark.

[0069] If the watermark embedded in a video is not secure enough, the watermark information can be illegally obtained. For example, in digital media copyright protection, if the watermark is not secure, pirates can easily extract the watermark information, remove or tamper with the watermark, and then distribute the pirated content. This makes it impossible for copyright owners to prove copyright ownership through watermarks, leading to rampant copyright infringement.

[0070] In some scenarios where information authenticity needs to be verified, such as electronic document signatures or watermarks for authentication, if the watermark security is compromised, attackers can forge fake documents or identity information bearing legitimate watermarks. For example, in electronic contracts, attackers could tamper with the contract content and then use the obtained watermark information to make it appear legitimate, posing serious legal and economic risks to commercial transactions and other activities. Watermarks are also often used to ensure data integrity. If the watermark is insecure, attackers can modify the data content without detection because they can control the extraction and modification of the watermark. For example, in medical data storage, if watermarked medical images are illegally modified and the image content is tampered with, serious consequences such as misdiagnosis can occur.

[0071] Therefore, the process of generating random private keys and verifying their security increases the difficulty of obtaining valid private keys. Because not every randomly generated private key can be used, only private keys that, after hash value verification, fall within a pre-set secure hash value range are considered safe for use. This judgment mechanism makes it difficult for attackers to easily guess or generate valid private keys. For example, consider hash values ​​generated using a complex encryption algorithm, such as SHA-256. Its output hash values ​​are highly discrete and unpredictable. Generating a private key within the pre-set secure hash value range would require an attacker to make numerous attempts, which is computationally expensive.

[0072] Only the legitimate recipient with the corresponding private key k can decrypt the encrypted noise parameters. This means that the watermark information is effectively protected, just like a safe with a special keyhole that can only be opened with a key of a specific shape. For illegal users who do not have the private key, even if they obtain the encrypted noise parameters, they cannot correctly decrypt and extract the watermark.

[0073] The process of regenerating a new private key and re-verifying it when the generated private key hash value is not within the safe range further increases the time and computational cost for attackers to obtain valid private keys. If an attacker wants to obtain a valid private key through brute force, they must restart the verification process every time they generate a private key that does not meet the requirements, which greatly reduces their chances of success.

[0074] Automatically adjust encryption strategies based on the video's content type. For example, for action films, due to the rapid visual changes and large amounts of information, the noise intensity can be appropriately increased or the encryption key update frequency can be adjusted to enhance watermark security. For animated films, a finer encryption granularity can be adopted, with encryption parameters set for different animated characters or scenes. Optimize encryption based on video quality parameters such as resolution and frame rate. For example, for high-resolution videos, increase the accuracy of encryption calculations to prevent the increased risk of watermark information leakage due to high resolution. For low-frame-rate videos, adjust the number of encryption algorithm iterations to balance computing resource consumption with watermark security.

[0075] S103: performing size-blocking processing on the first scene change frame in the first video to obtain a plurality of size image blocks.

[0076] The visual saliency value of the pixel of the first scene change frame in the first video is calculated, and the first scene change frame in the first video is divided into blocks of different sizes according to the visual saliency value to obtain a plurality of image blocks of different sizes. The calculation formula is as follows:

[0077]

[0078]

[0079] in, is the visual saliency value of the pixel point in the first scene change frame. The visual saliency value reflects the degree of prominence of the pixel point in the entire image of the first scene change frame to the human eye. The larger the value, the more attention the pixel point can attract. c is the color channel, including the red channel R, the green channel G, and the blue channel B. The coordinates are The pixel value of the pixel on the color channel c, x is the horizontal coordinate position index of the pixel in the first scene change frame, y is the vertical coordinate position index of the pixel in the first scene change frame, is the mean value of the pixel on color channel c, is the Gaussian kernel, is a Gaussian kernel function, and the dynamic size of the image block is determined by n. is the horizontal and vertical length of the dynamic size of the image block, is the maximum visual saliency value of the frame, is the average value of the entire frame.

[0080] To achieve adaptive segmentation, we first need to calculate the visual saliency value of the first scene change frame. For each pixel in the first scene change frame, we perform the calculation on the red (R), green (G), and blue (B) color channels. It represents the difference between the pixel point in a certain color channel and the mean value of the channel. The larger the difference, the more significant the pixel point. Then, through the Gaussian kernel These differences are weighted, The smoothing effect and edge preservation can be better balanced. Finally, the results of the three color channels are added together to obtain the visual significance value of the pixel. The calculated visual saliency value This is used to determine the dynamic size of the image blocks. Here, we use a value of 8 as the base, and adjust the block size based on the ratio of the frame's maximum saliency value to the full-frame average. For example, if a frame contains a highly salient object, such as a bright headlight, resulting in a large maximum saliency value and a relatively small full-frame average, the calculated value of n will be increased accordingly, meaning a larger block size will be used for the salient area. Conversely, in non-salient areas, the block size will be relatively smaller. This achieves the purpose of dynamic block segmentation, reduces computational effort, and improves the efficiency of watermark embedding and extraction.

[0081] S104 . Perform discrete cosine transform on the luminance component of the first scene change frame in the first video to obtain a first DC coefficient and a first low-frequency AC coefficient.

[0082] For an image of a given size, the discrete cosine transform (DCT) performs calculations on every pixel. When both corresponding parameters are zero, the DCT process combines the values ​​of all pixels in the image. Simply put, the value of every pixel in the image is taken into account, a specific summation is performed, and then some processing is performed based on the image size. The result is the DC coefficient. The DC coefficient reflects the average brightness of the image as a whole, providing a general numerical representation of the overall brightness of the image.

[0083] The calculation process is also based on the discrete cosine transform. When setting specific parameter values, here they are all set to 0.1. In actual applications, this can be changed as needed to calculate the value of each pixel in the image. This calculation process involves using a specific cosine function to weight the value of each pixel, and then summing all the weighted pixel values. The final result is the AC coefficient. The AC coefficient reflects not the overall brightness of the image, but rather the texture, details, and other aspects of the image. It can help us understand the characteristics of the image from another perspective. The following will explain in detail how the first DC coefficient and the first low-frequency AC coefficient are calculated through the discrete cosine transform through specific formulas.

[0084] The formula for calculating the first DC coefficient and the first low-frequency AC coefficient by discrete cosine transform is as follows:

[0085]

[0086]

[0087] when and When F(0, 0) is called the DC coefficient, that is, the first DC coefficient, the formula is as follows:

[0088]

[0089] In this application, the and When F(0.1, 0.1) is called the AC coefficient, that is, the first low-frequency AC coefficient, the formula is as follows:

[0090]

[0091] Among them, M represents the number of pixels in the horizontal direction of the image, and N represents the number of pixels in the vertical direction of the image. u represents the frequency component in the horizontal direction, v Represents the frequency component in the vertical direction, different ( u , v ) value determines the position of the coefficient in the frequency domain. 、 is the frequency coordinate u 、 v The relevant weighting coefficient, Represents the image in the spatial domain coordinates ( x , y ) is the pixel value at .

[0092] S105 , processing the first DC coefficient, the first low-frequency AC coefficient, and the encrypted quick response code matrix to obtain a second DC coefficient and a second low-frequency AC coefficient.

[0093] The calculation formula is as follows:

[0094]

[0095]

[0096] in, is the second DC coefficient, is the second lowest frequency AC coefficient, To encrypt the quick response code matrix, is the first DC coefficient, that is , is the first low-frequency AC coefficient, that is Since there may be multiple low-frequency AC coefficients in a size image block, the kth low-frequency AC coefficient is selected as the first low-frequency AC coefficient, where k is an integer. is the first correlation coefficient related to the significance value of the image block of size, is the second correlation coefficient related to the significance value of the image block of size .

[0097] The DC coefficient represents the average brightness information of the image block. Here, the elements of the encrypted quick response code matrix are combined with the DC coefficient. is a coefficient related to the visual saliency value of the image patch of size, . The first i Row, No. j The saliency metric of the column block is used to quantify the visual appeal of the region. This means that the more visually significant the block, the greater the modification of the DC coefficient, thus more effectively embedding the watermark in the visually sensitive area. In addition to the DC coefficient, the low-frequency AC coefficient is also modulated. The low-frequency AC coefficient contains the main texture information of the image. Similarly, Combined with the low frequency AC coefficient, is a coefficient related to the visual saliency value, In this way, the watermark information can be embedded into the video frame more covertly while ensuring the image quality.

[0098] The watermark is concentrated in visually sensitive areas (PSNR > 48dB), which are often feature-rich and information-rich parts of the image. Embedding the watermark in this area leverages this rich information, allowing it to better blend with the image and enhance its robustness. Compared to embedding watermarks in less sensitive areas with limited information and features, watermarks embedded in visually sensitive areas are more difficult to destroy or remove by various image processing operations. This improves the stability and detectability of the watermark during subsequent use and dissemination of the image, ensuring that the watermark can continue to perform its functions for copyright protection and content authentication. Furthermore, while preserving the watermark information, the impact on image quality is minimized, achieving a PSNR value greater than 48dB, making the watermark virtually imperceptible to the human eye.

[0099] S106 , performing inverse discrete cosine transform on the image blocks of multiple sizes according to the second DC coefficient and the second low-frequency AC coefficient to embed watermark information and obtain a first video with a watermark.

[0100] The calculation formula for the inverse discrete cosine transform is as follows:

[0101]

[0102] The discrete cosine transform (DCT) converts an image from the spatial domain to the frequency domain, while the inverse discrete cosine transform (IDCT) converts the image from the frequency domain back to the spatial domain. By performing the IDCT on the modified blocks, the coefficients containing the watermark information in the frequency domain are converted back to the spatial domain, resulting in the Y component containing the watermark. This step converts the watermark information from frequency domain coefficients to the image's luminance component, allowing it to be embedded in the luminance portion of the first scene change frame, thus achieving watermark embedding.

[0103] Since the width or height of the Y component has been adjusted before, after selecting the Y component of the scene change frame, it is determined whether its width or height can be r If it is not divisible, then the Y component has been enlarged or reduced. Here, the size of the Y component needs to be adjusted back to the original size. This step is to ensure the overall size and structural integrity of the first scene change frame, so that the first scene change frame after embedding the watermark is visually consistent with the original video frame size without deformation. For example, if the Y component width or height is enlarged to a size that can be r Now we need to scale it back to the original width and height values ​​according to a certain ratio.

[0104] S107 , processing the first video with the watermark according to the second DC coefficient and the second low-frequency AC coefficient to obtain a target estimated value of the encrypted quick response code matrix.

[0105] The calculation formula for the target estimated value of the encrypted quick response code matrix is ​​as follows:

[0106]

[0107]

[0108]

[0109] in, is the first part of the estimated value of the encrypted quick response code matrix, is the second part of the estimated value of the encrypted quick response code matrix, is the target estimate of the encrypted fast response code matrix, is the first correlation coefficient related to the significance value of the image block of size, is the second correlation coefficient related to the significance value of the image block of size, is the second DC coefficient, is the second lowest frequency AC coefficient, is the first DC coefficient, that is , is the first low-frequency AC coefficient, that is .

[0110] S108. Decrypt the target estimated value of the encrypted quick response code matrix to obtain a target quick response code matrix.

[0111] Use the corresponding decryption algorithm, such as the elliptic curve cryptography (ECC) decryption algorithm, to Decryption is performed. This step removes the encryption performed on the quick response code matrix during the watermark embedding process, restoring its original encoded form. Furthermore, if interference factors such as noise were introduced during the watermark embedding process, noise cancellation operations such as these will need to be performed in conjunction with relevant noise parameters. For example, the watermarking scheme may include quantum noise injection. In this case, the effects of this noise on the quick response code matrix need to be reversed based on the principles of noise generation and superposition. After decryption and noise cancellation operations, the target quick response code matrix is ​​obtained.

[0112] S109: Decode the target quick response code matrix to obtain target watermark information, thereby extracting the watermark information.

[0113] Since the method of decoding the target quick response code matrix is ​​an existing technology, it will not be described in detail here. By decoding, the target watermark information is obtained, and the watermark information is extracted.

[0114] In order to further reduce the computational complexity of embedding and extracting watermark information, a watermark extraction optimization model is established. First, the confidence of the target watermark information is calculated. Reliability evaluation is performed based on the confidence of the target watermark information. If the watermark extraction action is terminated and the confidence of the target watermark information is higher than the first confidence threshold, the target watermark information is considered reliable, and a reward value is given to the watermark extraction optimization model. If the confidence of the target watermark information is higher than the first confidence threshold but the watermark extraction action is continued, the extraction action is considered a waste of resources, and a penalty value is given to the watermark extraction optimization model. The confidence value is output. ∈[0,1].

[0115] When extracting the watermark, each frame in the video stream is first decoded by the quick response code matrix. Due to the influence of noise and various factors in the video processing process, the reliability of the decoding result needs to be evaluated by confidence. The value range is between 0 and 1, where 0 means the decoding result is completely unreliable and 1 means the decoding result is very reliable. For example, when the noise interference of a frame's quick response code matrix is ​​small during decoding, and the decoded information matches the expected quick response code matrix format and error correction information well, the confidence level of the frame is 1. Will be close to 1; on the contrary, if there is severe noise interference or decoding errors, Close to 0. State in reinforcement learning The confidence of the current frame and the historical mean of the confidence of the previous frame This status information reflects the real-time situation and overall trend of the confidence level during decoding. For example, if the confidence level of the current frame is High and historical average It is also relatively high, indicating that the current decoding process is relatively stable and reliable.

[0116] Intelligent extraction termination decision by action To achieve, action There are only two values: 0 means continuing to decode the next frame, and 1 means terminating the decoding process and assuming that reliable watermark information has been successfully extracted.

[0117] The calculation formula of the reward and penalty functions is as follows:

[0118]

[0119] Reward Function Used to guide reinforcement learning algorithms to make optimal decisions. = 1 and the decoding is successful, a higher reward value of 10 is given, which encourages the algorithm to terminate in time when it can successfully decode, saving computing resources. =0 and the current frame confidence When the confidence level is >0.99, a penalty of -1 is applied, as continuing to decode when the confidence level is already very high is a waste of resources. In other cases, the reward is 0. By continuously adjusting the action, the algorithm maximizes the cumulative reward in the long term, thus achieving intelligent extraction termination.

[0120] Based on the above content, a first video and a quick response code matrix are obtained, and a first scene change frame in the first video is selected; the quick response code matrix is ​​processed to obtain an encrypted quick response code matrix; the first scene change frame in the first video is subjected to size block processing to obtain multiple size image blocks, and the size can be dynamically adjusted according to the first scene change frame of the first video, and then the block processing is performed to reduce the amount of calculation caused by the fixed size and improve the accuracy of subsequent watermark information extraction; then a discrete cosine transform is performed according to the brightness component of the first scene change frame in the first video to obtain a first DC coefficient and a first low-frequency AC coefficient; the first DC coefficient , the first low-frequency AC coefficient and the encrypted quick response code matrix are processed to obtain the second DC coefficient and the second low-frequency AC coefficient; based on the second DC coefficient and the second low-frequency AC coefficient, a plurality of image blocks of different sizes are subjected to inverse discrete cosine transform to embed the watermark information and obtain the first video with the watermark; based on the second DC coefficient and the second low-frequency AC coefficient, the first video with the watermark is processed to obtain the target estimated value of the encrypted quick response code matrix; the target estimated value of the encrypted quick response code matrix is ​​decrypted to obtain the target quick response code matrix, the target quick response code matrix is ​​decoded to obtain the target watermark information, and the watermark information is extracted. It can be seen that the present application provides a method for extracting video watermarks based on a quick response code, which can dynamically adjust the size of the embedded watermark, does not need to process all video frames according to a fixed size, reduces the amount of calculation, improves the extraction efficiency, and also adds encryption processing to improve the security of the watermark.

[0121] The embodiment of the present application also provides a device for extracting video watermark based on a quick response code, such as Figure 2 As shown, this figure is a schematic diagram of a video watermark extraction device based on a quick response code provided by an embodiment of the present application, the device includes: an acquisition module 201, an embedding module 202 and an extraction module 203;

[0122] The acquisition module 201 is configured to acquire a first video and a quick response code matrix, select a first scene change frame in the first video, and process the quick response code matrix to obtain an encrypted quick response code matrix.

[0123] The embedding module 202 is used to perform size-blocking processing on the first scene change frame in the first video to obtain multiple size image blocks; perform discrete cosine transform based on the brightness component of the first scene change frame in the first video to obtain a first DC coefficient and a first low-frequency AC coefficient; process the first DC coefficient, the first low-frequency AC coefficient and the encrypted quick response code matrix to obtain a second DC coefficient and a second low-frequency AC coefficient; perform inverse discrete cosine transform on the multiple size image blocks based on the second DC coefficient and the second low-frequency AC coefficient to embed watermark information and obtain a first video with a watermark.

[0124] The extraction module 203 is configured to process the first video with the watermark according to the second DC coefficient and the second low-frequency AC coefficient to obtain a target estimated value of the encrypted quick response code matrix; decrypt the target estimated value of the encrypted quick response code matrix to obtain a target quick response code matrix; and decode the target quick response code matrix to obtain target watermark information, thereby extracting the watermark information.

[0125] In some possible implementations, the embedding module 202 is specifically used to calculate the visual saliency value of the pixel points of the first scene change frame in the first video based on the first scene change frame in the first video, and perform size block processing on the first scene change frame in the first video according to the visual saliency value to obtain multiple size image blocks.

[0126] In some possible implementations, the extraction module 203 is specifically configured to perform restoration processing on the second DC coefficient and the second low-frequency AC coefficient to obtain a first partial estimated value and a second partial estimated value of the encrypted quick response code matrix, and perform processing based on the first partial estimated value and the second partial estimated value of the encrypted quick response code matrix to obtain a target estimated value of the encrypted quick response code matrix.

[0127] In some possible implementations, the extraction module 203 is specifically configured to generate a dynamic noise matrix for the quick response code matrix using Bernoulli distribution, and process the quick response code matrix, the dynamic noise matrix, and a time attenuation factor to obtain an encrypted quick response code matrix.

[0128] In some possible implementations, the apparatus further includes:

[0129] The judgment module is used to encrypt the dynamic noise matrix according to the dynamic noise matrix and the time attenuation factor to obtain encrypted noise parameters.

[0130] During the encryption process, a random private key is generated, the security of the random private key is verified, a hash value is obtained based on the random private key, the hash value is compared with a preset secure hash value range to obtain a first judgment result, and if the first judgment result indicates that the hash value is within the preset secure hash value range, it is confirmed that the random private key is safe and usable, the encrypted noise parameters are decrypted based on the random private key, and the watermark information is extracted.

[0131] In some possible implementations, the judgment module is further configured to regenerate a new random private key and re-judge if the first judgment result indicates that the hash value is not within a preset secure hash value range until the new random private key is safely available.

[0132] The present application also provides a computing device. Figure 3 As shown, this figure is a schematic diagram of a computing device provided by an embodiment of the present application, wherein the computing device 400 includes a bus 401, a processor 402, a communication interface 403, and a memory 404. The processor 402, the memory 404, and the communication interface 403 communicate with each other via the bus 401.

[0133] The bus 401 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0134] The processor 402 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0135] The communication interface 403 is used for external communication. For example, when the computing device is a first switch, the communication interface 403 can be used for communication between the first switch and the first user terminal, or between the first switch and the second switch.

[0136] Memory 404 may include volatile memory, such as random access memory (RAM). Memory 404 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0137] The memory 404 stores executable codes, and the processor 402 executes the executable codes to perform the aforementioned method for extracting video watermarks based on quick response codes.

[0138] Embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk, or magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the aforementioned method for extracting video watermarks based on quick response codes.

[0139] The present application also provides a computer program product comprising one or more computer instructions that, when loaded and executed on a computing device, fully or partially generate the process or function described in the present application.

[0140] The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer or data center to another website, computer or data center via wired (e.g., coaxial cable, optical fiber) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0141] When the computer program product is executed by a computer, the computer executes any of the aforementioned methods for extracting video watermarks based on quick response codes. The computer program product may be a software installation package. When any of the aforementioned methods for extracting video watermarks based on quick response codes is needed, the computer program product may be downloaded and executed on the computer.

[0142] The descriptions of the processes or structures corresponding to the above figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.

[0143] The above description is only a specific implementation method of the present application, but the protection scope of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in the present application should be included in the protection scope of the present application.

Claims

1. A method for extracting video watermark based on quick response code, characterized in that: The method comprises: Obtaining a first video and a quick response code matrix, and selecting a first scene change frame in the first video; Processing the quick response code matrix to obtain an encrypted quick response code matrix; performing size-block processing on a first scene change frame in the first video to obtain a plurality of size image blocks; Performing discrete cosine transform on a luminance component of a first scene change frame in the first video to obtain a first DC coefficient and a first low-frequency AC coefficient; Processing the first DC coefficient, the first low-frequency AC coefficient, and the encrypted quick response code matrix to obtain a second DC coefficient and a second low-frequency AC coefficient; Performing inverse discrete cosine transform on the image blocks of multiple sizes according to the second DC coefficient and the second low-frequency AC coefficient to embed watermark information and obtain a first video with a watermark; Processing the first video with the watermark according to the second DC coefficient and the second low-frequency AC coefficient to obtain a target estimated value of the encrypted quick response code matrix; decrypting the target estimated value of the encrypted quick response code matrix to obtain a target quick response code matrix, decoding the target quick response code matrix to obtain target watermark information, and extracting the watermark information; The step of performing size-block processing on the first scene change frame in the first video to obtain a plurality of size image blocks includes: A visual saliency value of a pixel point of a first scene change frame in the first video is calculated based on the first scene change frame in the first video, and the first scene change frame in the first video is segmented into blocks of different sizes based on the visual saliency value to obtain a plurality of image blocks of different sizes.

2. The method according to claim 1, characterized in that The step of processing the first video with the watermark according to the second DC coefficient and the second low-frequency AC coefficient to obtain a target estimated value of the encrypted quick response code matrix includes: The second DC coefficient and the second low-frequency AC coefficient are restored to obtain a first partial estimated value and a second partial estimated value of the encrypted quick response code matrix, and the first partial estimated value and the second partial estimated value of the encrypted quick response code matrix are processed to obtain a target estimated value of the encrypted quick response code matrix.

3. The method according to claim 1, characterized in that The processing of the quick response code matrix to obtain an encrypted quick response code matrix includes: A dynamic noise matrix is ​​generated for a quick response code matrix using Bernoulli distribution, and the quick response code matrix, the dynamic noise matrix and a time attenuation factor are processed to obtain an encrypted quick response code matrix.

4. The method according to claim 3, characterized in that The method further comprises: Performing encryption processing on the dynamic noise matrix according to the dynamic noise matrix and the time attenuation factor to obtain encrypted noise parameters; During the encryption process, a random private key is generated, the security of the random private key is verified, a hash value is obtained based on the random private key, the hash value is compared with a preset secure hash value range to obtain a first judgment result, and if the first judgment result indicates that the hash value is within the preset secure hash value range, it is confirmed that the random private key is safe and usable, the encrypted noise parameters are decrypted based on the random private key, and the watermark information is extracted.

5. The method according to claim 4, characterized in that The method further comprises: If the first judgment result indicates that the hash value is not within the preset secure hash value range, a new random private key is regenerated and the judgment is repeated until the new random private key is safely available.

6. A video watermark extraction device based on a quick response code, characterized in that: The device comprises: An acquisition module is configured to acquire a first video and a quick response code matrix, select a first scene change frame in the first video, and process the quick response code matrix to obtain an encrypted quick response code matrix. An embedding module is configured to perform size-blocking processing on a first scene change frame in the first video to obtain a plurality of size image blocks; perform discrete cosine transform on a luminance component of the first scene change frame in the first video to obtain a first DC coefficient and a first low-frequency AC coefficient; process the first DC coefficient, the first low-frequency AC coefficient, and the encrypted quick response code matrix to obtain a second DC coefficient and a second low-frequency AC coefficient; perform inverse discrete cosine transform on the plurality of size image blocks according to the second DC coefficient and the second low-frequency AC coefficient to embed watermark information and obtain a first video with a watermark; the size-blocking processing on the first scene change frame in the first video to obtain a plurality of size image blocks includes: calculating visual saliency values ​​of pixels of the first scene change frame according to the first scene change frame in the first video, and size-blocking the first scene change frame in the first video according to the visual saliency values ​​to obtain a plurality of size image blocks; The extraction module is configured to process the first video with the watermark according to the second DC coefficient and the second low-frequency AC coefficient to obtain a target estimated value of the encrypted quick response code matrix; decrypt the target estimated value of the encrypted quick response code matrix to obtain a target quick response code matrix; and decode the target quick response code matrix to obtain target watermark information, thereby extracting the watermark information.

7. The device according to claim 6, characterized in that The embedding module is specifically used to calculate the visual saliency value of the pixel points of the first scene change frame in the first video based on the first scene change frame, and perform size block processing on the first scene change frame in the first video according to the visual saliency value to obtain multiple size image blocks.

8. A computing device, characterized in that including memory and processor; One or more computer programs are stored in the memory, and the one or more computer programs include instructions; when the instructions are executed by the processor, the computing device executes the method according to any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • A method and a device for generating a digital watermark image based on graphic codes

    AU2020104204A4

  • Fast robust watermarking method and system for color image and application of fast robust watermarking method

    CN113538200A