A content protection method, electronic device and storage medium
By using multi-level feature fusion and deep learning technology, watermarks are embedded layer by layer, solving the problem of watermarks being easily tampered with and lost in the copyright protection of AI-generated content, and achieving efficient and stable watermark protection.
Patent Information
- Application Number
- CN202511164750.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-20
AI Technical Summary
Existing watermarking technologies are insufficient to effectively protect the copyright of AI-generated content, are easily tampered with or lost, and lack compatibility and robustness when faced with multimodal features.
A multi-level feature fusion method is adopted to divide the image into multiple regions. Watermarks are embedded layer by layer through deep learning technology. Feature extraction and watermark generation are performed using deep residual networks and SE blocks. The hidden watermarking technology is combined to improve the watermark's resistance to attacks.
It significantly enhances the watermark's resistance to attacks, ensures the stability and reliability of the watermark in complex environments, improves the accuracy and robustness of watermark extraction, and achieves efficient protection of AI-generated content.
Smart Images

Figure CN120744887B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a method for generating content protection, an electronic device and a storage medium. BACKGROUND
[0002] AIGC (Artificial Intelligence Generated Content) is a new content creation method in which artificial intelligence learns a large amount of data to automatically generate various contents such as images and videos. At present, the protection of AIGC content is still a very tricky problem. The traditional watermarking technology is easy to be tampered with or lost, and it is difficult to effectively protect the copyright of the generated content, which cannot meet the needs of generated content protection. SUMMARY
[0003] In view of the above problems, the present application provides a method for generating content protection, an electronic device and a storage medium to achieve the purpose of improving the reliability and security of the watermark. The specific scheme is as follows:
[0004] The first aspect of the present application provides a method for generating content protection, comprising:
[0005] obtaining a to-be-processed image and a watermark to be implanted in the to-be-processed image;
[0006] performing feature fusion of at least one level on the to-be-processed image and the watermark based on an encoder to obtain first image feature data with the watermark fused, and in the feature fusion process of each level: splitting the to-be-processed image into a preset number of regions corresponding to the level, fusing second image feature data of each region and the watermark, and the number of regions included in the feature fusion of a high level is lower than the number of regions included in the feature fusion of a low level in adjacent levels;
[0007] performing fusion of the first image feature data and the to-be-processed image based on the encoder to obtain a protection image with the watermark implanted.
[0008] In a possible implementation, the fusing of the second image feature data of each region and the watermark includes:
[0009] mapping the watermark into the second image feature data according to an amplitude intensity corresponding to a style content of the to-be-processed image.
[0010] In a possible implementation, the performing of the feature fusion of at least one level on the to-be-processed image and the watermark based on the encoder to obtain the first image feature data with the watermark fused includes: in the feature fusion of each level:
[0011] perform feature extraction on each region based on an image feature extractor composed of a first deep residual network and at least one first SE block to obtain second image feature data of each region;
[0012] perform watermark conversion based on a watermark generator composed of a second deep residual network, a diffusion block, at least one transpose convolution layer and at least one second SE block to obtain generated watermarks adapted to each region;
[0013] superimpose each generated watermark and corresponding second image feature data.
[0014] In a possible implementation, the fusion of the first image feature data and the to-be-processed image based on the encoder to obtain a protection image in which the watermark is implanted includes:
[0015] perform processing on the first image feature data based on a third deep residual network to obtain third image feature data;
[0016] fuse the third image feature data and the to-be-processed image and input the fusion result to a fourth deep residual network for processing to obtain the protection image.
[0017] The second aspect of the present application provides a content protection generation method, including:
[0018] determine a target region in the to-be-decoded image according to a proportional relationship between the to-be-decoded image and the protection image;
[0019] perform decoding operation on each region in each level in the target region based on a decoder to obtain first watermarks of each region in each level;
[0020] fuse all first watermarks to obtain first decoded watermarks, and compare the first decoded watermarks with original watermarks corresponding to the protection image to obtain comparison result data.
[0021] In a possible implementation, the method further includes:
[0022] determine whether an attack is made according to watermark integrity represented by the comparison result data;
[0023] when it is determined that an attack is made, repair the to-be-decoded image according to an attack type;
[0024] perform decoding operation on each region in each level of the repaired to-be-decoded image based on the decoder to obtain second watermarks of each region in each level;
[0025] The second watermarks are fused to obtain second decoded watermarks, and the second decoded watermarks are compared with the original watermarks.
[0026] In a possible implementation, the first watermarks of each region in each level in the target region are obtained by performing decoding operations on each region in each level in the target region based on the decoder.
[0027] The first watermarks of each region in each level in the target region are obtained by performing decoding operations on each region in each level in the target region based on the decoder composed of the fifth deep residual network, the third SE block, and the anti-diffusion block.
[0028] In a possible implementation, the first decoded watermarks are obtained by fusing the first watermarks, and the comparison result data is obtained by comparing the first decoded watermarks with the original watermarks corresponding to the protection image.
[0029] The values of each position in the average value are compared with a threshold value, and values greater than the threshold value are set to 1, and values less than the threshold value are set to 0, to obtain the first decoded watermarks.
[0030] The values of each position in the average value are compared with a threshold value, and values greater than the threshold value are set to 1, and values less than the threshold value are set to 0, to obtain the first decoded watermarks.
[0031] The first decoded watermarks are compared with the original watermarks in terms of bit accuracy to obtain the comparison result data.
[0032] The third aspect of the present application provides a computer program product, which includes computer readable instructions, when the computer readable instructions run on an electronic device, the electronic device implements the content protection method of the first aspect or the second aspect.
[0033] The fourth aspect of the present application provides an electronic device, which includes at least one processor and a memory connected to the processor, wherein:
[0034] The memory is configured to store a computer program.
[0035] The processor is configured to execute the computer program, so that the electronic device can implement the content protection method of the first aspect or the second aspect.
[0036] The fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs, when the one or more computer programs are executed by an electronic device, the electronic device can implement the content protection method of the first aspect or the second aspect.
[0037] By the technical scheme, the content protection generation method provided by the application obtains first image feature data by performing feature fusion of at least one level on the obtained to-be-processed image and the watermark of the to-be-processed image based on an encoder, and in the feature fusion process of each level: the to-be-processed image is split into a preset number of regions corresponding to the level, the second image feature data of each region and the watermark are fused, the number of regions included in the feature fusion of a high level in adjacent levels of feature fusion is lower than the number of regions included in the feature fusion of a low level, and then the embedding of the watermark in multiple levels on the image features is realized. Finally, the first image feature data and the to-be-processed image are fused based on the encoder to obtain a protection image in which the watermark is implanted. By using the deep learning technology to inject the watermark into the image layer by layer, the complexity of the watermark in the image is effectively increased, and then the anti-attack ability of the watermark is significantly enhanced, so that the watermark can resist various common attacks such as noise, compression, geometric deformation, and the like, and the stability and reliability of the watermark in various complex environments are ensured. BRIEF DESCRIPTION OF DRAWINGS
[0038] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent by describing in detail the following specific embodiments with reference to the attached drawings. The same or similar components are denoted by the same or similar reference numerals throughout the drawings. It should be understood that the drawings are schematic, and the shapes and elements are not necessarily drawn to scale.
[0039] Figure 1 An architecture diagram of a content protection generation system provided by the application;
[0040] Figure 2 A flowchart of a content protection generation method provided by the application;
[0041] Figure 3 A schematic diagram of watermark hierarchical implantation provided by the application;
[0042] Figure 4 An architecture diagram of an encoder and a decoder provided by the application;
[0043] Figure 5 A flowchart of another content protection generation method provided by the application;
[0044] Figure 6 A specific implementation mode diagram of a content protection generation method provided by the application;
[0045] Figure 7 A structure diagram of an electronic device provided by the application. DETAILED DESCRIPTION
[0046] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0047] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0048] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0049] See Figure 1 , Figure 1 A schematic diagram of the architecture of a content generation protection system is shown. The system may include a terminal 100 and a server 200. The server 200 can provide the content generation protection method provided in this application embodiment to one or more terminals.
[0050] The terminal 100 may have an application for generating content input installed. The application and the webpage can provide an interface. The terminal 100 can receive relevant parameters entered by the user on the content input interface and send the parameters to the server 200. The server 200 can obtain the processing result based on the received parameters and return the processing result to the terminal 100.
[0051] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the processing result based on the received parameters on its own, without the need for the server to cooperate. This application embodiment is not limited to this.
[0052] The following description Figure 1 The product form of the mid-terminal 100;
[0053] The terminal 100 in this application embodiment can be a mobile phone, tablet computer, in-vehicle equipment, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., and this application embodiment does not impose any restrictions on it.
[0054] Terminal 100 may include a radio frequency unit, memory, input unit, display unit, camera (optional), audio circuitry (optional), speaker (optional), microphone (optional), headphone jack (optional), processor, external interface, power supply, and other components. Those skilled in the art will understand that the above-mentioned components are merely examples and do not constitute a limitation on the terminal or multifunctional device; it may include more or fewer components, or a combination of certain components, or different components.
[0055] The input unit can be used to receive input numeric or character information, and to generate key signal inputs related to user settings and function control of the portable multi-functional device. Specifically, the input unit may include a touchscreen (optional) and / or other input devices. Other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0056] Among them, the input device can receive input data, etc.
[0057] The display unit can be used to display information input by the user or information provided to the user, various menus of the terminal, interactive interfaces, file display, and / or playback of any multimedia file. In the embodiments of this application, the display unit can be used to display the interface for generating content input, processing results, etc.
[0058] The memory can be used to store software code related to the content protection method, the processor can execute the steps of generating the content protection method, and can also schedule other units (such as the input unit and display unit mentioned above) to achieve the corresponding functions.
[0059] This radio frequency unit (optional) can be used to receive and send signals during information transmission or calls.
[0060] In this embodiment of the application, the radio frequency unit can send data to the server 200 and receive the processing results sent by the server 200.
[0061] It should be understood that this radio frequency unit is optional and can be replaced with other communication interfaces, such as a network port.
[0062] Terminal 100 also includes a power source (such as a battery) for supplying power to the various components.
[0063] Terminal 100 also includes an external interface, which can be a standard Micro USB interface or a multi-pin connector, which can be used to connect terminal 100 to other devices for communication or to connect a charger to charge terminal 100.
[0064] Server 200 includes a bus, a processor, a communication interface, and memory. The processor, memory, and communication interface communicate with each other via the bus.
[0065] The memory can be used to store software code related to the content protection method, the processor can execute the steps of the chip's content protection method, and can also schedule other units to achieve the corresponding functions.
[0066] Currently, copyright protection is a pressing issue in the field of AIGC (AI-Generated Content). While traditional watermarking technologies can protect copyright to some extent, they have many shortcomings when dealing with AI-generated content. For example, watermarks are easily tampered with or lost, and they struggle to adapt to the multimodal characteristics of AI-generated content. Existing watermarking technologies, such as digital watermarking and steganography, while meeting copyright protection requirements to some extent, still suffer from poor compatibility and insufficient robustness when dealing with AIGC content. Early watermark embedding methods primarily operated at the pixel level, directly modifying pixel values in specific color spaces (such as RGB or YUV channels) of image or video frames to embed watermark information. However, these pixel-level embedding algorithms have been found to lack robustness and are susceptible to interference. Subsequently, transform domain embedding schemes were developed, including Fourier transform, DCT (Discrete Cosine Transform), wavelet transform, DTCWT (Dual-Tree Complex Wavelet Transform), and some were even combined with SVD (Singular Value Decomposition). However, these schemes still could not meet the robustness requirements of adversarial attacks.
[0067] To address the aforementioned problems, this application provides a method for generating content protection. The method for generating content protection according to this application will be described in detail below with reference to the accompanying drawings.
[0068] Reference Figure 2 , Figure 2This application provides a flowchart illustrating a content protection generation method as an embodiment of the present application. Figure 2 As shown in the embodiment of this application, a method for protecting generated content is provided. This method can be applied to the process of embedding watermarks into generated content and may include steps 201 to 203. These steps are described in detail below.
[0069] 201. Obtain the image to be processed and the watermark to be implanted into the image to be processed.
[0070] Specifically, the image to be processed here can be a directly acquired image or an image of each frame of video content. In order to meet the processing requirements of the encoder model, the image to be processed can be pre-processed to convert it into a format suitable for the encoder to process. For example, the acquired images can be uniformly processed into images with a resolution of 720P.
[0071] The watermark to be implanted in the corresponding image to be processed can be generated using hidden watermarking technology. Hidden watermarking technology, by embedding a hidden watermark containing a specific identifier, can track the dissemination path of digital content, locate unauthorized users, thereby improving content security, reducing the risk of data leakage, and not affecting the efficiency of content dissemination. It does not significantly affect the quality of the original content, ensuring its viewing experience and usability, and improving user experience. In specific implementation, a deterministic Bernoulli distribution can be used to generate a unique binary watermark sequence. By setting appropriate parameters, the randomness and uniqueness of the watermark can be controlled. While ensuring the randomness of the watermark, its parameter range is limited, enhancing the controllability and consistency of the watermark.
[0072] It is understood that those skilled in the art may use other methods to generate hidden watermarks, and no restrictions are imposed here.
[0073] 202. Based on the encoder, feature fusion of the image to be processed and the watermark at at least one level is performed to obtain the first image feature data incorporating the watermark. In the feature fusion process at each level: the image to be processed is divided into a preset number of regions corresponding to the level, and the second image feature data of each region is fused with the watermark. In the feature fusion of adjacent levels, the number of regions contained in the feature fusion of higher levels is lower than the number of regions contained in the feature fusion of lower levels.
[0074] Specifically, the generated content can be divided into multiple regions, each serving as an independent embedding unit. For example, in video, it can be segmented according to temporal and spatial dimensions. In images, consecutive pixel regions can be segmented. Within each region, adjacent blocks of the same size are designed to repeatedly embed the same bit information. Multiple adjacent regions can form a larger region. Embedding watermarks in this way enhances the redundancy and resistance to attacks; even if some regions are damaged, the watermark information can still be recovered from adjacent blocks.
[0075] For example, refer to Figure 3 As shown, the entire image is divided into 20 small regions. Based on this, four small regions are combined to form one large region, resulting in four large regions. Watermarks are then implanted in each small region and also in each large region. During implantation, the frequency features of each region are extracted using an encoder, and the watermark is diffused to a state adapted to each region. The watermark information is then fused with the image's frequency domain features. This method subtly embeds the watermark information into the image's frequency domain features, thereby improving the watermark's robustness and invisibility.
[0076] It is understood that those skilled in the art can adjust the number of the above-mentioned levels and the size of each region in each level as needed, without any restrictions.
[0077] 203. Based on the encoder, the first image feature data and the image to be processed are fused to obtain a watermarked protected image.
[0078] Specifically, based on the first image feature data incorporating the watermark information, an encoder further integrates the first image feature data into the pixels of the image to be processed, thereby obtaining a watermarked protected image. Through multi-level watermark structure design and deep integration of deep learning technology, the watermark's resistance to attacks is significantly enhanced, enabling it to withstand various common attacks such as noise, compression, and geometric deformation, ensuring the stability and reliability of the watermark in various complex environments. Simultaneously, the accuracy and robustness of watermark extraction are improved, providing a solid guarantee for the security of the generated content.
[0079] In one embodiment, to further ensure the quality of the generated content and avoid affecting the generated content, the second image feature data and watermark of each region are fused in the above embodiment, which may specifically include:
[0080] The watermark is mapped onto the second image feature data according to the amplitude intensity corresponding to the style content of the image to be processed.
[0081] Specifically, based on the style of the generated content (such as comics, realistic, etc.), its characteristics and attributes can be analyzed to determine parameters such as intensity. Realistic images with complex content are given higher intensity parameters, while comic-style images with relatively simpler content are given lower intensity parameters. Lower intensity maintains better image quality. Furthermore, in complex image scenarios, stronger hidden watermarks are less likely to be detected.
[0082] In practical implementation, the intensity parameter, such as amplitude intensity, mainly refers to the strength of the watermark signal, that is, the magnitude of the watermark value mapped to the pixels of the original image. Specifically, it controls the degree of modification to the image pixel values during watermark embedding. A larger intensity parameter makes the watermark more robust and easier to extract, but may reduce visual quality; a smaller intensity parameter makes the watermark more concealed and has less visual impact, but may reduce robustness and extractability. For example, in complex realistic images, the intensity parameter can be set to a higher value so that the watermark resists attacks such as noise, compression, and geometric distortion; while in simple cartoon-style images, the intensity parameter can be reduced to minimize the impact on image quality. Furthermore, in complex image conditions, a strong hidden watermark is also difficult to detect.
[0083] When identifying specific style content, a large number of content samples labeled as anime and realistic genres can be collected from resource libraries such as videos and images. These samples cover a wide variety of visual style features, providing a rich data foundation for model training. Then, a binary classifier model is trained using deep learning technology. During training, data augmentation techniques, such as random cropping, rotation, scaling, and adjusting brightness and contrast, are employed to increase the diversity of samples. Simultaneously, transfer learning techniques are used to apply the pre-trained model (ResNet-50) to the style classification task, improving the model's performance and efficiency. In practical applications, the AI-generated content to be identified is input into the trained binary classifier model, which can quickly and accurately distinguish between anime and realistic genres, providing a basis for setting subsequent watermark strength parameters.
[0084] In some specific embodiments, based on the encoder, feature fusion is performed on the image to be processed and the watermark at at least one level to obtain first image feature data incorporating the watermark. This may specifically include: in each level of feature fusion:
[0085] An image feature extractor based on a first deep residual network and at least one first SE block is used to extract features from each region to obtain second image feature data for each region.
[0086] A watermark generator, consisting of a second deep residual network, a diffusion block, at least one transposed convolutional layer, and at least one second SE block, converts the watermark into a generated watermark that is adapted to each region.
[0087] Each generated watermark is superimposed with a corresponding second image feature data.
[0088] Specifically, refer to Figure 4 As shown, the encoder consists of an image feature extractor at the top and a watermark generator at the bottom. The image feature extractor consists of a deep residual network and multiple SE blocks. The deep residual network, such as ResNet-50, extracts high-level semantic features of the image through multi-layer convolution operations, gradually reducing the spatial size of the image while increasing the number of channels to capture the complex features of the image. The deep residual network mainly consists of convolutional layers, batch normalization layers, and activation layers.
[0089] The SE block (Squeeze-and-Excitation block) employs multi-scale pooling operations, including global average pooling, local average pooling, and max pooling, to capture feature information at different scales. The multi-scale pooling results are then fused, and weights for each channel are generated through two fully connected layers and a gating mechanism. Finally, adaptive channel calibration is used to enhance important features and suppress unimportant features, thereby improving the model's ability to learn image features.
[0090] The watermark generator consists of a diffusion block, a deep residual network, multiple transpose convolutions, and multiple SE blocks. The diffusion block spreads the watermark throughout the entire image, making the watermark information more dispersed and improving its resistance to attacks. Multiple loss functions are selected, such as image quality loss, watermark extraction accuracy loss, adversarial loss, and perceptual loss. Furthermore, data augmentation operations are performed on the training data during training, such as random cropping, rotation, scaling, and adjusting brightness / contrast, to increase data diversity and improve the model's robustness.
[0091] Based on this, the first image feature data is processed using a subsequent third deep residual network to obtain the third image feature data.
[0092] The third image feature data and the image to be processed are then fused and input into the fourth deep residual network for further processing to obtain the protected image.
[0093] During the training of this model, a large number of AI-generated content samples are collected and generated, and their image complexity is labeled. Multiple loss functions are selected for the model, including image quality loss, watermark extraction accuracy loss, adversarial loss, and perceptual loss. Data augmentation operations are performed on the training data during training, such as random cropping, rotation, scaling, and adjusting brightness / contrast, to increase data diversity and improve the model's robustness. This enables the model to learn how to effectively embed and extract watermark information while preserving image quality.
[0094] Based on the same design concept, and referring to Figure 5 As shown, embodiments of this application also provide a method for protecting generated content, which can be used in the process of watermarking the acquired generated content, and may specifically include the following processing steps:
[0095] 501. Based on the proportional relationship between the image to be decoded and the protected image, determine the target region in the image to be decoded.
[0096] Specifically, considering that the generated content may be subject to attacks such as deformation and cropping, resulting in the image size not being consistent with the original image size, it is necessary to determine the area in the image to be decoded that is consistent with the protected image by the size ratio between the image to be decoded and the protected image, and then perform targeted watermark decoding operations to improve the accuracy and reliability of watermark extraction.
[0097] 502. Based on the decoder, perform decoding operations on each region in each level of the target region to obtain the first watermark of each region in each level.
[0098] Specifically, corresponding to the encoder in the above embodiments, a suitable decoder can be used here to extract the watermark layer by layer, so as to obtain the watermark of each region in each layer.
[0099] 503. Perform fusion processing on all the first watermarks to obtain the first decoded watermark, and compare the first decoded watermark with the original watermark corresponding to the protected image to obtain the comparison result data.
[0100] Specifically, this can be achieved by summing the watermarks and taking the average. The value at each position in the average is compared to a threshold, with values greater than the threshold set to 1 and values less than the threshold set to 0, resulting in the first decoded watermark. This first decoded watermark is then compared to the original watermark in terms of bit accuracy to obtain the comparison result data.
[0101] For example, for Figure 3The two-tiered watermarking method shown comprises 20 small regions and 4 large regions, with watermark extraction performed on each small region. For each large region consisting of 4 small regions, a total of 24 watermark information can be extracted. Next, these 24 watermark information are fused using a maximum voting method to improve extraction accuracy and attack resistance. The specific implementation process is as follows: each binary watermark sequence is averaged, and then the average result for each bit is binarized using a threshold of 0.5 (i.e., compared with the average result for each bit using 0.5) to obtain the final result for each watermark bit, thus achieving watermark information fusion. Even when the test input resolution differs from the training image resolution, the test image can be cropped into multiple blocks, watermarks can be repeatedly embedded, and all extracted watermarks can be voted on during decoding to obtain the final extraction result. This enhances the accuracy and robustness of watermark extraction, improving the success rate of watermark extraction under attack.
[0102] In other embodiments, to accurately identify the attacks received and provide a reliable reference for subsequent targeted defense, the following processing steps may also be included:
[0103] Step 11: Determine whether an attack has occurred based on the watermark integrity represented by the comparison results data.
[0104] Step 12: When an attack is confirmed, repair the image to be decoded according to the type of attack.
[0105] Step 13: Based on the decoder, perform decoding operations on each region in each layer of the repaired image to be decoded to obtain the second watermark for each region in each layer.
[0106] Step 14: Perform fusion processing on all the second watermarks to obtain the second decoded watermark, and compare the second decoded watermark with the original watermark.
[0107] Specifically, to address different types of attacks, it's necessary to determine the attack type and intensity. This can be achieved by comparing the basic dimensions of the attacked image with those of the original image. If the dimensions of the two images differ only slightly, the attack intensity is low, and direct watermark extraction can be attempted. However, if the dimensions differ significantly, it indicates a high-intensity attack, potentially involving large-scale cropping or other geometric deformation attacks. In this case, feature matching techniques can be used to restore the image to its original dimensions before watermark extraction. This approach effectively improves the completeness and success rate of watermark extraction, ensuring accurate recovery of the watermark information. After restoring the image to its original dimensions, further watermark extraction based on the decoder and comparison with the original watermark can be performed.
[0108] As a specific implementation of step 502 above, the decoder performs decoding operations on each region in each layer of the target region to obtain the first watermark of each region in each layer, which may specifically include:
[0109] The decoder, which is composed of the fifth deep residual network, the third SE block, and the anti-diffusion block, performs decoding operations on each region in each layer of the target region to obtain the first watermark of each region in each layer.
[0110] Specifically, refer to Figure 4 The decoder structure shown on the right reconstructs the image or extracts watermark information from features through multiple deconvolution operations, gradually restoring the spatial size of the image while reducing the number of channels, ultimately generating an output with the same size as the input image. SE blocks are added to the model to enhance its ability to learn image features, especially in the frequency domain. The anti-diffusion block employs the opposite processing procedure to the diffusion block, extracting the watermark at multiple levels.
[0111] As a specific application of the above-mentioned content protection method for watermark extraction, refer to Figure 6 As shown, the specific processing steps may include the following:
[0112] First, determine if the image to be decoded is 720P. If it is, proceed with the subsequent decoding operations. If it is not 720P, calculate the ratio between the current image size and the size of a 720P image, and then perform the image decoding operation.
[0113] When the resolution is set to 720P, each layer of the image area is decoded according to the designed image area size to obtain all decoded watermarks.
[0114] If it is determined that it is not 720P, adjust the size of the area to be decoded according to the ratio of the original image to 720P, and then perform layer decoding.
[0115] Voting is conducted on the watermark information for each region and different levels to obtain the extracted watermark.
[0116] A complete watermark is determined when the watermark integrity reaches 90%.
[0117] If the watermark extraction integrity does not reach 90%, check for various attacks in sequence, such as interception, deformation, and noise addition. After performing the corresponding repairs, return to execute the above decoding operation.
[0118] If it is determined that the image has not been attacked and the watermark extraction integrity does not reach 90%, it is determined that there is no watermark in the image.
[0119] The content protection method provided in this application can achieve efficient and covert watermark embedding without significantly affecting image quality. It can automatically adjust watermark parameters according to different content styles, ensuring the robustness and invisibility of the watermark. Simultaneously, through multi-level watermark structure design and deep learning model optimization, the watermark's resistance to attacks is significantly enhanced, enabling it to withstand various common attacks such as noise, compression, and geometric deformation. Furthermore, combining the maximum voting method with multi-channel information improves the accuracy and robustness of watermark extraction. This invention also achieves full-process automation from content generation to copyright protection, greatly improving content production efficiency and copyright management effectiveness, and providing solid technical support for the healthy development of the AIGC industry.
[0120] This application also provides an electronic device in its embodiments. (See reference...) Figure 7 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0121] like Figure 7 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. When the electronic device is powered on, the RAM 703 also stores various programs and data required for the operation of the electronic device. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0122] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, memory cards, hard drives, etc.; and communication devices 709. Communication device 709 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have instead.
[0123] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the content protection methods provided in this application.
[0124] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the content protection methods provided in this application.
[0125] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0126] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0127] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0128] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A method for generating content protection, characterized in that, include: Acquire the image to be processed and the watermark to be implanted into the image to be processed; Based on the encoder, feature fusion is performed on the image to be processed and the watermark at at least one level to obtain the first image feature data incorporating the watermark. In the feature fusion process at each level: the image to be processed is divided into a preset number of regions corresponding to the level, and the second image feature data of each region is fused with the watermark. In the feature fusion of adjacent levels, the number of regions included in the feature fusion of higher levels is lower than the number of regions included in the feature fusion of lower levels. The encoder fuses the first image feature data and the image to be processed to obtain a protected image with the watermark embedded. The step of performing feature fusion at least one level on the image to be processed and the watermark based on the encoder to obtain first image feature data incorporating the watermark includes: in each level of feature fusion: An image feature extractor based on a first deep residual network and at least one first SE block is used to extract features from each region to obtain second image feature data for each region. A watermark generator based on a second deep residual network, a diffusion block, at least one transposed convolutional layer, and at least one second SE block converts the watermark into a generated watermark adapted to each region. Each generated watermark is superimposed with a corresponding second image feature data; The process involves fusing the first image feature data and the image to be processed based on the encoder to obtain a protected image with the watermark embedded, including: The first image feature data is processed based on the third deep residual network to obtain the third image feature data; The third image feature data and the image to be processed are fused and then input into the fourth deep residual network for processing to obtain the protected image.
2. The content protection method according to claim 1, characterized in that, The step of fusing the second image feature data of each region with the watermark includes: The watermark is mapped onto the second image feature data according to the amplitude intensity corresponding to the style content of the image to be processed.
3. A method for generating content protection, characterized in that, include: The target region in the image to be decoded is determined based on the ratio between the image to be decoded and the protected image. Based on the decoder, each region in each level of the target region is decoded to obtain the first watermark of each region in each level; All the first watermarks are fused to obtain the first decoded watermark, and the first decoded watermark is compared with the original watermark corresponding to the protected image to obtain the comparison result data. Specifically, the decoding operation is performed on each region in each level of the target region based on the decoder to obtain the first watermark of each region in each level, including: The decoder, which is composed of a fifth deep residual network, a third SE block, and an anti-diffusion block, performs decoding operations on each region in each layer of the target region to obtain a first watermark for each region in each layer.
4. The content protection method according to claim 3, characterized in that, Also includes: Based on the watermark integrity represented by the comparison results, it is determined whether an attack has occurred. Upon confirmation of an attack, the image to be decoded is repaired according to the type of attack. Based on the decoder, a decoding operation is performed on each region in each layer of the repaired image to be decoded to obtain a second watermark for each region in each layer; All second watermarks are fused to obtain a second decoded watermark, and the second decoded watermark is compared with the original watermark.
5. The content protection method according to claim 3, characterized in that, The process involves fusing all the first watermarks to obtain a first decoded watermark, and then comparing the first decoded watermark with the original watermark corresponding to the protected image to obtain comparison result data, including: Add up all the watermarks and take the average. The value at each position in the average value is compared with a threshold, and values greater than the threshold are set to 1, and values less than the threshold are set to 0, to obtain the first decoded watermark; The first decoded watermark and the original watermark are compared in terms of bit accuracy to obtain the comparison result data.
6. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the content protection method as described in any one of claims 1 to 5.
7. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the content protection method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Method for generating 3D image watermark based on AIGC and model framework
CN119205479A
Watermark detection method, device, terminal and storage medium
WO2021129466A1