Image protection method and device, electronic equipment and storage medium
By generating and superimposing protection noise, the problem of image watermarks being easily cracked is solved, efficient image protection and quality retention are achieved, and image copyright protection is enhanced.
Patent Information
- Application Number
- CN202510528960.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, image watermarks are easily cracked, affecting the image display effect and poor protection effect, and cannot effectively protect image copyright.
By obtaining the user's local protected text encoding and trained text adapter, text conditions are generated and inputted into the server's noise generation model to generate protection noise, superimposing it on the image to be protected according to the user's preset ratio to form the protected image.
The protection effect of the image and the quality of the image after protection are improved, and the cracking difficulty is increased. Users can adjust the relationship between protection intensity and image quality to improve user experience.
Smart Images

Figure CN120355590A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to image processing technologies, and in particular, to an image protection method, apparatus, electronic device, and storage medium. Background Art
[0002] With the rapid development of image generation technologies, many advanced large-scale image generation models can effectively understand the semantic features of images through deep learning and generate synthetic images that are almost indistinguishable from real images. This image generation technology has brought great convenience to the art creation and design industries, but it has also raised copyright protection issues for image content.
[0003] In the prior art, image protection can be achieved by adding watermarks to images. However, watermarks can affect the display effect of images, and watermarks are easily cracked, resulting in poor protection effects for images. Summary of the Invention
[0004] The present application provides an image protection method, apparatus, electronic device, and storage medium to improve the protection effect of images and the quality of protected images.
[0005] In a first aspect, an embodiment of the present application provides an image protection method, which includes:
[0006] Obtain an image to be protected;
[0007] Obtain a protection text encoding and a trained text adapter stored locally by the user, and input the protection text encoding into the trained text adapter to obtain a text condition;
[0008] Input the image to be protected and the text condition into a trained noise generation model stored in the server to obtain protection noise; the trained noise generation model and the trained text adapter are a pair of models in a matched training model group;
[0009] Overlay the protection noise and the image to be protected according to a preset ratio of the user to obtain a protected image.
[0010] In a second aspect, an embodiment of the present application further provides an image protection apparatus, which includes:
[0011] An image-to-be-protected acquisition module, configured to obtain an image to be protected;
[0012] A text condition acquisition module, configured to obtain a protection text encoding and a trained text adapter stored locally by the user, and input the protection text encoding into the trained text adapter to obtain a text condition;
[0013] A protection noise acquisition module is configured to input the image to be protected and text conditions into a trained noise generation model stored in a server to obtain protection noise; the trained noise generation model and the trained text adapter are a pair of models in a model group that have been trained to match each other.
[0014] A protection noise superposition module is configured to superpose the protection noise and the image to be protected according to a preset ratio of a user to obtain a protected image.
[0015] Thirdly, an embodiment of the present application further provides an electronic device, which includes:
[0016] One or more processors;
[0017] A storage device configured to store one or more programs;
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement any image protection method provided by the embodiment of the present application.
[0019] Fourthly, an embodiment of the present application further provides a storage medium including computer-executable instructions, and the computer-executable instructions are used to execute any image protection method provided by the embodiment of the present application when being executed by a computer processor.
[0020] The present application obtains an image to be protected; obtains a protected text encoding and a trained text adapter stored locally by a user, inputs the protected text encoding into the trained text adapter to obtain text conditions; the text conditions can change semantic features of the image to be protected and improve the difficulty of being cracked; inputs the image to be protected and the text conditions into a trained noise generation model stored in a server to obtain protection noise; by storing the trained noise generation model and the trained text adapter remotely, the difficulty of cracking the model group is improved, and the protection effect on the image to be protected is improved. At the same time, since the trained noise generation model and the trained text adapter are a pair of models in a model group that have been trained to match each other, a user can privately own a group of protection models, providing proprietary image protection for the user, further improving the protection effect and cracking difficulty of the model group; superposes the protection noise and the image to be protected according to a preset ratio of the user to obtain a protected image, and the user can adjust the relationship between the protection intensity and the image quality through the preset ratio, select a satisfactory protection effect, and improve the user experience. Therefore, through the technical solution of the present application, the problems that watermarks affect the display effect of images, are easily cracked, and have a poor protection effect on images are solved, and the effects of improving the protection effect on images and the quality of protected images are achieved. Description of the Drawings
[0021] Figure 1It is a flowchart of an image protection method in Embodiment 1 of the present application;
[0022] Figure 2 It is a flowchart of an image protection method in Embodiment 2 of the present application;
[0023] Figure 3 It is a flowchart of an image protection method in Embodiment 3 of the present application;
[0024] Figure 4 It is a schematic structural diagram of an image protection device in Embodiment 4 of the present application;
[0025] Figure 5 It is a schematic structural diagram of an electronic device in Embodiment 5 of the present application. Detailed implementation manners
[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0027] It should be noted that the terms "first" and "second" in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] Embodiment 1
[0029] Figure 1 It is a flowchart of an image protection method provided in Embodiment 1 of the present application. This embodiment is applicable to the situation of protecting images. This method can be executed by an image protection device, which can be implemented by software and / or hardware and is specifically configured in an electronic device for displaying images to be protected, such as a mobile phone.
[0030] See Figure 1 The image protection method shown, which specifically includes the following steps:
[0031] S100. Obtain the image to be protected.
[0032] The image to be protected can be an image that the user designates to be protected to prevent it from being stolen and used by unauthorized users, thus infringing on the user's copyright. The user can set the image to be protected according to their needs, and obtain the image set by the user as the image to be protected.
[0033] S110. Obtain the protected text encoding stored locally by the user and the trained text adapter, and input the protected text encoding into the text adapter to obtain a text condition.
[0034] The protected text encoding can be the encoding after compiling the protected text preset by the user, and is used to generate protection noise for the image to be protected. The protected text can be a preset protection word. It should be noted that the protection word can be an image object displayed on the image to generate a text condition that has a protective meaning for the subsequent image to be protected. For example, the protected text can be a burning house. Exemplarily, the trained text encoder can be used to encode the protected text set by the user to obtain the protected text encoding, and the protected text encoding is stored locally by the user.
[0035] The protected text encoding can be used to train the text adapter so that the text condition generated by the text adapter can maximize the coding difference from the image to be protected, thereby maximizing the change in the semantic information of the image to be protected and increasing the difficulty of the subsequent protected image being learned. The specific process of training the text adapter with the protected text encoding can be seen in the subsequent embodiments and will not be elaborated here.
[0036] The trained text adapter can be used to process the protected text encoding to obtain a text condition. The text adapter can be a lightweight deep learning model. The text adapter can be specified by the user, and this application does not make specific limitations in this regard. Exemplarily, the text adapter can be an artificial neural network model. For example, the text adapter can be a Multilayer Perceptron (MLP). The text adapter can be trained in advance with image samples and protected text encoding to obtain the trained text adapter, and the trained text adapter is stored locally by the user.
[0037] After obtaining the image to be protected, obtain the protected text encoding and the trained text adapter from the user's local storage. Input the protected text encoding into the text adapter, and the output is a text condition. The text condition can be used to generate protection noise to change the semantic information of the image to be protected, thereby effectively preventing the image generation model from learning the image to be protected and improving the protection effect of the protection noise generated by the text condition on the image to be protected.
[0038] S120, inputting the image to be protected and the text conditions into the trained noise generation model stored in the server to obtain the protection noise; the trained noise generation model and the trained text adapter are a pair of model groups after matching training.
[0039] The trained noise generation model is used to generate protection noise for the image to be protected according to the text condition. Exemplarily, the image to be protected and the text condition are input into the trained noise generation model, and the protection noise is output. Exemplarily, the noise generation model can be a deep learning model. For example, the noise generation model can be a Conditional Resnet model (a professional term, a deep learning model) or a Conditional Unet model (a professional term, a deep learning model).
[0040] The trained noise generation model and the trained text adapter are a pair of model groups that have been matched and trained. It can be understood that the parameter modifications of the noise generation model and the text adapter during the training process are matched, that is, when the noise generation model is trained through image samples, the training goal is to minimize the output protection noise and the input image, and the parameters of the noise generation model and the text adapter are updated at the same time, so that the trained noise generation model and the trained text adapter are a pair of matched model groups, which can be understood as a private set of image protection models for users, which increases the difficulty of deciphering the model and improves the protection effect of the protected image.
[0041] At the same time, by storing the trained noise generation model in the server and storing the protection text encoding and the trained text adapter locally on the user, off-site storage is achieved. The user-defined protection text encoding and the trained text adapter are stored locally, reducing the risk of all protection models being obtained and cracked by potential intruders at the same time, improving the security of the generated protection noise, and improving the protection effect of the protection noise.
[0042] S130: Superimpose the protection noise and the image to be protected according to a ratio preset by the user to obtain a protected image.
[0043] The preset ratio may be a mixing ratio of the protection noise and the image to be protected that is preset by the user, and is used to mix the protection noise and the image to be protected. The protection noise is the same size as the image to be protected. After obtaining the protection noise, the protection noise and the pixel value at the corresponding position in the image to be protected may be weighted and summed according to the preset ratio, and the weighted pixel value obtained is the pixel value at the corresponding position in the protected image.
[0044] Through the preset ratio of the user, the relationship between the protection intensity and the image quality can be adjusted. This customized part can not only provide exclusive protection for each user, but also enable the user to select a satisfactory protection effect, improving the visual effect of the protected image.
[0045] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data comply with the relevant laws, regulations, and standards in the relevant regions.
[0046] In recent years, with the rapid development of image generation technology, many advanced and large-scale image generation models have been able to effectively understand the semantic features of images through deep learning and generate synthetic images that are almost indistinguishable from real images. Image generation technology has brought great convenience to the art creation and design industries, but it has also raised major issues regarding the copyright protection of image content producers (such as painters, photographers, and designers). Exemplarily, when an image content producer publicly releases their work on the Internet, the publicly released image is easily crawled by crawler programs and used as fine-tuning data for image generation models, enabling the fine-tuned image generation model to not only learn the style of the image but also potentially copy and reproduce these styles uncontrollably, resulting in the infringement of the copyright of the image content producer. In particular, users (potential infringers) who use the fine-tuned generation model can generate images with the original author's style at extremely low cost, seriously threatening the originality of the creator and the exclusivity of the work.
[0047] Although publicly releasing images on the Internet by image content producers is a necessary means to promote their works and attract audiences, this behavior also makes their works vulnerable to unauthorized use. Therefore, there is an urgent need for an effective image protection mechanism to ensure that the image works publicly released by image content producers are properly protected in the network environment and to prevent their styles and creations from being copied or tampered with by unauthorized image generation models.
[0048] In the prior art, image protection mostly adopts methods of adding watermarks and noise. However, adding watermarks to images cannot protect the images from being learned by the above-mentioned image generation models. Adding noise is usually a deep learning-based protection method, and often an integrated online model is used to add noise to pictures. Since the online model also has the risk of being cracked, there is also a risk that the online model can be obtained by potential infringers.
[0049] The technical solution of this embodiment is to obtain the image to be protected; obtain the protected text encoding and the trained text adapter stored locally by the user, input the protected text encoding into the trained text adapter to obtain the text condition; the text condition can change the semantic features of the image to be protected and improve the difficulty of being cracked; input the image to be protected and the text condition into the trained noise generation model stored in the server to obtain the protected noise; store the trained noise generation model and the trained text adapter remotely to improve the difficulty of cracking the model group and the protection effect on the image to be protected. At the same time, since the trained noise generation model and the trained text adapter are a pair of model groups after matching training, the user can have a private set of protection models to provide proprietary image protection for the user, further improving the protection effect and cracking difficulty of the model group; superimpose the protected noise and the image to be protected according to the preset ratio of the user to obtain the protected image. The user can adjust the relationship between the protection intensity and the image quality through the preset ratio, select the satisfactory protection effect, and improve the user experience. Therefore, through the technical solution of this application, the problems that the watermark affects the display effect of the image, is easily cracked, and has a poor protection effect on the image are solved, and the protection effect on the image and the quality of the protected image are improved.
[0050] Embodiment 2
[0051] Figure 2 The flowchart of an image protection method provided by the second embodiment of this application. The technical solution of this embodiment is further refined on the basis of the above technical solution.
[0052] Further, before obtaining the protected text encoding and the trained text adapter stored locally by the user, append "obtain the protected text, image sample, trained text encoder, trained image encoder matching the trained text encoder, initial text adapter, and initial noise generation model specified by the user; input the protected text into the trained text encoder to obtain the protected text encoding, and input the image sample into the trained image encoder to obtain the image encoding; train the initial text adapter through the protected text encoding and the image encoding to obtain the first text adapter; train the first text adapter and the initial noise generation model through the protected text encoding and the image sample to obtain the trained noise generation model and the trained text adapter" to obtain the trained noise generation model and the trained text adapter.
[0053] See Figure 2 An image protection method shown below includes:
[0054] S200. Obtain the image to be protected.
[0055] S210. Obtain the protected text, image sample, trained text encoder, trained image encoder that matches the trained text encoder, initial text adapter, and initial noise generation model that matches the initial text adapter.
[0056] The protected text, image sample, trained text encoder, trained image encoder that matches the trained text encoder, initial text adapter, and initial noise generation model that matches the initial text adapter can all be specified by the user. Through several user-selectable parts, this customized part can not only provide exclusive protection for each user, but also enable the user to select the protection effect they are satisfied with, improving the visual effect of the protected image.
[0057] The image sample can be an image of the same style or the same author as the image to be protected, so that the output of the trained text adapter can maximize the semantic difference from the image to be protected, and the output of the trained noise generation model can minimize the image difference from the image to be protected.
[0058] The trained text encoder and the trained image encoder are a set of matching model groups. For example, the trained text encoder and the trained image encoder based on CLIP or BLIP. Multiple sets of matching trained text encoders and trained image encoders can be provided to the user in advance, and the user can select one of them. Multiple sets of matching initial text adapters and initial noise generation models can be determined by professional technicians based on experience or experiments, and this application does not make specific limitations. Specifically, multiple sets of matching trained text encoders and trained image encoders can be displayed through a display interface, and the user can be prompted to select from the displayed options. After obtaining the user's selection, the corresponding trained text encoder and trained image encoder are obtained.
[0059] The initial text adapter and the initial noise generation model that matches the initial text adapter are also a set of matching model groups. Multiple sets of matching initial text adapters and initial noise generation models can be provided to the user in advance, and the user can select one of them. Multiple sets of matching initial text adapters and initial noise generation models can be determined by professional technicians based on experience or experiments, and this application does not make specific limitations. Specifically, multiple sets of matching text adapters and noise generation models can be displayed through a display interface, and the user can be prompted to select from the displayed options. After obtaining the user's selection, the corresponding text adapter and noise generation model are obtained as the initial text adapter and initial noise generation model.
[0060] Optionally, if the protected text, image sample, initial text adapter, trained text encoder, and trained image encoder specified by the user are not obtained, jump to S250. That is, if the protected text, image sample, initial text adapter, and the matching trained text encoder and trained image encoder specified by the user are not obtained, directly obtain the protected text encoding and the trained text adapter from the user's local storage. It can be understood that if the user does not specify a new protected text, image sample, initial text adapter, and the matching trained text encoder and trained image encoder, directly obtain the protected text encoding and the trained text adapter from the user's local storage. If the user specifies at least one of the new protected text, image sample, initial text adapter, and the matching trained text encoder and trained image encoder, replace the original part with the newly specified part by the user, and train according to S220 - S240 to obtain the updated protected text encoding, trained text adapter, and trained noise generation model, and store the updated protected text encoding and trained text adapter in the user's local storage, and store the updated trained noise generation model on the server. It should be noted that since the initial text adapter and the initial noise generation model matching the initial text adapter are a set of matching model groups, when the initial text adapter remains unchanged, the initial noise generation model matching the initial text adapter also remains unchanged.
[0061] Exemplarily, for the paintings of a certain author, if the protected text content remains unchanged, it can be considered that the text and image to be protected by the user remain unchanged. Then, after one training, the protected text encoding and the trained text adapter stored in the user's local storage can be reused. When the protected text or the image to be protected changes, it is necessary to retrain the initial text adapter and the initial noise generation model matching the initial text adapter according to the protected text, image sample, initial text adapter, trained text encoder, and trained image encoder specified by the user to obtain the protected text encoding, trained text adapter, and trained noise generation model.
[0062] S220: Input the protected text into the trained text encoder to obtain the protected text encoding, and input the image sample into the trained image encoder to obtain the image encoding.
[0063] Input the protected text into the trained text encoder, and the trained text encoder encodes the protected text and outputs the protected text encoding. Input the image sample into the trained image encoder, and the trained image encoder encodes the image sample and outputs the image encoding.
[0064] S230. Train the initial text adapter using the protected text encoding and the image encoding to obtain a first text adapter.
[0065] Training the initial text adapter can involve inputting the protected text encoding into the initial text adapter, determining the semantic difference between the output and the image encoding through the loss function of the initial text adapter, and performing feedback learning based on maximizing the semantic difference between the output and the image encoding to adjust the model parameters of the initial text adapter, thereby achieving the training of the initial text adapter. This can improve the semantic difference between the output of the trained first text adapter and the image to be protected, increase the difficulty of cracking the protected image subsequently, and enhance the image protection effect. After the initial text adapter converges, use the initial text adapter at this time as the first text adapter.
[0066] Exemplarily, a corresponding loss function of the initial text adapter can be preset, input the protected text encoding into the initial text adapter, the loss function determines the loss value based on the semantic difference between the data processed by the initial text adapter and the image encoding, and performs feedback learning based on the loss value to update the model parameters of the initial text adapter. Through feedback learning using the loss function, the semantic difference between the data processed by the initial text adapter and the image encoding is maximized.
[0067] Train the initial text adapter using the protected text encoding and the image encoding. After the initial text adapter converges, obtain a first text adapter, improve the semantic difference between the output of the first text adapter and the image to be protected, avoid the protected image being parsed and then regenerated, and can enhance the protection effect on the image to be protected.
[0068] In an alternative embodiment, training the initial text adapter using the protected text encoding and the image encoding to obtain a first text adapter includes: inputting the protected text encoding into the initial text adapter to obtain an intermediate text condition; adjusting the model parameters of the initial text adapter through a first loss function to maximize the difference between the intermediate text condition and the image encoding; and obtaining the first text adapter after the initial text adapter converges.
[0069] The intermediate text condition is the output obtained after inputting the protected text encoding into the initial text adapter. The first loss function can be the loss function used during the training process of the initial text adapter, and is used to adjust the model parameters of the matching initial text adapter during the backpropagation process of the initial text adapter to maximize the semantic difference between the intermediate text condition and the image encoding. By using the first loss function to maximize the semantic difference between the intermediate text condition and the image encoding, the semantic difference between the protection noise of the trained noise generation model and the original image can be enhanced, the degree of change in the semantics of the input image to be protected can be increased, and the protection effect on the image to be protected can be improved.
[0070] Exemplarily, the first loss function can be shown as the following formula:
[0071]
[0072] where i' is the image encoding and t' is the output intermediate text condition.
[0073] According to the pre-defined model convergence condition, it is determined whether the trained initial text adapter converges. After the trained initial text adapter converges, the text adapter at this time is used as the first text adapter. The convergence condition can be that the first loss function is stable or the parameters of the model are stable. This application does not make specific limitations on this.
[0074] By inputting the protected text encoding into the initial text adapter, an intermediate text condition is obtained; the model parameters of the initial text adapter are adjusted through the first loss function to maximize the difference between the intermediate text condition and the image encoding; after the initial text adapter converges, a first text adapter is obtained, so that the first text adapter can maximize the change in the semantic information of the image to be protected and improve the protection effect.
[0075] S240. Train the first text adapter and the initial noise generation model through the protected text encoding and the image samples to obtain a trained noise generation model and a trained text adapter.
[0076] The output obtained by inputting the protected text encoding into the first text adapter can be used as the text noise for image generation in the initial noise generation model. The text noise and the image samples are input into the initial noise generation model. The image difference between the input and output of the initial noise generation model is determined through the loss function of the initial noise generation model, and based on minimizing the image difference between the input and output, feedback learning is performed to adjust the model parameters of the first text adapter and the initial noise generation model, realizing the matching training of the first text adapter and the initial noise generation model. In this way, the image difference between the protected noise generated by the trained noise generation model and the image to be protected can be reduced, and the quality of the protected image obtained by subsequently superimposing the protected noise and the image to be protected according to the preset ratio of the user can be improved. After the first text adapter and the initial noise generation model converge, a trained noise generation model and a trained text adapter are obtained.
[0077] S250. Obtain the protected text encoding stored locally by the user and the trained text adapter, and input the protected text encoding into the trained text adapter to obtain a text condition.
[0078] S260. Input the image to be protected and the text condition into the trained noise generation model stored in the server to obtain protected noise.
[0079] S270. Superimpose the protection noise and the image to be protected according to the preset ratio of the user to obtain the protected image.
[0080] For the technical solution of this embodiment, obtain the protection text specified by the user, the image sample, the trained text encoder, the trained image encoder matching the trained text encoder, the initial text adapter, and the initial noise generation model matching the initial text adapter. Having models and protection text related to the user's specification realizes the user's personalized customization, improves the personalization of the subsequent entire protection model, and increases the difficulty of deciphering the protected image obtained through the trained model; input the protection text into the trained text encoder to obtain the protection text encoding, and input the image sample into the trained image encoder to obtain the image encoding; train the initial text adapter through the protection text encoding and the image encoding to obtain the first text adapter; train the first text adapter and the initial noise generation model through the protection text encoding and the image sample to obtain the trained noise generation model and the trained text adapter. Since the training process is matched, it increases the difficulty of deciphering the protected image generated by the trained noise generation model and the trained text adapter, and improves the protection effect on the image to be protected.
[0081] Embodiment III
[0082] Figure 3 It is a flowchart of an image protection method provided by Embodiment III of this application. The technical solution of this embodiment is further refined on the basis of the above technical solution.
[0083] Further, training the first text adapter and the initial noise generation model through the protection text encoding and the image sample to obtain the trained noise generation model and the trained text adapter is refined as "input the protection text encoding into the first text adapter to obtain the training text condition; input the training text condition and the image sample into the initial noise generation model to train the initial noise generation model and the first text adapter; after the initial noise generation model converges, obtain the trained noise generation model and the trained text adapter", so as to obtain the trained noise generation model and the trained text adapter.
[0084] See Figure 3 An image protection method shown, including:
[0085] S300. Obtain the image to be protected.
[0086] S310. Obtain the protected text specified by the user, the image sample, the trained text encoder, the trained image encoder that matches the trained text encoder, the initial text adapter, and the initial noise generation model that matches the initial text adapter.
[0087] S320. Input the protected text into the trained text encoder to obtain the protected text encoding, and input the image sample into the trained image encoder to obtain the image encoding.
[0088] S330. Train the initial text adapter with the protected text encoding and the image encoding to obtain the first text adapter.
[0089] S340. Input the protected text encoding into the first text adapter to obtain the training text condition.
[0090] Input the protected text into the first text adapter, and the output is the training text condition. The training text condition can be used as an input during the training process of the initial noise generation model to train the initial noise generation model.
[0091] S350. Input the training text condition and the image sample into the initial noise generation model to train the initial noise generation model and the first text adapter.
[0092] The image sample can be an image used to train the initial noise generation model. Exemplarily, the image sample can be provided by the user. The training process of the initial noise generation model and the first text adapter can be as follows: Set the loss function based on the training purpose, input the training text condition and the image sample into the initial noise generation model, perform feedback learning based on the loss function, and update the model parameters of the matching first text adapter and initial noise generation model to achieve the training of the initial noise generation model and the first text adapter.
[0093] In an optional embodiment, inputting the training text condition and the image sample into the initial noise generation model to train the initial noise generation model and the first text adapter includes: Inputting the training text condition and the image sample into the initial noise generation model to obtain the training noise; Adjusting the model parameters of the matching first text adapter and initial noise generation model through the second loss function to minimize the difference between the input image sample and the corresponding training noise, so as to achieve the training of the initial noise generation model.
[0094] The training noise is the output obtained after inputting the training text condition and the image sample into the initial noise generation model. The second loss function can be the loss function used in the training process of the initial noise generation model, and is used to adjust the model parameters of the matching first text adapter and the initial noise generation model during the backpropagation process, so as to minimize the difference between the input image sample and the corresponding training noise. By minimizing the semantic difference between the input image sample and the corresponding training noise through the second loss function, the protection noise output by the trained noise generation model can be reduced, the degree of change in the semantics of the input image to be protected can be reduced, and the protection effect on the image to be protected can be improved. Exemplarily, the second loss function of the initial noise generation model can be a regression loss function. For example, the second loss function of the initial noise generation model can adopt the Mean Square Error (MSE).
[0095] The second loss function can be shown as the following formula:
[0096]
[0097] Where, i is the input, p is the output, k represents the position where the pixels and channels in the output match those in the input, and n represents the size of the input. For example, n can be 224*224*3.
[0098] By inputting the training text condition and the image sample into the initial noise generation model, training noise is obtained; by adjusting the model parameters of the matching first text adapter and the initial noise generation model through the second loss function, the difference between the input image sample and the corresponding training noise is minimized, so as to realize the training of the initial noise generation model. By minimizing the difference between the image sample input to the initial noise generation model and the corresponding training noise, a kind of noise that hardly affects visual perception can be imposed on the original image to change the semantic information of the image, accurately reproduce the style of the original image, and improve the quality of the protected image.
[0099] S360. After the initial noise generation model converges, a trained noise generation model and a trained text adapter are obtained.
[0100] According to the predefined model convergence condition, it is judged whether the trained initial noise generation model converges. After the trained initial noise generation model converges, the model at this time is used as the trained noise generation model, and the trained noise generation model is stored in the server. The convergence condition can be that the second loss function is stable or the parameters of the model are stable, and the present application does not make specific limitations on this.
[0101] In an optional embodiment, after obtaining the trained noise generation model and the trained text adapter, the method further includes: storing the trained noise generation model in the server, and storing the protected text encoding and the trained text adapter locally on the user side.
[0102] The training of the first text adapter and the initial noise generation model that match each other is performed on the server, reducing the hardware requirements for the client. After obtaining the trained noise generation model and the trained text adapter, the protected text encoding and the trained text adapter can be sent to the user side for storage through the user identifier, and the trained noise generation model can be identified through the user identifier to achieve the matching between the trained noise generation model and the trained text adapter.
[0103] By storing the trained noise generation model in the server and storing the protected text encoding and the trained text adapter locally on the user side, off-site storage of the trained noise generation model, the protected text encoding, and the trained text adapter is realized, increasing the difficulty of cracking after the model is stolen and improving the security of image protection.
[0104] The image protection method proposed by the present invention is an image semantic protection method that can be stored off-site. This method changes the semantic information of the image by applying a kind of noise that hardly affects visual perception to the original image, thereby effectively preventing unauthorized users from reproducing the image based on some general image generation models. Further, the general image generation model cannot accurately reproduce the style of the original image, so as to achieve the purpose of protecting the copyright and creative rights of the image content producer. Specifically, in this application, a lightweight trained text adapter is stored locally on the user side; this adapter can only generate a protected picture for the user's work in combination with the remaining model parts stored online, enabling each user to construct and privately own a protection model online. This design makes the protected model more difficult to be obtained and cracked by potential infringers.
[0105] S370: Obtain the protected text encoding and the trained text adapter stored locally on the user side, and input the protected text encoding into the trained text adapter to obtain a text condition.
[0106] S380: Input the image to be protected and the text condition into the trained noise generation model stored in the server to obtain protected noise.
[0107] S390: Superimpose the protected noise and the image to be protected according to the preset ratio of the user to obtain the protected image.
[0108] In the technical solution of this embodiment, by inputting the protected text encoding into the first text adapter, a training text condition is obtained; the training text condition and the image sample are input into the initial noise generation model to train the initial noise generation model and the first text adapter; after the initial noise generation model converges, a trained noise generation model and a trained text adapter are obtained. Since the first text adapter and the initial noise generation model are trained simultaneously, the trained text adapter and the trained noise generation model are matched, improving the protection effect of the protected noise output by the trained noise generation model, increasing the cracking difficulty of the model, and improving the protection effect of image protection.
[0109] Embodiment 4
[0110] Figure 4 The following is a schematic structural diagram of an image protection device provided in Embodiment 4 of the present application. This embodiment is applicable to the situation of protecting images. The specific structure of the image protection device is as follows:
[0111] The to-be-protected image acquisition module 400 is used to acquire the to-be-protected image;
[0112] The text condition acquisition module 410 is used to acquire the protected text encoding stored locally by the user and the trained text adapter, and input the protected text encoding into the trained text adapter to obtain the text condition;
[0113] The protected noise acquisition module 420 is used to input the to-be-protected image and the text condition into the trained noise generation model stored in the server to obtain the protected noise; the trained noise generation model and the trained text adapter are a pair of model groups after matching training;
[0114] The protected noise superposition module 430 is used to superpose the protected noise and the to-be-protected image according to the preset ratio of the user to obtain the protected image.
[0115] The technical solution of this embodiment is to obtain the image to be protected; obtain the protected text encoding and the trained text adapter stored locally by the user, input the protected text encoding into the trained text adapter to obtain the text condition; the text condition can change the semantic features of the image to be protected and increase the difficulty of being cracked; input the image to be protected and the text condition into the trained noise generation model stored in the server to obtain the protection noise; by storing the trained noise generation model and the trained text adapter remotely, the difficulty of cracking the model group is increased, and the protection effect on the image to be protected is improved. At the same time, since the trained noise generation model and the trained text adapter are a pair of model groups after matching training, the user can have a private set of protection models, providing proprietary image protection for the user, further improving the protection effect and cracking difficulty of the model group; superimpose the protection noise and the image to be protected according to the preset ratio of the user to obtain the protected image. The user can adjust the relationship between the protection intensity and the image quality through the preset ratio, select the satisfactory protection effect, and improve the user experience. Therefore, through the technical solution of this application, the problems that the watermark affects the display effect of the image, the watermark is easily cracked, and the protection effect on the image is poor are solved, and the protection effect on the image and the quality of the protected image are improved.
[0116] Optionally, the image protection device further includes:
[0117] A user-specified data acquisition module, configured to acquire the protected text, image samples, trained text encoder, trained image encoder matching the trained text encoder, initial text adapter, and initial noise generation model matching the initial text adapter specified by the user;
[0118] A data encoding module, configured to input the protected text into the trained text encoder to obtain the protected text encoding, and input the image samples into the trained image encoder to obtain the image encoding;
[0119] A first text adapter determination module, configured to train the initial text adapter through the protected text encoding and the image encoding to obtain a first text adapter;
[0120] An image protection model training module, configured to train the first text adapter and the initial noise generation model through the protected text encoding and the image samples to obtain a trained noise generation model and a trained text adapter.
[0121] Optionally, the first text adapter determination module includes:
[0122] An intermediate text condition determination unit, configured to input the protected text encoding into the initial text adapter to obtain an intermediate text condition;
[0123] A model parameter adjustment unit for adjusting the model parameters of the initial text adapter through a first loss function to maximize the difference between the intermediate text condition and the image encoding;
[0124] A first text adapter determination unit for obtaining a first text adapter after the initial text adapter converges.
[0125] Optionally, the image protection model training module includes:
[0126] A training text condition determination unit for inputting the protected text encoding into the first text adapter to obtain a training text condition;
[0127] A model training unit for inputting the training text condition and the image sample into the initial noise generation model to train the initial noise generation model and the first text adapter;
[0128] A trained model determination unit for obtaining a trained noise generation model and a trained text adapter after the initial noise generation model converges.
[0129] Optionally, the image protection device further includes:
[0130] A model storage module for storing the trained noise generation model in the server and storing the protected text encoding and the trained text adapter locally on the user side.
[0131] Optionally, the model training module includes:
[0132] A training noise determination unit for inputting the training text condition and the image sample into the initial noise generation model to obtain training noise;
[0133] A model parameter adjustment unit for adjusting the model parameters of the matching first text adapter and the initial noise generation model through a second loss function to minimize the difference between the input image sample and the corresponding training noise, so as to train the initial noise generation model.
[0134] The image protection device provided by the embodiments of the present application can execute the image protection method provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the image protection method.
[0135] According to the embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.
[0136] Embodiment 5
[0137] Figure 5 It is a schematic structural diagram of an electronic device provided by Embodiment 5 of the present application, asFigure 5 As shown, the electronic device includes a processor 500, a memory 510, an input device 520, and an output device 530; the number of processors 500 in the electronic device may be one or more, Figure 5 and one processor 500 is taken as an example herein; the processor 500, the memory 510, the input device 520, and the output device 530 in the electronic device may be connected through a bus or other means, Figure 5 and connection through a bus is taken as an example herein.
[0138] The memory 510, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the image protection method in the embodiments of the present application (for example, the to-be-protected image acquisition module 400, the text condition acquisition module 410, the protection noise acquisition module 420, and the protection noise superposition module 430). The processor 500 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 510, that is, implements the above-mentioned image protection method.
[0139] The memory 510 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 510 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 510 may further include a memory remotely set relative to the processor 500, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0140] The input device 520 can be used to receive input character information and generate key signal inputs related to the user settings and function controls of the electronic device. The output device 530 may include a display device such as a display screen.
[0141] Embodiment Six
[0142] Embodiment 6 of the present application also provides a storage medium containing computer-executable instructions. The computer-executable instructions are used to execute an image protection method when executed by a computer processor. The method includes: obtaining an image to be protected; obtaining a protected text encoding and a trained text adapter stored locally by the user, inputting the protected text encoding into the trained text adapter to obtain a text condition; inputting the image to be protected and the text condition into a trained noise generation model stored in the server to obtain protected noise; the trained noise generation model and the trained text adapter are a pair of model groups that have been matched and trained; superimposing the protected noise and the image to be protected according to a preset ratio of the user to obtain a protected image.
[0143] Of course, for a storage medium containing computer-executable instructions provided by the embodiments of the present application, the computer-executable instructions are not limited to the method operations described above, and can also execute relevant operations in the image protection method provided by any embodiment of the present application.
[0144] Through the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and necessary general hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk, or optical disc of a computer, etc., including several instructions to enable an electronic device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0145] It should be noted that in the embodiments of the above image protection device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present application.
[0146] Note that the above is only the preferred embodiment of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, more other equivalent embodiments can be included, and the scope of the present application is determined by the scope of the appended claims.
Claims
1. An image protection method, characterized in that, Including: Obtain the image to be protected; Obtain the protected text encoding and the trained text adapter stored locally by the user, and input the protected text encoding into the trained text adapter to obtain a text condition; Input the image to be protected and the text condition into the trained noise generation model stored in the server to obtain protected noise; the trained noise generation model and the trained text adapter are a pair of model groups that have been matched and trained; Superimpose the protected noise and the image to be protected according to the preset ratio of the user to obtain the protected image.
2. The method according to claim 1, wherein Before obtaining the protected text encoding and the trained text adapter stored locally by the user, it further includes: Obtain the protected text specified by the user, the image sample, the trained text encoder, the trained image encoder matched with the trained text encoder, the initial text adapter, and the initial noise generation model matched with the initial text adapter; Input the protected text into the trained text encoder to obtain the protected text encoding, and input the image sample into the trained image encoder to obtain the image encoding; Train the initial text adapter through the protected text encoding and the image encoding to obtain the first text adapter; Train the first text adapter and the initial noise generation model through the protected text encoding and the image sample to obtain the trained noise generation model and the trained text adapter.
3. The method according to claim 2, characterized in that, The training of the initial text adapter through the protected text encoding and the image encoding to obtain the first text adapter includes: Input the protected text encoding into the initial text adapter to obtain an intermediate text condition; Adjust the model parameters of the initial text adapter through the first loss function to maximize the difference between the intermediate text condition and the image encoding; After the initial text adapter converges, obtain the first text adapter.
4. The method according to claim 2, wherein The training of the first text adapter and the initial noise generation model through the protected text encoding and the image sample to obtain the trained noise generation model and the trained text adapter includes: Input the protected text encoding into the first text adapter to obtain the training text condition; Input the training text condition and the image sample into the initial noise generation model to train the initial noise generation model and the first text adapter; After the initial noise generation model converges, obtain the trained noise generation model and the trained text adapter.
5. The method according to claim 4, wherein After obtaining the trained noise generation model and the trained text adapter, it further includes: Store the trained noise generation model in the server, and store the protected text encoding and the trained text adapter locally by the user.
6. The method according to claim 4, wherein The input of the training text condition and the image sample into the initial noise generation model to train the initial noise generation model and the first text adapter includes: Input the training text condition and the image sample into the initial noise generation model to obtain training noise; Adjust the model parameters of the first text adapter and the initial noise generation model that match through a second loss function to minimize the difference between the input image sample and the corresponding training noise, so as to realize the training of the initial noise generation model.
7. An image protection device, characterized in that, It includes: A to-be-protected image acquisition module, configured to acquire a to-be-protected image; A text condition acquisition module, configured to acquire a protected text encoding stored locally by the user and a trained text adapter, and input the protected text encoding into the trained text adapter to obtain a text condition; A protected noise acquisition module, configured to input the to-be-protected image and the text condition into a trained noise generation model stored in the server to obtain protected noise; the trained noise generation model and the trained text adapter are a pair of model groups after being matched and trained; A protected noise superposition module, configured to superpose the protected noise and the to-be-protected image according to a preset ratio of the user to obtain a protected image.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the image protection method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the image protection method according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the image protection method according to any one of claims 1-6.