Image processing method and apparatus, device and storage medium
The original image is subjected to noise addition and denoising through the diffusion model, and the disturbance value is calculated in combination with the noise prediction loss, and the image is disturbed to generate anti-edited images. This solves the problem of poor image editing protection effect in the prior art and improves the difficulty of image feature learning and editing.
Patent Information
- Application Number
- PCT/CN2024/114614
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-07
- Filing Date
- 2024-08-26
- Publication Date
- 2025-06-26
AI Technical Summary
In the prior art, when preventing pictures from being edited, the style converter cannot effectively represent the picture style, resulting in the possibility of the picture being edited cannot be effectively reduced, and the method is only applicable to specific scenarios.
The original image is added and denoised by the forward and reverse diffusion network of the diffusion model, the noise prediction loss between the predicted noise and the sampling noise is determined, the perturbation value is calculated, and the original image is disturbed to generate an anti-edited image.
The diffusion model learns image features in anti-edited images and increases the editing difficulty of anti-edited images, making it difficult for the diffusion model to accurately extract image features.
Smart Images

Figure CN2024114614_26062025_PF_FP_ABST
Abstract
Description
Image processing method, device, equipment and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 7, 2023, application number 202311468616.7, and application name “Image processing method, device, equipment and storage medium”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of image processing technology, and in particular to an image processing method, apparatus, device, and storage medium. Background Art
[0003] With the emergence of many open-source large-scale text-image models, the threshold for editing images posted by users on the Internet has gradually become lower. Therefore, in order to protect images from being edited arbitrarily, it is necessary to reduce the possibility of image tampering through anti-editing methods.
[0004] In related technologies, a style converter is used to maximize the gradient of the style converter so that the latent space representing the style of the original image is different from the styles of other images, thereby preventing the image style from being edited.
[0005] However, if the style converter cannot well represent the style of the image, it cannot effectively reduce the possibility of the image being edited. It can be seen that the method of using the style converter to prevent the image from being edited can only be applied to specific scenarios and cannot effectively reduce the possibility of the image being edited.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide an image processing method, apparatus, device, and storage medium that can increase the difficulty of learning image features in an anti-editing image using a diffusion model, and increase the difficulty of editing an anti-editing image using a diffusion model. The technical solution is as follows:
[0008] In one aspect, an embodiment of the present application provides an image processing method, which is executed by a computer device and includes:
[0009] Get the original image;
[0010] According to the sampling noise, the original image is subjected to noise processing by a forward diffusion network of a diffusion model to obtain a noisy image;
[0011] performing denoising on the noisy image through a reverse diffusion network of the diffusion model, and determining predicted noise based on the denoised image obtained by the denoising;
[0012] determining a disturbance value according to a noise prediction loss between the predicted noise and the sampled noise;
[0013] The original image is disturbed according to the disturbance value to obtain an anti-editing image.
[0014] On the other hand, an embodiment of the present application provides an image processing apparatus, which is deployed on a computer device and includes:
[0015] An acquisition module, used to acquire the original image;
[0016] A noise processing module is used to perform noise processing on the original image according to the sampling noise through a forward diffusion network of a diffusion model to obtain a noisy image;
[0017] a noise prediction module, configured to perform denoising on the noisy image using a reverse diffusion network of the diffusion model, and determine predicted noise based on the denoised image obtained by the denoising process;
[0018] a disturbance determination module, configured to determine a disturbance value according to a noise prediction loss between the predicted noise and the sampled noise;
[0019] The first disturbance processing module is used to perform disturbance processing on the original image according to the disturbance value to obtain an anti-editing image.
[0020] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory; the memory stores at least one instruction, and the at least one instruction is used to be executed by the processor to implement the image processing method described in the above aspect.
[0021] On the other hand, an embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the image processing method described in the above aspects.
[0022] In another aspect, embodiments of the present application provide a computer program product comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image processing method provided in various optional implementations of the aforementioned aspects.
[0023] In an embodiment of the present application, after acquiring an original image, the original image is first subjected to noise processing by a forward diffusion network of a diffusion model based on the sampling noise to obtain a noisy image. The noisy image is then denoised by a backward diffusion network of the diffusion model, and the predicted noise is determined based on the denoised image obtained by the denoising process. The process of denoising and denoising the original image by the diffusion model is essentially a process of learning image features from the original image. The noise prediction loss between the predicted noise and the sampling noise can reflect the uncertainty of the diffusion model in its current state regarding the noise. Therefore, a perturbation value can be determined based on the noise prediction loss between the predicted noise and the sampling noise. This perturbation value is intended to simulate or amplify the uncertainty of the diffusion model in predicting noise. In this way, the original image can be perturbated based on the perturbation value to obtain an anti-editing image, thereby introducing additional, unpredictable noise into the anti-editing image. That is, each pixel in the anti-editing image is subjected to signal interference, thereby increasing the difficulty of the diffusion model in learning the image features of the anti-editing image. Accordingly, when using the trained diffusion model to edit the anti-editing image, since each pixel in the anti-editing image is subject to signal interference, the diffusion model cannot accurately extract the image features in the anti-editing image, thereby increasing the difficulty of editing the anti-editing image using the diffusion model. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0025] FIG1 shows a schematic structural diagram of a diffusion model provided by an exemplary embodiment of the present application;
[0026] FIG2 shows a schematic diagram of the LoRA structure provided by an exemplary embodiment of the present application;
[0027] FIG3 shows an image schematic diagram for LoRA fine-tuning provided by an exemplary embodiment of the present application;
[0028] FIG4 is a schematic diagram showing the result of image editing using a diffusion model fine-tuned by LoRA, provided by an exemplary embodiment of the present application;
[0029] FIG5 is a schematic diagram showing an implementation environment provided by an exemplary embodiment of the present application;
[0030] FIG6 shows a flowchart of an image processing method provided by an exemplary embodiment of the present application;
[0031] FIG7 shows a flowchart of an image processing method provided by another exemplary embodiment of the present application;
[0032] FIG8 shows a comparison of visual effects between an anti-editing image and an original image provided by an exemplary embodiment of the present application;
[0033] FIG9 is a schematic diagram showing the results of learning and editing an anti-editing image using a diffusion model provided by an exemplary embodiment of the present application;
[0034] FIG10 is a schematic diagram showing a process of obtaining an anti-editing image through M rounds of perturbation processing according to an exemplary embodiment of the present application;
[0035] FIG11 shows a structural block diagram of an image processing apparatus provided by an exemplary embodiment of the present application;
[0036] FIG12 shows a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0038] In related technologies, in order to implement image editing, a computer device usually trains a diffusion model using an image set having target image features, so that the diffusion model can use the learned target image features to edit other images.
[0039] Optionally, the diffusion model may include a text-to-image latent space model (SD, Stable Diffusion), a text-to-image cascade model (Deep Floyd IF), etc.
[0040] In some embodiments, the diffusion model can learn image features by adding noise and denoising the image, so the diffusion model needs to implement the adding noise and denoising of the image through a forward process and a reverse process, that is, the diffusion model adds noise to the image through a forward diffusion network and denoises the noisy image through a reverse diffusion network.
[0041] Optionally, the process of adding noise to the image by the diffusion model can be expressed as dx t =f(t)x t dt+g(t)dw, the denoising process can be expressed as Where w represents the standard Wiener process (also known as Brownian motion), f(t) is x t Drift Coefficient, g(T) is xt The diffusion coefficient, x t is the image sample corresponding to time t, represents the standard Wiener process when time flows backward from T to 0, q(x t ) is x t The data distribution, Represents the data distribution score.
[0042] In the process of adding noise to the image, the diffusion time is a continuous time variable t∈[0, T]. The data distribution corresponding to time 0 can be expressed as x0~q(x0), and the data distribution corresponding to time T can be expressed as x T ~q(x T In order to learn the image features, the computer device needs to sample the real samples x0~q(x0) through the discretization inverse process.
[0043] Schematically, as shown in FIG1 , the computer device may first perform noise processing on the image through the forward diffusion network 102 in the diffusion model 101 to obtain a noisy image, and then perform denoising processing on the noisy image through the backward diffusion network 103 .
[0044] In one possible implementation, the computer device can use a denoising score matching method to predict the network To estimate, in this case, the objective function Among them, s θ In order to predict the network, in the process of training the diffusion model based on the objective function, when the objective function converges to the optimal point, it satisfies The computer equipment thus realizes the optimized training of the diffusion model.
[0045] To apply the diffusion model to generate images for a specific domain, the computer device needs to further fine-tune the diffusion model. In some embodiments, the computer device can fine-tune the diffusion model using LoRA (Low-Rank Adaptation), a low-cost fine-tuning structure for large language models.
[0046] Optionally, when fine-tuning all parameters of the diffusion model, LoRA can fit the image data set by converting the W matrix, which originally requires a large amount of parameter adjustment, into two small matrices A and B, and according to the formula f(X)=W′X+ΔW′X=W′X+(A′B)′X.
[0047] Schematically, as shown in Figure 2, A is a matrix mapped from dimension d to dimension r, and B is a matrix mapped from dimension r to dimension d, wherein the computer device can initialize matrix A through random Gaussian distribution, initialize matrix B through zero matrix, and only train matrix A and matrix B during training, so that after the training is completed, matrix B is multiplied by matrix A and the pre-trained model parameters are merged as the fine-tuned model parameters.
[0048] In one possible implementation, during the fine-tuning phase of the diffusion model, the computer device can train the diffusion model using a small number of domain-specific image sets. Optionally, the model training objective function during the fine-tuning phase can be expressed as Among them, c represents the text description corresponding to the image.
[0049] Schematically, as shown in FIG1 , a computer device can use a neural network module to predict the noise of a denoised image in the denoising process according to a simple description text corresponding to the image, thereby training a diffusion model based on the predicted noise and the objective function.
[0050] Schematically, Figure 3 shows a set of images used for LoRA fine-tuning. After fine-tuning the diffusion model using the image set shown in Figure 3, the computer device can use the image features learned by the diffusion model to edit the image shown in 401 in Figure 4, thereby obtaining the image shown in 402 in Figure 4, which includes the generated image output by the diffusion model fine-tuned based on different fine-tuning weights.
[0051] In an embodiment of the present application, in order to increase the difficulty of learning the image features of the original image by the diffusion model and prevent the original image from being learned and edited by the diffusion model, in the process of denoising and noise-adding processing of the original image using the diffusion model, the predicted noise is obtained by performing noise prediction on the denoised image, and then the disturbance value is determined according to the noise prediction loss between the predicted noise and the sampling noise, and then the original image is disturbed with the disturbance value, that is, the anti-editing image corresponding to the original image can be obtained, so that it is difficult for the diffusion model to learn the image features from the anti-editing image, thereby increasing the difficulty of the diffusion model in editing the anti-editing image.
[0052] In a possible implementation, considering that in order to prevent the image features of the original image from being learned and edited by the diffusion model in the embodiment of the present application, a disturbance value is determined according to the noise prediction loss in the process of fine-tuning the diffusion model using the original image, and the original image is subjected to reverse disturbance processing according to the disturbance value, thereby obtaining an anti-editing image. Therefore, the objective function of the process can be determined as Among them, δ is the disturbance value, s θis the prediction network, t is the sampling time, c is the text description corresponding to the image, x0 is the original image, x t ′ is the noise image corresponding to the sampling time, λ t is the adjustment coefficient, ∈ is Gaussian noise.
[0053] That is, by fine-tuning the diffusion model so that J DSM During the minimization process, the embodiment of the present application needs to determine the maximum disturbance value δ, so as to perform disturbance processing on the original image with the disturbance value, thereby increasing the anti-editing effectiveness of the anti-editing image.
[0054] In a possible implementation, considering that different fine-tuning forms may be used in the process of fine-tuning the diffusion model, that is, different prediction network s may be generated after different fine-tuning forms. θ Therefore, in order to improve the efficiency of determining the disturbance value, the objective function can also be approximated, that is, assuming that the diffusion model has been optimized, the objective function can be approximately expressed as in, It represents the prediction network in the diffusion model after pre-training.
[0055] Optionally, the prediction network in the diffusion model can be a score network, a noise network, or a v-network. Different prediction networks depict the same information during the training objective function, and all depict the direction in which the probability density of data in different dimensions grows fastest. The noise network and the score network satisfy That is, the output data of the noise network and the score network can be converted to each other based on this linear relationship.
[0056] Please refer to Figure 5, which shows a schematic diagram of an implementation environment provided by one embodiment of the present application. The implementation environment includes a terminal 520 and a server 540. The terminal 520 and the server 540 communicate data via a communication network. Optionally, the communication network can be a wired network or a wireless network, and the communication network can be at least one of a local area network, a metropolitan area network, and a wide area network.
[0057] Terminal 520 is an electronic device installed with an application program having an image processing function. The image processing function can be a function of a native application in the terminal or a function of a third-party application. The electronic device can be a smartphone, tablet computer, personal computer, wearable device, or vehicle-mounted terminal, etc. FIG. 5 illustrates terminal 520 as a personal computer, but this is not intended to be limiting.
[0058] Server 540 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. In the embodiment of the present application, server 540 can be a backend server for an application with image processing capabilities.
[0059] In one possible implementation, as shown in FIG5 , data exchange occurs between a server 540 and a terminal 520. After the terminal 520 acquires the original image, the terminal 520 transmits the original image to the server 540. The server 540 then performs noise processing on the original image using a forward diffusion network of a diffusion model based on the sampling noise to obtain a noisy image. The server then performs denoising processing on the noisy image using a backward diffusion network of the diffusion model. The server then performs noise prediction on the denoised image obtained by the denoising process to obtain predicted noise. Furthermore, a disturbance value is determined based on the noise prediction loss between the sampling noise and the predicted noise. The disturbance value is then transmitted to the terminal 520, which then performs disturbance processing on the original image based on the disturbance value to obtain an anti-editing image.
[0060] Please refer to FIG6 , which shows a flowchart of an image processing method provided by an exemplary embodiment of the present application. This embodiment uses the method applied to a computer device (including a terminal 520 and / or a server 540) as an example for explanation. The method includes the following steps:
[0061] Step 601: Acquire the original image.
[0062] Optionally, the original image refers to a two-dimensional digital matrix composed of pixels.
[0063] Optionally, the original image may be red-green-blue (RGB) image data, that is, each pixel corresponds to three color channels; the original image may also be image data obtained by merging the channels of an RGB image and normalizing the image to a floating-point number between 0 and 1.
[0064] Step 602 : Based on the sampling noise, the original image is subjected to noise addition processing by a forward diffusion network of a diffusion model to obtain a noisy image.
[0065] In some embodiments, after acquiring the original image, the computer device may perform noise addition processing on the original image through a forward diffusion network of a diffusion model according to the sampling noise, thereby obtaining a noisy image.
[0066] In one possible implementation, a computer device encodes an original image through a codec module in a diffusion model and compresses it from a pixel space to a latent space, thereby performing noise processing on the original image based on the sampling noise using a forward diffusion network to obtain a noisy image.
[0067] Optionally, sampling noise refers to a signal interference applied to the original image, which may cause changes in image information or pixel brightness in the original image. Optionally, sampling noise may be various types of noise, such as Gaussian noise, salt and pepper noise, etc. When the sampling noise is Gaussian noise, each pixel in the noisy image is a pixel to which noise is applied.
[0068] Optionally, the noisy image is an image to which signal interference is applied, and compared with the original image, image information or pixel brightness in the noisy image has changed.
[0069] Optionally, the process of adding noise to the original image through the forward diffusion network is the process of applying signal interference to each pixel in the original image using the sampling noise, wherein the original image can be represented as x0, the sampling noise can be represented as ∈, and the diffusion time can be represented as T, so that the computer device applies noise to each pixel in the original image based on the sampling noise within the continuous diffusion time T through the forward diffusion network, and the noisy image x can be obtained. T , and the noisy image x T It can be a pure noise image.
[0070] Step 603: De-noising the noisy image using the inverse diffusion network of the diffusion model, and determining predicted noise based on the de-noised image obtained by the de-noising process.
[0071] In some embodiments, after the original image is denoised to obtain a noisy image, the computer device may further utilize a reverse diffusion network within the diffusion model to denoise the noisy image, thereby obtaining a denoised image, in order to enable the diffusion model to learn image features within the original image. Optionally, the reverse diffusion network may be a U-net network.
[0072] Optionally, the process of denoising the noisy image by using the back diffusion network is the process of removing the signal interference applied to each pixel in the noisy image. In this process, in order to restore the original image, the back diffusion network needs to learn the image features and predict the noise applied to the noisy image to achieve denoising of the noisy image. The diffusion time in the denoising process is also T, that is, from x T Backward diffusion process to x0.
[0073] In some embodiments, during the process of reverse denoising a noisy image, in order to ensure that the denoised image and the original image have the same data distribution, the computer device also needs to perform noise prediction on the denoised image through a neural network during the denoising process for subsequent loss calculation.
[0074] Optionally, the predicted noise refers to noise data obtained by predicting signal interference applied to the noisy image. The prediction process can be implemented through a prediction network during the process of denoising the noisy image.
[0075] In one possible implementation, the computer device may predict the denoised image using a prediction network to obtain predicted noise. Optionally, the prediction network may be a score network, a noise network, or a V-network, etc., which is not limited in this embodiment of the present application.
[0076] Step 604: Determine a disturbance value according to the noise prediction loss between the predicted noise and the sampled noise.
[0077] In some embodiments, since the process of adding noise and denoising the original image by the diffusion model is essentially a process of learning the image features of the original image, the computer device can calculate the noise prediction loss based on the predicted noise and the sampling noise, so that based on the noise prediction loss, the diffusion model can fully learn the image features of the original image.
[0078] In an embodiment of the present application, in order to prevent the image features of the original image from being learned by the diffusion model, the computer device needs to determine the disturbance value based on the noise prediction loss after determining the noise prediction loss, and then use the disturbance value to make it difficult for the diffusion model to learn the image features of the original image.
[0079] Optionally, the noise prediction loss refers to the result of performing a norm calculation on a noise difference between the predicted noise and the sampled noise, where the noise difference is the difference between the actual signal interference value applied to each pixel in the image and the predicted signal interference value.
[0080] Optionally, the disturbance value is used to indicate the adjustment amount of the RGB value of each pixel in the original image, and can be a disturbance signal applied to each pixel in the original image. Based on the disturbance signal, the image features represented by each pixel in the original image can be protected.
[0081] Optionally, in the process of training the diffusion model based on the noise prediction loss, the smaller the noise prediction loss, the more effective the training of the diffusion model, and the easier it is for the diffusion model to learn image features. In order to prevent the diffusion model from learning the image features of the original image, or increase the difficulty of the diffusion model in learning the features of the original image, the computer device needs to reversely determine the disturbance value applied to the original image based on the noise prediction loss, wherein the noise prediction loss and the disturbance value are positively correlated, and the larger the noise prediction loss, the larger the disturbance value.
[0082] Step 605 : Perform disturbance processing on the original image according to the disturbance value to obtain an anti-editing image.
[0083] In some embodiments, after determining the disturbance value corresponding to the original image, the computer device can perform disturbance processing on the original image according to the disturbance value to obtain an anti-editing image, thereby making it difficult for the diffusion model to learn the image features in the original image through the anti-editing image.
[0084] Optionally, the disturbance value has the same data expression form as the original image, and the disturbance value is in matrix form. Each disturbance value in the matrix corresponds to each pixel point in the image, that is, the process of perturbing the original image according to the disturbance value is the process of applying disturbance to each pixel point of the original image.
[0085] Optionally, the process of perturbation of the original image is the process of applying a perturbation signal to each pixel in the original image, and the perturbation values corresponding to different pixels may be different. Thus, when the diffusion model uses noise addition and denoising to learn the image features of the anti-editing image, the perturbation value applied to each pixel in the anti-editing image will affect the noise prediction result of the diffusion model, thereby increasing the difficulty of the diffusion model in learning the image features of the anti-editing image.
[0086] In summary, in the embodiment of the present application, after acquiring the original image, the original image is first subjected to noise processing by the forward diffusion network of the diffusion model according to the sampling noise to obtain a noisy image, and then the noisy image is denoised by the reverse diffusion network of the diffusion model, and the predicted noise is determined based on the denoised image obtained by the denoising process. The process of the diffusion model performing noise processing and denoising on the original image is essentially a process of learning the image features of the original image, and the noise prediction loss between the predicted noise and the sampling noise can reflect the uncertainty of the diffusion model in the current state of the noise. Therefore, the perturbation value can be determined based on the noise prediction loss between the predicted noise and the sampling noise. The perturbation value is intended to simulate or amplify the uncertainty of the diffusion model in predicting noise. In this way, the original image can be perturbated according to the perturbation value to obtain an anti-editing image, thereby introducing additional, unpredictable noise into the anti-editing image, that is, each pixel in the anti-editing image is subjected to signal interference, thereby increasing the difficulty of the diffusion model in learning the image features of the anti-editing image. Accordingly, when using the trained diffusion model to edit the anti-editing image, since each pixel in the anti-editing image is subject to signal interference, the diffusion model cannot accurately extract the image features in the anti-editing image, thereby increasing the difficulty of editing the anti-editing image using the diffusion model.
[0087] In some embodiments, in order to improve the accuracy of determining the disturbance value, the computer device may also determine the sampling moment by randomly and uniformly sampling time within the diffusion duration, thereby performing noise prediction based on the denoised image corresponding to the sampling moment, and further determining the disturbance value.
[0088] Please refer to FIG7 , which shows a flowchart of an image processing method provided by an exemplary embodiment of the present application. This embodiment uses the method applied to a computer device (including a terminal 520 and / or a server 540) as an example for description. The method includes the following steps:
[0089] Step 701: Acquire an original image.
[0090] Step 702: Sample the standard Gaussian noise to obtain sampled noise.
[0091] In a possible implementation, the computer device performs noise sampling on noise that satisfies a standard Gaussian distribution, thereby obtaining sampled noise, namely, Gaussian noise.
[0092] Alternatively, the noise sampling process can be expressed as Where ∈ is Gaussian noise.
[0093] Step 703 : Noise the original image using a forward diffusion network of a diffusion model according to the sampling noise and the diffusion time to obtain a noisy image.
[0094] In some embodiments, the computer device may pre-set a diffusion time corresponding to the diffusion model, and the noisy images obtained after different diffusion times have different noise distributions.
[0095] In a possible implementation, the computer device performs noise processing on the original image through a forward diffusion network of a diffusion model according to the sampling noise and the diffusion time, thereby obtaining a noisy image.
[0096] Optionally, the diffusion time can be expressed as T, and the noise addition process can be expressed as q(x t |x t-1 ), x t That is, the corresponding noisy image when the noise duration is t, t∈[0, T], x t It can be expressed as x t =a t x0+σ t ∈, where a t and σ t Indicates the discretization during the noise addition process, x0 is the original image, and ∈ is Gaussian noise.
[0097] In one possible embodiment, before adding noise to the original image using the diffusion model, in order to improve the accuracy of determining the disturbance value, the computer device may first train the original diffusion model using a sample image set to obtain a trained diffusion model, and then add noise to the original image using the trained diffusion model.
[0098] Optionally, in order to enable the prediction network in the diffusion model to accurately predict the noise of the denoised image based on the text description, the sample image set also needs to include at least sample images with similar image features to the original image, so that the diffusion model can fully learn the image features.
[0099] For example, when the original image is an image of a kitten, the sample image set includes at least several sample images with cat features.
[0100] Step 704: Sample the diffusion duration to obtain a sampling time.
[0101] In a possible implementation, in order to perform noise prediction based on the denoised image corresponding to the random denoising duration during the denoising process, the computer device may further sample the diffusion duration to determine the sampling moment.
[0102] Optionally, the process of determining the sampling time can be expressed as t~U(T), where T is the diffusion time.
[0103] Optionally, during the denoising process, the computer device may perform a single sampling of the diffusion duration to obtain a single sampling moment; or may perform multiple samplings of the diffusion duration to obtain multiple sampling moments.
[0104] Step 705: De-noising the noisy image through a back-diffusion network to obtain a de-noised image corresponding to the sampling moment.
[0105] In some embodiments, after determining the sampling moment, the computer device may determine the denoised image corresponding to the sampling moment during the process of denoising the noisy image through the back diffusion network.
[0106] In a possible implementation, the sampling time is t, and the computer device may determine the image that has undergone denoising for a time period of Tt as the denoised image.
[0107] In a possible implementation, when the diffusion duration is sampled once, the computer device obtains a denoised image corresponding to a single sampling moment; when the diffusion duration is sampled multiple times, the computer device obtains a denoised image corresponding to each sampling moment.
[0108] In some embodiments, in order to enable the diffusion model to learn specific image features, when using the back diffusion network to denoise the noisy image, the computer device can also combine image text features to denoise the noisy image.
[0109] In a possible implementation, the computer device performs denoising on the noisy image through a back diffusion network according to the text description of the original image, thereby obtaining a denoised image corresponding to the sampling moment.
[0110] Optionally, the computer device inputs the text description into a back-diffusion network of the diffusion model. The back-diffusion network can then perform denoising on pixels in different regions of the noisy image based on the text description, thereby generating a denoised image corresponding to the sampling time. For example, if the text description reads "There is a cat in the middle of the image," the back-diffusion network can perform finer denoising on the middle region of the image.
[0111] Optionally, the computer device may perform text feature extraction on the original image through a BLIP model (Bootstrapping Language-Image Pretraining, a multimodal model for unified understanding and generation), thereby obtaining a text description corresponding to the original image.
[0112] Step 706: Perform noise prediction on the denoised image to obtain predicted noise.
[0113] In a possible implementation, after determining the denoised image corresponding to the sampling moment, the computer device may perform noise prediction on the denoised image through a prediction network to obtain predicted noise.
[0114] Optionally, the computer device may predict the noise of the denoised image by using a noise network in a diffusion model, and output the predicted noise through the noise network. The process can be expressed as in, is the noise network when the diffusion model is trained to the optimal value, c is the text description corresponding to the original image, x t is the denoised image corresponding to the sampling time t.
[0115] In one possible implementation, when the diffusion duration is sampled once, the computer device obtains a denoised image corresponding to a single sampling moment, and performs noise prediction on the single denoised image to obtain a single predicted noise; when the diffusion duration is sampled multiple times, the computer device obtains denoised images corresponding to each sampling moment, and performs noise prediction on each denoised image to obtain corresponding predicted noises.
[0116] In a possible implementation, considering that the purpose of obtaining a denoised image in the embodiment of the present application is to predict the noise of the denoised image and thus further determine the noise prediction loss, that is, in the embodiment of the present application, it is not necessarily necessary to perform denoising for the entire diffusion time T to obtain a noisy image. Therefore, in order to improve the image processing efficiency, the computer device may also first sample the diffusion time to determine the sampling time, and then directly perform denoising on the original image according to the sampling time through the forward diffusion network of the diffusion model, thereby obtaining a noisy image x t , and pass the noise network according to the noise image x t Noise prediction is performed to obtain predicted noise.
[0117] Step 707 : performing norm calculation on the noise difference between the predicted noise and the sampled noise to obtain the noise prediction loss.
[0118] In some embodiments, after obtaining the predicted noise corresponding to the denoised image, the computer device may calculate the noise difference between the predicted noise and the sampled noise, and obtain the noise prediction loss by performing a norm calculation on the noise difference. Norm calculation is an important concept in mathematics, primarily used to measure the size or length of vectors, matrices, or other mathematical objects. There are various types of norm calculations, such as single-norm calculations and double-norm calculations.
[0119] In one possible implementation, the computer device may calculate the noise prediction loss by performing a two-norm calculation on the noise difference. Alternatively, the process may be expressed as in, is the prediction noise, ∈ is the sampling noise, then the noise difference is
[0120] In one possible implementation, when the diffusion time is sampled once, after determining the predicted noise of a single denoised image, the computer device may perform a norm calculation on the noise difference between the predicted noise and the sampled noise to obtain a noise prediction loss; when the diffusion time is sampled multiple times, after obtaining the predicted noise of each denoised image, the computer device may perform a norm calculation on the noise difference between the predicted noise and the sampled noise respectively, and obtain the noise prediction loss by averaging.
[0121] Step 708: Perform gradient calculation on the noise prediction loss based on the original image to obtain a disturbance value.
[0122] In one possible implementation, to determine the perturbation value, the computer device may perform a gradient calculation on the noise prediction loss based on the original image, thereby obtaining the perturbation value corresponding to the original image. Calculating the gradient of the noise prediction loss may refer to calculating a partial derivative of the noise prediction loss with respect to its independent variable.
[0123] Alternatively, if the noise prediction loss is represented as Loss and its independent variable is x0, the process of calculating the gradient of the noise prediction loss can be expressed as Among them, x0 is the original image and δ is the perturbation value.
[0124] Step 709 : Perform disturbance processing on the original image according to the disturbance value to obtain an anti-editing image.
[0125] In some embodiments, in order to ensure that the anti-editing image has the same visual effect as the original image and avoid visual differences caused by the disturbance processing of the original image according to the disturbance value, the computer device can also set a disturbance threshold, so that when it is determined based on the disturbance value and the disturbance threshold that the disturbance threshold condition is met, the original image is disturbed according to the disturbance value to obtain the anti-editing image; when it is determined based on the disturbance value and the disturbance threshold that the disturbance threshold condition is not met, the original image is disturbed according to the disturbance threshold to obtain the anti-editing image.
[0126] Optionally, the disturbance threshold condition can be expressed as ||x0+δ-x0|| p <δ0, where δ0 is the disturbance threshold and δ is the disturbance value.
[0127] In a possible implementation, in order to improve the anti-editing degree of the anti-editing image and effectively protect the image features of the original image, the computer device can also perform multiple rounds of disturbance processing on the original image according to the disturbance value to obtain the anti-editing image.
[0128] Schematically, the right side of Figure 8 shows the original image 803, and the left side shows the anti-editing image 801 obtained by perturbing the original image 803. It can be seen that there is little visual difference between the anti-editing image 801 and the original image 803. The center of Figure 8 shows the image 802 after anti-editing image 801 has been edited using the diffusion model. It can be seen that the edited image 802 is essentially identical to the anti-editing image 801, indicating that the diffusion model has difficulty editing the anti-editing image 801.
[0129] Schematically, as shown in Figure 9, 901 in Figure 9 is a set of images that have been processed for anti-editing based on different fine-tuning weights, and 902 in Figure 9 is a result image after the anti-editing image in 901 is edited by the diffusion model based on different fine-tuning weights. It can be seen that there is no successfully edited image in the result image, that is, the diffusion model cannot edit the image that has been processed for anti-editing.
[0130] In the above embodiment, the sampling time is determined by sampling the diffusion time, and the noise network is used to predict the noise of the denoised image corresponding to the sampling time to obtain the predicted noise. Then, the disturbance value is determined based on the noise prediction loss between the predicted noise and the sampling noise, thereby improving the accuracy of the disturbance value determination and effectively increasing the difficulty of editing the anti-editing image using the diffusion model.
[0131] In some embodiments, during the process of the diffusion model performing image feature learning on the original image, the diffusion model needs to perform denoising processing on the original image through multiple rounds of noise addition processing, so as to perform optimization training based on the noise prediction loss of each round. Therefore, accordingly, in an embodiment of the present application, in order to improve the anti-editing quality of the anti-editing image, the computer device can determine the disturbance value based on the noise prediction loss of each round, thereby performing M rounds of disturbance processing on the original image, and determining the disturbed image obtained by the Mth round of disturbance processing as the anti-editing image.
[0132] In a possible implementation, during the M rounds of disturbance processing, the i-th round of disturbance processing may include the following steps (not shown in the figure):
[0133] Step 709a: Based on the i-th sampling noise, the i-1th disturbed image is subjected to noise addition processing through the forward diffusion network to obtain the i-th noisy image. The i-1th disturbed image is the disturbed image obtained after the i-1th round of disturbance processing.
[0134] In one possible implementation, when performing the i-th round of perturbation processing, the computer device performs the i-th round of noise sampling on the noise that conforms to the standard Gaussian distribution to obtain the i-th sampled noise, and then, based on the i-th sampled noise, performs noise processing on the i-1th perturbed image through the forward diffusion network of the diffusion model to obtain the i-th noisy image, wherein the i-1th perturbation image is the perturbation image obtained after the i-1th round of perturbation processing, and i is a positive integer greater than 1 and less than or equal to M.
[0135] Step 709b: De-noising the i-th noisy image through a back-diffusion network, and determining the i-th predicted noise based on the i-th denoised image obtained by de-noising.
[0136] In a possible implementation, after obtaining the i-th noisy image, the computer device can perform denoising on the i-th noisy image through a back diffusion network to obtain the i-th denoised image, and then perform noise prediction on the i-th denoised image through a noise network to obtain the i-th predicted noise.
[0137] Step 709c: determining an i-th disturbance value according to an i-th noise prediction loss between the i-th predicted noise and the i-th sampled noise.
[0138] In one possible implementation, after obtaining the i-th predicted noise, the computer device can obtain the i-th noise prediction loss by performing a norm calculation on the noise difference between the i-th predicted noise and the i-th sampled noise, and in order to improve the accuracy of determining the disturbance value in each round, the computer device can perform a gradient calculation on the i-th noise prediction loss based on the i-1-th disturbance image to obtain the i-th disturbance value.
[0139] Optionally, the i-1th perturbation image can be expressed as x i-1 , the process of gradient calculation for noise prediction loss can be expressed as
[0140] Step 709d: If it is determined based on the i-th disturbance value and the disturbance threshold that the disturbance threshold condition is met, the i-1-th disturbance image is disturbed according to the i-th disturbance value to obtain the i-th disturbance image.
[0141] In some embodiments, in order to ensure that the anti-editing image has the same visual effect as the original image and avoid visual differences caused by the disturbance processing of the original image according to the disturbance value, the computer device can also set a disturbance threshold and perform conditional judgment on each round of disturbance value during each round of disturbance processing. Therefore, when it is determined that the disturbance threshold condition is met based on the i-th disturbance value and the disturbance threshold, the computer device then performs disturbance processing on the i-1th disturbed image according to the i-th disturbance value to obtain the i-th disturbed image.
[0142] Optionally, the purpose of setting a perturbation threshold condition is to ensure that the perturbed image obtained after each round of perturbation processing has minimal visual differences from the original image and has the same visual effect as the original image as much as possible. Specifically, when the original image is perturbed based on the perturbation threshold, the perturbed image can maintain the same visual effect as the original image; when the original image is perturbed based on a perturbation value greater than the perturbation threshold, the perturbed image cannot maintain the same visual effect as the original image.
[0143] In some embodiments, the computer device may pre-simulate the perturbation processing of the (i-1)th disturbed image using the (i)th disturbance value to obtain the (i)th candidate disturbed image, thereby judging whether the perturbation threshold condition is met based on the visual difference between the (i)th candidate disturbed image and the original image.
[0144] In a possible implementation, the computer device may first perform perturbation processing on the (i-1)th disturbed image using the (i)th disturbance value to obtain the (i)th candidate disturbed image, thereby determining the (i)th disturbance difference based on the (i)th candidate disturbed image and the original image, and then performing norm calculation on the (i)th disturbance difference to obtain the (i)th disturbance norm, wherein the (i)th disturbance difference is used to characterize the visual difference between the (i)th candidate disturbed image and the original image, thereby comparing the (i)th disturbance norm with the disturbance threshold. When the (i)th disturbance norm is not greater than the disturbance threshold, the computer device may determine the (i)th candidate disturbed image as the (i)th disturbed image, and the (i)th candidate disturbed image is obtained by perturbation processing on the (i-1)th disturbed image using the (i)th disturbance value, thereby achieving, when the disturbance threshold condition is met, performing perturbation processing on the (i-1)th disturbed image according to the (i)th disturbance value to obtain the (i)th disturbed image.
[0145] Optionally, the i-th perturbation norm may be the binary norm of the i-th perturbation difference value, or the infinite norm of the i-th perturbation difference value. That is, the perturbation threshold condition may be that the binary norm of the i-th perturbation difference value is not greater than the perturbation threshold value, or that the infinite norm of the i-th perturbation difference value is not greater than the perturbation threshold value. When the perturbation threshold condition is that the infinite norm of the i-th perturbation difference value is not greater than the perturbation threshold value, the anti-editing image has a smaller visual difference from the original image.
[0146] Step 709e: If it is determined based on the i-th disturbance value and the disturbance threshold that the disturbance threshold condition is not satisfied, the original image is disturbed according to the disturbance threshold to obtain the i-th disturbed image.
[0147] In one possible implementation, to ensure that each round of perturbed images maintains the same visual effect as the original image, when it is determined based on the i-th perturbation value and the perturbation threshold that the perturbation threshold condition is not met, the computer device needs to directly perturb the original image according to the perturbation threshold to obtain the i-th round of perturbed images.
[0148] In one possible implementation, the computer device first performs a perturbation process on the (i-1)th disturbed image using the (i)th disturbance value to obtain the (i)th candidate disturbed image, thereby determining the (i)th disturbance difference based on the (i)th candidate disturbed image and the original image, and then performing a norm calculation on the (i)th disturbance difference to obtain the (i)th disturbance norm. Therefore, when the (i)th disturbance norm is greater than the disturbance threshold, the computer device needs to perform a perturbation process on the original image according to the disturbance threshold to obtain the (i)th disturbed image.
[0149] Optionally, in the process of calculating the norm of the i-th disturbance difference, the computer device may calculate the 2-norm value of the i-th disturbance difference, or the infinite-norm value, which is not limited in the embodiment of the present application.
[0150] In a possible implementation, in order to improve the consistency of visual effects between the anti-editing image and the original image, the computer device may calculate the infinite norm of the i-th perturbation difference to obtain the i-th perturbation norm.
[0151] Optionally, the process of determining whether the i-th disturbance value satisfies the disturbance threshold condition can be expressed as Clip(||x0 i-1 +δ i -x0|| p ,δ0), where x0 is the original image, δ i is the i-th disturbance value, x0 i-1 is the perturbation image of the i-1th round, x0 i-1 +δ i Then is the candidate perturbation image of the i-th round, δ0 is the perturbation threshold, and p can take the infinite norm. i-1 +δ i -x0|| p When x0 is not greater than δ0, i =x0 i-1 +δ i ; in || x0 i-1 +δ i -x0|| p When x0 is greater than δ0, i =x0+δ0.
[0152] In the above embodiment, M rounds of perturbation processing are performed on the original image, and after the i-th perturbation value is determined in each round of perturbation processing, the i-th perturbation value is subjected to a perturbation threshold condition judgment, thereby determining the perturbed image in each round of perturbation processing. By performing multiple rounds of perturbation processing on the original image and using the perturbed image obtained in the final round of perturbation processing as the anti-editing image, the anti-editing degree of the anti-editing image is improved, and the visual difference between the anti-editing image and the original image is reduced.
[0153] Please refer to FIG10 , which shows a schematic flow chart of obtaining an anti-editing image through M rounds of perturbation processing according to an exemplary embodiment of the present application.
[0154] As shown in FIG10 , first, the computer device inputs the original image 1001 into the diffusion model 1002. After the denoising and denoising processes of the diffusion model 1002, a first denoised image 1003 can be obtained. Then, by performing noise prediction on the first denoised image 1003, a first predicted noise can be obtained. Then, based on the first sampling noise and the first predicted noise, a first noise prediction loss 1004 can be calculated to determine a first disturbance value 1005.
[0155] Secondly, in order to reduce the visual difference between the original image and the anti-editing image, the computer device also needs to determine whether the disturbance threshold condition is met based on the first disturbance value 1005 and the disturbance threshold 1006. If the disturbance threshold condition is met, the original image 1001 is disturbed by the first disturbance value 1005 to obtain a first disturbed image 1007; if the disturbance threshold condition is not met, the original image 1001 is disturbed by the disturbance threshold 1006 to obtain the first disturbed image 1007.
[0156] Then, after obtaining the first disturbed image 1007 , the computer device continues to input the first disturbed image 1007 into the diffusion model 1002 to perform multiple rounds of disturbance processing that are the same as the first round of disturbance processing.
[0157] Finally, after the Mth round of noise prediction obtains the Mth noise prediction loss 1008, the computer device can determine the Mth disturbance value 1009 according to the Mth noise prediction loss 1008, and thereby determine whether the disturbance threshold condition is met based on the Mth disturbance value 1009 and the disturbance threshold 1006. If the disturbance threshold condition is met, the M-1th disturbance image is disturbed with the Mth disturbance value 1009 to obtain the anti-editing image 1010; if the disturbance threshold condition is not met, the original image 1001 is disturbed with the disturbance threshold 1006 to obtain the anti-editing image 1010.
[0158] In one possible implementation, a computer device determines an anti-editing image set corresponding to an original image set containing N original images based on a Monte Carlo simulation, using a diffusion model and the anti-editing algorithm proposed in the embodiments of the present application.
[0159] Among them, the number of Monte Carlo simulations is M, and the diffusion model is The original image set is (I1,I2,...,I n ), the anti-editing image set is (I1′,I2′,...,I n ′), so the algorithm flow can be expressed as:
[0160] In some embodiments, considering that the diffusion model may only edit and learn the image features of a certain area in the image during the process of learning and editing image features, for example, the diffusion model may only focus on learning the features of the foreground person in the image, or the features of a certain object in the foreground, etc., therefore, in order to improve the efficiency of generating anti-editing images, the computer device can also perform perturbation processing on the original image according to the degree of anti-editing required for different image areas in the original image.
[0161] In one possible implementation, a computer device may divide the original image into image regions based on image features of the original image to obtain multiple original image regions, wherein different original image regions correspond to different anti-editing levels, for example, the anti-editing level corresponding to the foreground image region is greater than the anti-editing level corresponding to the background image region.
[0162] The computer device can then determine the disturbance values corresponding to different original image areas based on the noise prediction loss between the predicted noise and the sampled noise, and the anti-editing degree corresponding to different original image areas, and thus perform disturbance processing on the original image according to the disturbance values corresponding to different original image areas to obtain an anti-editing image.
[0163] In a possible implementation, the computer device may further indicate the degree of anti-editing of each original image region by setting an anti-editing level, for example, the anti-editing level of the foreground image region is high, and the anti-editing level of the background image region is low.
[0164] In another possible implementation, different levels of anti-editing can be achieved by adjusting different perturbation values for different original image regions. For example, after determining a perturbation value δ based on the noise prediction loss, the computer device can directly perturb the foreground image region based on the perturbation value δ, while perturbating the background image region based on half the perturbation value, 0.5δ.
[0165] In another possible implementation, different levels of anti-editing can also be manifested by performing different rounds of perturbation processing on different original image regions. For example, after determining a perturbation value for each round, the computer device can perform perturbation processing on the foreground image region based on the perturbation value for each round, and perform perturbation processing on the background image region based on the perturbation value every five rounds.
[0166] In one possible implementation, in order to prevent editing while ensuring that the visual effect of the original image after conversion into an anti-editing image is not affected as much as possible, the computer device can also perform local perturbation processing on the original image. For example, after determining the disturbance value based on the noise prediction loss between the predicted noise and the sampling noise, the computer device can perform disturbance processing on the foreground image area of the original image based on the disturbance value corresponding to the foreground image area in the disturbance value, that is, the computer device can set the disturbance value of the pixel point corresponding to the background image area in the disturbance matrix to zero, thereby performing disturbance processing on the original image according to the set disturbance value to obtain an anti-editing image.
[0167] In the above embodiment, the original image is divided into image areas to obtain different original image areas, and then the disturbance value corresponding to each original image area is determined according to the anti-editing degree corresponding to the different original image areas, and then the original image is disturbed according to the disturbance value to obtain an anti-editing image. This achieves the goal of reducing the visual difference between the original image and the anti-editing image while performing the anti-editing processing, improving the anti-editing processing efficiency, and optimizing the image quality of the anti-editing image.
[0168] Please refer to FIG11 , which shows a structural block diagram of an image processing device provided by an exemplary embodiment of the present application. The device includes:
[0169] An acquisition module 1101 is used to acquire an original image;
[0170] A noise processing module 1102 is configured to perform noise processing on the original image according to the sampling noise through a forward diffusion network of a diffusion model to obtain a noisy image;
[0171] a noise prediction module 1103 configured to perform denoising on the noisy image using a reverse diffusion network of the diffusion model, and determine predicted noise based on the denoised image obtained by the denoising process;
[0172] a disturbance determination module 1104 for determining a disturbance value according to a noise prediction loss between the predicted noise and the sampled noise;
[0173] The first disturbance processing module 1105 is configured to perform disturbance processing on the original image according to the disturbance value to obtain an anti-editing image.
[0174] Optionally, the noise processing module 1102 is configured to:
[0175] Sampling standard Gaussian noise to obtain the sampled noise;
[0176] According to the sampling noise and the diffusion time, the original image is subjected to noise processing by the forward diffusion network to obtain the noisy image.
[0177] Optionally, the noise prediction module 1103 is configured to:
[0178] Sampling the diffusion time to obtain a sampling moment;
[0179] Performing denoising on the noisy image through the back diffusion network to obtain the denoised image corresponding to the sampling moment;
[0180] Noise prediction is performed on the denoised image to obtain the predicted noise.
[0181] Optionally, the noise prediction module 1103 is further configured to:
[0182] According to the text description of the original image, the noisy image is denoised by the back diffusion network to obtain the denoised image corresponding to the sampling moment.
[0183] Optionally, the disturbance determination module 1104 is configured to:
[0184] performing a norm calculation on a noise difference between the predicted noise and the sampled noise to obtain the noise prediction loss;
[0185] A gradient calculation is performed on the noise prediction loss based on the original image to obtain the disturbance value.
[0186] Optionally, the first disturbance processing module 1105 is configured to:
[0187] According to the disturbance value, the original image is subjected to M rounds of disturbance processing, and the disturbed image obtained by the M-th round of disturbance processing is determined as the anti-editing image, where M is a positive integer.
[0188] Optionally, the i-th round of disturbance processing in the M rounds of disturbance processing includes:
[0189] Based on the i-th sampling noise, the i-1-th perturbed image is subjected to a noise addition process through the forward diffusion network to obtain an i-th noisy image, where the i-1-th perturbed image is a perturbed image obtained after the i-1-th round of perturbation processing, and i is a positive integer greater than 1 and less than or equal to M;
[0190] Performing denoising on the i-th noisy image through the back diffusion network, and determining an i-th predicted noise based on the i-th denoised image obtained by denoising;
[0191] determining an i-th disturbance value according to an i-th noise prediction loss between the i-th predicted noise and the i-th sampled noise;
[0192] If it is determined based on the i-th disturbance value and the disturbance threshold that the disturbance threshold condition is satisfied, performing disturbance processing on the (i-1)-th disturbed image according to the i-th disturbance value to obtain an i-th disturbed image;
[0193] If it is determined based on the i-th disturbance value and the disturbance threshold that the disturbance threshold condition is not satisfied, the original image is disturbed according to the disturbance threshold to obtain the i-th disturbed image.
[0194] Optionally, the first disturbance processing module 1105 is further configured to:
[0195] Performing a norm calculation on a noise difference between the i-th predicted noise and the i-th sampled noise to obtain the i-th noise prediction loss;
[0196] Performing a gradient calculation on the i-th noise prediction loss based on the i-1-th disturbance image to obtain the i-th disturbance value.
[0197] Optionally, the device further includes:
[0198] a second disturbance processing module, configured to perform disturbance processing on the (i-1)th disturbed image according to the (i)th disturbance value to obtain an (i)th candidate disturbed image;
[0199] a difference determination module, configured to determine an i-th disturbance difference based on the i-th candidate disturbed image and the original image, wherein the i-th disturbance difference is used to represent a visual difference between the i-th candidate disturbed image and the original image;
[0200] A norm calculation module, configured to perform norm calculation on the i-th disturbance difference to obtain an i-th disturbance norm;
[0201] The first disturbance processing module 1105 is configured to:
[0202] If the i-th disturbance norm is not greater than the disturbance threshold, determining the i-th candidate disturbed image as the i-th disturbed image;
[0203] The first disturbance processing module 1105 is further configured to:
[0204] If the i-th disturbance norm is greater than the disturbance threshold, the original image is disturbed according to the disturbance threshold to obtain the i-th disturbed image.
[0205] Optionally, the device further includes:
[0206] The training module is used to train the original diffusion model through a sample image set to obtain the trained diffusion model, wherein the sample image set at least includes sample images having similar image features to the original image.
[0207] Optionally, the device further includes:
[0208] an image division module, configured to divide the original image into image regions according to image features of the original image, to obtain a plurality of original image regions, wherein different original image regions among the plurality of original image regions correspond to different anti-editing levels;
[0209] The disturbance determination module 1104 is configured to:
[0210] determining disturbance values corresponding to the different original image regions respectively according to a noise prediction loss between the predicted noise and the sampled noise, and anti-editing degrees corresponding to the different original image regions;
[0211] The first disturbance processing module 1105 is further configured to:
[0212] The original image is disturbed according to the disturbance values corresponding to the different original image regions to obtain the anti-editing image.
[0213] In summary, in the embodiment of the present application, after acquiring the original image, the original image is first subjected to noise processing by the forward diffusion network of the diffusion model according to the sampling noise to obtain a noisy image, and then the noisy image is denoised by the reverse diffusion network of the diffusion model, and the predicted noise is determined based on the denoised image obtained by the denoising process. The process of the diffusion model performing noise processing and denoising on the original image is essentially a process of learning the image features of the original image, and the noise prediction loss between the predicted noise and the sampling noise can reflect the uncertainty of the diffusion model in the current state of the noise. Therefore, the perturbation value can be determined based on the noise prediction loss between the predicted noise and the sampling noise. The perturbation value is intended to simulate or amplify the uncertainty of the diffusion model in predicting noise. In this way, the original image can be perturbated according to the perturbation value to obtain an anti-editing image, thereby introducing additional, unpredictable noise into the anti-editing image, that is, each pixel in the anti-editing image is subjected to signal interference, thereby increasing the difficulty of the diffusion model in learning the image features of the anti-editing image. Accordingly, when using the trained diffusion model to edit the anti-editing image, since each pixel in the anti-editing image is subject to signal interference, the diffusion model cannot accurately extract the image features in the anti-editing image, thereby increasing the difficulty of editing the anti-editing image using the diffusion model.
[0214] It should be noted that the apparatus provided in the above embodiments is merely exemplified by the division of the above functional modules. In actual applications, the above functions can be distributed among different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The implementation process is detailed in the method embodiments and will not be repeated here.
[0215] It should be noted that before collecting original images and other related user data, and during the process of collecting original images and other related user data, this application can display a prompt interface, pop-up window or output voice prompt information. The prompt interface, pop-up window or voice prompt information is used to remind the user that their relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining user-related data after obtaining the user's confirmation operation on the prompt interface or pop-up window. Otherwise (that is, when the user's confirmation operation on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining user-related data are terminated, that is, the user's relevant data is not obtained.
[0216] Please refer to Figure 12, which shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Specifically, the computer device 1200 includes a central processing unit (CPU) 1201, a system memory 1204 including a random access memory 1202 and a read-only memory 1203, and a system bus 1205 connecting the system memory 1204 and the CPU 1201. The computer device 1200 also includes a basic input / output system (I / O system) 1206 that facilitates information transmission between various components within the computer, and a mass storage device 1207 for storing an operating system 1213, application programs 1214, and other program modules 1215.
[0217] The basic input / output system 1206 includes a display 1208 for displaying information and an input device 1209, such as a mouse and keyboard, for user input. Both the display 1208 and the input device 1209 are connected to the central processing unit 1201 via an input / output controller 1210 connected to the system bus 1205. The basic input / output system 1206 may also include an input / output controller 1210 for receiving and processing input from a variety of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1210 also provides output to a display screen, printer, or other types of output devices.
[0218] The mass storage device 1207 is connected to the central processing unit 1201 via a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1207 and its associated computer-readable media provide non-volatile storage for the computer device 1200. In other words, the mass storage device 1207 may include a computer-readable medium (not shown) such as a hard disk or drive.
[0219] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cassette, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage medium is not limited to the above-mentioned ones. The above-mentioned system memory 1204 and mass storage device 1207 can be collectively referred to as memory.
[0220] The memory stores one or more programs, and the one or more programs are configured to be executed by one or more central processing units 1201. The one or more programs contain instructions for implementing the above-mentioned methods. The central processing unit 1201 executes the one or more programs to implement the methods provided by the above-mentioned various method embodiments.
[0221] According to various embodiments of the present application, the computer device 1200 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 1200 may be connected to the network 1211 via the network interface unit 1212 connected to the system bus 1205, or the network interface unit 1212 may be used to connect to other types of networks or remote computer systems (not shown).
[0222] An embodiment of the present application further provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the image processing method described in the above embodiment.
[0223] Optionally, the computer-readable storage medium may include: ROM, RAM, solid state drives (SSDs) or optical disks, etc. Among them, RAM may include resistance random access memory (ReRAM) and dynamic random access memory (DRAM).
[0224] The present invention provides a computer program product comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image processing method described in the above embodiment.
[0225] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0226] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. An image processing method, the method being executed by a computer device, the method comprising: Get the original image; According to the sampling noise, the original image is subjected to noise processing by a forward diffusion network of a diffusion model to obtain a noisy image; De-noising the noisy image by using a reverse diffusion network of the diffusion model, and determining predicted noise based on the de-noised image obtained by the de-noising process; determining a disturbance value according to a noise prediction loss between the predicted noise and the sampled noise; The original image is disturbed according to the disturbance value to obtain an anti-editing image.
2. The method according to claim 1, wherein the process of adding noise to the original image through a forward diffusion network of a diffusion model according to the sampling noise to obtain a noisy image comprises: Sampling standard Gaussian noise to obtain the sampled noise; According to the sampling noise and the diffusion time, the original image is subjected to noise processing through the forward diffusion network to obtain the noisy image.
3. The method according to claim 2, wherein the denoising process is performed on the noisy image by using the inverse diffusion network of the diffusion model, and the predicted noise is determined according to the denoised image obtained by the denoising process, comprising: Sampling the diffusion time to obtain a sampling time; Performing denoising on the noisy image through the back diffusion network to obtain the denoised image corresponding to the sampling moment; Noise prediction is performed on the denoised image to obtain the predicted noise.
4. The method according to claim 3, wherein the denoising process is performed on the noisy image by the back diffusion network to obtain the denoised image corresponding to the sampling time, comprising: According to the text description of the original image, the noisy image is denoised through the back diffusion network to obtain the denoised image corresponding to the sampling moment.
5. The method according to any one of claims 1 to 4, wherein determining the disturbance value according to the noise prediction loss between the predicted noise and the sampled noise comprises: Performing norm calculation on the noise difference between the predicted noise and the sampled noise to obtain the noise prediction loss; The noise prediction loss is gradient calculated based on the original image to obtain the disturbance value.
6. The method according to any one of claims 1 to 5, wherein the perturbation process is performed on the original image according to the perturbation value to obtain the anti-editing image, comprising: According to the disturbance value, the original image is subjected to M rounds of disturbance processing, and a disturbance image obtained by the M-th round of disturbance processing is determined as the anti-editing image, where M is a positive integer.
7. The method according to claim 6, wherein the i-th round of disturbance processing in the M rounds of disturbance processing comprises: Based on the i-th sampling noise, the i-1-th disturbance image is subjected to noise processing through the forward diffusion network to obtain the i-th noisy image, wherein the i-1-th disturbance image is a disturbance image obtained after the i-1-th round of disturbance processing, and i is a positive integer greater than 1 and less than or equal to M; Denoising the i-th noisy image by using the back diffusion network, and determining the i-th predicted noise according to the i-th denoised image obtained by denoising; Determining an i-th disturbance value according to an i-th noise prediction loss between the i-th predicted noise and the i-th sampled noise; If it is determined based on the i-th disturbance value and the disturbance threshold that the disturbance threshold condition is met, performing disturbance processing on the i-1th disturbance image according to the i-th disturbance value to obtain an i-th disturbance image; If it is determined based on the i-th disturbance value and the disturbance threshold that the disturbance threshold condition is not satisfied, the original image is disturbed according to the disturbance threshold to obtain the i-th disturbed image.
8. The method according to claim 7, wherein determining the i-th disturbance value according to the i-th noise prediction loss between the i-th predicted noise and the i-th sampled noise comprises: Performing norm calculation on a noise difference between the i-th predicted noise and the i-th sampled noise to obtain the i-th noise prediction loss; A gradient calculation is performed on the i-th noise prediction loss based on the i-1th disturbance image to obtain the i-th disturbance value.
9. The method according to claim 7 or 8, further comprising: Performing disturbance processing on the (i-1)th disturbance image according to the (i)th disturbance value to obtain an (i)th candidate disturbance image; Based on the i-th candidate disturbed image and the original image, determine an i-th disturbance difference, where the i-th disturbance difference is used to represent a visual difference between the i-th candidate disturbed image and the original image; Performing norm calculation on the i-th disturbance difference to obtain the i-th disturbance norm; If it is determined based on the i-th disturbance value and the disturbance threshold that the disturbance threshold condition is satisfied, performing disturbance processing on the i-1-th disturbance image according to the i-th disturbance value to obtain the i-th disturbance image, comprising: If the i-th disturbance norm is not greater than the disturbance threshold, determining the i-th candidate disturbed image as the i-th disturbed image; If it is determined based on the i-th disturbance value and the disturbance threshold that the disturbance threshold condition is not satisfied, performing disturbance processing on the original image according to the disturbance threshold to obtain the i-th disturbance image, comprising: If the i-th disturbance norm is greater than the disturbance threshold, the original image is disturbed according to the disturbance threshold to obtain the i-th disturbed image.
10. The method according to any one of claims 1 to 9, further comprising: The original diffusion model is trained by using a sample image set to obtain the trained diffusion model, wherein the sample image set at least includes sample images having similar image features to the original image.
11. The method according to any one of claims 1 to 10, further comprising: According to the image features of the original image, the original image is divided into image regions to obtain a plurality of original image regions, wherein different original image regions among the plurality of original image regions correspond to different anti-editing degrees; The determining of the disturbance value according to the noise prediction loss between the predicted noise and the sampled noise comprises: Determining disturbance values corresponding to the different original image regions respectively according to a noise prediction loss between the predicted noise and the sampled noise, and anti-editing degrees corresponding to the different original image regions; The step of performing a disturbance process on the original image according to the disturbance value to obtain an anti-editing image includes: The original image is disturbed according to the disturbance values corresponding to the different original image regions to obtain the anti-editing image.
12. An image processing apparatus, the apparatus being deployed on a computer device, the apparatus comprising: An acquisition module, used for acquiring the original image; A noise processing module, used for performing noise processing on the original image through a forward diffusion network of a diffusion model according to sampling noise to obtain a noisy image; A noise prediction module, configured to perform denoising on the noisy image through a reverse diffusion network of the diffusion model, and determine predicted noise based on the denoised image obtained by the denoising; a disturbance determination module, configured to determine a disturbance value according to a noise prediction loss between the predicted noise and the sampled noise; The first disturbance processing module is used to perform disturbance processing on the original image according to the disturbance value to obtain an anti-editing image.
13. A computer device, comprising a processor and a memory; the memory stores at least one instruction, and the at least one instruction is used to be executed by the processor to implement the image processing method according to any one of claims 1 to 11.
14. A computer-readable storage medium, wherein at least one instruction is stored in the computer-readable storage medium, and the at least one instruction is loaded and executed by a processor to implement the image processing method according to any one of claims 1 to 11.
15. A computer program product, comprising computer instructions, wherein the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the image processing method as described in any one of claims 1 to 11.