Image processing method, device, equipment, storage medium, product and system

By training a target diffusion model and combining it with fine-tuning and preservation models, the problem of imperfections generated by the diffusion model was solved, thus improving image quality and conformity.

CN120876348APending Publication Date: 2025-10-31BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410528363.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-28
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Images generated by existing diffusion models have flaws, such as blurred watermarks at the edges of the images, which affect image quality.

Method used

By training a target diffusion model, combining a fine-tuning model and a preservation model, the fine-tuning model learns the ability to generate flawless images during training, while the model's generation ability remains unchanged. The generation process is optimized to avoid flaws, thus obtaining the target diffusion model.

Benefits of technology

The quality of the generated images has been improved, ensuring that the images match the target description text and reducing defects, resulting in better generation effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876348A_ABST
    Figure CN120876348A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method, device, equipment, storage medium, product and system, and relates to the technical field of image processing, the method comprises the following steps: processing an original image and a target description text through a target diffusion model to obtain a target image, the target diffusion model being obtained based on a plurality of sample image training basic models, the basic model comprises a fine tuning model and a holding model, the holding model remains unchanged in the training process, and the fine tuning model is used for learning the image generation capability of the holding model in the training process so as to obtain a target diffusion model after training. The generation capability of a basic model in a non-defective area can be reserved as much as possible, generation of flaws is avoided as much as possible, and therefore a target diffusion model with better image generation quality is obtained. Therefore, a target image which better conforms to the target description text and is better in quality can be obtained based on the target diffusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and in particular to an image processing method, apparatus, device, storage medium, product and system. Background Technology

[0002] Diffusion models can be used for image generation and editing tasks. They can process original graphics and descriptive text to obtain corresponding effect images. However, images generated using diffusion models may contain imperfections that affect their quality. For example, blurry watermarks may appear at the edges of the image. Summary of the Invention

[0003] To overcome the problems existing in related technologies, this disclosure provides an image processing method, apparatus, device, storage medium, product, and system to improve the quality of images generated by diffusion models.

[0004] According to a first aspect of the present disclosure, an image processing method is provided, comprising: Acquire the original image to be processed and the target description text; The original image and the target description text are processed by a target diffusion model to obtain a target image. The target diffusion model is obtained by training a base model based on multiple sample images. The base model includes a fine-tuning model and a maintenance model. The maintenance model remains unchanged during training. The fine-tuning model is used to learn the image generation capability of the maintenance model during training so as to obtain the target diffusion model after training.

[0005] Optionally, the target diffusion model is trained through the following steps: The multiple sample images are divided into a training set and a test set; The base model is trained and optimized in multiple rounds using the training set, so that the fine-tuned model learns the image generation capability of the model during the training process. If the fine-tuned model converges, training is stopped, and a candidate diffusion model is obtained. The candidate diffusion model is optimized through multiple rounds of testing using the test set to obtain the target diffusion model.

[0006] Optionally, the step of performing multiple rounds of training and optimization on the base model using the training set includes: The base model is trained iteratively in multiple rounds using the training set. For any round of training iterations, determine the fine-tuning loss for the fine-tuning model and the retention loss for the retention model in that round of training iterations. The fine-tuning loss and the maintenance loss are used to optimize the fine-tuning model corresponding to this round of iterative training.

[0007] Optionally, determining the fine-tuning loss for the fine-tuned model and the retention loss for the retention model in any given round of training includes: For any round of training iteration, obtain the target sample image corresponding to that round of training. The target sample image carries a target label, which includes sample description text. The target sample image is processed to obtain a target mask image visible in a preset area; The target sample image, the target mask image, and the target label are processed using the fine-tuning model to obtain the fine-tuning result; The target sample image, the target mask image, and the target label are processed using the preservation model to obtain the preservation result; Based on the fine-tuning results, the fine-tuning loss corresponding to the fine-tuned model in this round of iterative training is obtained, and based on the fine-tuning results and the holding results, the holding loss corresponding to the holding model in this round of iterative training is obtained.

[0008] Optionally, processing the target sample image to obtain a target mask image visible in a preset region includes: Generate a mask of the same size as the target sample image, wherein the mask of the preset region of the mask is 0, and the mask of the other regions of the preset region of the mask is 1; Based on the mask and the target sample image, a target mask image visible in the preset area is obtained.

[0009] Optionally, the fine-tuning result includes a first random accumulated noise, a first predicted noise, and a predicted defect detection result; The process of processing the target sample image, the target mask image, and the target label using the fine-tuning model to obtain the fine-tuning result includes: The target sample image, the target mask image, and the target label are processed by the fine-tuning model to obtain the first sample image code corresponding to the target sample image, the first mask image code corresponding to the mask image, and the first descriptive text code corresponding to the target label. The target sample image is subjected to noise addition for a preset number of steps to obtain the first random accumulated noise. The first random accumulated noise is the accumulated noise corresponding to multiple steps after a random number of steps during the process of the fine-tuning model adding noise to the target sample image for a preset number of steps. The first predicted noise is obtained by processing the first noise map encoding, the first descriptive text encoding, the first mask encoding, and the random step number through the fine-tuning model. The first noise map encoding is obtained by processing the first sample map encoding with the first random accumulated noise. The target diffusion map is obtained by processing the prediction map encoding through the fine-tuning model. The prediction map encoding is obtained by processing the first noise map encoding through the first prediction noise. The target diffusion map is processed by the target defect detection model to obtain the predicted defect detection result.

[0010] Optionally, the retention result includes a second prediction noise; The process of processing the target sample image, the target mask image, and the target label using the preservation model to obtain the preservation result includes: The target sample image, the target mask image, and the target label are processed by the preservation model to obtain the second sample image code corresponding to the target sample image, the second mask image code corresponding to the mask image, and the second descriptive text code corresponding to the target label. The target sample image is subjected to noise addition for a preset number of steps to obtain a second random accumulated noise. The second random accumulated noise is the accumulated noise corresponding to multiple steps after the random number of steps during the process of the preservation model adding noise to the target sample image for a preset number of steps. The second predicted noise is obtained by processing the second noise map encoding, the second descriptive text encoding, the second mask encoding, and the random step number through the preservation model. The second noise map encoding is obtained by processing the second sample map encoding with the second random accumulated noise.

[0011] Optionally, obtaining the fine-tuning loss corresponding to the fine-tuned model in this round of iterative training based on the fine-tuning result, and obtaining the retention loss corresponding to the retained model in this round of iterative training based on the fine-tuning result and the retention result, includes: The product of the absolute value of the difference between the first random accumulated noise and the first predicted noise and the predicted defect detection result is determined as the fine-tuning loss corresponding to the fine-tuning model in this round of iterative training. The product of the absolute value of the difference between the first predicted noise and the second predicted noise and the expected result is determined as the holding loss corresponding to the model in this round of iterative training. The expected result is the value obtained by subtracting the predicted defect detection result from 1.

[0012] According to a second aspect of the present disclosure, an image processing apparatus is provided, comprising: The acquisition module is configured to acquire the raw image to be processed and the target description text; The first acquisition module is configured to process the original image and the target description text through a target diffusion model to obtain a target image. The target diffusion model is obtained by training a base model based on multiple sample images. The base model includes a fine-tuning model and a maintenance model. The maintenance model remains unchanged during training. The fine-tuning model is used to learn the image generation capability of the maintenance model during training so as to obtain the target diffusion model after training.

[0013] According to a third aspect of the present disclosure, an electronic device is provided, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to execute the steps of the image processing method provided in the first aspect of this disclosure.

[0014] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the steps of the image processing method provided in the first aspect of the present disclosure.

[0015] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the image processing method provided in the first aspect of the present disclosure.

[0016] According to a sixth aspect of the present disclosure, a chip system is provided, the chip system including a processing unit and an interface circuit, the processing unit obtaining program instructions through the interface circuit, the program instructions being executed by the processing unit, the processing unit being used to implement the steps of the image processing method provided in the first aspect of the present disclosure when executing.

[0017] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: The target diffusion model processes the original image and target description text to obtain the target image. This model is trained on a base model using multiple sample images. The base model includes a fine-tuning model and a preservation model. The preservation model remains unchanged during training, while the fine-tuning model learns the image generation capabilities of the preservation model during training, resulting in the final target diffusion model. Because the base model includes both fine-tuning and preservation models, the fine-tuning model continuously learns to generate flawless sample images during training, while the preservation model continuously learns its image generation capabilities. This approach aims to retain the base model's ability to generate images in non-flawed areas as much as possible, while minimizing the generation of flawed images, thus resulting in a target diffusion model with better image generation quality. Ultimately, this allows for the production of target images that better match the target description text and are of higher quality.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0020] Figure 1 This is a flowchart illustrating an image processing method according to an exemplary embodiment.

[0021] Figure 2 This is a flowchart illustrating a method for optimizing a base model according to an exemplary embodiment.

[0022] Figure 3 This is a flowchart illustrating a method for determining loss according to an exemplary embodiment.

[0023] Figure 4 This is a flowchart illustrating another method for determining loss according to an exemplary embodiment.

[0024] Figure 5 This is a data processing flow during the training process of a target diffusion model, as illustrated in an exemplary embodiment.

[0025] Figure 6 This is a block diagram illustrating an image processing apparatus according to an exemplary embodiment.

[0026] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment.

[0027] Figure 8 This is a block diagram illustrating a chip system according to an exemplary embodiment. Detailed Implementation

[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0029] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0030] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.

[0031] Diffusion models have developed rapidly in recent years and are widely used in image generation and editing tasks. The original diffusion model requires 1000 iterations to generate a satisfactory image. Therefore, research on large diffusion models has largely focused on reducing the number of iterations. Existing research has achieved near-lossless reduction of the iteration count to 8 steps or even lower. Regarding the generated image quality, diffusion models almost inevitably produce some imperfections, such as blurry watermarks at the edges of the image.

[0032] To address the aforementioned technical problems, this disclosure provides an image processing method, apparatus, device, storage medium, product, and system. Since the base model includes a fine-tuning model and a retention model, the fine-tuning model continuously learns its ability to generate flawless sample images during training, while the retention model continuously learns its image generation capabilities. This allows the base model to retain its generation capabilities in non-flawed areas as much as possible, minimizing the generation of flawed images, thereby obtaining a target diffusion model with better image generation quality. This enables the generation of target images that better match the target description text and have higher quality based on the target diffusion model.

[0033] Figure 1 This is a flowchart illustrating an image processing method according to an exemplary embodiment, such as... Figure 1 As shown, this method can be used in a terminal or server and may include the following steps.

[0034] In step S101, the original image to be processed and the target description text are obtained.

[0035] In this embodiment, the original image to be processed may be a noisy image or a blurry, low-quality image. The target descriptive text is used to describe the features of the image to be generated, such as the image's subject and style.

[0036] In step S102, the original image and target description text are processed by the target diffusion model to obtain the target image. The target diffusion model is obtained by training a base model based on multiple sample images. The base model includes a fine-tuning model and a maintenance model. The maintenance model remains unchanged during training. The fine-tuning model is used to learn the image generation capability of the maintenance model during training so as to obtain the target diffusion model after training.

[0037] In this embodiment, a base model is first trained based on multiple sample images to obtain a target diffusion model that can generate images that closely match the target description and are free of defects. The sample images are defect-free images. The base model can include a fine-tuning model and a preservation model, where the pre-trained fine-tuning model and preservation model can be identical. For example, a copy of the pre-trained diffusion model can be used as the preservation model, and another copy as the fine-tuning model. During training, the parameters of the preservation model are fixed, while the parameters of the fine-tuning model are adjustable. During training, the fine-tuning model continuously learns its ability to generate defect-free sample images by optimizing its parameters, and continuously learns the image generation ability of the preservation model. This ensures that the base model's generation ability in defect-free areas is preserved as much as possible, while minimizing the generation of defects, resulting in a target diffusion model with better image generation quality. This allows for the generation of target images that better match the target description and are of higher quality.

[0038] In one possible implementation, the target diffusion model can be trained through the following steps: Multiple sample images are divided into training and test sets. The base model is trained and optimized multiple times using the training set so that the fine-tuned model learns to maintain the model's image generation ability during training. Training is stopped when the fine-tuned model converges, and a candidate diffusion model is obtained. The candidate diffusion model is tested and optimized multiple times using the test set to obtain the target diffusion model.

[0039] In this implementation, taking watermarks as an example, watermarks are a type of text, and any mature text detection model can be chosen to detect them. The diffusion model can be trained on the Laion dataset, where each image is accompanied by corresponding descriptive text, and some images contain watermarks. The text detection model can be used to filter out images containing text in the Laion dataset, leaving clean images as sample images. Multiple sample images can be divided into training and testing sets. First, the base model is trained multiple times using the training set, and after each round of training, the parameters of the fine-tuned model in the base model are optimized. This allows the fine-tuned model to continuously learn its ability to generate flawless sample images during training, while maintaining its image generation capabilities. Training stops when the fine-tuned model converges, resulting in a candidate diffusion model, at which point training on the corresponding training set ends. Convergence of the fine-tuned model is defined as the loss no longer changing or changing within a small preset range after multiple rounds of training. Then, the candidate diffusion model is tested and optimized again using the testing set to obtain a target diffusion model with better image generation quality.

[0040] Figure 2 This is a flowchart illustrating a method for optimizing a base model according to an exemplary embodiment, such as... Figure 2 As shown, in one possible implementation, the base model is trained and optimized through multiple rounds using a training set, which may include the following steps: In step S201, the base model is trained iteratively multiple times using the training set.

[0041] In this embodiment, the base model is trained iteratively through a training set. For each iteration, a certain number of sample images are acquired and input into the base model, outputting the corresponding prediction results.

[0042] In step S202, for any round of iterative training, the fine-tuning loss corresponding to the fine-tuning model and the holding loss corresponding to the holding model are determined in that round of iterative training.

[0043] In this embodiment, the fine-tuning loss corresponding to the fine-tuning model and the retention loss corresponding to the retention model in this training round can be determined based on the outputs of the fine-tuning model and the retention model in this training round.

[0044] In step S203, the fine-tuned model corresponding to this round of iterative training is optimized by fine-tuning the loss and maintaining the loss.

[0045] In this implementation, the total loss for each training round can be determined based on the fine-tuning loss and the retention loss. For example, the sum of the fine-tuning loss and the retention loss can be used to determine the total loss. Based on the total loss, adjustment parameters are determined to optimize the parameters of the fine-tuned model, enabling it to output more accurate images. By iteratively optimizing the base model using the training set, and optimizing the fine-tuned model based on its fine-tuning loss and the retention loss of the retention model during training, the fine-tuned model can not only learn image generation capabilities during training but also inherit the image generation capabilities of the retention model. This allows the trained target diffusion model to minimize image generation defects during image generation.

[0046] Figure 3 This is a flowchart illustrating a method for determining loss according to an exemplary embodiment, such as... Figure 3 As shown, in one possible implementation, for any round of iterative training, determining the fine-tuning loss corresponding to the fine-tuning model and the retention loss corresponding to the retention model in that round of iterative training may include: In step S301, for any round of iterative training, the target sample image corresponding to that round of training is obtained. The target sample image carries a target label, which includes sample description text.

[0047] In this embodiment, for any given training iteration, the target sample image used in that iteration is obtained, and this target sample image has a corresponding target label. The target label may include sample descriptive text, which is used to describe the features of the target sample image, such as the theme and style of the image.

[0048] In step S302, the target sample image is processed to obtain a target mask image visible in a preset area.

[0049] In this embodiment, each target sample image is processed separately. The target sample image can be extracted to obtain a portion of its content. The preset region can be a region of a preset size located in the center of the target sample image. A mask can be used to extract the target sample image, resulting in a target mask image where the preset region is visible.

[0050] In step S303, the target sample image, target mask image, and target label are processed by the fine-tuning model to obtain the fine-tuning result.

[0051] In step S304, the target sample image, target mask image, and target label are processed by the preservation model to obtain the preservation result.

[0052] In step S305, based on the fine-tuning results, the fine-tuning loss corresponding to the fine-tuned model in this round of iterative training is obtained, and based on the fine-tuning results and the holding results, the holding loss corresponding to the holding model in this round of iterative training is obtained.

[0053] In this embodiment, considering that the defects in the generated image mainly occur in the edge region of the generated image, the central region can be set as the preset region, and a mask image can be generated to obtain the fine-tuning result and the preservation result, and obtain the corresponding fine-tuning loss and preservation loss, so as to learn the image generation capability that does not generate defects in the edge region and minimize the generation of defects.

[0054] In one possible implementation, the method for processing the target sample image to obtain a target mask image visible in a preset region can be as follows: Generate a mask of the same size as the target sample image. The mask of the preset region of the mask is 0, and the mask of other regions of the preset region of the mask is 1. Based on the mask and the target sample image, obtain the target mask image that is visible in the preset region.

[0055] In this embodiment, a mask of the same size as the target sample image is generated. For example, if the target sample image is 512 pixels in both length and width, a mask of the same size can be generated, with the mask value set to 0 for the central 256×256 region and the mask value set to 1 for the surrounding regions. Then, the target mask image can be obtained based on the formula: Target mask image = Target sample image × (1 - Mask). The target mask image is such that the predetermined central region is visible, while the surrounding regions are invisible. This is to maintain generative capability while suppressing defects during training.

[0056] In one possible implementation, the fine-tuning result includes a first random accumulated noise, a first predicted noise, and a predicted defect detection result; The fine-tuning process involves processing the target sample image, target mask image, and target label using a fine-tuning model to obtain the fine-tuning result. This may include the following steps: The target sample image, target mask image, and target label are processed by a fine-tuning model to obtain the first sample image code, the first mask code, and the first descriptive text code corresponding to the target label. Noise is added to the target sample image for a preset number of steps to obtain the first random accumulated noise. This first random accumulated noise is the accumulated noise corresponding to multiple steps after a random number of steps during the noise addition process. The first noise image code, the first descriptive text code, the first mask code, and the random steps are processed by the fine-tuning model to obtain the first predicted noise. The first noise image code is obtained by processing the first sample image code with the first random accumulated noise. The predicted image code is processed by the fine-tuning model to obtain the target diffusion map. The predicted image code is obtained by processing the first noise image code with the first predicted noise. The target diffusion map is processed by a target defect detection model to obtain the predicted defect detection result.

[0057] In this embodiment, the fine-tuning model may include a first text encoder, a first image encoder, a first semantic denoiser, and a first image decoder. For the fine-tuning model, the first image encoder can process the target sample image and the target mask image respectively to obtain the first sample image code corresponding to the target sample image and the first mask code corresponding to the mask image. The first text encoder can process the target label to obtain the first descriptive text code corresponding to the target label. Then, a noise additive is used to add noise to the target sample image for a preset number of steps. This preset number of steps can be 1000 steps, and a Gaussian distribution can be gradually added to the target sample image. A random step number t is randomly selected from 0 to 1000 steps, for example, t is 500, and equivalent noise accumulation is performed on the subsequent steps, i.e., from 500 to 1000 steps, to obtain the first randomly accumulated noise.

[0058] Then, the first sample image encoding is noise-added using a first random accumulated noise to obtain a first noise image encoding. The first noise image encoding, the first descriptive text encoding, the first mask encoding, and the random step count are then processed by a first semantic denoiser to obtain first prediction noise, which is the prediction result for the t-step accumulated noise. The first noise image encoding is then denoised using the first prediction noise to obtain a prediction image encoding. The prediction image encoding is then processed by a first image decoder to obtain a target diffusion map. Finally, a target defect detection model is used to detect defects in the target diffusion map to obtain the predicted defect detection result. If the defect is a watermark, the target diffusion map may include a text detection model. The predicted defect detection result may include the detected defect region, which can be scaled to the same size as the first sample image encoding. The target defect detection model can be obtained by training a basic detection model based on multiple sample images carrying the target defect.

[0059] In this embodiment, by determining a random number of steps and obtaining accumulated noise in each training process, a first predicted noise can be predicted, and a target diffusion map can be obtained based on the first predicted noise. By using a random number of steps in each training process, multiple training iterations can cover multiple steps from 0 to 1000, making the predicted first predicted noise more accurate and thus generating a more accurate target diffusion map.

[0060] In one possible implementation, the target defect detection model can be trained through the following steps: Multiple training images are acquired, each carrying a labeled tag representing the target defect. The basic detection model is then trained iteratively through these images. After each iteration, the defect detection result is obtained. The detection loss is determined based on the result and the labeled tag of the training image. The basic detection model is then optimized based on the detection loss. Training is stopped once the basic detection model converges, resulting in the target defect detection model.

[0061] In this embodiment, the basic detection model is trained through multiple rounds of iterative training using multiple sample training images after the target defect is annotated, thereby obtaining a target defect detection model capable of detecting the target defect. The basic detection model can be a DETR (Detection with Transformer) model.

[0062] In one possible implementation, the preservation result includes a second prediction noise; the preservation result is obtained by processing the target sample image, target mask image, and target label through a preservation model, which may include the following steps: By preserving the model and processing the target sample image, target mask image, and target label respectively, the second sample image code, the second mask code, and the second descriptive text code corresponding to the target sample image, are obtained. The model then adds noise to the target sample image for a preset number of steps to obtain second random accumulated noise. This second random accumulated noise is the accumulated noise corresponding to multiple steps after a random number of steps during the preserving model's noise addition process. Finally, the model processes the second noise image code, the second descriptive text code, the second mask code, and the random number of steps to obtain second predicted noise. The second noise image code is obtained by processing the second sample image code with the second random accumulated noise.

[0063] In this embodiment, the preservation model may include a second text encoder, a second image encoder, a second semantic denoiser, and a second image decoder. For the preservation model, the second image encoder can process the target sample image and the target mask image separately to obtain the second sample image code corresponding to the target sample image and the second mask code corresponding to the mask image. The second text encoder can process the target label to obtain the second descriptive text code corresponding to the target label. Then, a noise additive is used to add noise to the target sample image for a preset number of steps, which can be 1000 steps, progressively adding a Gaussian distribution to the target sample image. Within steps 0 to 1000, a random step number t, the same as that of the fine-tuning model, is used, for example, 500. For the multiple steps after the random step number, i.e., steps 500 to 1000, equivalent noise accumulation is performed to obtain the second randomly accumulated noise.

[0064] Then, the second sample image encoding is noise-added using a second random accumulated noise to obtain the second noise image encoding. The second noise image encoding, the second descriptive text encoding, the second mask encoding, and the random step count are then processed by a second semantic denoiser to obtain the second prediction noise.

[0065] In this embodiment, by determining the same number of random steps as in the fine-tuning training during each training process and obtaining accumulated noise, a second prediction noise can be predicted. By using a random number of steps in each training process, multiple training iterations can cover multiple steps from 0 to 1000, thereby making the predicted second prediction noise more accurate.

[0066] Figure 4 This is a flowchart illustrating another method for determining loss according to an exemplary embodiment, such as... Figure 4 As shown, in one possible implementation, the fine-tuning loss corresponding to the fine-tuned model in this round of iterative training is obtained based on the fine-tuning result, and the retention loss corresponding to the retained model in this round of iterative training is obtained based on the fine-tuning result and the retention result. This may include the following steps: In step S401, the product of the absolute value of the difference between the first random accumulated noise and the first predicted noise and the predicted defect detection result is determined as the fine-tuning loss corresponding to the fine-tuning model in this round of iterative training.

[0067] In step S402, the product of the absolute value of the difference between the first prediction noise and the second prediction noise and the expected result is determined as the holding loss corresponding to the model in this round of iterative training. The expected result is the value obtained by subtracting the predicted defect detection result from 1.

[0068] In this embodiment, the formula for calculating the fine-tuning loss can be: Fine-tuning loss = |first prediction noise –first random accumulated noise| × prediction defect detection result.

[0069] The formula for calculating retention loss can be: Retention loss = |First prediction noise – Second prediction noise| × (1 – Predicted defect detection result).

[0070] After the above training, the converged and fine-tuned model can inherit and maintain the model's generative ability without generating any more defects.

[0071] Figure 5 This is a data processing flow during the training process of a target diffusion model, as illustrated in an exemplary embodiment, such as... Figure 5 As shown, the target sample image, target mask image, and target label can be processed by fine-tuning the model to obtain the fine-tuning result, and the target sample image, target mask image, and target label can be processed by maintaining the model to obtain the maintenance result.

[0072] Figure 6 This is a block diagram illustrating an image processing apparatus according to an exemplary embodiment. (Refer to...) Figure 6 The image processing device 600 includes an acquisition module 601 and a first acquisition module 602.

[0073] The acquisition module 601 is configured to acquire the original image to be processed and the target description text; The first obtaining module 602 is configured to process the original image and the target description text through a target diffusion model to obtain a target image. The target diffusion model is obtained by training a base model based on multiple sample images. The base model includes a fine-tuning model and a maintenance model. The maintenance model remains unchanged during training. The fine-tuning model is used to learn the image generation capability of the maintenance model during training so as to obtain the target diffusion model after training.

[0074] Optionally, the image processing apparatus 600 further includes: The classification module is configured to divide the plurality of sample images into a training set and a test set; The training module is configured to perform multiple rounds of training optimization on the base model using the training set, so that the fine-tuned model learns the image generation capability of the model during training. The second acquisition module is configured to stop training and obtain a candidate diffusion model when the fine-tuned model converges. The third acquisition module is configured to perform multiple rounds of testing and optimization on the candidate diffusion model using the test set to obtain the target diffusion model.

[0075] Optionally, the training module includes: The training submodule is configured to perform multiple rounds of iterative training on the base model using the training set; The determination submodule is configured to determine the fine-tuning loss for the fine-tuning model and the retention loss for the retention model in any given round of training iterations. The optimization submodule is configured to optimize the fine-tuned model corresponding to the round of iterative training using the fine-tuning loss and the maintenance loss.

[0076] Optionally, the determining submodule includes: The acquisition unit is configured to acquire the target sample image corresponding to any round of training iteration, wherein the target sample image carries a target label, and the target label includes sample description text. The first obtaining unit is configured to process the target sample image to obtain a target mask image visible in a preset area; The second obtaining unit is configured to process the target sample image, the target mask image, and the target label through the fine-tuning model to obtain the fine-tuning result; The third obtaining unit is configured to process the target sample image, the target mask image, and the target label through the preservation model to obtain the preservation result; The fourth obtaining unit is configured to obtain the fine-tuning loss corresponding to the fine-tuned model in the current round of iterative training based on the fine-tuning result, and to obtain the holding loss corresponding to the holding model in the current round of iterative training based on the fine-tuning result and the holding result.

[0077] Optionally, the first obtaining unit includes: The generation subunit is configured to generate a mask of the same size as the target sample image, wherein the mask of a preset region of the mask is 0, and the mask of other regions of the preset region of the mask is 1; The first obtaining subunit is configured to obtain a target mask image visible in the preset region based on the mask and the target sample image.

[0078] Optionally, the fine-tuning result includes a first random accumulated noise, a first predicted noise, and a predicted defect detection result; The second obtaining unit includes: The second obtaining subunit is configured to process the target sample image, the target mask image, and the target label respectively through the fine-tuning model to obtain the first sample image code corresponding to the target sample image, the first mask image code corresponding to the mask image, and the first descriptive text code corresponding to the target label; The third obtaining subunit is configured to add noise to the target sample image for a preset number of steps to obtain the first random accumulated noise. The first random accumulated noise is the accumulated noise corresponding to multiple steps after a random number of steps during the process of adding noise to the target sample image by the fine-tuning model for a preset number of steps. The fourth obtaining subunit is configured to process the first noise map encoding, the first descriptive text encoding, the first mask encoding, and the random step number through the fine-tuning model to obtain the first predicted noise. The first noise map encoding is obtained by processing the first sample map encoding with the first random accumulated noise. The fifth obtaining subunit is configured to process the prediction map encoding through the fine-tuning model to obtain the target diffusion map, wherein the prediction map encoding is obtained by processing the first noise map encoding through the first prediction noise; The sixth obtaining subunit is configured to process the target diffusion map through the target defect detection model to obtain the predicted defect detection result.

[0079] Optionally, the retention result includes a second prediction noise; The third obtaining unit includes: The seventh obtaining subunit is configured to process the target sample image, the target mask image, and the target label respectively through the preservation model to obtain the second sample image code corresponding to the target sample image, the second mask image code corresponding to the mask image, and the second descriptive text code corresponding to the target label; The eighth obtaining subunit is configured to add noise to the target sample image for a preset number of steps to obtain a second random accumulated noise. The second random accumulated noise is the accumulated noise corresponding to multiple steps after the random number of steps during the process of adding noise to the target sample image by the holding model for a preset number of steps. The ninth obtaining subunit is configured to process the second noise map encoding, the second descriptive text encoding, the second mask encoding, and the random step number through the preservation model to obtain the second predicted noise. The second noise map encoding is obtained by processing the second sample map encoding with the second random accumulated noise.

[0080] Optionally, the fourth obtaining unit includes: The first determining subunit is configured to determine the product of the absolute value of the difference between the first random accumulated noise and the first predicted noise and the predicted defect detection result as the fine-tuning loss corresponding to the fine-tuning model in this round of iterative training. The second determining subunit is configured to determine the holding loss corresponding to the model in this round of iterative training by multiplying the absolute value of the difference between the first predicted noise and the second predicted noise with the expected result, wherein the expected result is the value obtained by subtracting the predicted defect detection result from 1.

[0081] Regarding the image processing apparatus 600 in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0082] This disclosure also provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the steps of the image processing method provided in this disclosure.

[0083] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment. For example, the electronic device 700 may be a mobile phone, a computer, a tablet device, a personal digital assistant, etc.

[0084] Reference Figure 7 The electronic device 700 may include one or more of the following components: a processing component 702, a first memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output interface 712, a sensor component 714, and a communication component 716.

[0085] Processing component 702 typically controls the overall operation of electronic device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 702 may include one or more first processors 720 to execute instructions to complete all or part of the steps of the image processing method described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.

[0086] The first memory 704 is configured to store various types of data to support the operation of the electronic device 700. Examples of this data include instructions for any application or method operating on the electronic device 700, contact data, phonebook data, messages, pictures, videos, etc. The first memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0087] Power supply component 706 provides power to various components of electronic device 700. Power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 700.

[0088] Multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0089] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when electronic device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in first memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.

[0090] Input / output interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, start buttons, and lock buttons.

[0091] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of electronic device 700. For example, sensor assembly 714 can detect the on / off state of electronic device 700, the relative positioning of components such as the display and keypad of electronic device 700, changes in position of electronic device 700 or a component of electronic device 700, the presence or absence of user contact with electronic device 700, orientation or acceleration / deceleration of electronic device 700, and temperature changes of electronic device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0092] Communication component 716 is configured to facilitate wired or wireless communication between electronic device 700 and other devices. Electronic device 700 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0093] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the image processing method described above.

[0094] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a first memory 704 including instructions, which can be executed by a first processor 720 of an electronic device 700 to complete the image processing method described above. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0095] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a programmable device, the computer program having a code portion for performing the image processing method described above when executed by the programmable device.

[0096] Figure 8 This is a block diagram illustrating a chip system according to an exemplary embodiment. Some embodiments of this disclosure also provide a chip system, such as... Figure 8 As shown, the chip system includes at least one second processor 801 and at least one interface circuit 802. The second processor 801 and the interface circuit 802 are interconnected via lines. For example, the interface circuit 802 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 802 can be used to send signals to other devices (e.g., the second processor 801). Exemplarily, the interface circuit 802 can read instructions stored in the memory and send those instructions to the second processor 801. When the instructions are executed by the second processor 801, the image processing device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete components, and some embodiments of this disclosure do not specifically limit this.

[0097] In some embodiments of this disclosure, the interface circuit 802 can acquire data, program instructions, and / or information from the internal storage area of ​​the chip system; it can also acquire data, program instructions, and / or information from outside the chip system.

[0098] Optionally, the chip system also includes a second memory 803 for storing necessary computer programs and data.

[0099] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented through hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the described functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.

[0100] In the above detailed description, terms such as "center," "upper," "lower," "left," and "right" indicate direction or positional relationship. Since components of the described device can be positioned in multiple different orientations, these directional terms are for illustrative purposes and not restrictive. It should be understood that other aspects can be utilized and structural or logical changes can be made without departing from the concept of this disclosure. Therefore, the following detailed description should not be considered limiting.

[0101] It should be understood that, unless otherwise specifically indicated, features of various embodiments of this disclosure described herein can be combined with each other.

[0102] Although terms such as “first,” “second,” and “third” may be used herein to describe various components, parts, regions, layers, or sections, these components, parts, regions, layers, or sections are not limited to these terms. Rather, these terms are used only to distinguish one component, part, region, layer, or section from another. Therefore, without departing from the teachings of the examples described herein, the first component, part, region, layer, or section mentioned in the examples may also be referred to as the second component, part, region, layer, or section. Furthermore, the terms “first” and “second” are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as “first” or “second” may explicitly or implicitly include at least one of that feature. In the description herein, “a plurality” means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0103] Furthermore, the term “exemplary” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “exemplary” is not necessarily to be construed as advantageous compared to other aspects or designs. Rather, the use of the term “exemplary” is intended to present the concept in a concrete manner. As used herein, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless otherwise specified or clear from the context, “X applies A or B” is intended to mean any of the natural inclusive arrangements. That is, “X applies A or B” satisfies any of the foregoing instances if X applies A; X applies B; or both X applies A and B. Additionally, unless otherwise specified or clear from the context to refer to the singular form, the articles “a” and “an” as used in this application and the appended claims are generally understood to mean “one or more.”

[0104] Similarly, although this disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding the specification and drawings. This disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terminology used to describe such components is intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if structurally not equivalent to the disclosed structure. Furthermore, although specific features of this disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations, as may be desired and advantageous for any given or particular application. Moreover, with regard to the terms “comprising,” “owning,” “having,” “having,” or variations thereof as used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term “including.”

[0105] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.

[0106] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. An image processing method, characterized in that, include: Acquire the original image to be processed and the target description text; The original image and the target description text are processed by a target diffusion model to obtain a target image. The target diffusion model is obtained by training a base model based on multiple sample images. The base model includes a fine-tuning model and a maintenance model. The maintenance model remains unchanged during training. The fine-tuning model is used to learn the image generation capability of the maintenance model during training so as to obtain the target diffusion model after training.

2. The image processing method according to claim 1, characterized in that, The target diffusion model is trained through the following steps: The multiple sample images are divided into a training set and a test set; The base model is trained and optimized in multiple rounds using the training set, so that the fine-tuned model learns the image generation capability of the model during the training process. If the fine-tuned model converges, training is stopped, and a candidate diffusion model is obtained. The candidate diffusion model is optimized through multiple rounds of testing using the test set to obtain the target diffusion model.

3. The image processing method according to claim 2, characterized in that, The step of performing multiple rounds of training and optimization on the base model using the training set includes: The base model is trained iteratively in multiple rounds using the training set. For any round of training iterations, determine the fine-tuning loss for the fine-tuning model and the retention loss for the retention model in that round of training iterations. The fine-tuning loss and the maintenance loss are used to optimize the fine-tuning model corresponding to this round of iterative training.

4. The image processing method according to claim 3, characterized in that, For any given round of training iterations, determining the fine-tuning loss for the fine-tuned model and the retention loss for the retained model in that round of training includes: For any round of training iteration, obtain the target sample image corresponding to that round of training. The target sample image carries a target label, which includes sample description text. The target sample image is processed to obtain a target mask image visible in a preset area; The target sample image, the target mask image, and the target label are processed using the fine-tuning model to obtain the fine-tuning result; The target sample image, the target mask image, and the target label are processed using the preservation model to obtain the preservation result; Based on the fine-tuning results, the fine-tuning loss corresponding to the fine-tuned model in this round of iterative training is obtained, and based on the fine-tuning results and the holding results, the holding loss corresponding to the holding model in this round of iterative training is obtained.

5. The image processing method according to claim 4, characterized in that, The process of processing the target sample image to obtain a target mask image visible in a preset region includes: Generate a mask of the same size as the target sample image, wherein the mask of the preset region of the mask is 0, and the mask of the other regions of the preset region of the mask is 1; Based on the mask and the target sample image, a target mask image visible in the preset area is obtained.

6. The image processing method according to claim 4, characterized in that, The fine-tuning results include the first random accumulated noise, the first predicted noise, and the predicted defect detection results; The process of processing the target sample image, the target mask image, and the target label using the fine-tuning model to obtain the fine-tuning result includes: The target sample image, the target mask image, and the target label are processed by the fine-tuning model to obtain the first sample image code corresponding to the target sample image, the first mask image code corresponding to the mask image, and the first descriptive text code corresponding to the target label. The target sample image is subjected to noise addition for a preset number of steps to obtain the first random accumulated noise. The first random accumulated noise is the accumulated noise corresponding to multiple steps after a random number of steps during the process of the fine-tuning model adding noise to the target sample image for a preset number of steps. The first predicted noise is obtained by processing the first noise map encoding, the first descriptive text encoding, the first mask encoding, and the random step number through the fine-tuning model. The first noise map encoding is obtained by processing the first sample map encoding with the first random accumulated noise. The target diffusion map is obtained by processing the prediction map encoding through the fine-tuning model. The prediction map encoding is obtained by processing the first noise map encoding through the first prediction noise. The target diffusion map is processed by the target defect detection model to obtain the predicted defect detection result.

7. The image processing method according to claim 6, characterized in that, The retention result includes a second prediction noise; The process of processing the target sample image, the target mask image, and the target label using the preservation model to obtain the preservation result includes: The target sample image, the target mask image, and the target label are processed by the preservation model to obtain the second sample image code corresponding to the target sample image, the second mask image code corresponding to the mask image, and the second descriptive text code corresponding to the target label. The target sample image is subjected to noise addition for a preset number of steps to obtain a second random accumulated noise. The second random accumulated noise is the accumulated noise corresponding to multiple steps after the random number of steps during the process of the preservation model adding noise to the target sample image for a preset number of steps. The second predicted noise is obtained by processing the second noise map encoding, the second descriptive text encoding, the second mask encoding, and the random step number through the preservation model. The second noise map encoding is obtained by processing the second sample map encoding with the second random accumulated noise.

8. The image processing method according to claim 7, characterized in that, The step of obtaining the fine-tuning loss corresponding to the fine-tuned model in this round of iterative training based on the fine-tuning result, and obtaining the retention loss corresponding to the retained model in this round of iterative training based on the fine-tuning result and the retention result, includes: The product of the absolute value of the difference between the first random accumulated noise and the first predicted noise and the predicted defect detection result is determined as the fine-tuning loss corresponding to the fine-tuning model in this round of iterative training. The product of the absolute value of the difference between the first predicted noise and the second predicted noise and the expected result is determined as the holding loss corresponding to the model in this round of iterative training. The expected result is the value obtained by subtracting the predicted defect detection result from 1.

9. An image processing apparatus, characterized in that, include: The acquisition module is configured to acquire the raw image to be processed; The first acquisition module is configured to perform noise reduction processing on the original image using a target diffusion model to obtain a target image. The target diffusion model is obtained by training a base model based on multiple sample images. The base model includes a fine-tuning model and a maintenance model. The maintenance model remains unchanged during training. The fine-tuning model is used to learn the image generation capability of the maintenance model during training so as to obtain the target diffusion model after training.

10. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to perform the steps of the image processing method according to any one of claims 1 to 8 when executing.

11. A computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the steps of the image processing method according to any one of claims 1 to 8.

12. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the steps of the image processing method according to any one of claims 1 to 8.

13. A chip system, characterized in that, The chip system includes a processing unit and an interface circuit. The processing unit obtains program instructions through the interface circuit, and the program instructions are executed by the processing unit. The processing unit is used to perform the steps of the image processing method as described in any one of claims 1 to 8.