Image processing method and apparatus, and computing device cluster
Patent Information
- Application Number
- PCT/CN2026/079967
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-02-25
- Publication Date
- 2026-09-24
Smart Images

Figure CN2026079967_24092026_PF_FP_ABST
Abstract
Description
An image processing method, apparatus, and computing device cluster
[0001] This application claims priority to Chinese Patent Application No. 202510329404.3, filed on March 19, 2025, entitled “An Image Processing Method, Apparatus and Computing Device Cluster”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence technology, and in particular to an image processing method, apparatus, and computing device cluster. Background Technology
[0003] Image super-resolution (SR) technology aims to recover a high-quality image from a low-quality one. It is a crucial technique in computer vision image processing and has numerous applications in real life. For example, in medicine, SR technology enhances the quality of images from magnetic resonance (MR) and computed tomography (CT) scans to improve diagnostic accuracy through clearer structures. In mobile phone photography, SR technology increases image resolution to provide more detail, which is especially important when using telephoto lenses with poor imaging quality. In film and entertainment, SR technology improves the picture quality of movies and games, providing viewers with a better visual experience. In practical applications, many factors contribute to image quality degradation, such as motion blur, JPEG compression, and noise. This degradation is known as image degradation. Due to the influence of different imaging devices, such as different mobile phones, camera modules, and image signal processing pathways, the degradation patterns vary. Therefore, how to recover high-quality images to improve the effectiveness of image super-resolution is a pressing technical problem that needs to be solved. Summary of the Invention
[0004] This application provides an image processing method, apparatus, and computing device cluster that can accurately extract degradation features that characterize degradation information from an image and input the degradation features as conditions into a diffusion model, so that the diffusion model can accurately perceive the degradation information in the image when processing the image, thereby recovering a high-quality image and improving the image super-resolution results.
[0005] In a first aspect, this application provides an image processing method, comprising: acquiring a first image; extracting features from degradation information in the first image to obtain degradation features; and inputting the first image and first data into a diffusion model to obtain a second image, wherein the first data includes degradation features, and the image quality of the second image is higher than that of the first image.
[0006] In the image processing method provided in this application, feature extraction is performed on the degradation information in the first image to obtain the degradation features in the first image. The extracted degradation features are then used as conditions to input into the diffusion model, so that the diffusion model can accurately perceive the degradation information of the image during the processing of the first image, thereby recovering a high-quality image and improving the super-resolution effect of the image.
[0007] In this application, a neural network model capable of extracting degradation features can be pre-trained, enabling the model to accurately extract degradation features from images under different degradation modes. Then, by using the trained neural network model to extract degradation information from the first image, the degradation features of the first image can be accurately obtained. Furthermore, since degradation features can characterize the degradation information of the first image, using these features as conditional inputs to the diffusion model allows the diffusion model to accurately perceive the degradation information of the first image, thereby recovering a high-quality second image from the first image and improving the image super-resolution effect.
[0008] In one possible implementation, the first data further includes content features. Before inputting the first image and the first data into the diffusion model, the method further includes: extracting features from the content of the first image to obtain content features. In this way, the diffusion model can perceive both the degradation information and the content information of the first image during the image restoration process, thereby accurately restoring the content information in the first image and improving the image super-resolution effect.
[0009] First, a neural network model capable of extracting content features is pre-trained. Then, the content information in the first image is extracted using the trained neural network model, thus accurately obtaining the content features of the first image.
[0010] In one possible implementation, a first image and first data are input into a diffusion model to obtain a second image, including: injecting degradation features into content features to obtain fused features; and inputting the first image and fused features into the diffusion model to obtain the second image. This allows the degradation features and content features to be fused together, enabling the diffusion model to perceive the correlation between them, helping it better understand the degradation process, thereby improving the super-resolution effect.
[0011] In one possible implementation, degradation features are injected into content features to obtain fused features. This includes: performing a dot product operation on the degradation features using a convolution kernel to obtain modulation parameters; and performing a convolution operation on the content features using the modulation parameters to obtain fused features. In this way, the content features can be adjusted based on the degradation features, thereby fusing the degradation features and content features together.
[0012] In one possible implementation, a third image and a fourth image are acquired, wherein the image quality of the third image is lower than that of the fourth image, and the content in the third image is the same as that in the fourth image. The third image is input into a neural network model to obtain a fifth image. The neural network model includes a first encoder and a first decoder. The first encoder is used to extract features from the content of the third image, and the first decoder is used to generate the fifth image based on the features extracted by the first encoder. A first loss is calculated based on the difference between the fourth and fifth images. The neural network model is trained with the goal of minimizing the first loss. After training, the first encoder in the above neural network model is used to extract features from the content of the first image. This allows for the training of a first encoder for extracting content features. This first encoder can focus on recovering high-quality content from degraded, low-quality images, thus ignoring the degradation itself and extracting the core content of the image.
[0013] In one possible implementation, multiple sixth images with different content under the same degradation mode are acquired, and multiple seventh images with the same content under different degradation modes are acquired. The multiple sixth images are input into a second encoder for degradation information extraction to obtain degradation features for each sixth image. The multiple seventh images are also input into the second encoder for degradation information extraction to obtain degradation features for each seventh image. A second loss is calculated based on the differences between the degradation features of the multiple sixth images, a third loss is calculated based on the differences between the degradation features of the multiple seventh images, and a fourth loss is calculated based on the differences between the degradation features of the sixth and seventh images. The second encoder is trained with the goal of minimizing the second and third losses and maximizing the fourth loss. The trained second encoder is then used to extract features from the degradation information in the first image. In this way, by using the trained second encoder to extract degradation features, similar representations can be extracted from multiple images under the same degradation mode, while different degradation representations can be extracted from multiple images taken in different scenes and with different cameras (different degradation models). This ensures that image content is accurately preserved under various degradation conditions, while accurately representing and distinguishing degradation patterns.
[0014] In one possible implementation, the seventh image is input into the first encoder for feature extraction to obtain the content features of the seventh image; the seventh image is then input into the second encoder for feature extraction to obtain the degradation features of the seventh image; the content features and degradation features of the seventh image are input into the second decoder for processing to obtain the eighth image; a loss is calculated based on the difference between the seventh and eighth images to obtain the sixth loss; the parameters in the first encoder and the second decoder are updated with the goal of minimizing the sixth loss. In this way, by first training the second encoder and the second encoder separately, and then training them together, both encoders can accurately extract the corresponding features.
[0015] In one possible implementation, the first encoder has a progressive downsampling structure; the second encoder has a structure combining progressive downsampling and pooling layers, where the pooling layers are used to convert the features obtained from the last downsampling layer into a global description. Since degradation is often global, this approach can more accurately represent the degradation characteristics of an image.
[0016] In one possible implementation, the diffusion model is either a latent diffusion model or a denoised diffusion probability model.
[0017] Secondly, this application provides an image processing apparatus, comprising: an acquisition module and a processing module. The acquisition module is used to acquire a first image. The processing module is used to extract features from degradation information in the first image to obtain degradation features. The processing module is further used to input the first image and first data into a diffusion model to obtain a second image, wherein the first data includes degradation features, and the image quality of the second image is higher than that of the first image.
[0018] In one possible implementation, the first data further includes the content features; before inputting the first image and the first data into the diffusion model, the processing module is further configured to: extract features from the content in the first image to obtain the content features.
[0019] In one possible implementation, inputting the first data into the diffusion model to obtain the second image includes: injecting the degradation features into the content features to obtain fusion features; inputting the first image and the fusion features into the diffusion model to obtain the second image.
[0020] In one possible implementation, injecting the degradation features into the content features to obtain fused features includes: performing a dot product operation on the degradation features using a convolution kernel to obtain modulation parameters; and performing a convolution operation on the content features using the modulation parameters to obtain the fused features.
[0021] In one possible implementation, the processing module is further configured to: acquire a third image and a fourth image, wherein the image quality of the third image is lower than that of the fourth image, and the content in the third image is the same as that in the fourth image; input the third image into a neural network model to obtain a fifth image, wherein the neural network model includes a first encoder and a first decoder, the first encoder is used to extract features from the content of the third image, and the first decoder is used to generate the fifth image based on the features extracted by the first encoder; calculate a loss based on the difference between the fourth image and the fifth image to obtain a first loss; train the neural network model with the objective of minimizing the first loss, wherein the first encoder in the trained neural network model is used to extract features from the content of the first image.
[0022] In one possible implementation, the processing module is further configured to: acquire multiple sixth images with different content under the same degradation mode, and acquire multiple seventh images with the same content under different degradation modes; input the multiple sixth images to a second encoder for degradation information extraction to obtain degradation features of each sixth image; input the multiple seventh images to the second encoder for degradation information extraction to obtain degradation features of each seventh image; calculate a loss based on the differences between the degradation features of the multiple sixth images to obtain a second loss, and calculate a loss based on the differences between the degradation features of the multiple seventh images to obtain a third loss, and calculate a loss based on the differences between the degradation features of the multiple sixth images and the degradation features of the multiple seventh images to obtain a fourth loss; train the second encoder with the goal of minimizing the second loss and the third loss, and maximizing the fourth loss, wherein the trained second encoder is used to extract features from the degradation information in the first image.
[0023] In one possible implementation, the processing module is further configured to: input the seventh image into the first encoder for feature extraction to obtain the content features of the seventh image; input the seventh image into the second encoder for feature extraction to obtain the degradation features of the seventh image; input the content features and degradation features of the seventh image into the second decoder for processing to obtain the eighth image; perform loss calculation based on the difference between the seventh image and the eighth image to obtain the sixth loss; and update the parameters in the second encoder and the second decoder with the goal of minimizing the sixth loss.
[0024] In one possible implementation, the first encoder has a progressive downsampling structure; the second encoder has a structure that combines progressive downsampling with a pooling layer, wherein the pooling layer is used to convert the features obtained from the final downsampling into a global description.
[0025] In one possible implementation, the diffusion model is either a latent diffusion model or a denoised diffusion probability model.
[0026] Thirdly, this application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method described in the first aspect or any possible implementation of the first aspect.
[0027] Fourthly, this application provides a computer-readable storage medium including computer program instructions that, when executed by a cluster of computing devices, perform the method described in the first aspect or any possible implementation thereof. Exemplarily, the computing device cluster may include one or more computing devices.
[0028] Fifthly, this application provides a computer program product containing instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method described in the first aspect or any possible implementation thereof. Exemplarily, the cluster of computing devices may include one or more computing devices.
[0029] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0030] Figure 1 is a schematic diagram of the architecture of an image processing system provided in an embodiment of this application;
[0031] Figure 2 is a schematic diagram of a degradation extraction module provided in an embodiment of this application;
[0032] Figure 3 is a schematic diagram of the architecture of another image processing system provided in an embodiment of this application;
[0033] Figure 4 is a schematic diagram of the training process of a content extraction module and a degradation extraction module provided in an embodiment of this application;
[0034] Figure 5 is a schematic diagram of the architecture of another image processing system provided in an embodiment of this application;
[0035] Figure 6 is a schematic diagram of another degradation extraction module provided in an embodiment of this application;
[0036] Figure 7 is a schematic diagram of user interaction with a cloud computing platform provided in an embodiment of this application;
[0037] Figure 8 is a flowchart illustrating an image processing method provided in an embodiment of this application;
[0038] Figure 9 is a schematic diagram comparing image processing effects provided in an embodiment of this application;
[0039] Figure 10 is a schematic diagram of the structure of an image processing device provided in an embodiment of this application;
[0040] Figure 11 is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0041] Figure 12 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;
[0042] Figure 13 is a schematic diagram of another computing device cluster provided in an embodiment of this application. Detailed Implementation
[0043] In this article, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The symbol " / " in this article indicates that the related objects are in an "or" relationship; for example, A / B means A or B.
[0044] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. For example, "first response message" and "second response message," etc., are used to distinguish different response messages, not to describe a specific order of response messages.
[0045] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0046] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple processing units means two or more processing units, multiple elements means two or more elements, etc.
[0047] First, the relevant technical terms used in this application will be introduced.
[0048] (1) Degradation types and degradation characteristics
[0049] Image degradation types refer to the classification of various causes or processes that lead to a decrease in image quality. For example, degradation types may include one or more of blur degradation, distortion degradation, and noise degradation. Image degradation characteristics are a specific description of the degradation effects exhibited during the degradation process, encompassing various phenomena, manifestations, and inherent properties that occur during degradation. In other words, "degradation type" is a classification description of the causes of image degradation, while "degradation characteristics" are a detailed characterization of the degradation effects, including the degree of blur, noise type, and distortion form. Together, they constitute a comprehensive understanding of image degradation phenomena.
[0050] (2) Image quality
[0051] Image quality is a multi-dimensional concept, which may include, but is not limited to, one or more of the following: sharpness, resolution, color, brightness, noise, sharpness, contrast, and artifacts. For example, when some of an image's parameters are improved while others are reduced, judging image quality can be done by comprehensively considering the impact of each parameter on practical applications. For instance, the key to color is to realistically reproduce the scene, avoiding oversaturation or distortion; brightness needs to be moderate—too high will be glaring and lose detail, while too low will affect sharpness. Therefore, image quality can be evaluated by combining subjective perception with objective indicators, and the priority of each parameter should be weighed according to specific needs to achieve the best visual effect.
[0052] Next, the technical solution provided in this application will be introduced.
[0053] Image quality degradation is typically caused by a variety of factors, a phenomenon known as image degradation. In practical applications, degradation can be caused by various factors, such as motion blur, JPEG compression, and various types of noise, all of which can impair image quality. Due to the influence of different imaging devices, such as different mobile phones, different camera modules, and the entire image signal processor (ISP) pathway behind them, the various degradation modes are not the same. Therefore, for super-resolution technology to effectively recover high-quality images in various scenarios, the extraction of image degradation characteristics is particularly important.
[0054] In super-resolution imaging, a vision-language model can be employed to improve image inpainting quality. The core idea of this approach is to combine visual and linguistic information to enhance inpainting quality. Especially in complex and diverse natural scenes, by introducing linguistic descriptions, the model can better understand the image content and the inpainting target, thus maintaining semantic consistency and visual realism during the inpainting process. This method allows users to control the inpainting process through language commands, improving the flexibility and accuracy of the inpainting. However, limited by the degradation patterns of the constructed training data, this method has limited effectiveness and cannot accurately capture more nuanced degradation representations. For various reasons, simple descriptions using language cannot accurately express the degradation characteristics of real-world scenes.
[0055] Alternatively, some hand-designed operators are used to simulate degradation to obtain a large amount of training data. This training data is then used to train the neural network model. This approach has achieved good results in some scenarios, but unresponsiveness still occurs in some unseen scenarios.
[0056] In view of this, embodiments of this application provide an image processing method that can effectively decouple the content and degradation characteristics in an image, thereby accurately representing the degradation features of a real scene and improving the effect of super-resolution technology.
[0057] For example, Figure 1 shows a schematic diagram of the architecture of an image processing system provided in an embodiment of this application. As shown in Figure 1, the image processing system 100 mainly includes: a degradation extraction module 110 and a diffusion model 120.
[0058] The degradation extraction module 110 is mainly used to extract degradation features from the low-quality image to be processed. These degradation features characterize the degradation information of the low-quality image, such as degradation type and / or degradation characteristics. In this embodiment, the design goal of the degradation extraction module 110 is that, under the same degradation mode, multiple images should yield similar representations. For images captured in different scenes and by different cameras, multiple images should yield different degradation representations. This helps to distinguish different types of degradation and accurately describe the characteristics of each type of degradation.
[0059] The training method for the degradation extraction module 110 is as follows: First, different degradation models are constructed to simulate different degradation patterns. Then, multiple sets of positive samples are constructed using images with different content under the same degradation pattern, and multiple sets of negative samples are constructed using images with the same content under different degradation patterns, to obtain multiple sets of positive and negative sample training data pairs. Finally, the degradation extraction module 110 is trained by contrast using the positive and negative sample training data pairs, thereby enabling the degradation extraction module 110 to effectively extract degradation features. The goal of this training method is to create a distance between different degradation patterns in the feature space, while bringing images under the same degradation pattern closer together, thus clearly distinguishing and representing various degradations. In this way, by comparing multiple combinations of positive and negative sample training data pairs, the degradation extraction module 110 can learn the commonalities and differences of degradation patterns. For example, as shown in Figure 4, in part ②, the encoder can be understood as the degradation extraction module 110. In part ② of Figure 4, the encoder (or "second encoder") can be trained by contrastive learning. For example, the training process for the second encoder can be as follows: First, acquire multiple sixth images with different content under the same degradation mode, and acquire multiple seventh images with the same content under different degradation modes, thus constructing positive and negative sample training data pairs. Then, input the multiple sixth images into the second encoder to extract degradation information, obtaining the degradation features of each sixth image; and input the multiple seventh images into the second encoder to extract degradation information, obtaining the degradation features of each seventh image. Next, calculate the loss based on the differences between the degradation features of each sixth image to obtain the second loss; calculate the loss based on the differences between the degradation features of each seventh image to obtain the third loss; and calculate the loss based on the differences between the degradation features of each sixth image and the degradation features of each seventh image to obtain the fourth loss. Finally, train the second encoder with the goal of minimizing the second and third losses and maximizing the fourth loss. In this way, the distance between different degradation modes can be increased in the feature space, while the distance between images under the same degradation mode can be decreased, thereby clearly distinguishing and representing various degradations. In some embodiments, as shown in FIG2, the degradation extraction module 110 may include: multiple cascaded residual blocks, multiple convolutional layers, and at least one pooling layer. Each residual block may include at least one convolutional layer. In any residual block, after performing convolution processing on the data input to that residual block through the convolutional layer in that residual block, the processing result and the input data can be used to perform residual calculation to obtain the processing result of that residual block. For example, the convolutional layer in the residual block may be a convolutional layer with a kernel size of 3×3 and 64 output channels. For example, the input to the first residual block may be a low-quality image to be processed.Referring again to Figure 2, after processing through cascaded residual blocks, the results can be further processed by convolutional layers between the pooling layer (ada-pool) and the residual blocks. These convolutional layers can be, but are not limited to, convolutional layers with a kernel size of 3×3 and 256 output channels. The pooling layer converts the features output by the convolutional layer into a 1x1x2048 dimension, becoming a global description of the image. For example, the global description can be a comprehensive representation of the overall image information, encompassing factors such as the main objects, background, spatial relationships between them, and overall color distribution. Since degradation is often global, this method can more accurately express the degradation characteristics of the image. Furthermore, by processing the pooled features through convolutional layers following the pooling layer, degradation features can be extracted. For example, the convolutional layer following the pooling layer can be, but is not limited to, a convolutional layer with a kernel size of 1×1 and 2048 output channels. It should be understood that the number of residual blocks and convolutional layers in the degradation extraction module 110 can be determined according to the actual situation and is not limited here. Alternatively, the output of the pooling layer can also be used as degradation features, which can be determined according to the actual situation and is not limited here. In some embodiments, in Figure 2, the processing before the pooling layer can be understood as performing stepwise downsampling on the image. Furthermore, in Figure 2, the structure of the degradation extraction module 110 can be understood as a combination of stepwise downsampling and pooling layers, where the pooling layer is used to convert the features obtained from the final downsampling into a global description.
[0060] The diffusion model 120 is primarily used to repair low-quality images to be processed, based on the degradation features extracted by the degradation extraction module 110, in order to obtain high-quality images. The degradation features can guide the super-resolution process. For example, degradation features can guide the diffusion model 120 to focus on and repair the degraded parts of the image. The super-resolution method achieved by the diffusion model 120 is mainly based on simulating the process of an image going from sharp to blurry (or from high quality to low quality), and then recovering the high-quality details of the image by reversing this process. Since degradation features can characterize the degradation type and / or degradation characteristics of the low-quality image to be processed, the diffusion model 120 can accurately understand the degradation information of the low-quality image to be processed, thereby accurately guiding the super-resolution process and improving the image super-resolution effect. For example, the diffusion model 120 can be obtained through pre-training. For example, the diffusion model 120 can be, but is not limited to, a latent diffusion model (LDM) or a denoising diffusion probabilistic model (DDPM).
[0061] As can be seen from the above description of the image processing system 100, this system extracts degradation features from the low-quality image to be processed, enabling accurate representation of the image degradation features during super-resolution processing. When the extracted degradation features are accurate, they are used as conditional inputs to the diffusion model, allowing the model to accurately perceive the degradation information of the low-quality image and thus accurately recover a high-quality image, improving the super-resolution effect. Furthermore, to better recover image content during super-resolution processing and further improve the super-resolution effect, content features of the low-quality image to be processed can be added to the diffusion model 120 during processing. Based on this concept, this application provides another image processing system. The following describes this other image processing system.
[0062] For example, Figure 3 shows a schematic diagram of the architecture of another image processing system provided in an embodiment of this application. As shown in Figure 3, the image processing system 200 mainly includes: a content extraction module 210, a degradation extraction module 220, and a diffusion model 230. Among them, the degradation extraction module 220 is the same as the degradation extraction module 110 described in Figure 1, and can be referred to the relevant description in Figure 1 above, which will not be repeated here.
[0063] The content extraction module 210 is mainly used to extract content features from the low-quality image to be processed. These content features can be used to characterize the content in the low-quality image. For example, the content features can include one or more of the following: subject object features, semantic information features, salient region features, scene layout, and contextual features. Specifically, subject object features can include the shape and contour details, key structural components, etc., of objects in the image; semantic information features can include object category labels, action and state descriptions, etc., of objects in the image; salient region features can include the salience of color and texture, visual focus information, etc., of objects in the image; scene layout can include the spatial arrangement of objects in the image and their geometric relationships; and contextual features can include the semantic relationships between different parts of the image.
[0064] In this embodiment, the design goal of the content extraction module 210 is to ensure that content features remain unchanged under different degradation conditions. This means that regardless of the quality degradation the image has undergone, its core content and information should be effectively preserved.
[0065] The content extraction module 210 is trained by training a low-quality to high-quality reconstruction model to extract the content features of an image. This model focuses on recovering high-quality content from degraded low-quality images, ignoring the degradation itself and focusing on extracting the core content of the image. The main goal of the content extraction module 210 is to ignore the degradation details of the image and focus on extracting and preserving the content of the image. For example, the content extraction module 210 can be the feature extraction part of the trained model. For example, as shown in Figure 4, in part ①, the low-quality to high-quality reconstruction model mainly consists of an encoder (or "first encoder") and a decoder (or "first decoder"). The training dataset of this reconstruction model contains: low-quality images with the same content under different degradation conditions, and high-quality images corresponding to each low-quality image. This reconstruction model can convert low-quality images with the same content under different degradation conditions into the same high-quality image. For example, the training process of the reconstruction model (or "neural network model") shown in part ① can be: first, acquire the third image and the fourth image. The image quality of the third image is lower than that of the fourth image, and the content in the third image is the same as that in the fourth image. For example, images of different quality but the same content can refer to images with the same content (such as scene, subject, etc.), but differences in detail, sharpness, color representation, noise, etc. Then, the third image is input into the neural network model to obtain the fifth image. Next, a loss is calculated based on the difference between the fourth and fifth images to obtain the first loss. Finally, the neural network model is trained with the goal of minimizing the first loss. In the reconstruction model shown in Part ①, the encoder can be used to extract content features from low-quality images, and the decoder is used to process the content features to obtain high-quality images. The content extraction module 210 can be the encoder in Part ①. For example, in Figure 4, the encoder in Part ① can be a progressive downsampling structure, where the encoder performs feature processing through progressive downsampling, and the features obtained from the last downsampling can be, but are not limited to, content features. The decoder in Part ① can be a progressive upsampling structure.
[0066] Because contrastive learning cannot encompass all degradation information, this reconstruction can reconstruct both the content and the degradation, thus ensuring that the degradation encoder can encompass all degradation information.
[0067] Furthermore, referring to Figure 4, to enable the encoder in Part ② to extract degradation features under most or all degradation modes, the same image can be processed separately in Parts ① and ② to extract the content features and degradation features of the image. Then, in Part ③, a decoder is used to process the extracted content features and degradation features to reconstruct the original image. Further, in Part ③, based on the differences between the reconstructed image and the original image, the parameters in the encoder and decoder used in Part ② can be adjusted to ensure that the reconstructed image approximates or equals the original image, and that the encoder in Part ② can extract more accurate degradation features, thus enabling it to extract degradation features under most or all degradation modes. For example, in Part ③, the decoder is designed to reconstruct the original low-quality image (i.e., the original image) using the extracted content features and degradation features. This step ensures that the extracted features not only accurately describe the content and degradation but are also effectively used for image reconstruction. The main objective of Part ③ is to ensure the completeness of the information in the extracted content features and degradation features. Through joint learning (combining content features and degradation features to learn the original image), the decoder can simultaneously understand and utilize image content and degradation information, thereby more accurately restoring image quality during reconstruction. Thus, through the collaborative work of these three parts, image quality can be processed and improved more effectively, ensuring accurate preservation of image content under various degradation conditions, while accurately representing and distinguishing degradation patterns. For example, in Figure 4, the structure of the decoder in part ③ can be the same as or different from the decoder in part ①, depending on the actual situation; no limitation is made here. Furthermore, to improve training efficiency, the decoder used in part ③ can be the decoder trained in part ①, i.e., joint training. For example, the process of jointly training the second encoder in part ② and the decoder in part ③ in part ③ can be as follows: First, acquire the seventh image. Then, input the seventh image into the first encoder for feature extraction to obtain the content features of the seventh image, and input the seventh image into the second encoder for feature extraction to obtain the degradation features of the seventh image. Next, the content features and degradation features of the seventh image are input into the decoder (or "second decoder") in Part ③ for processing, resulting in the eighth image. Then, a loss is calculated based on the difference between the seventh and eighth images, yielding the fifth loss. Finally, the parameters, such as weights, in the second encoder and second decoder are updated with the goal of minimizing the fifth loss. For example, the parameters of the first encoder can be frozen during the training process in Part ③, but this is not limited to the case where the first encoder's parameters are frozen.
[0068] In the image processing system 200, the diffusion model 230 can be used to repair low-quality images to be processed, based on the degradation features extracted by the degradation extraction module 220 and the content features extracted by the content extraction module 210, to obtain high-quality images. Both degradation features and content features can be used to guide the super-resolution process. For example, degradation features can guide the diffusion model 230 to focus on and repair the degraded parts of the image, while content features can guide the diffusion model 230 to focus on and repair the content parts of the image. Since degradation features facilitate the diffusion model 230's understanding of the degradation information of the low-quality image to be processed, and content features facilitate its understanding of the content of the low-quality image to be processed, the diffusion model 230 can understand both degradation details and image content during the super-resolution process, thereby improving the super-resolution effect.
[0069] As can be seen from the above description of the image processing system 200, this system extracts degradation features and content features from the low-quality image to be processed, enabling accurate representation of the degradation features and content features during the super-resolution process, thereby improving the super-resolution effect. Furthermore, to better understand the relationship between degradation features and content features during the super-resolution process and further improve the super-resolution effect, degradation features can be injected into content features, and the content features with injected degradation features can be input into the diffusion model 120 for processing. Based on this concept, this application provides another image processing system. The following describes another image processing system provided by this application embodiment.
[0070] For example, Figure 5 shows a schematic diagram of the architecture of another image processing system provided in an embodiment of this application. As shown in Figure 5, the image processing system 300 mainly includes: a content and degradation modulation module 310, a degradation extraction module 320, a content extraction module 330, and a diffusion model 340. The degradation extraction module 320 is the same as the degradation extraction module 110 in Figure 1, and can be referred to the relevant description in Figure 1 above, which will not be repeated here. The content extraction module 330 is the same as the content extraction module 210 in Figure 3, and can be referred to the relevant description in Figure 3 above, which will not be repeated here.
[0071] The content and degradation modulation module 310 is mainly used to inject degradation features into content features, thereby integrating degradation information into the image content and more effectively improving the diffusion model 340's perception and utilization of degradation and content. For example, the features output by the content and degradation modulation module 310 can be features that fuse degradation features and content features, i.e., features combining degradation features and content features. For example, the content and degradation modulation module 310 can inject degradation features into content features to obtain features that fuse degradation features and content features. In some embodiments, as shown in FIG6, the content and degradation modulation module 310 may include multiple cascaded modulation layers. Each modulation layer may include a convolutional layer, an activation function (leaky ReLU), and a modulator. In any modulation layer, the content features input to the modulation layer or the output of the previous modulation layer can be convolved by the convolutional layer; then, the convolutional result can be processed by the activation function; finally, the modulation function processes the activation function's processing result and the degradation features to obtain the final processing result of the modulation layer. In this modulation layer, the input data of the convolutional layer is the content features, while the input data of the convolutional layers in other modulation layers is the output of the previous modulation layer. For example, the convolutional layers in the modulation layers can be, but are not limited to, convolutional layers with a kernel size of 3×3 and 64 output channels. After processing by the last modulation layer, the processing result of that modulation layer can be convolved by another convolutional layer to obtain the final processing result of the content and degradation modulation module 310. Referring again to Figure 6, the modulator in any modulation layer can consist of a series of convolutional layers. When combined with degradation features, the modulator is designed to receive degradation features as input and generate a set of modulation parameters. These modulation parameters are then used to adjust the content features to reflect the impact of the degradation process. The modulator adjusts the image content by combining degradation features with content features. In this way, its processing can be dynamically adjusted according to degradation information. The main advantage of introducing a modulator is that it allows the model to handle different types of degradation more flexibly. By explicitly incorporating degradation information into the model, the modulator can help the model better understand the degradation process and perform more effective image restoration accordingly.For example, continuing to refer to Figure 6, the working process of the modulator in any modulation layer can first process the degradation features and content features separately through convolutional layers; then, a point multiplication is performed between a convolutional kernel and the result of the processing of degradation features by the convolutional layer to obtain the modulation parameters, where the kernel can be obtained through pre-training; finally, a convolution operation is performed on the point multiplication result and the result of the processing of content features by the convolutional layer to modulate the content features of the image, thereby obtaining the processing result of the modulator.
[0072] In the image processing system 300, the diffusion model 340 can process the low-quality image to be processed based on the features output by the content and degradation modulation module 310 to obtain a high-quality image. Since the features output by the content and degradation modulation module 310 integrate degradation features and content features, and include the relationship between the two, the diffusion model 340 can understand degradation details, the content in the image, and the relationship between degradation details and content during the super-resolution process, thereby improving the super-resolution effect.
[0073] As can be seen from the above description of the image processing system 300, this system improves the super-resolution processing effect by fusing the degradation features and content features extracted from the low-quality image to be processed. This allows for the accurate representation of the degradation features and content features of the image, as well as the relationship between them, during the super-resolution process. Furthermore, in the image processing system 300, under new application scenarios, the parameters in the content and degradation modulation module 310, degradation extraction module 320, or content extraction module 330 can be fine-tuned based on a small number of images, allowing the system to adapt to new scenarios. Similarly, in the image processing system 200, under new application scenarios, the parameters in the degradation extraction module 220 or content extraction module 210 can also be fine-tuned based on a small number of images, enabling the system to adapt to new scenarios.
[0074] In some embodiments, each component of the image processing system described above can be configured on a cloud computing platform, for example, deployed on at least one instance such as a virtual machine or container, so that the cloud computing platform can provide image processing services. Of course, each component of the image processing system described above can also be configured on nodes other than the cloud computing platform, for example, deployed in at least one data center or on at least one server, depending on the actual situation, and is not limited here. The cloud computing platform can provide pages related to public cloud services for users to remotely access public cloud services. In this embodiment, users can pre-purchase image processing services on the cloud computing platform. For ease of understanding, the interaction between the user and the cloud computing platform is described below. As shown in Figure 7, the interaction between the user and the cloud computing platform mainly includes: the user logs into the cloud computing platform 700 through a client webpage, selects and purchases image processing services on the cloud computing platform 700, and after purchase, the user can process images on the cloud computing platform 700 based on the functions provided by the image processing services. The cloud computing platform 700 is mainly used to manage the infrastructure running the image processing services. For example, the infrastructure running the image processing services may include multiple data centers located in different regions, each data center including multiple servers. Data centers can provide basic resources for image processing services, such as computing and storage resources. Therefore, when users purchase and use image processing services, they primarily pay for the resources they use. When using image processing services, users can input low-quality images to be processed through the configuration interface, application programming interface, or user interaction interface provided by the cloud computing platform 700. The cloud computing platform 700 then processes the low-quality images according to the user input and returns the processing results to the user. In some embodiments, the components of the image processing system described above can also be configured on a local server; the specific configuration depends on the actual situation and is not limited here.
[0075] The specific implementation process of the above image processing system will be introduced below.
[0076] For example, Figure 8 shows a schematic flowchart of an image processing method provided in an embodiment of this application. It is understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. For example, this method can be executed by an image processing device, which can be implemented by software and / or hardware, and can be configured in, but is not limited to, an electronic device or a server; typically, it can be configured on a server. For ease of description, a cloud computing platform will be used as the execution subject below. As shown in Figure 8, the image processing method may include the following steps:
[0077] S801, Obtain the first image.
[0078] In this embodiment, an image upload portal can be provided on the client associated with the cloud computing platform. Users can select a first image through the image upload portal on the client and confirm the upload of the first image to the cloud computing platform. The client associated with the cloud computing platform can be a desktop application, mobile application, web application, or web-based application, etc. After the client transmits the user-inputted first image to the cloud computing platform, the cloud computing platform can obtain the first image.
[0079] S802. Extract features from the degradation information in the first image to obtain degradation features.
[0080] In this embodiment, the cloud computing platform can extract features from the degradation information in the first image using the degradation extraction module 110 described in Figure 1, to obtain degradation features. For example, the degradation information of the first image may include degradation type and / or degradation characteristics. As a possible implementation, when extracting degradation features, the first image can be downsampled step-by-step, and the features obtained from the last downsampling can be pooled to convert the features obtained from the last downsampling into a global description of the first image, thus obtaining the degradation features. Since degradation is often global, this method can more accurately express the degradation characterization of the image.
[0081] S803. Input the first image and the first data containing degradation features into the diffusion model to obtain the second image, wherein the image quality of the second image is higher than that of the first image.
[0082] In this embodiment, the cloud computing platform can input a first image and first data containing degradation features into a diffusion model, so that the diffusion model can repair the first image based on the degradation features to obtain a second image. The second image has a higher image quality than the first image. For example, the diffusion model used here can be, but is not limited to, the diffusion model 120 described in Figure 1 above.
[0083] Since degradation features can characterize the degradation information of the first image, by using degradation features as a conditional input to the diffusion model, the diffusion model can accurately perceive the degradation information of the first image, thereby accurately recovering a high-quality second image from the first image and improving the image super-resolution effect.
[0084] In some embodiments, in order for the diffusion model to understand the content of the first image during image restoration, the image processing method shown in FIG8 can also extract features from the content of the first image using the content extraction module 210 described in FIG3, etc., to obtain content features. Furthermore, the first data in S803 of FIG8 also includes the content features of the first image. In this way, the diffusion model can perceive both the degradation information and the content of the first image during image restoration, thereby accurately restoring the content of the first image and improving the image super-resolution effect.
[0085] Furthermore, inputting the first image and first data containing degradation features and content features into the diffusion model can include: first, injecting degradation features into content features using the content and degradation modulation module 310 described in Figure 5 to obtain fused features; then, inputting the first image and fused features into the diffusion model. For example, when injecting degradation features into content features, a dot product operation can be performed on the degradation features using a pre-learned convolution kernel to obtain modulation parameters; then, convolution operations can be performed on the content features using the modulation parameters to inject degradation features into the content features, thereby obtaining fused features. In this way, during subsequent processing, the diffusion model can understand both degradation details and image content through the first data, as well as the relationship between degradation details and content, thereby improving the super-resolution effect.
[0086] The above is an introduction to the image processing method provided in this embodiment. As shown in Figure 9, the method provided in this embodiment significantly improves the resolution and clarity of images processed by the method on real-world mobile phone data compared to the most advanced (state-of-the-art, SOTA) methods commonly used in the industry (such as scaling-up image restoration, SUPIR).
[0087] It is understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. In addition, the various embodiments and features described in the above embodiments can be combined according to actual conditions, and the combined solutions are still within the protection scope of this application.
[0088] Next, based on the methods in the above embodiments, an image processing apparatus provided in this application will be described.
[0089] For example, Figure 10 shows a schematic diagram of an image processing apparatus provided in an embodiment of this application. As shown in Figure 10, the image processing apparatus 1000 includes: an acquisition module 1001, used to acquire a first image; a processing module 1002, used to extract features from degradation information in the first image to obtain degradation features; and the processing module 1002 is further used to input the first image and first data into a diffusion model to obtain a second image, wherein the first data includes degradation features, and the image quality of the second image is higher than that of the first image.
[0090] In some embodiments, the processing module 1002 is further configured to: extract features from the content in the first image to obtain content features; wherein the first data also includes content features.
[0091] In some embodiments, before inputting the first data into the diffusion model, the processing module 1002 is further configured to: inject degradation features into content features to fuse degradation features and content features.
[0092] In some embodiments, when the processing module 1002 injects degradation features into the content features, it specifically performs the following: performs a dot product operation on the degradation features using a convolution kernel to obtain modulation parameters; and performs a convolution operation on the content features using the modulation parameters to inject the degradation features into the content features.
[0093] In some embodiments, the processing module 1002 is further configured to: acquire a third image and a fourth image, wherein the image quality of the third image is lower than that of the fourth image, and the content in the third image is the same as the content in the fourth image; input the third image into a neural network model to obtain a fifth image, wherein the neural network model includes a first encoder and a first decoder, the first encoder is used to extract features from the content in the third image, and the first decoder is used to generate the fifth image based on the features extracted by the first encoder; calculate a loss based on the difference between the fourth image and the fifth image to obtain a first loss; train the neural network model with the goal of minimizing the first loss, wherein the first encoder in the trained neural network model is used to extract features from the content in the first image.
[0094] In some embodiments, the processing module 1002 is further configured to: acquire multiple sixth images with different content under the same degradation mode, and acquire multiple seventh images with the same content under different degradation modes; input the multiple sixth images into a second encoder to extract degradation information and obtain degradation features of each sixth image; input the multiple seventh images into the second encoder to extract degradation information and obtain degradation features of each seventh image; calculate a loss based on the difference between the degradation features of the multiple sixth images to obtain a second loss, and calculate a loss based on the difference between the degradation features of the multiple seventh images to obtain a third loss, and calculate a loss based on the difference between the degradation features of the multiple sixth images and the degradation features of the multiple seventh images to obtain a fourth loss; train the second encoder with the goal of minimizing the second loss and the third loss, and maximizing the fourth loss, wherein the trained second encoder is used to extract features from the degradation information in the first image.
[0095] In some embodiments, the processing module 1002 is further configured to: input the seventh image into the first encoder for feature extraction to obtain the content features of the seventh image; input the seventh image into the second encoder for feature extraction to obtain the degradation features of the seventh image; input the content features and degradation features of the seventh image into the second decoder for processing to obtain the eighth image; perform loss calculation based on the difference between the seventh image and the eighth image to obtain the sixth loss; and update the parameters in the second encoder and the second decoder with the goal of minimizing the sixth loss.
[0096] In some embodiments, the first encoder has a progressive downsampling structure; the second encoder has a structure combining progressive downsampling and pooling layers, wherein the pooling layer is used to convert the features obtained from the last downsampling stage into a global description.
[0097] In some embodiments, the diffusion model is a latent diffusion model or a denoised diffusion probability model.
[0098] In some embodiments, both the acquisition module 1001 and the processing module 1002 shown in FIG. 10 can be implemented in software or in hardware. For example, the implementation of the acquisition module 1001 will be described below. Similarly, the implementation of the processing module 1002 can refer to the implementation of the acquisition module 1001.
[0099] As an example of a software functional unit, module 1001 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, module 1001 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0100] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0101] As an example of a hardware functional unit, the acquisition module 1001 may include at least one computing device, such as a server. Alternatively, the acquisition module 1001 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0102] The multiple computing devices included in the acquisition module 1001 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the acquisition module 1001 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the acquisition module 1001 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0103] It should be noted that, in other embodiments, the acquisition module 1001 can be used to execute any step in the image processing method described in the above embodiments, and the processing module 1002 can also be used to execute any step in the image processing method described in the above embodiments. Furthermore, the steps implemented by the acquisition module 1001 and the processing module 1002 can be specified as needed. By having the acquisition module 1001 and the processing module 1002 respectively implement different steps in the image processing method described in the above embodiments, all the functions of the image processing device 1000 shown in FIG10 can be achieved.
[0104] This application also provides a computing device 1100. As shown in FIG11, the computing device 1100 includes: a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate with each other via the bus 1102. The computing device 1100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1100.
[0105] Bus 1102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 11, but this does not imply that there is only one bus or one type of bus. Bus 1104 can include pathways for transmitting information between various components of computing device 1100 (e.g., memory 1106, processor 1104, communication interface 1108).
[0106] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0107] The memory 1106 may include volatile memory, such as random access memory (RAM). The processor 1104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0108] The memory 1106 stores executable program code, and the processor 1104 executes the executable program code to implement the functions of the acquisition module 1001 and the processing module 1002 shown in FIG. 10, thereby implementing the image processing method described in the above embodiments. That is, the memory 1106 stores instructions for executing the image processing method described in the above embodiments.
[0109] Alternatively, the memory 1106 may store executable code, which the processor 1104 executes to implement the functions of the image processing apparatus 1000 shown in FIG. 10, thereby implementing the image processing method described in the above embodiments. That is, the memory 1106 stores instructions for executing the image processing method described in the above embodiments.
[0110] The communication interface 1108 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1100 and other devices or communication networks.
[0111] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0112] As shown in Figure 12, the computing device cluster includes at least one computing device 1100. The memory 1106 of one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the image processing method described in the above embodiments.
[0113] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the image processing method described in the above embodiments. In other words, a combination of one or more computing devices 1100 can jointly execute instructions for executing the image processing method described in the above embodiments.
[0114] It should be noted that the memory 1106 in different computing devices 1100 within the computing device cluster can store different instructions, which are used to execute some of the functions of the image processing apparatus 1000 shown in FIG. 10. That is, the instructions stored in the memory 1106 in different computing devices 1100 can implement the functions of one or more modules in the acquisition module 1001 and the processing module 1002.
[0115] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 13 illustrates one possible implementation. As shown in Figure 13, two computing devices 1100A and 1100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1106 in computing device 1100A stores instructions for executing the functions of the acquisition module 1001. Simultaneously, the memory 1106 in computing device 1100B stores instructions for executing the functions of the processing module 1002.
[0116] It should be understood that the functions of computing device 1100A shown in Figure 13 can also be performed by multiple computing devices 1100. Similarly, the functions of computing device 1100B can also be performed by multiple computing devices 1100.
[0117] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to the connection method of the computing device cluster described in Figures 12 and 13. The difference is that the memory 1106 of one or more computing devices 1100 in this computing device cluster can store the same instructions for executing the methods in the above embodiments.
[0118] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the aforementioned image processing method. In other words, a combination of one or more computing devices 1100 can jointly execute the instructions for executing the aforementioned image processing method.
[0119] It should be understood that each step of the above method embodiments can be accomplished by hardware logic circuits or software instructions in a processor.
[0120] Based on the methods in the above embodiments, this application provides a computer-readable storage medium including computer program instructions. When executed by a cluster of computing devices including at least one computing device, the computer program instructions cause the cluster of computing devices to perform the methods in the above embodiments. Exemplarily, the computer-readable storage medium can be any available medium that the computing device can store, or a data storage device such as a data center containing one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives).
[0121] Based on the methods in the above embodiments, this application provides a computer program product containing instructions that, when executed by a cluster of computing devices containing at least one computing device, cause the cluster of computing devices to perform the methods in the above embodiments.
[0122] It is understood that the processor in the embodiments of this application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor may be a microprocessor or any conventional processor.
[0123] The method steps in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.
[0124] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0125] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
Claims
1. An image processing method, characterized in that, The method includes: Get the first image; Feature extraction is performed on the degradation information in the first image to obtain degradation features; The first image and the first data are input into a diffusion model to obtain a second image; wherein the first data includes the degradation features, and the image quality of the second image is higher than that of the first image.
2. The method according to claim 1, characterized in that, The first data also includes content features, and before inputting the first image and the first data into the diffusion model, it further includes: Feature extraction is performed on the content of the first image to obtain content features.
3. The method according to claim 2, characterized in that, The first image and the first data are input into the diffusion model to obtain the second image, including: The degradation features are injected into the content features to obtain the fused features; The first image and the fusion features are input into the diffusion model to obtain the second image.
4. The method according to claim 3, characterized in that, The step of injecting the degradation features into the content features to obtain fused features includes: The degradation features are multiplied by a convolution kernel to obtain modulation parameters; The content features are convolved using the modulation parameters to obtain the fused features.
5. The method according to any one of claims 2-4, characterized in that, The method further includes: Acquire a third image and a fourth image, wherein the image quality of the third image is lower than that of the fourth image, and the content in the third image is the same as that in the fourth image; The third image is input into a neural network model to obtain a fifth image. The neural network model includes a first encoder and a first decoder. The first encoder is used to extract features from the content of the third image, and the first decoder is used to generate the fifth image based on the features extracted by the first encoder. A first loss is obtained by calculating the loss based on the difference between the fifth image and the fourth image; The neural network model is trained with the goal of minimizing the first loss, wherein the first encoder in the trained neural network model is used to extract features from the content of the first image.
6. The method according to claim 5, characterized in that, The method further includes: Acquire multiple sixth images with different content under the same degradation mode, and acquire multiple seventh images with the same content under different degradation modes; Multiple sixth images are input into a second encoder to extract degradation information, thereby obtaining the degradation features of each sixth image; Multiple seventh images are input into the second encoder to extract degradation information, thereby obtaining the degradation features of each seventh image; A second loss is obtained by calculating the loss based on the differences between the degradation features of multiple sixth images; a third loss is obtained by calculating the loss based on the differences between the degradation features of multiple seventh images; and a fourth loss is obtained by calculating the loss based on the differences between the degradation features of multiple sixth images and the degradation features of multiple seventh images. The second encoder is trained with the goal of minimizing the second loss and the third loss, and maximizing the fourth loss, wherein the trained second encoder is used to extract features from the degradation information in the first image.
7. The method according to claim 6, characterized in that, Also includes: The seventh image is input into the first encoder for feature extraction to obtain the content features of the seventh image; The seventh image is input into the second encoder for feature extraction to obtain the degradation features of the seventh image; The content features of the seventh image and the degradation features of the sixth image are input into the second decoder for processing to obtain the eighth image; The fifth loss is obtained by calculating the loss based on the difference between the seventh image and the eighth image; With the goal of minimizing the fifth loss, the parameters in the second encoder and the second decoder are updated.
8. The method according to claim 6 or 7, characterized in that, The first encoder is a step-by-step downsampling structure; the second encoder is a structure combining step-by-step downsampling and a pooling layer, wherein the pooling layer is used to convert the features obtained from the last downsampling stage into a global description.
9. The method according to any one of claims 1-8, characterized in that, The diffusion model is either a potential diffusion model or a denoised diffusion probability model.
10. An image processing apparatus, characterized in that, include: The acquisition module is used to acquire the first image; The processing module is used to extract features from the degradation information in the first image to obtain degradation features; The processing module is further configured to input the first image and the first data into a diffusion model to obtain a second image, wherein the first data includes the degradation features, and the image quality of the second image is higher than that of the first image.
11. The apparatus according to claim 10, characterized in that, The first data also includes the content features, and before inputting the first image and the first data into the diffusion model, the processing module is further configured to: Feature extraction is performed on the content of the first image to obtain content features.
12. The apparatus according to claim 11, characterized in that, The step of inputting the first data into the diffusion model to obtain the second image includes: The degradation features are injected into the content features to obtain the fused features; The second image can be obtained by inputting the first image and the fusion features into the diffusion model.
13. The apparatus according to claim 12, characterized in that, The step of injecting the degradation features into the content features to obtain fused features includes: The degradation features are multiplied by a convolution kernel to obtain modulation parameters; The content features are convolved using the modulation parameters to obtain the fused features.
14. The apparatus according to any one of claims 11-13, characterized in that, The processing module is further configured to: Acquire a third image and a fourth image, wherein the image quality of the third image is lower than that of the fourth image, and the content in the third image is the same as that in the fourth image; The third image is input into a neural network model to obtain a fifth image. The neural network model includes a first encoder and a first decoder. The first encoder is used to extract features from the content of the third image, and the first decoder is used to generate the fifth image based on the features extracted by the first encoder. A first loss is obtained by calculating the loss based on the difference between the fourth image and the fifth image. The neural network model is trained with the goal of minimizing the first loss, wherein the first encoder in the trained neural network model is used to extract features from the content of the first image.
15. The apparatus according to claim 14, characterized in that, The processing module is further configured to: Acquire multiple sixth images with different content under the same degradation mode, and acquire multiple seventh images with the same content under different degradation modes; Multiple sixth images are input into a second encoder to extract degradation information, thereby obtaining the degradation features of each sixth image; Multiple seventh images are input into the second encoder to extract degradation information, thereby obtaining the degradation features of each seventh image; A second loss is obtained by calculating the loss based on the differences between the degradation features of multiple sixth images; a third loss is obtained by calculating the loss based on the differences between the degradation features of multiple seventh images; and a fourth loss is obtained by calculating the loss based on the differences between the degradation features of multiple sixth images and the degradation features of multiple seventh images. The second encoder is trained with the goal of minimizing the second loss and the third loss, and maximizing the fourth loss, wherein the trained second encoder is used to extract features from the degradation information in the first image.
16. The apparatus according to claim 15, characterized in that, The processing module is further configured to: The seventh image is input into the first encoder for feature extraction to obtain the content features of the seventh image; The seventh image is input into the second encoder for feature extraction to obtain the degradation features of the seventh image; The content features and degradation features of the seventh image are input into the second decoder for processing to obtain the eighth image; A sixth loss is obtained by calculating the loss based on the difference between the seventh and eighth images; With the goal of minimizing the sixth loss, the parameters in the second encoder and the second decoder are updated.
17. The apparatus according to claim 15 or 16, characterized in that, The first encoder is a step-by-step downsampling structure; the second encoder is a structure combining step-by-step downsampling and a pooling layer, wherein the pooling layer is used to convert the features obtained from the last downsampling stage into a global description.
18. The apparatus according to any one of claims 10-17, characterized in that, The diffusion model is either a potential diffusion model or a denoised diffusion probability model.
19. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-9.
20. A computer-readable storage medium, characterized in that, The method includes computer program instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method as described in any one of claims 1-9, wherein the cluster of computing devices includes at least one computing device.
21. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1-9, wherein the computing device cluster includes at least one computing device.