Image enhancement method and device
By enhancing the initial image features and predicting the degraded feature of the compressed image, and combining N different scale features for image enhancement, the problem of insufficient image structure in the prior art is solved, and the enhancement effect is more in line with the real image, which is suitable for machine vision.
Patent Information
- Application Number
- CN202510196880.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-06
AI Technical Summary
The existing compressed image enhancement methods have shortcomings in the structural aspects of visual content and contextual relationships and cannot be applied to machine vision.
By performing noise enhancement processing on the initial image features, the target noise image features are obtained; the original compressed image is predicted to obtain N different scale features; these features are used as image enhancement conditions, and image enhancement processing is performed to obtain the target enhanced image.
Effectively pay attention to image structure information and guide image restoration. The restored target enhancement image has a data distribution and image structure that is more in line with the real image, which is suitable for machine vision.
Smart Images

Figure CN120107136A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and more specifically to an image enhancement method and device. Background Art
[0002] Image compression technology aims to reduce redundant information in image data through mathematical transformations and coding strategies, while retaining information that is critical to the human visual system and machine vision system. However, the compression process will inevitably introduce compression distortion, such as block effects and image blur.
[0003] In related technologies, image enhancement is performed by using the mapping relationship between compressed images and uncompressed images using pixel-level loss constraints. However, this only focuses on the reconstruction of pixel values, while sacrificing the structural nature of the image in terms of visual content and contextual relationships. Therefore, existing compressed image enhancement methods are not suitable for machine vision. Summary of the invention
[0004] In view of the above problems, the present disclosure provides an image enhancement method and device.
[0005] According to a first aspect of the present disclosure, an image enhancement method is provided, comprising: performing noise enhancement processing on initial image features to obtain target noise image features; performing degradation feature prediction processing on the original compressed image to obtain N different scale features of the original compressed image, wherein each scale feature represents an image degradation type feature; using the N different scale features of the original compressed image as image enhancement conditional features, performing image enhancement processing on the target noise image features to obtain a target enhanced image.
[0006] According to an embodiment of the present disclosure, the N different-scale features include M first-scale features and L second-scale features; performing degradation feature prediction processing on the original compressed image to obtain N different-scale features of the original compressed image, including: performing convolution on the m-1th first-scale feature to obtain the mth first-scale feature, and obtaining M first-scale features, wherein, when m is equal to 1, the first first-scale feature is obtained by convolving the original compressed image, M≥1, 1≤m≤M; performing degradation feature prediction processing on the i-1th second-scale feature to obtain the i-th second-scale feature, and obtaining L second-scale features, wherein, when i is equal to 1, the first second-scale feature is obtained by performing degradation feature prediction processing on the Mth first-scale feature, and L≥1, 1≤i≤L.
[0007] According to an embodiment of the present disclosure, a degradation feature prediction process is performed on the i-1th second-scale feature to obtain the i-th second-scale feature, including: convolving the i-1th second-scale feature to obtain a convolution feature; based on an attention mechanism, performing deep feature extraction on the convolution feature to obtain an intermediate feature; and performing structural feature extraction on the intermediate feature to obtain the i-th second-scale feature.
[0008] According to an embodiment of the present disclosure, structural feature extraction is performed on the intermediate features to obtain the i-th second-scale feature, including: performing fusion normalization processing on the intermediate features to obtain normalized features; and performing structural feature extraction on the normalized features to obtain the i-th second-scale feature.
[0009] According to an embodiment of the present disclosure, intermediate features are fused and normalized to obtain normalized features, including: performing instance normalization on the intermediate features to obtain first normalized features; performing batch normalization on the intermediate features to obtain second normalized features; performing attention prediction on the third normalized features to obtain a first attention weight, wherein the third normalized feature is obtained by adding the first normalized feature and the second normalized feature; multiplying the first attention weight and the first normalized feature to obtain a fourth normalized feature; multiplying the first attention weight and the second normalized feature to obtain a fifth normalized feature; and performing feature fusion on the fourth normalized feature and the fifth normalized feature to obtain a normalized feature.
[0010] According to an embodiment of the present disclosure, structural feature extraction is performed on the normalized feature to obtain the i-th second-scale feature, including: determining a prompt parameter corresponding to the normalized feature; performing attention calculation on the normalized feature to obtain a second attention weight; multiplying the second attention weight and the prompt parameter to obtain a target prompt parameter; and fusing the normalized feature and the target prompt parameter to obtain the i-th second-scale feature.
[0011] According to an embodiment of the present disclosure, multiple different scale features of the original compressed image are used as image enhancement condition features, and image enhancement processing is performed on the target noise image features to obtain a target enhanced image, including: down-sampling the t-th step denoising image features to obtain the t-th step down-sampled denoising image features; up-sampling the t-th step down-sampled denoising image features and multiple different scale features to obtain the t+1-th step denoising image features, and obtain the T-th step denoising image features; decoding the T-th step denoising image features to obtain the target enhanced image.
[0012] According to an embodiment of the present disclosure, upsampling is performed on the t-th step down-sampled denoised image feature and multiple different-scale features to obtain the t+1-th step denoised image feature, including: fusing the k-1-th sub-upsampled denoised image feature and the k-th scale feature to obtain the fused k-1-th sub-upsampled denoised image feature; upsampling is performed on the fused k-1-th sub-upsampled denoised image feature to obtain the k-th sub-upsampled denoised image feature, and the N-th sub-upsampled denoised image feature is obtained, wherein, when k is equal to 1, the k-1-th sub-upsampled denoised image feature is obtained by fusing and upsampling the t-th step down-sampled denoised image feature and the 1st scale feature, 1≤k-1≤N, N=M+L; and the N-th sub-upsampled denoised image feature is determined as the t+1-th step denoised image feature.
[0013] According to an embodiment of the present disclosure, noise enhancement processing is performed on the initial image features to obtain target noise image features, including: performing noise enhancement processing on the p-1th step noise image features to obtain the p+1th step noise image features, and obtaining the Pth step noise image features, wherein, when p is equal to 1, the first step noise image features are obtained by performing noise enhancement processing on the initial image features; and determining the Pth step noise image features as the target noise image features.
[0014] A second aspect of the present disclosure provides an image enhancement device, comprising: a first processing module, used to perform noise enhancement processing on initial image features to obtain target noise image features; a second processing module, used to perform degradation feature prediction processing on the original compressed image to obtain N different scale features of the original compressed image, wherein each scale feature represents an image degradation type feature; and a third processing module, used to use the N different scale features of the original compressed image as image enhancement condition features, perform image enhancement processing on the target noise image features, and obtain a target enhanced image.
[0015] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0016] The fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the above computer program or instructions are executed by a processor.
[0017] The fifth aspect of the present disclosure further provides a computer program product, including a computer program or instructions, which implement the steps of the above method when the above computer program or instructions are executed by a processor.
[0018] According to the embodiments of the present disclosure, the initial image features are subjected to noise enhancement processing to obtain target noise image features; the original compressed image is subjected to degradation feature prediction processing to obtain N different scale features of the original compressed image, wherein each scale feature represents the image degradation type feature; the N different scale features of the original compressed image are used as image enhancement condition features, and the target noise image features are subjected to image enhancement processing to obtain a target enhanced image. Since the original compressed image is subjected to degradation suppression prediction to obtain N different scale features, the problem of ignoring image structure information in the related art is avoided, so that the image structure information is further focused on, and the N different scale features of the original compressed image are used as image enhancement condition features to guide image restoration, and the restored target enhanced image has a data distribution and image structure that is more consistent with the real image. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0020] Figure 1 The application scenario diagram of the image enhancement method according to the embodiment of the present disclosure is schematically shown;
[0021] Figure 2 A flowchart of an image enhancement method according to an embodiment of the present disclosure is schematically shown;
[0022] Figure 3 A schematic diagram of structural feature extraction according to an embodiment of the present disclosure is schematically shown;
[0023] Figure 4 A schematic diagram of an image enhancement model according to an embodiment of the present disclosure is schematically shown;
[0024] Figure 5 A schematic diagram of a degradation feature prediction module according to an embodiment of the present disclosure is schematically shown;
[0025] Figure 6 The structure block diagram of the image enhancement device according to the embodiment of the present disclosure is schematically shown;
[0026] Figure 7 A block diagram of an electronic device suitable for implementing an image enhancement method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0027] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0028] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0029] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0030] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0031] In related technologies, image enhancement is performed by using the mapping relationship between compressed images and uncompressed images constrained by pixel-level loss. For example, a four-layer convolutional neural network architecture is used to perform the key functions of feature extraction, feature enhancement, feature mapping, and image reconstruction. However, this technology only focuses on the reconstruction of pixel values, but sacrifices the structural nature of the image in visual content and contextual relationships. These features are crucial for machine vision tasks such as target detection. Therefore, existing compressed image enhancement methods are not suitable for machine vision.
[0032] In view of this, an embodiment of the present disclosure provides an image enhancement method, comprising: performing noise enhancement processing on initial image features to obtain target noise image features; performing degradation feature prediction processing on the original compressed image to obtain N different scale features of the original compressed image, wherein each scale feature represents an image degradation type feature; using multiple different scale features of the original compressed image as image enhancement conditional features, performing image enhancement processing on the target noise image features to obtain a target enhanced image.
[0033] Figure 1 The application scenario diagram of the image enhancement method according to the embodiment of the present disclosure is schematically shown.
[0034] like Figure 1 As shown, the application scenario according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0035] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0036] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0037] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0038] It should be noted that the image enhancement method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the image enhancement device provided in the embodiment of the present disclosure can generally be set in the server 105. The image enhancement method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the image enhancement device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0039] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0040] The following will be based on Figure 1 The scene described by Figure 2~Figure 5 The image enhancement method of the disclosed embodiment is described in detail.
[0041] Figure 2 The flowchart of the image enhancement method according to the embodiment of the present disclosure is schematically shown.
[0042] like Figure 2 As shown, the image enhancement method of this embodiment includes operations S210 to S230.
[0043] In operation S210, noise enhancement processing is performed on the initial image features to obtain target noise image features.
[0044] According to an embodiment of the present disclosure, the initial image feature can be obtained by encoding the initial image using an encoder to reduce the computational complexity. This process can be formalized as: Mapped to a low-dimensional latent space, As the initial image features, wherein the initial image may be a high-definition image in any posture, and the encoder may be composed of a plurality of residual blocks and downsampling layers.
[0045] According to an embodiment of the present disclosure, the noise enhancement processing of the initial image features may be a step-by-step noise adding process, which gradually introduces noise in multiple time steps T, and eventually makes the target noise image features conform to the noise characteristics of the Gaussian distribution.
[0046] In operation S220, degradation feature prediction processing is performed on the original compressed image to obtain N different scale features of the original compressed image, wherein each scale feature represents a feature of an image degradation type.
[0047] According to an embodiment of the present disclosure, the original compressed image may be obtained by compressing the above initial image.
[0048] According to an embodiment of the present disclosure, the degradation feature prediction processing is performed on the original compressed image, and the Unet neural network is used to predict and extract the structured features of the original compressed image that may be lost during the compression process to obtain N different scale features, wherein the image degradation type feature represented by each scale feature may be multiple types of structured features that may be lost during the compression process of the original compressed image, and may be the structural features of the image in the visual content and contextual relationship, for example, the shape, texture, edge information, semantic relationship, etc. of the image.
[0049] In operation S230, N different scale features of the original compressed image are used as image enhancement condition features, and image enhancement processing is performed on the target noise image features to obtain a target enhanced image.
[0050] According to an embodiment of the present disclosure, image enhancement processing is performed on target noise image features, which may be to gradually denoise and reconstruct the target noise image features to obtain a clear image, i.e., a target enhanced image, wherein N different scale features are introduced in the above process to guide the restored target enhanced image to have a data distribution and image structure that is more consistent with the initial image.
[0051] According to the embodiments of the present disclosure, the initial image features are subjected to noise enhancement processing to obtain target noise image features; the original compressed image is subjected to degradation feature prediction processing to obtain N different scale features of the original compressed image, wherein each scale feature represents the image degradation type feature; the N different scale features of the original compressed image are used as image enhancement condition features, and the target noise image features are subjected to image enhancement processing to obtain a target enhanced image. Since the original compressed image is subjected to degradation suppression prediction to obtain N different scale features, the problem of ignoring image structure information in the related art is avoided, so that the image structure information is further focused on, and the N different scale features of the original compressed image are used as image enhancement condition features to guide image restoration, and the restored target enhanced image has a data distribution and image structure that is more consistent with the real image.
[0052] According to an embodiment of the present disclosure, the N different-scale features include M first-scale features and L second-scale features; performing degradation feature prediction processing on the original compressed image to obtain N different-scale features of the original compressed image, including: performing convolution on the m-1th first-scale feature to obtain the mth first-scale feature, and obtaining M first-scale features, wherein, when m is equal to 1, the first first-scale feature is obtained by convolving the original compressed image, M≥1, 1≤m≤M; performing degradation feature prediction processing on the i-1th second-scale feature to obtain the i-th second-scale feature, and obtaining L second-scale features, wherein, when i is equal to 1, the first second-scale feature is obtained by performing degradation feature prediction processing on the Mth first-scale feature, and L≥1, 1≤i≤L.
[0053] According to an embodiment of the present disclosure, convolution is performed on the m-1th first-scale feature to obtain the m-th first-scale feature. Feature extraction may be performed on the m-1th first-scale feature through convolution to obtain shallow-level image features.
[0054] For example, if M=3, the original compressed image is convolved to obtain the first first-scale feature, the first first-scale feature is convolved to obtain the second first-scale feature, the second first-scale feature is convolved to obtain the third first-scale feature, and the first first-scale feature, the second first-scale feature and the third first-scale feature are used as M first-scale features.
[0055] According to an embodiment of the present disclosure, performing degradation feature prediction processing on the i-1th second scale feature to obtain the i-th second scale feature may be to perform deep extraction on structural information about the image in the i-1th second scale feature to obtain the i-th second scale feature.
[0056] For example, if L=3, the third first scale feature is subjected to degraded feature prediction processing to obtain the first second scale feature, the first second scale feature is subjected to degraded feature prediction processing to obtain the second second scale feature, the second second scale feature is subjected to degraded feature prediction processing to obtain the third second scale feature, and the first second scale feature, the second second scale feature and the third second scale feature are used as L second scale features.
[0057] According to the embodiments of the present disclosure, the values of M and L can be set according to actual needs, and those skilled in the art will not make any limitation to this.
[0058] According to an embodiment of the present disclosure, a degradation feature prediction process is performed on the i-1th second-scale feature to obtain the i-th second-scale feature, including: convolving the i-1th second-scale feature to obtain a convolution feature; based on an attention mechanism, performing deep feature extraction on the convolution feature to obtain an intermediate feature; and performing structural feature extraction on the intermediate feature to obtain the i-th second-scale feature.
[0059] According to an embodiment of the present disclosure, convolution is performed on the i-1th second-scale feature, feature extraction is performed on the i-1th second-scale feature, and then a large convolution feature is obtained after being processed by a residual block, wherein the residual block includes two convolution layers and one residual connection to ensure information transmission and feature extraction.
[0060] According to an embodiment of the present disclosure, based on the attention mechanism, deep feature extraction is performed on the convolutional features, the global information of the image is processed, and the information of different modalities is fused to obtain intermediate features.
[0061] According to an embodiment of the present disclosure, structural features are extracted from the intermediate features to obtain structural features such as shape, texture, edge information, and semantic relationship, and obtain the i-th second-scale feature.
[0062] According to an embodiment of the present disclosure, structural feature extraction is performed on the intermediate features to obtain the i-th second-scale feature, including: performing fusion normalization processing on the intermediate features to obtain normalized features; and performing structural feature extraction on the normalized features to obtain the i-th second-scale feature.
[0063] According to an embodiment of the present disclosure, the intermediate features are fused and normalized to extract the image personalized features of the intermediate features, so as to distinguish the personalized differences between different original compressed images and adapt to images of different styles.
[0064] According to the embodiments of the present disclosure, structural feature extraction is performed on the normalized features, so that structural information such as context information of the normalized features can be effectively extracted to obtain the i-th second scale feature.
[0065] According to an embodiment of the present disclosure, intermediate features are fused and normalized to obtain normalized features, including: performing instance normalization on the intermediate features to obtain first normalized features; performing batch normalization on the intermediate features to obtain second normalized features; performing attention prediction on the third normalized features to obtain a first attention weight, wherein the third normalized feature is obtained by adding the first normalized feature and the second normalized feature; multiplying the first attention weight and the first normalized feature to obtain a fourth normalized feature; multiplying the first attention weight and the second normalized feature to obtain a fifth normalized feature; and performing feature fusion on the fourth normalized feature and the fifth normalized feature to obtain a normalized feature.
[0066] According to an embodiment of the present disclosure, batch normalization is to process different compression distortions contained in different original compressed images in the same batch, thereby ensuring that the network can effectively process multiple compression distortions at the same time.
[0067] According to an embodiment of the present disclosure, the instance normalization branch focuses on preserving the unique style of each original compressed image itself to prevent it from being interfered by other original compressed images in the same batch.
[0068] According to an embodiment of the present disclosure, after the first normalized feature and the second normalized feature are added, an attention map is generated through an attention prediction layer to determine the position of high importance in the third normalized feature obtained by adding, and the attention map is used as the first attention weight. The attention prediction layer is composed of two convolutions.
[0069] According to an embodiment of the present disclosure, the fourth normalized feature and the fifth normalized feature are fused by respectively performing dot multiplication of the first attention weight with the first normalized feature and the second normalized feature and then adding them together, and fusing the fused feature of the fourth normalized feature and the fifth normalized feature with the intermediate feature to obtain the normalized feature.
[0070] According to the embodiments of the present disclosure, the obtained normalized features can also be further subjected to channel attention operations through an enhancement module (Squeeze-and-Excitation Networks, SE), which can focus more on more important features in the normalized features and realize the network's powerful generalization capability for complex compression distortions.
[0071] According to the embodiments of the present disclosure, instance normalization and batch normalization are performed on intermediate features to help the model process different compression distortions in the same batch, and the features corresponding to different original compressed images in the same batch do not interfere with each other, thereby improving the stability and accuracy of the model while ensuring the processing efficiency of the model.
[0072] According to an embodiment of the present disclosure, structural feature extraction is performed on the normalized feature to obtain the i-th second-scale feature, including: determining a prompt parameter corresponding to the normalized feature; performing attention calculation on the normalized feature to obtain a second attention weight; multiplying the second attention weight and the prompt parameter to obtain a target prompt parameter; and fusing the normalized feature and the target prompt parameter to obtain the i-th second-scale feature.
[0073] According to an embodiment of the present disclosure, a prompt parameter corresponding to the normalized feature is determined, and the prompt parameter may be a learnable parameter such as the type and degree of degradation.
[0074] According to an embodiment of the present disclosure, based on the attention mechanism, global average pooling is performed on the normalized features to obtain a set of channel attention weights, namely, the second attention weights.
[0075] According to an embodiment of the present disclosure, the second attention weight is multiplied by the prompt parameter to obtain a target prompt parameter suitable for normalizing the feature, which is used to identify and extract key information in the feature.
[0076] According to an embodiment of the present disclosure, after the normalized feature is fused and convolved with the target prompt parameter, the obtained fused feature is fused again with the normalized feature to obtain the i-th second-scale feature.
[0077] According to an embodiment of the present disclosure, the normalized features and the target hint parameters are fused to embed degradation information in the normalized features, and then structural information such as context information associated with the original compressed image is extracted from the features after convolution and fusion as the second scale feature, which is used as a guiding condition for denoising in the subsequent image enhancement process.
[0078] Figure 3 A schematic diagram of structural feature extraction according to an embodiment of the present disclosure is schematically shown.
[0079] According to the embodiment of the present disclosure, the above operation can be performed as follows: Figure 3 As shown in the figure, the intermediate features are input into the batch normalization and instance normalization layers respectively to obtain the first normalized features and the second normalized features; the first normalized features and the second normalized features are fused and processed by the attention prediction layer to obtain the first attention weight; the first attention weight is dot-multiplied with the first normalized features and the second normalized features and then added, the result is multiplied with the intermediate features and then processed by the SE layer to obtain the normalized features; the normalized features are processed by the global average pooling layer, the convolution layer and the Softmax activation function layer to obtain the second attention weight; the second attention weight is multiplied and convolved with the prompt parameters to obtain the target prompt parameters; after the normalized features are fused and convolved with the target prompt parameters, the obtained fused features are fused with the normalized features to obtain the i-th second-scale features.
[0080] According to the embodiments of the present disclosure, based on the above Figure 3 The operation shown above obtains the second scale feature As shown in formula (1),
[0081] (1)
[0082] in, Indicates intermediate features; represents batch normalization and instance normalization layers; represents the attention prediction layer; represents the first attention weight; represents the enhancement layer; represents the normalized features; Indicates the attention calculation for the normalized features; and They represent convolutional layers with kernel sizes of 1×1 and 3×3 respectively; represents the activation function; P represents the prompt parameter, represents the target prompt parameter; represents the second scale feature; Represents an element-wise multiplication operation.
[0083] According to an embodiment of the present disclosure, if four different scale features are extracted, including one first scale feature and three second scale features, it can be as shown in formula (2):
[0084] (2)
[0085] in, represents the original compressed image; Represents a convolution layer with a convolution kernel size of 3×3; represents the residual block; Represents the ransformer block; Respectively represent the first scale feature of the first scale, the second scale feature of the second scale, the second scale feature of the third scale, and the second scale feature of the fourth scale; Represents the degradation feature prediction submodule.
[0086] According to an embodiment of the present disclosure, multiple different scale features of the original compressed image are used as image enhancement condition features, and image enhancement processing is performed on the target noise image features to obtain a target enhanced image, including: down-sampling the t-th step denoising image features to obtain the t-th step down-sampled denoising image features; up-sampling the t-th step down-sampled denoising image features and multiple different scale features to obtain the t+1-th step denoising image features, and obtain the T-th step denoising image features; decoding the T-th step denoising image features to obtain the target enhanced image.
[0087] According to an embodiment of the present disclosure, the denoised image features of the t-th step are downsampled to extract high-level features to obtain the downsampled denoised image features of the t-th step, wherein the downsampling block performing the downsampling process is composed of multiple residual blocks and a downsampling layer, and the downsampling layer is implemented using a convolution with a convolution kernel size of 3×3 and a step size of 2; the residual block enables the network to be built deeper, and at the same time, the time embedding information is embedded in the model, and each residual block contains two convolution layers and a residual connection to ensure the transmission of information and the extraction of features.
[0088] According to an embodiment of the present disclosure, upsampling processing is performed on the downsampled denoised image features and multiple different-scale features of the t-th step, and the resolution is gradually restored, wherein the upsampling block performing the upsampling processing is composed of multiple residual blocks and an upsampling layer, and the upsampling layer uses a bilinear interpolation algorithm to amplify the spatial resolution of the downsampled denoised image features and multiple different-scale features of the t-th step, and then a convolution layer is used to adjust the number of channels of the downsampled denoised image features and multiple different-scale features of the t-th step and restore the details.
[0089] According to an embodiment of the present disclosure, the above-mentioned upsampling block and downsampling block are connected through an intermediate block, and the t-th step downsampling denoised image features are input to the upsampling block through the intermediate block. Among them, the intermediate block includes a residual block and a spatial attention (Transformer) block. The Transformer block is composed of group normalization, convolution, basic Transformer block and convolution to ensure the effective fusion of local features and global features. The basic Transformer block contains self-attention, cross-attention and feedforward network, which can process the global information of the t-th step downsampling denoised image features and fuse information of different modalities.
[0090] According to the embodiments of the present disclosure, after T denoising processes, the noise is gradually removed. During the denoising process, the denoising process is performed through the Unet neural network workflow, the noise residual is predicted, and the target noise image features are gradually reconstructed into clear image features. The clear image features are reconstructed into a pixel-level image through a decoder to obtain a target enhanced image, thereby completing image enhancement.
[0091] For example, when T=3, the target noise image features are downsampled to obtain downsampled denoised image features of the target noise image features, and the downsampled denoised image features of the target noise image features and multiple different-scale features are upsampled to obtain the first-step denoised image features; the first-step denoised image features are downsampled to obtain the first-step downsampled denoised image features, and the first-step downsampled denoised image features and multiple different-scale features are upsampled to obtain the second-step denoised image features; the second-step denoised image features are downsampled to obtain the second-step downsampled denoised image features, and the second-step downsampled denoised image features and multiple different-scale features are upsampled to obtain the third-step denoised image features, and the third-step denoised image features are used as the T-th step denoised image features, and are decoded to obtain the target enhanced image.
[0092] According to an embodiment of the present disclosure, upsampling is performed on the t-th step down-sampled denoised image feature and multiple different-scale features to obtain the t+1-th step denoised image feature, including: fusing the k-1-th sub-upsampled denoised image feature and the k-th scale feature to obtain the fused k-1-th sub-upsampled denoised image feature; upsampling is performed on the fused k-1-th sub-upsampled denoised image feature to obtain the k-th sub-upsampled denoised image feature, and the N-th sub-upsampled denoised image feature is obtained, wherein, when k is equal to 1, the k-1-th sub-upsampled denoised image feature is obtained by fusing and upsampling the t-th step down-sampled denoised image feature and the 1st scale feature, 1≤k-1≤N, N=M+L; and the N-th sub-upsampled denoised image feature is determined as the t+1-th step denoised image feature.
[0093] According to an embodiment of the present disclosure, the k-1th sub-upsampling denoised image feature and the k-th scale feature are fused, and when the fused k-1th sub-upsampling denoised image feature is upsampled, the image feature can be denoised based on the structural information condition given by the k-th scale feature to obtain the k-th sub-upsampling denoised image feature.
[0094] For example, when N=3, the t-th step down-sampling denoising image feature and the first scale feature are fused and up-sampled to obtain the fused t-th step down-sampling denoising image feature, and the fused t-th step down-sampling denoising image feature is up-sampled to obtain the first sub-upsampling denoising image feature; the first sub-upsampling denoising image feature and the second scale feature are fused and up-sampled to obtain the fused first sub-upsampling denoising image feature, and the fused first sub-upsampling denoising image feature is up-sampled to obtain the second sub-upsampling denoising image feature; the second sub-upsampling denoising image feature and the third scale feature are fused and up-sampled to obtain the fused second sub-upsampling denoising image feature, and the fused second sub-upsampling denoising image feature is up-sampled to obtain the third sub-upsampling denoising image feature, and the third sub-upsampling denoising image feature is used as the N-th sub-upsampling denoising image feature, and is determined as the t+1-th step denoising image feature.
[0095] According to an embodiment of the present disclosure, based on the above operation, the denoised image features of the t+1th step are determined, and then the denoised image features of the Tth step are determined from the denoised image features of the t+1th step, and are decoded to obtain a target enhanced image.
[0096] According to the embodiment of the present disclosure, the above is to obtain the target enhanced image The image enhancement processing can be shown as formula (3):
[0097] (3)
[0098] in, represents the target noise image characteristics; T represents the total number of denoising steps; Represents the features after denoising at step t+1; Represents features of different scales; represents the target enhanced image; D represents the decoder; , and Represent the upsampling block, the intermediate block and the upsampling block respectively, Represents the Unet neural network at the tth step in the noise addition process, including , ADE represents the degenerate feature prediction, Represents the original compressed image.
[0099] According to an embodiment of the present disclosure, noise enhancement processing is performed on the initial image features to obtain target noise image features, including: performing noise enhancement processing on the p-1th step noise image features to obtain the p+1th step noise image features, and obtaining the Pth step noise image features, wherein, when p is equal to 1, the first step noise image features are obtained by performing noise enhancement processing on the initial image features; and determining the Pth step noise image features as the target noise image features.
[0100] According to an embodiment of the present disclosure, the noise enhancement submodule performs an image noise addition process through a Unet neural network workflow, wherein the target noise image feature is obtained. The operation can be shown as formula (4):
[0101] (4)
[0102] in, represents the initial image; E represents the encoder; represents the initial image features, p represents the time step, and the maximum value is ; They represent upsampling block, intermediate block and upsampling block respectively; Represents the Unet neural network at the pth step in the noise addition process, including .
[0103] Figure 4 A schematic diagram of an image enhancement model according to an embodiment of the present disclosure is schematically shown; Figure 5 The schematic diagram of the degradation feature prediction module according to the embodiment of the present disclosure is schematically shown.
[0104] According to an embodiment of the present disclosure, using Figure 4 The image enhancement model shown executes the above-mentioned image enhancement method, wherein the image enhancement model includes a noise enhancement module, a degradation feature prediction module and an image enhancement module. The noise enhancement module includes an encoder and a noise enhancement submodule, and the image enhancement module includes a downsampling submodule, an upsampling submodule and a decoder.
[0105] The initial image P1 is decoded by the decoder to obtain the initial image features, and the noise enhancement submodule is used to perform noise enhancement on the initial image features to obtain the target noise image features. The degradation feature prediction module performs degradation feature prediction on the original compressed image P2 to obtain N different scale features, and the image enhancement module performs multi-step image enhancement processing, wherein the downsampling submodule is used to downsample the input features, and the upsampling submodule is used to upsample the features output by the downsampling submodule and N different scale features. The features obtained after completing the preset number of steps of image enhancement processing are The decoder is used to decode the target enhanced image P3.
[0106] Among them, the degradation feature prediction module is as follows Figure 5 As shown, the degradation feature prediction module includes a convolution submodule and multiple degradation feature prediction submodules. The convolution submodule is used to extract features of the original compressed image to obtain a first scale feature (A) of the first scale, the first degradation feature prediction submodule is used to perform degradation feature prediction on the first scale feature (A) of the first scale to obtain a second scale feature (B) of the second scale, the second degradation feature prediction submodule is used to perform degradation feature prediction on the second scale feature (B) of the second scale to obtain a second scale feature (C) of the third scale, and the third degradation feature prediction submodule is used to perform degradation feature prediction on the second scale feature (C) of the third scale to obtain a second scale feature (D) of the fourth scale.
[0107] Furthermore, the upsampling submodule includes a first upsampling unit, a second upsampling unit, a third upsampling unit and a fourth upsampling unit.
[0108] In the image enhancement processing of the t-th step, after the downsampling submodule outputs the downsampled denoised image features of the t-th step, the downsampled denoised image features of the t-th step and the first scale features (A) of the first scale are fused and upsampled in the first upsampling unit (1) to obtain the first sub-upsampled denoised image features; the first sub-upsampled denoised image features and the second scale features (B) of the second scale are fused and upsampled in the second upsampling unit (2) to obtain the second sub-upsampled denoised image features; the second sub-upsampled denoised image features and the second scale features (C) of the third scale are fused and upsampled in the third upsampling unit (3) to obtain the third sub-upsampled denoised image features; and the third sub-upsampled denoised image features and the second scale features (D) of the fourth scale are fused and upsampled in the fourth upsampling unit (4) to obtain the upsampled denoised image features of the t-th step.
[0109] According to an embodiment of the present disclosure, the number of noise addition steps is consistent with the number of denoising steps. When training the image enhancement model, the deviation between the noise added and the noise removed in each step can be calculated, and back propagation can be performed to update the image enhancement model parameters so that its prediction error on the test set is minimized.
[0110] Specifically, the optimizer (Adaptive Moment Estimation, Adam) can be selected, the initial learning rate is set to 0.00001, the total number of iterations is designed to be 160000, and the image enhancement model with the best loss in this process is saved. The test data set is applied to obtain the enhanced image, and the Faster RCNN object detection network is used to verify the performance of the obtained image enhancement model.
[0111] Based on the above image enhancement method, the present disclosure also provides an image enhancement device. Figure 6 The device is described in detail.
[0112] Figure 6 The structural block diagram of the image enhancement device according to an embodiment of the present disclosure is schematically shown.
[0113] like Figure 6 As shown, the image enhancement device 600 of this embodiment includes a first processing module 610 , a second processing module 620 and a third processing module 630 .
[0114] The first processing module 610 is used to perform noise enhancement processing on the initial image features to obtain target noise image features. In one embodiment, the first processing module 610 can be used to perform the operation S210 described above, which will not be described in detail here.
[0115] The second processing module 620 is used to perform degradation feature prediction processing on the original compressed image to obtain N different scale features of the original compressed image, wherein each scale feature represents the image degradation type feature. In one embodiment, the second processing module 620 can be used to perform the operation S220 described above, which will not be repeated here.
[0116] The third processing module 630 is used to use the N different scale features of the original compressed image as image enhancement condition features to perform image enhancement processing on the target noise image features to obtain a target enhanced image. In one embodiment, the third processing module 630 can be used to perform the operation S230 described above, which will not be repeated here.
[0117] According to the embodiments of the present disclosure, the initial image features are subjected to noise enhancement processing to obtain target noise image features; the original compressed image is subjected to degradation feature prediction processing to obtain N different scale features of the original compressed image, wherein each scale feature represents the image degradation type feature; the N different scale features of the original compressed image are used as image enhancement condition features, and the target noise image features are subjected to image enhancement processing to obtain a target enhanced image. Since the original compressed image is subjected to degradation suppression prediction to obtain N different scale features, the problem of ignoring image structure information in the related art is avoided, so that the image structure information is further focused on, and the N different scale features of the original compressed image are used as image enhancement condition features to guide image restoration, and the restored target enhanced image has a data distribution and image structure that is more consistent with the real image.
[0118] According to an embodiment of the present disclosure, any multiple modules of the first processing module 610, the second processing module 620 and the third processing module 630 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the first processing module 610, the second processing module 620 and the third processing module 630 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in any appropriate combination of any of them. Alternatively, at least one of the first processing module 610, the second processing module 620 and the third processing module 630 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding function can be executed.
[0119] According to an embodiment of the present disclosure, the second processing module 620 includes a convolution submodule and a first processing submodule.
[0120] The convolution submodule is used to convolve the m-1th first-scale feature to obtain the mth first-scale feature and obtain M first-scale features, wherein when m is equal to 1, the first first-scale feature is obtained by convolving the original compressed image, M≥1, 1≤m≤M.
[0121] The first processing submodule is used to perform degradation feature prediction processing on the i-1th second scale feature to obtain the i-th second scale feature and obtain L second scale features, wherein when i is equal to 1, the first second scale feature is obtained by performing degradation feature prediction processing on the Mth first scale feature, and L≥1, 1≤i≤L.
[0122] According to an embodiment of the present disclosure, the first processing submodule includes a convolution unit, a first extraction unit and a second extraction unit.
[0123] The convolution unit is used to perform convolution on the i-1th second-scale feature to obtain a convolution feature.
[0124] The first extraction unit is used to perform deep feature extraction on the convolution feature based on the attention mechanism to obtain the intermediate feature.
[0125] The second extraction unit is used to extract structural features from the intermediate features to obtain the i-th second scale feature.
[0126] According to an embodiment of the present disclosure, the second extraction unit includes a processing subunit and an extraction subunit.
[0127] The processing subunit is used to perform fusion and normalization processing on the intermediate features to obtain normalized features.
[0128] The extraction subunit is used to extract the structural features of the normalized features to obtain the i-th second-scale feature.
[0129] According to an embodiment of the present disclosure, the processing subunit includes an instance normalization component, a batch normalization component, a prediction component, a first multiplication component, a second multiplication component and a first fusion component.
[0130] The instance normalization component is used to perform instance normalization on the intermediate features to obtain the first normalized features.
[0131] Batch normalization component, used to batch normalize the intermediate features to obtain the second normalized features
[0132] A prediction component is used to perform attention prediction on the third normalized feature to obtain a first attention weight, wherein the third normalized feature is obtained by adding the first normalized feature and the second normalized feature.
[0133] The first multiplication component is used to multiply the first attention weight and the first normalized feature to obtain a fourth normalized feature.
[0134] The second multiplication component is used to multiply the first attention weight and the second normalized feature to obtain a fifth normalized feature.
[0135] The first fusion component is used to perform feature fusion on the fourth normalized feature and the fifth normalized feature to obtain a normalized feature.
[0136] According to an embodiment of the present disclosure, the extraction subunit includes a determination component, a calculation component, a third multiplication component and a second fusion component.
[0137] A determination component is used to determine a prompt parameter corresponding to the normalized feature.
[0138] A calculation component is used to perform attention calculation on the normalized features to obtain a second attention weight.
[0139] The third multiplication component is used to multiply the second attention weight and the prompt parameter to obtain the target prompt parameter.
[0140] The second fusion component is used to fuse the normalized features and the target prompt parameters to obtain the i-th second scale feature.
[0141] According to an embodiment of the present disclosure, the third processing module 630 includes a second processing submodule, a third processing submodule and a decoding submodule.
[0142] The second processing submodule is used to perform down-sampling processing on the denoised image features of the t-th step to obtain the down-sampled denoised image features of the t-th step.
[0143] The third processing submodule is used to upsample the downsampled denoised image features of the t-th step and multiple different-scale features to obtain the denoised image features of the t+1-th step and obtain the denoised image features of the T-th step.
[0144] The decoding submodule is used to decode the denoised image features in the Tth step to obtain the target enhanced image.
[0145] According to an embodiment of the present disclosure, the third processing submodule includes a fusion unit, a processing unit and a determination unit.
[0146] The fusion unit is used to fuse the k-1th sub-upsampling denoised image feature and the kth scale feature to obtain the fused k-1th sub-upsampling denoised image feature.
[0147] The processing unit is used to perform upsampling processing on the fused k-1th sub-upsampling denoised image feature to obtain the kth sub-upsampling denoised image feature and the Nth sub-upsampling denoised image feature, wherein when k is equal to 1, the k-1th sub-upsampling denoised image feature is obtained by fusing and upsampling the tth step downsampling denoised image feature and the first scale feature, 1≤k-1≤N, N=M+L.
[0148] The determining unit is used to determine the Nth sub-upsampled denoised image feature as the t+1th step denoised image feature.
[0149] According to an embodiment of the present disclosure, the first processing module 610 includes a noise adding submodule and a determining submodule.
[0150] The noise adding submodule is used to perform noise enhancement processing on the noisy image features of the p-1th step to obtain the noisy image features of the p+1th step and obtain the noisy image features of the Pth step, wherein, when p is equal to 1, the noisy image features of the first step are obtained by performing noise enhancement processing on the initial image features.
[0151] The determination submodule is used to determine the noisy image features in the Pth step as the target noisy image features.
[0152] Figure 7 A block diagram of an electronic device suitable for implementing an image enhancement method according to an embodiment of the present disclosure is schematically shown.
[0153] like Figure 7As shown, the electronic device according to the embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage part 708 to the random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (for example, an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include an onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present disclosure.
[0154] In RAM 703, various programs and data required for the operation of the electronic device are stored. The processor 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to the embodiment of the present disclosure by executing the programs in ROM 702 and / or RAM 703. It should be noted that the program may also be stored in one or more memories other than ROM 702 and RAM 703. The processor 701 may also perform various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0155] According to an embodiment of the present disclosure, the electronic device may further include an input / output (I / O) interface 705, which is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that a computer program read therefrom is installed into the storage portion 708 as needed.
[0156] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0157] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 702 and / or RAM 703 described above and / or one or more memories other than ROM 702 and RAM 703.
[0158] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the image enhancement method provided by the embodiment of the present disclosure.
[0159] The above functions defined in the system / device of the embodiment of the present disclosure are performed when the computer program is executed by the processor 701. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0160] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 709, and / or installed from the removable medium 711. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0161] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.
[0162] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).
[0163] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0164] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present disclosure.
[0165] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. An image enhancement method, characterized in that: The method comprises: Perform noise enhancement processing on the initial image features to obtain target noise image features; Performing degradation feature prediction processing on the original compressed image to obtain N different scale features of the original compressed image, wherein each scale feature represents a feature of the image degradation type; The N different scale features of the original compressed image are used as image enhancement condition features, and image enhancement processing is performed on the target noise image features to obtain a target enhanced image.
2. The method according to claim 1, characterized in that: The N different scale features include M first scale features and L second scale features; The degradation feature prediction process is performed on the original compressed image to obtain N different scale features of the original compressed image, including: Convolving the m-1th first scale feature to obtain the mth first scale feature, and obtaining M first scale features, wherein when m is equal to 1, the first first scale feature is obtained by convolving the original compressed image, M≥1, 1≤m≤M; Performing a degraded feature prediction process on the i-1th second scale feature to obtain an ith second scale feature, and obtaining L second scale features, wherein, when i is equal to 1, the first second scale feature is obtained by performing a degraded feature prediction process on the Mth first scale feature, and L≥1, 1≤i≤L.
3. The method according to claim 2, characterized in that The performing degradation feature prediction processing on the i-1th second scale feature to obtain the i-th second scale feature includes: Convolving the (i-1)th second-scale feature to obtain a convolution feature; Based on the attention mechanism, deep feature extraction is performed on the convolutional features to obtain intermediate features; Structural features are extracted from the intermediate features to obtain the i-th second scale feature.
4. The method according to claim 3, characterized in that The extracting structural features from the intermediate features to obtain the i-th second scale feature includes: Performing fusion normalization processing on the intermediate features to obtain normalized features; Perform structural feature extraction on the normalized feature to obtain the i-th second scale feature.
5. The method according to claim 4, characterized in that The fusing and normalizing the intermediate features to obtain normalized features includes: Performing instance normalization on the intermediate feature to obtain a first normalized feature; Performing batch normalization on the intermediate features to obtain second normalized features; Performing attention prediction on a third normalized feature to obtain a first attention weight, wherein the third normalized feature is obtained by adding the first normalized feature and the second normalized feature; Multiplying the first attention weight and the first normalized feature to obtain a fourth normalized feature; Multiplying the first attention weight and the second normalized feature to obtain a fifth normalized feature; The fourth normalized feature and the fifth normalized feature are subjected to feature fusion to obtain the normalized feature.
6. The method according to claim 4, characterized in that The extracting structural features from the normalized features to obtain the i-th second scale feature includes: determining a prompt parameter corresponding to the normalized feature; Performing attention calculation on the normalized feature to obtain a second attention weight; Multiplying the second attention weight and the prompt parameter to obtain a target prompt parameter; The normalized feature and the target prompt parameter are fused to obtain the i-th second scale feature.
7. The method according to claim 1, characterized in that The method of using the N different scale features of the original compressed image as image enhancement condition features and performing image enhancement processing on the target noise image features to obtain a target enhanced image includes: Down-sampling the denoised image features of the t-th step to obtain the down-sampled denoised image features of the t-th step; Upsampling the downsampled denoised image features of the t-th step and the plurality of the different-scale features to obtain the denoised image features of the t+1-th step, and to obtain the denoised image features of the T-th step; The T-th step denoising image features are decoded to obtain the target enhanced image.
8. The method according to claim 7, characterized in that The upsampling of the t-th step down-sampled denoised image features and the plurality of the different-scale features to obtain the t+1-th step denoised image features comprises: Fusing the k-1th sub-upsampling denoised image feature and the kth scale feature to obtain a fused k-1th sub-upsampling denoised image feature; Performing upsampling processing on the fused k-1th sub-upsampled denoised image feature to obtain the kth sub-upsampled denoised image feature, and obtaining the Nth sub-upsampled denoised image feature, wherein when k is equal to 1, the k-1th sub-upsampled denoised image feature is obtained by fusing and upsampling the tth step downsampled denoised image feature and the first scale feature, 1≤k-1≤N, N=M+L; The Nth sub-upsampling denoised image feature is determined as the t+1th step denoised image feature.
9. The method according to claim 1, characterized in that: The performing noise enhancement processing on the initial image features to obtain target noise image features includes: Performing noise enhancement processing on the noisy image feature of the p-1th step to obtain the noisy image feature of the p+1th step, and obtaining the noisy image feature of the Pth step, wherein, when p is equal to 1, the noisy image feature of the first step is obtained by performing noise enhancement processing on the initial image feature; The P-th step noisy image feature is determined as the target noise image feature.
10. An image enhancement device, characterized in that: The device comprises: The first processing module is used to perform noise enhancement processing on the initial image features to obtain target noise image features; A second processing module is used to perform degradation feature prediction processing on the original compressed image to obtain N different scale features of the original compressed image, wherein each scale feature represents a degradation type feature of the image; The third processing module is used to use the N different scale features of the original compressed image as image enhancement condition features, perform image enhancement processing on the target noise image features, and obtain a target enhanced image.