Image Enhancement Method, Apparatus, Device, and Readable Storage Medium
By adopting multi-stage deep semantic feature processing and downsampling convolution processing based on sliding window mechanisms of different scales in the image enhancement model, the problem of poor image enhancement effect in multi-weather environments is solved, and efficient and real-time image enhancement effect is achieved.
Patent Information
- Application Number
- CN202311487671.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-11-09
AI Technical Summary
The existing image enhancement model has poor image enhancement effect in multi-weather environments, and the model structure and calculation complexity are high, which affects real-time performance.
Multi-stage deep semantic feature processing and downsampling convolution processing based on sliding window mechanisms of different scales are used to build an image enhancement model through the encoder-decoder structure, reducing model complexity and improving image enhancement effect.
It effectively improves the image enhancement effect, reduces the model structure and calculation complexity, improves real-time performance, and can achieve accurate image recovery in a variety of harsh weather environments.
Smart Images

Figure CN117575971B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and particularly relates to an image enhancement method, device, equipment and readable storage medium. Background Art
[0002] With the wide popularization of visual sensors and the rapid development of artificial intelligence technology, the automated visual processing technology and methods for outdoor scenarios have become a research hotspot and a research difficulty in the industrial and academic fields. Among them, the image enhancement technology for various different environments is one of the most rapidly developing and widely used technologies in the fields of artificial intelligence and intelligent driving, and has important research and application values; this technology aims to process the images captured in different environments such as bad weather through an algorithm model to output high-quality images after removing weather noises (such as rain, snow, etc.) in the images.
[0003] In related technologies, the current image enhancement models usually can only enhance images in a single weather environment. However, the actual weather environment often contains multiple weather types, so that the types and distributions of weather noises are diverse and complex. Therefore, the current image enhancement models are not suitable for enhancing images in multi-weather environments, and there is a problem of poor image enhancement effect; in addition, although a small number of current technologies can initially achieve image enhancement under multi-weather conditions, the model structure and computational complexity are relatively high, and the computing power resources consumption is relatively high, thus affecting the real-time performance of image enhancement. Summary of the Invention
[0004] The present application provides an image enhancement method, device, equipment and readable storage medium, which can effectively improve the image enhancement effect while reducing the model structure complexity, computational complexity and computing power resources consumption.
[0005] In a first aspect, an embodiment of the present application provides an image enhancement method, and the image enhancement method includes:
[0006] Performing multi-stage deep semantic feature processing and downsampling convolution processing on the image to be enhanced based on a sliding window mechanism of different scales to output a deep feature tensor;
[0007] Performing upsampling processing on the deep feature tensor to output an enhanced image;
[0008] Wherein, a sliding window mechanism of a first stage is constructed based on a moving window mechanism.
[0009] Combined with the first aspect, in an implementation manner, the performing multi-stage deep semantic feature processing on the image to be enhanced based on a sliding window mechanism of different scales includes:
[0010] Perform the first-stage deep semantic feature processing on the image to be enhanced based on the sliding window mechanism of the first scale;
[0011] Perform the second-stage deep semantic feature processing on the output result of the first stage based on the sliding window mechanism of the second scale;
[0012] Perform the third-stage deep semantic feature processing on the output result of the second stage based on the sliding window mechanism of the third scale;
[0013] Perform the fourth-stage deep semantic feature processing on the output result of the third stage based on the sliding window mechanism of the first scale;
[0014] Wherein, the third scale is greater than the second scale and the second scale is greater than the first scale.
[0015] Combined with the first aspect, in one embodiment, the performing the first-stage deep semantic feature processing on the image to be enhanced based on the sliding window mechanism of the first scale includes:
[0016] Perform regularization processing on the image to be enhanced based on the sliding window mechanism of the first scale to obtain a first regularization tensor;
[0017] Perform equal-dimensional mapping processing on the first regularization tensor to obtain a feature tensor;
[0018] Perform regularization processing on the image to be enhanced and the feature tensor to obtain a second regularization tensor;
[0019] Perform equal-dimensional transformation on the second regularization tensor to obtain a target tensor, and repeat the above steps once based on the target tensor to obtain the final target tensor.
[0020] Combined with the first aspect, in one embodiment, the performing upsampling processing on the deep feature tensor includes:
[0021] Perform multiple upsampling processes on the deep feature tensor;
[0022] For each upsampling process, perform deconvolution processing on the result of the previous upsampling process to obtain a deconvolution tensor;
[0023] Perform first convolution processing on the deconvolution tensor to obtain a first convolution tensor;
[0024] Map the first convolution tensor through an activation function to obtain an activation tensor;
[0025] Perform second convolution processing on the activation tensor to obtain a second convolution tensor,
[0026] Perform a residual connection on the second convolutional tensor, the transposed convolutional tensor, and the downsampling result corresponding to the current upsampling processing scale, and use the residual result and the downsampling result corresponding to the next upsampling processing scale as the input for the next upsampling processing;
[0027] Among them, during the first upsampling process, the object of the transposed convolutional process is the depth feature tensor.
[0028] Combined with the first aspect, in an implementation manner, the downsampling convolutional process is implemented by a two-dimensional convolution with a stride of 2, a convolutional kernel of 7, and an output channel dimension that is twice the input channel dimension.
[0029] In a second aspect, an embodiment of the present application provides an image enhancement device, and the image enhancement device includes: an encoder and a decoder;
[0030] The encoder is used to perform multi-stage deep semantic feature processing and downsampling convolutional processing on the image to be enhanced based on a sliding window mechanism of different scales, so as to output a depth feature tensor;
[0031] The decoder is used to perform upsampling processing on the depth feature tensor to output an enhanced image;
[0032] Among them, each stage of the encoder includes a transformer module and a downsampling module connected in series. The transformer modules in different stages include sliding window multi-head self-attention modules of different scales. One of the sliding window multi-head self-attention modules in the first stage is constructed based on a moving window mechanism. The transformer module and the downsampling module are used to perform deep semantic feature processing and downsampling convolutional processing on the image to be enhanced.
[0033] Combined with the second aspect, in an implementation manner, the transformer module includes two attention units connected in series, and each attention unit includes a first-layer regularization module, a sliding window multi-head self-attention module, a second-layer regularization module, and a multi-layer perceptron connected in sequence;
[0034] The first-layer regularization module is used to perform regularization processing on the input tensor to obtain a first regularized tensor;
[0035] The sliding window multi-head self-attention module is used to perform an equal-dimensional mapping process on the first regularized tensor to obtain a feature tensor;
[0036] The second-layer regularization module is used to perform regularization processing on the input tensor and the feature tensor to obtain a second regularized tensor;
[0037] The multi-layer perceptron is used to perform an equal-dimensional transformation on the second regularized tensor to obtain a target tensor.
[0038] In combination with the second aspect, in one implementation, the sliding window scale of the sliding window multi-head self-attention module in the first stage is the first scale; the sliding window scale of the sliding window multi-head self-attention module in the second stage is the second scale; the sliding window scale of the sliding window multi-head self-attention module in the third stage is the third scale; the sliding window scale of the sliding window multi-head self-attention module in the fourth stage is the first scale; wherein, the third scale is greater than the second scale and the second scale is greater than the first scale.
[0039] In combination with the second aspect, in one implementation, the decoder is specifically configured to:
[0040] Perform multiple upsampling processes on the depth feature tensor;
[0041] For each upsampling process, perform deconvolution processing on the result of the previous upsampling process to obtain a deconvolution tensor;
[0042] Perform first convolution processing on the deconvolution tensor to obtain a first convolution tensor;
[0043] Map the first convolution tensor through an activation function to obtain an activation tensor;
[0044] Perform second convolution processing on the activation tensor to obtain a second convolution tensor,
[0045] Perform residual connection on the second convolution tensor, the deconvolution tensor, and the downsampling result corresponding to the current upsampling process scale, and use the residual result and the downsampling result corresponding to the next upsampling process scale as the input for the next upsampling process;
[0046] Wherein, during the first upsampling process, the object of the deconvolution processing is the depth feature tensor.
[0047] In combination with the second aspect, in one implementation, the downsampling module includes a two-dimensional convolution with a stride of 2, a convolution kernel of 7, and an output channel dimension that is twice the input channel dimension.
[0048] In a third aspect, an embodiment of the present application provides an image enhancement device, which includes a processor, a memory, and an image enhancement program stored on the memory and executable by the processor. When the image enhancement program is executed by the processor, the steps of the aforementioned image enhancement method are implemented.
[0049] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which an image enhancement program is stored. When the image enhancement program is executed by a processor, the steps of the aforementioned image enhancement method are implemented.
[0050] The beneficial effects brought by the technical solution provided in the embodiment of the present application include:
[0051] By performing multi-stage deep semantic feature processing and downsampling convolution processing on the image to be enhanced, the feature learning ability of the model is improved, and thus it can robustly generalize various environmental noises. The learned features not only cover the detailed information in the original image but also filter out environmental noise information such as weather in the original image to a certain extent, thereby realizing accurate restoration of images in various harsh weather and other environmental scenarios; and through a hybrid window mechanism of different scales and only using a moving window mechanism in the first stage to efficiently implement a hierarchical model architecture, the model structure and computational complexity are effectively reduced and the high computing power consumption of a global-based model architecture is avoided, while the quality of image enhancement can be significantly improved. Description of the Drawings
[0052] Figure 1 It is a schematic flowchart of the embodiment of the image enhancement method of the present application;
[0053] Figure 2 It is a schematic structural diagram of the encoder involved in the solution of the embodiment of the present application;
[0054] Figure 3 It is a schematic structural diagram of the Transformer module involved in the solution of the embodiment of the present application;
[0055] Figure 4 It is a schematic specific structural diagram of the encoder involved in the solution of the embodiment of the present application;
[0056] Figure 5 It is a schematic structural diagram of the deconvolution residual module involved in the solution of the embodiment of the present application;
[0057] Figure 6 It is a schematic hardware structural diagram of the image enhancement device involved in the solution of the embodiment of the present application. Detailed Embodiments
[0058] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0059] First, some technical terms in the present application are explained to facilitate the understanding of the present application by those skilled in the art.
[0060] Transformer: A deep learning model widely used in the field of natural language processing, such as machine translation, text classification, question answering systems, etc.
[0061] Swin Transformer: A deep learning model based on Transformer.
[0062] Shifiting window: A moving window mechanism, which is a sliding window mechanism proposed by Swin Transformer. This mechanism realizes efficient hierarchical feature expression by non-uniformly partitioning the global features and then calculating the multi-head self-attention for the divided local regions respectively.
[0063] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0064] In the first aspect, the embodiments of this application provide an image enhancement method.
[0065] In one embodiment, referring to Figure 1 , Figure 1 is the schematic flowchart of the embodiment of the image enhancement method of this application. As Figure 1 shown, the image enhancement method includes:
[0066] Step S10: Perform multi-stage deep semantic feature processing and downsampling convolution processing on the image to be enhanced based on the sliding window mechanism of different scales to output a deep feature tensor; wherein, the sliding window mechanism of the first stage is constructed based on the moving window mechanism.
[0067] Exemplarily, it should be understood that the current image enhancement methods for various harsh weather and other environmental scenarios mainly face two aspects of problems: 1) The applicable scenarios and image enhancement effects are limited, that is, the current methods mostly propose corresponding model algorithms for specific low-quality single scenarios (i.e., the same type of climate environment, such as rainy days, foggy days, low-light scenarios, etc.), and the enhancement effects of these model algorithms will be greatly weakened in low-quality scenarios mixed with multiple harsh weathers; 2) It is difficult to balance the real-time performance of the model and the enhancement effect. Currently, a small number of technologies have begun to specifically enhance the images captured in multiple low-quality scenarios. However, in order to enhance the generalization ability of the model, complex model structures are often required, which leads to high computational complexity and high consumption of computing power resources, making it difficult to meet the real-time performance of the application.
[0068] To solve the above problems, in this embodiment, an image enhancement model will be constructed based on the encoder-decoder structure, and there is also a skip connection between the encoder and the decoder, so that the encoder and the decoder together form a U-shaped Net model structure, that is, the encoder and the decoder form an end-to-end image enhancement model in a series and skip connection manner. Among them, each stage of the encoder includes a cascaded Transformer module and a downsampling module in series. The Transformer modules in different stages include sliding window multi-head self-attention modules of different scales, and one of the sliding window multi-head self-attention modules in the first stage is constructed based on the moving window mechanism. It can be seen that this embodiment will use a hierarchical Transformer model based on the hybrid window mechanism as the backbone network of the encoder and combine the downsampling module to continuously downsample the image to be enhanced to generate a depth feature tensor.
[0069] Since Transformer is good at capturing the dependencies between long-range features and can fully learn the global features of image data, while downsampling convolution has more advantages in capturing local invariant features, this embodiment makes full use of Transformer and downsampling convolution to perform multi-receptive field and multi-level feature learning on the image, which not only improves the feature learning ability of the encoder to learn the deep semantic features of the image, so that the features output by the encoder cover both the detailed information in the original image and filter out the environmental noise information such as weather in the original image to a certain extent, but also reduces the complexity of the model structure and the model calculation complexity.
[0070] Specifically, refer to Figure 2 As shown, the encoder includes a convolutional layer for initially downsampling the image to be enhanced (the convolutional kernel can be set to 7 and the padding stride can be set to 2) and four processing stages (i.e., Stage1 to Stage4). Each stage is composed of 1 Transformer module and 1 downsampling module in series, and different sliding window mechanisms are used in the Transformer modules in different stages. The output of each stage will be used as the input of the next stage; for example, for any input image, its image matrix data can be represented in the form of a tensor as Img∈R 3×H×W , where H and W respectively represent the height and width of the input image, then the output of the i-th Stage can be represented as d represents the channel dimension, d i =2 i , H i =H / 2 i , W i =W / 2 i, where \(i = 1, 2, 3, 4\), FM4 will be passed as the input of the decoder to the decoder for upsampling, and finally an enhanced image is obtained; among them, FM1, FM2, and FM3 are input into the corresponding transposed convolution residual module of the decoder through skip connections.
[0071] Among them, in this embodiment, the scales of the sliding window multi-head self-attention modules in the Transformer modules at different stages are different, so as to efficiently implement a hierarchical Transformer architecture based on the hybrid window mechanism, enabling different sliding window segmentation and fusion strategies to be used in different stages, thereby enhancing the relationship between different sliding windows. Since the hybrid window mechanism has strong generalization ability, it can effectively avoid the high computing power consumption of the global-based Transformer, and at the same time can significantly improve the quality of image enhancement, thus efficiently and accurately generating multi-semantic features of the image. And this hybrid window mechanism can be generalized to be used in multiple classic models, and has high promotion and application value.
[0072] In addition, although the Shifiting window mechanism can effectively enhance the information correlation between different local regions, it will inevitably increase the computational complexity of the model. In this embodiment, the moving window mechanism is only used in one of the sliding window multi-head self-attention modules of the Transformer module in the first stage, so as to obtain the long-range dependencies relationship of features to the greatest extent with a relatively small computational complexity, thereby effectively reducing the computational complexity of the model.
[0073] It can be seen that in this embodiment, the encoder with sliding window mechanisms of different scales in the image enhancement model performs multi-stage deep semantic feature processing and downsampling convolution processing on the image to be enhanced. While outputting the deep feature tensor, the image enhancement effect can be effectively improved, and the structural complexity, computational complexity and computing power resource consumption of the model can be reduced.
[0074] Step S20: Perform upsampling processing on the deep feature tensor to output the enhanced image.
[0075] Exemplarily, in this embodiment, the learned features (i.e., the deep feature tensor) are upsampled by a decoder to restore the image to its original size, and environmental noises such as weather are eliminated in this process, and then an enhanced image with the same spatial scale and channel dimension as the image to be enhanced is output. It can be seen that in this embodiment, the Transformer in the encoder is mixed with downsampling convolution for multi-receptive field and multi-level deep semantic feature learning to improve the feature learning ability of the encoder, and thus can robustly generalize various environmental noises, so that the features output by the encoder not only cover the detailed information in the original image, but also filter the environmental noise information such as weather in the original image to a certain extent, thereby realizing the accurate restoration of images in various harsh weather and other environmental scenarios; and through the mixed window mechanism of different scales and only using the moving window mechanism in the first stage to efficiently implement the hierarchical Transformer architecture, the model structure and computational complexity can be effectively reduced and the high computing power consumption of the global-based Transformer can be avoided, and at the same time, the quality of image enhancement can be significantly improved.
[0076] Further, in one embodiment, the downsampling convolution process is implemented by a two-dimensional convolution with a stride of 2, a convolution kernel of 7, and an output channel dimension that is twice the input channel dimension.
[0077] Exemplarily, in this embodiment, the downsampling convolution process will be implemented by a two-dimensional convolution with a stride of 2, a convolution kernel of 7, and an output channel dimension that is twice the input channel dimension, that is, the downsampling module can be implemented by a two-dimensional convolution filter with a convolution kernel size of 7, a stride of 2, and an output channel dimension that is twice the input channel dimension; it should be noted that the specific values of the above convolution kernel size and stride are only presented in the embodiment, and they can be adjusted larger or smaller according to actual needs. Therefore, for any input feature tensor In ∈ R d×H×W in terms of, the downsampling module will map the tensor dimension output by the Transformer module to while the Transformer module will map In ∈ R d×H×W to a tensor with the same dimension of R d×H×w ; it can be seen that each stage Stage reduces the scale of the input feature tensor to 1 / 4 of the original, and the channel dimension is expanded to twice the original.
[0078] Further, in one embodiment, the multi-stage deep semantic feature processing of the image to be enhanced based on the sliding window mechanism of different scales includes:
[0079] Performing the first-stage deep semantic feature processing on the image to be enhanced based on the sliding window mechanism of the first scale;
[0080] Perform the second-stage deep semantic feature processing on the output result of the first stage based on the sliding window mechanism of the second scale;
[0081] Perform the third-stage deep semantic feature processing on the output result of the second stage based on the sliding window mechanism of the third scale;
[0082] Perform the fourth-stage deep semantic feature processing on the output result of the third stage based on the sliding window mechanism of the first scale;
[0083] Wherein, the third scale is greater than the second scale and the second scale is greater than the first scale.
[0084] Exemplarily, referring to Figure 3 shown, in this embodiment, the deep semantic feature processing of each stage will be completed by a Transformer module, and each Transformer module is composed of two attention units connected in series. Among them, each attention unit is composed of a first-layer regularization module, a sliding window multi-head self-attention module, a second-layer regularization module, and a multi-layer perceptron connected in series in sequence. It should be understood that for any input tensor In ∈ R d×H×W , the sliding window multi-head self-attention module first evenly divides the input tensor into windows of size h×w, and then performs the classical multi-head attention mechanism processing within each window. Among them, the calculation formula and calculation process of the classical multi-head attention mechanism are as follows:
[0085]
[0086] Wherein, Q, K, and V represent three linear transformations (i.e., the query vector Q, the key vector K, and the value vector V), and x represents the input feature map; since the spatial dimension of the input and output tensors of the multi-head attention module does not change, the output tensor dimension of the multi-head attention module within each window is h×w and the channel dimension is d, thus obtaining output tensors of d×h×w; then, by using the dimension transformation operation, an output tensor of dimension d×H×W can be obtained. After the above steps, the sliding window multi-head self-attention module maps the input tensor In ∈ R d×H×W to a tensor with the same dimension of R d×H×W .
[0087] In this embodiment, the sliding window scales of the sliding window multi-head self-attention modules in different stages will be set to different sizes; it should be noted that the specific values of the first scale, the second scale, and the third scale can be determined according to actual needs, as long as the third scale is greater than the second scale and the second scale is greater than the first scale, for example, the first scale h = w = 4, the second scale h = w = 16, and the third scale h = w = 8.
[0088] Therefore, for Stage1, a sliding window with h = w = 4 is used to divide the feature map, and the shifting window mechanism is used in the second sliding window multi-head self-attention module in the Transformer module to maximize the acquisition of long-range dependencies of features with relatively low computational complexity.
[0089] For Stage2, a sliding window with h = w = 16 is used to divide the feature map, and the shifting window mechanism is not used in either of the two sliding window multi-head self-attention modules in this Stage2.
[0090] For Stage3, a sliding window with h = w = 8 is used to divide the feature map, and the shifting window mechanism is not used in either of the two sliding window multi-head self-attention modules in this Stage3, so as to ensure that the features within each window have the same receptive field as the features within the window in Stage2.
[0091] For Stage4, a sliding window with h = w = 4 is used to divide the feature map, and the shifting window mechanism is not used in either of the two sliding window multi-head self-attention modules in this Stage4, so as to ensure that the features within each window have the same receptive field as the features within the window in Stage3.
[0092] It can be seen that the different sliding window designs used in the above 4 Stages constitute a hybrid sliding window mechanism. Based on the architecture shown in Figure 2 this mechanism enables the use of different sliding window mechanisms in the Transformer modules of different stages, and thus an encoder as shown in Figure 4 can be formed. In summary, in this embodiment, the moving window mechanism is only used in Stage1 of the encoder to further enhance the semantic correlation between local regions and reduce the computational complexity and model complexity of the encoder; at the same time, the balance between the feature expression ability and computing power consumption of the encoder is achieved through the controllable receptive field perception of the hybrid sliding window in subsequent Stages.
[0093] Further, in one embodiment, the sliding window mechanism based on the first scale performs the first-stage deep semantic feature processing on the image to be enhanced, including:
[0094] Performing regularization processing on the image to be enhanced based on the sliding window mechanism of the first scale to obtain a first regularization tensor;
[0095] Performing equal-dimensional mapping processing on the first regularization tensor to obtain a feature tensor;
[0096] Performing regularization processing on the image to be enhanced and the feature tensor to obtain a second regularized tensor;
[0097] Perform an equal-dimensional transformation on the second regularized tensor to obtain a target tensor, and repeat the above steps once based on the target tensor to obtain a final target tensor.
[0098] For example, it should be understood that in this embodiment, the deep semantic feature processing at each stage is completed by the attention unit in the Transformer module. The working principle of the attention unit is as follows: Input tensor IT∈R d×h×w First, the first layer regularization module is used to regularize the tensor elements to obtain the first regularized tensor; then the sliding window multi-head self-attention module performs equal-dimensional mapping on the first regularized tensor to obtain the feature tensor IT1∈R d×h×w ; Then use the residual structure to perform residual connection on IT+IT1 to obtain tensor IT2 and use it as the input of the second layer regularization module; the second layer regularization regularizes the tensor elements of IT2 to obtain the second regularized tensor; the second regularized tensor is input into a multi-layer perceptron with equal dimensional transformation to obtain the target tensor IT3∈R d×h×w ; Then use the residual structure to perform residual connection on IT2+IT3 to obtain tensor IT4 and use it as the input of the next attention unit. At this point, the workflow of an attention unit ends, and the above steps are repeated in the next attention unit to obtain the output result IT5 of the Transformer module; since the whole process does not change the dimension of the tensor, the dimension of IT5 is still R d×H×w However, after IT5 is processed by the downsampling module, its dimension becomes
[0099] It should be noted that since the principles and processes of deep semantic feature processing in each stage are similar, for the sake of simplicity of description, the following embodiments will take the deep semantic feature processing process of the first stage as an example for explanation. In this embodiment, the input tensor is first regularized by the first layer regularization module to obtain a first regularized tensor; then the first regularized tensor is subjected to equal-dimensional mapping processing by the sliding window multi-head self-attention module to obtain a feature tensor; then the enhanced image and feature tensor are regularized by the second layer regularization module to obtain a second regularized tensor; then the second regularized tensor is transformed into an equal-dimensional shape by a multi-layer perceptron to obtain a target tensor, and the above steps are repeated once based on the target tensor to obtain the final target tensor.
[0100] Furthermore, in one embodiment, the upsampling of the depth feature tensor includes:
[0101] Perform multiple upsampling processes on the depth feature tensor;
[0102] For each upsampling process, perform deconvolution on the result of the previous upsampling process to obtain a deconvolution tensor;
[0103] Perform a first convolution on the deconvolution tensor to obtain a first convolution tensor;
[0104] Map the first convolution tensor through an activation function to obtain an activation tensor;
[0105] Perform a second convolution on the activation tensor to obtain a second convolution tensor,
[0106] Perform a residual connection on the second convolution tensor, the deconvolution tensor, and the downsampling result corresponding to the current upsampling process scale, and use the residual result and the downsampling result corresponding to the next upsampling process scale as the input for the next upsampling process;
[0107] Among them, during the first upsampling process, the object of the deconvolution process is the depth feature tensor.
[0108] Exemplarily, in this embodiment, a classical deconvolution residual module can be used and a decoder can be constructed in a cascaded form, and multiple upsampling processes of the depth feature tensor can be implemented through this decoder. Among them, the entire decoder can be composed of 4 cascaded deconvolution residual modules for restoring the output features of the encoder to the input image scale. See Figure 5 As shown, each deconvolution residual module includes a deconvolution layer, a first convolution layer, an activation function, a second convolution layer, and a residual block connected in series in sequence. The residual block is used to perform a residual connection on the tensor output by the first convolution layer, the tensor output by the second convolution layer, and the tensor output by the encoder. Since the encoder in this embodiment has excellent generalization ability and outstanding image feature expression ability, even using a general deconvolution residual module as the decoder can achieve an outstanding image enhancement effect.
[0109] It is understandable that, except that the input of the first deconvolutional residual module is the output of the entire encoder, that is, during the first upsampling process, the object of the deconvolutional process is the depth feature tensor, the inputs of the other three deconvolutional residual modules consist of two parts: one is the output of the superior deconvolutional residual module, and the other is the output of the corresponding stage in the encoder; for example, the input of the second deconvolutional residual module in the decoder comes from the output of the first deconvolutional residual module and the output of Stage 3 of the encoder; the input of the third deconvolutional residual module in the decoder comes from the output of the second deconvolutional residual module and the output of Stage 2 of the encoder; the input of the fourth deconvolutional residual module in the decoder comes from the output of the third deconvolutional residual module and the output of Stage 1 of the encoder.
[0110] Since the working principle of each upsampling process is the same, for the sake of simplicity of description, the following embodiments will take the second upsampling process as an example for illustration: For any input tensor In ∈ R d×H×W , the second deconvolutional residual module first maps In to a tensor Out1 with dimensions d×2H×2W through a deconvolution with a stride of 2 and a convolutional kernel of 4; then Out1 is mapped to after convolution with a convolutional kernel of 3, a stride of 1, and an output channel dimension of After Out2 is mapped through the activation function, it is then processed by convolution with a convolutional kernel of 3, a stride of 1, and an output channel dimension of to obtain Finally, Out1 + Out3 + Out en is used as the output of this deconvolutional residual module, where Out en is the output of Stage2 of the encoder.
[0111] In summary, this embodiment proposes a Transformer module based on a hybrid sliding window to form the basic structure of an image encoder to map an image into a deep feature map tensor, and then uses a deconvolutional structure as the decoder to perform dimensional reduction on the image, thereby obtaining an image after removing environmental noises such as climate. It should be noted that when constructing the above image enhancement model with an encoder-decoder architecture, model training is also required. Specifically, the input tensor R i in the training data pair (R i , Gt i ) is input into the model, and after being processed by the encoder-decoder, the output result tensor Out i is obtained. Then, the L1 loss function is used to calculate the difference between Out i and the output tensor Gt iThe loss between them is calculated, and then, according to the SGD (Stochastic Gradient Descent) algorithm, backpropagation is performed on the network to optimize the model parameters. After a certain number of training iterations to obtain an image enhancement model that meets the requirements, the training is stopped.
[0112] It should be noted that the image enhancement processing method provided in this embodiment can not only be used for image enhancement processing in various harsh weather environments, but also for image enhancement processing in other scenarios.
[0113] In a second aspect, an embodiment of the present application also provides an image enhancement device.
[0114] In one embodiment, the image enhancement device includes: an encoder and a decoder;
[0115] The encoder is used to perform multi-stage deep semantic feature processing and downsampling convolution processing on the image to be enhanced based on a sliding window mechanism of different scales, so as to output a deep feature tensor;
[0116] The decoder is used to perform upsampling processing on the deep feature tensor to output the enhanced image;
[0117] Among them, each stage of the encoder includes a cascaded transformer module and a downsampling module. The transformer modules in different stages include sliding window multi-head self-attention modules of different scales. One of the sliding window multi-head self-attention modules in the first stage is constructed based on a moving window mechanism. The deep semantic feature processing and downsampling convolution processing are performed on the image to be enhanced through the transformer module and the downsampling module.
[0118] In this embodiment, the Transformer in the encoder is mixed with downsampling convolution to perform multi-receptive field and multi-level deep semantic feature learning, so as to improve the feature learning ability of the encoder. Furthermore, it can robustly generalize various climate noises, so that the features output by the encoder not only cover the detailed information in the original image, but also filter the weather noise information in the original image to a certain extent, thereby realizing the accurate restoration of images in various harsh weather scenarios. And through the mixed window mechanism of different scales and only using the moving window mechanism in one of the sliding window multi-head self-attention modules in the first stage, a hierarchical Transformer architecture is efficiently implemented, so as to effectively reduce the model structure and computational complexity and avoid the high computing power consumption of the global-based Transformer. At the same time, the quality of image enhancement can be significantly improved.
[0119] Further, in one embodiment, the transformer module includes two attention units connected in series, and each attention unit includes a first layer regularization module, a sliding window multi-head self-attention module, a second layer regularization module, and a multi-layer perceptron connected in sequence;
[0120] The first layer regularization module is used to perform regularization processing on the input tensor to obtain a first regularized tensor;
[0121] The sliding window multi-head self-attention module is used to perform equal-dimensional mapping processing on the first regularized tensor to obtain a feature tensor;
[0122] The second layer regularization module is used to perform regularization processing on the input tensor and the feature tensor to obtain a second regularized tensor;
[0123] The multi-layer perceptron is used to perform equal-dimensional transformation on the second regularized tensor to obtain a target tensor.
[0124] Further, in one embodiment, the sliding window scale of the sliding window multi-head self-attention module in the first stage is the first scale; the sliding window scale of the sliding window multi-head self-attention module in the second stage is the second scale; the sliding window scale of the sliding window multi-head self-attention module in the third stage is the third scale; the sliding window scale of the sliding window multi-head self-attention module in the fourth stage is the first scale; wherein, the third scale is greater than the second scale and the second scale is greater than the first scale.
[0125] Further, in one embodiment, the decoder is used to perform multiple upsampling processes on the depth feature tensor. The decoder includes four transposed convolution residual modules connected in series. Each transposed convolution residual module includes a transposed convolution layer, a first convolution layer, an activation function, a second convolution layer, and a residual block connected in sequence. For each upsampling process, the transposed convolution layer is used to perform transposed convolution processing on the result of the previous upsampling process to obtain a transposed convolution tensor; the first convolution layer is used to perform first convolution processing on the transposed convolution tensor to obtain a first convolution tensor; the activation function is used to map the first convolution tensor to obtain an activation tensor; the second convolution layer is used to perform second convolution processing on the activation tensor to obtain a second convolution tensor; the residual block is used to perform residual connection on the second convolution tensor, the transposed convolution tensor, and the downsampling result corresponding to the scale of the current upsampling process, and use the residual result and the downsampling result corresponding to the scale of the next upsampling process as the input of the next upsampling process; wherein, the object of the transposed convolution processing of the first transposed convolution residual module is the depth feature tensor.
[0126] Further, in one embodiment, the downsampling module includes a two-dimensional convolution with a stride of 2, a convolution kernel of 7, and an output channel dimension that is twice the input channel dimension.
[0127] In summary, the codec structure in this embodiment has a relatively simple model structure and low computing power consumption. At the same time, the encoder has excellent feature learning ability and can robustly generalize various climate noises, and thus can achieve image denoising and enhancement in various harsh weather scenarios, so as to accurately restore images with rain, snow, fog, etc. at the same time, thereby completing the precise restoration of images.
[0128] Among them, the function implementation of each module in the above image enhancement device corresponds to each step in the above image enhancement method embodiment, and its function and implementation process will not be elaborated here one by one.
[0129] In a third aspect, an embodiment of the present application provides an image enhancement device, which can be a device with data processing functions such as a personal computer (PC), a laptop computer, a server, etc.
[0130] Refer to Figure 6 , Figure 6 which is a schematic hardware structure diagram of the image enhancement device involved in the embodiment of the present application. In the embodiment of the present application, the image enhancement device may include a processor, a memory, a communication interface, and a communication bus.
[0131] Among them, the communication bus can be of any type and is used to interconnect the processor, the memory, and the communication interface.
[0132] The communication interface includes input / output (I / O) interfaces, physical interfaces, and logical interfaces, etc., which are used to implement the interconnection of internal devices of the image enhancement device, as well as interfaces for implementing the interconnection between the image enhancement device and other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display screen (Display), a keyboard (Keyboard), etc.
[0133] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0134] The processor can be a general-purpose processor, which can call the image enhancement program stored in the memory and execute the image enhancement method provided by the embodiments of the present application. For example, the general-purpose processor can be a central processing unit (CPU). Among them, the method executed when the image enhancement program is called can refer to the various embodiments of the image enhancement method of the present application, which will not be elaborated here.
[0135] Those skilled in the art can understand that Figure 6 the hardware structure shown in does not constitute a limitation to the present application, and may include more or fewer components than shown, or combine certain components, or arrange different components.
[0136] In a fourth aspect, the embodiments of the present application also provide a computer-readable storage medium.
[0137] An image enhancement program is stored on the readable storage medium of the present application. When the image enhancement program is executed by a processor, the steps of the image enhancement method as described above are implemented.
[0138] Among them, the method implemented when the image enhancement program is executed can refer to the various embodiments of the image enhancement method of the present application, which will not be elaborated here.
[0139] It should be noted that the serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0140] In the description of the specification, claims and the above-mentioned drawings of this application, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. Descriptions such as "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit that "first", "second" and "third" are different types.
[0141] In the description of the embodiments of this application, "exemplary", "for example" or "for instance" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary", "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example" or "for instance" is intended to present relevant concepts in a specific manner.
[0142] In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "a plurality of" means two or more than two.
[0143] In some processes described in the embodiments of this application, there are multiple operations or steps that appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of this application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.
[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to enable a terminal device to execute the methods described in the various embodiments of this application.
[0145] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.
Claims
1. An image enhancement method, characterized in that, The image enhancement method includes: The encoder performs multi-stage deep semantic feature processing and downsampling convolution processing on the image to be enhanced based on the sliding window mechanism of different scales, so as to output a deep feature tensor; The decoder performs upsampling processing on the deep feature tensor to output the enhanced image; Among them, each stage of the encoder includes a cascaded transformer module. The transformer modules in different stages include sliding window multi-head self-attention modules of different scales. One of the sliding window multi-head self-attention modules in the first stage is constructed based on the moving window mechanism, and the sliding window multi-head self-attention modules in other stages are not constructed through the moving window mechanism. The sliding window scale of the sliding window multi-head self-attention module in the first stage is the first scale; the sliding window scale of the sliding window multi-head self-attention module in the second stage is the second scale; the sliding window scale of the sliding window multi-head self-attention module in the third stage is the third scale; the sliding window scale of the sliding window multi-head self-attention module in the fourth stage is the first scale; the third scale is greater than the second scale and the second scale is greater than the first scale.
2. The image enhancement method according to claim 1, characterized in that, The multi-stage deep semantic feature processing of the image to be enhanced based on the sliding window mechanism of different scales includes: Performing the first-stage deep semantic feature processing on the image to be enhanced based on the sliding window mechanism of the first scale; Performing the second-stage deep semantic feature processing on the output result of the first stage based on the sliding window mechanism of the second scale; Performing the third-stage deep semantic feature processing on the output result of the second stage based on the sliding window mechanism of the third scale; Performing the fourth-stage deep semantic feature processing on the output result of the third stage based on the sliding window mechanism of the first scale.
3. The image enhancement method according to claim 2, characterized in that, The performing the first-stage deep semantic feature processing on the image to be enhanced based on the sliding window mechanism of the first scale includes: Performing regularization processing on the image to be enhanced based on the sliding window mechanism of the first scale to obtain a first regularization tensor; Performing equal-dimensional mapping processing on the first regularization tensor to obtain a feature tensor; Performing regularization processing on the image to be enhanced and the feature tensor to obtain a second regularization tensor; Performing equal-dimensional transformation on the second regularization tensor to obtain a target tensor, and repeating the above steps once based on the target tensor to obtain the final target tensor.
4. The image enhancement method according to claim 1, characterized in that, The upsampling processing of the deep feature tensor includes: Performing multiple upsampling processes on the deep feature tensor; For each upsampling process, performing deconvolution processing on the result of the previous upsampling process to obtain a deconvolution tensor; Performing the first convolution processing on the deconvolution tensor to obtain a first convolution tensor; Mapping the first convolution tensor through an activation function to obtain an activation tensor; Performing the second convolution processing on the activation tensor to obtain a second convolution tensor, Performing residual connection on the second convolution tensor, the deconvolution tensor, and the downsampling result corresponding to the current upsampling scale, and using the residual result and the downsampling result corresponding to the next upsampling scale as the input of the next upsampling process; Among them, during the first upsampling process, the object of the transposed convolution process is the depth feature tensor.
5. The image enhancement method according to claim 1, characterized in that: The downsampling convolution process is implemented by a two-dimensional convolution with a stride of 2, a convolution kernel of 7, and an output channel dimension that is twice the input channel dimension.
6. An image enhancement device, characterized in that, The image enhancement device includes: an encoder and a decoder; The encoder is used to perform multi-stage deep semantic feature processing and downsampling convolution processing on the image to be enhanced based on a sliding window mechanism of different scales, so as to output a depth feature tensor; The decoder is used to perform upsampling processing on the depth feature tensor to output the enhanced image; Among them, each stage of the encoder includes a transformer module and a downsampling module connected in series. The transformer modules in different stages include sliding window multi-head self-attention modules of different scales. One of the sliding window multi-head self-attention modules in the first stage is constructed based on the moving window mechanism, and the sliding window multi-head self-attention modules in other stages are not constructed through the moving window mechanism. The deep semantic feature processing and downsampling convolution processing of the image to be enhanced are performed through the transformer module and the downsampling module; Among them, the sliding window scale of the sliding window multi-head self-attention module in the first stage is the first scale; The sliding window scale of the sliding window multi-head self-attention module in the second stage is the second scale; The sliding window scale of the sliding window multi-head self-attention module in the third stage is the third scale; The sliding window scale of the sliding window multi-head self-attention module in the fourth stage is the first scale; Among them, the third scale is greater than the second scale and the second scale is greater than the first scale.
7. The image enhancement device according to claim 6, characterized in that: The transformer module includes two attention units connected in series. Each attention unit includes a first-layer regularization module, a sliding window multi-head self-attention module, a second-layer regularization module, and a multi-layer perceptron connected in sequence; The first-layer regularization module is used to perform regularization processing on the input tensor to obtain a first regularized tensor; The sliding window multi-head self-attention module is used to perform equal-dimension mapping processing on the first regularized tensor to obtain a feature tensor; The second-layer regularization module is used to perform regularization processing on the input tensor and the feature tensor to obtain a second regularized tensor; The multi-layer perceptron is used to perform equal-dimension transformation on the second regularized tensor to obtain a target tensor.
8. An image enhancement device, characterized in that, The image enhancement device includes a processor, a memory, and an image enhancement program stored on the memory and executable by the processor. When the image enhancement program is executed by the processor, the steps of the image enhancement method according to any one of claims 1 to 5 are implemented.
9. A computer-readable storage medium, characterized in that, An image enhancement program is stored on the computer-readable storage medium. When the image enhancement program is executed by a processor, the steps of the image enhancement method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Image denoising method, system and device and storage medium
CN116012266A