A wavelet transform vision enhancement method for robots in low-light environments
By adopting the wavelet transform vision enhancement method in a low-light environment, combined with the detail recovery and noise denoising module, the problem of high calculation overhead of low-light image enhancement and noise amplification will lead to image quality degradation in the prior art, and efficient image enhancement and detail retention effects are achieved.
Patent Information
- Application Number
- CN202510205714.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The existing low-light image enhancement methods have problems such as high computational overhead, noise amplification will lead to image quality degradation and blurred details under complex lighting conditions.
The robot wavelet transform vision enhancement method in low-light environments is adopted, and the details recovery module of space channel attention and the noise estimation denoising module are combined to adaptively enhance and denoising the low-light images. This method improves the brightness and detail performance of the image through steps such as HSV conversion, illumination mapping, multi-layer wavelet transformation, detail recovery and noise denoising.
It effectively reduces computing overhead, improves the image's detail retention and noise suppression capabilities, and enhances the visibility and quality of low-light images.
Smart Images

Figure CN119693244B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of computer vision and image processing, and relates to a low-light image enhancement method, and specifically to a robot wavelet transform vision enhancement method in a low-light environment. Background Art
[0002] Low-light image enhancement is an important technology in the field of image processing. It improves the brightness and detail of low-light images and improves their visibility. This technology is widely used in many fields, such as monitoring systems, medical imaging, intelligent driving, and night scene photography. However, existing low-light image enhancement methods still face many challenges in dealing with complex lighting conditions. In particular, during the brightness enhancement process, image details are easily blurred and noise may be further amplified. Current low-light image enhancement methods mainly face the following problems:
[0003] (1) Traditional low-light image enhancement methods, especially those based on deep learning attention mechanisms or complex convolutional neural network models, often have high computational overhead and are difficult to run in real time on resource-constrained devices. For example, although attention mechanisms such as IGAB (Illumination-Guided Attention Block) can improve the enhancement effect, due to their high computational complexity, they are difficult to meet the requirements of real-time performance and efficiency in practical applications.
[0004] (2) Noise is a common problem in low-light images, especially in images taken at high ISO values. Traditional methods cannot effectively distinguish between real information and noise in the image during brightness enhancement, and often enhance the noise together with the image content, causing the noise to be further amplified, thus destroying the image quality. In particular, due to the lack of a dedicated noise estimation and suppression module, the problem of noise amplification cannot be effectively addressed.
[0005] (3) In low-light images, much high-frequency information (such as edges and textures) becomes blurred or lost due to insufficient lighting. In traditional low-light image enhancement methods, simply increasing the image brightness often causes image details to become blurred or even introduces artifacts. This is because during the enhancement process, the brightness of all areas of the image is uniformly increased, and the details of complex areas cannot be effectively restored. Although methods based on global operations or attention mechanisms (IGAB) can improve the overall brightness of the image, they perform poorly when processing image details and high-frequency information. The attention mechanism tends to focus on global features and tends to ignore local details in the image, such as edges and textures. This neglect causes the image to lose important high-frequency information during the enhancement process, ultimately resulting in blurred image details and unclear edges.
[0006] Therefore, the present invention is just produced based on the above shortcomings. Summary of the invention
[0007] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a wavelet transform vision enhancement method for robots in low-light environments. By combining the detail recovery module of spatial channel attention and the noise estimation and denoising module, an image enhancement method for adaptively enhancing and denoising low-light images is provided to enhance the robot's vision ability in low-light environments.
[0008] The present invention is achieved through the following technical solutions:
[0009] A robot wavelet transform vision enhancement method in a low-light environment, characterized by comprising the following steps:
[0010] Step S1, the acquired low-light image is used as an input image model input, and the brightness is changed by HSV conversion to obtain several roughly enhanced images with different brightness;
[0011] Step S2, the input image and the roughly enhanced image are spliced by channel and then passed through a convolution layer to generate an illumination map. At the same time, the input image and the roughly enhanced image are also input into a noise estimation and denoising module to generate a noise feature map, and are also input into a detail recovery module to extract high-frequency information;
[0012] Step S3, multiplying the input image and the illumination map element by element to generate a preliminary brightness enhanced image;
[0013] Step S4, the preliminary brightness enhanced image is further processed through the convolution layer to extract the feature map, and enters the multi-layer wavelet transform module for wavelet transform, the high-frequency information extracted by the detail recovery module is integrated into the wavelet transform, and the noise estimation and denoising module denoises the feature map after wavelet transform under the guidance of the noise feature map;
[0014] Step S5: The denoised feature map is reassembled into a large-size image by integrating image information layer by layer through up-convolution inverse wavelet transform in the multi-layer wavelet transform module. The reorganized image is denoised by the noise estimation and denoising module, and finally passes through the convolution layer to obtain the final enhanced image.
[0015] The robot wavelet transform vision enhancement method in a low-light environment as described above is characterized in that: the multi-layer wavelet transform module includes an encoder, a bottleneck layer and a decoder, the encoder has a plurality of wavelet blocks for the feature map to pass through in sequence and each of the wavelet blocks performs feature extraction and downsampling through wavelet transform to extract feature information of different scales, the bottleneck layer further integrates deep-level features and passes the processed feature map to the decoder, the decoder has a plurality of upsampling convolutional layers and the inverse wavelet transform in the upsampling convolutional layer integrates the image information layer by layer with the feature map output by the bottleneck layer.
[0016] The wavelet transform vision enhancement method for robots in low-light environments as described above is characterized in that the step of generating noise features by the noise estimation and denoising module in step S2 comprises: S21, the NEM module accepts two images, an input image and a roughly enhanced image, as input;
[0017] S22. Calculate the pixel difference between the original low-light image and the roughly enhanced image, convert the original image and the roughly enhanced image into grayscale images, and calculate the element-by-element absolute difference:
[0018]
[0019] in is the original image grayscale image, is the grayscale image of the enhanced image, is the noise difference map, which represents the noise difference between the roughly enhanced image and the original image;
[0020] S23. After estimating the noise difference map, a mapping relationship of noise distribution is established at different brightness levels. The mapping relationship function between brightness and noise is:
[0021]
[0022] in, Indicates brightness, Indicates that at a specific brightness The expected noise intensity under Indicates the noise intensity at this brightness;
[0023] S24, according to the fitted mapping function , generate a noise feature map for representing the noise distribution at different brightness levels, the relationship function is:
[0024]
[0025] in, represents the final noise feature map, which provides an estimate of the noise at each pixel intensity.
[0026] The wavelet transform vision enhancement method for robots in low-light environments as described above is characterized in that: in the step S2, the detail recovery module receives the input image and the roughly enhanced image, splices them into a feature map by channel, and then inputs them into the SCM. The SCM includes N spatial attention mechanisms and N channel cross-attention mechanisms. A convolutional layer for adjusting the size and dimension of the feature map is provided after the last channel cross-attention mechanism.
[0027] The wavelet transform vision enhancement method for robots in low-light environments as described above is characterized in that the feature map is processed using the spatial attention mechanism, including: inputting the feature map , respectively, the maximum pooling and average pooling are used for the feature map, the pooled feature map is concatenated by channel and then convolution is performed, and finally the attention score map is calculated by the sigmoid function :
[0028]
[0029]
[0030]
[0031] in" " represents the concatenation in the channel dimension, Represents the sigmoid activation function;
[0032] The overall spatial attention is:
[0033]
[0034] in Represents element-wise multiplication.
[0035] The wavelet transform vision enhancement method for robots in low-light environments as described above is characterized in that the channel cross-attention mechanism is used to process feature maps, including: inputting feature maps processed by the spatial attention mechanism and feature maps not processed by the spatial attention mechanism, performing maximum pooling and average pooling, performing a fully connected layer operation on the pooled feature maps, and performing dimension reduction, ReLU activation and dimension increase on the pooled features through a convolutional layer in the fully connected layer, generating two feature tensors for addition, and calculating the sum of the feature tensors as follows:
[0036]
[0037]
[0038] The two feature tensors are added together as the input to the weights module and the cross attention is calculated:
[0039]
[0040] The wavelet transform vision enhancement method for robots in low-light environments as described above is characterized in that: after wavelet transform, the image is decomposed into four sub-bands, wavelet transform DWT:
[0041]
[0042] in It is almost self-contained. is the horizontal detail subband, is the vertical detail subband, It is a diagonal detail subband;
[0043] The noise estimation and denoising module performs denoising operations on the four sub-bands, defined as NEM ( ), the denoising process can be expressed as:
[0044] NEM(DWT(Image))=( )
[0045] in is the denoised subband.
[0046] The wavelet transform vision enhancement method for robots in low-light environments as described above is characterized in that: the bottleneck layer is composed of a dimensionality reduction convolution layer that can reduce the dimension of a feature map with a high number of channels, a nonlinear activation function ReLU for improving the feature expression capability, and a dimensionality increase convolution layer that can restore the feature map reduced in dimension by the dimensionality reduction convolution layer to a feature map with a high number of channels, and the process is:
[0047]
[0048] in is the input feature map, and are the weights and biases of the dimension-reducing convolutional layer, and are the weights and biases of the up-dimensional convolutional layer, is the activation function.
[0049] The wavelet transform vision enhancement method for robots in low-light environments as described above is characterized in that the inverse wavelet transform in the upsampling convolution layer recombines the four sub-bands, and the image size is gradually restored after three upsampling convolution layers. The upsampling convolution layer is:
[0050] Reconstructed Image= NEM(IDWT( )
[0051] Where IDWT stands for inverse wavelet transform.
[0052] The wavelet transform vision enhancement method for robots in low-light environments as described above is characterized in that the roughly enhanced image in step S1 is obtained by converting the original image into the HSV color space and enhancing the brightness of its V channel.
[0053] Compared with the prior art, the present invention has the following advantages:
[0054] 1. The present invention introduces wavelet transform and takes advantage of its simultaneous analysis in time and frequency domains. After extracting high-frequency domain information at multiple scales in the encoder, the image is reconstructed through inverse wavelet transform in the decoder stage. In this process, the high-frequency image information of the detail recovery module is integrated. The inverse wavelet transform operation in the upsampling stage integrates and reconstructs the rich image information and illumination information, thereby improving detail retention, noise suppression and multi-scale information processing capabilities in low-light image enhancement.
[0055] 2. In order to reduce the phenomenon that the inherent noise of the image is amplified during the enhancement process and enhance the noise suppression ability of the wavelet transform, the present invention designs a noise estimation and denoising module. The module estimates the relationship between brightness and noise by making a difference map between the enhanced image and the original image, and then guides the wavelet transform to remove noise. Especially when the noise may be over-amplified while the brightness is enhanced, the noise estimation and denoising module can effectively remove the prominent noise performance, thereby improving the model performance.
[0056] 3. In order to solve the problem of high-frequency detail loss in wavelet transform, the present invention designs a detail supplement module. This module extracts deep feature information from the image and injects it into the decoder stage through a series of alternating structures of spatial attention and channel cross attention, making up for the information loss in the wavelet transform layer, thereby enhancing edge details and abstract features and achieving better image enhancement effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 This is the overall framework diagram of the model;
[0058] Figure 2 for the detail recovery module;
[0059] Figure 3 Schematic diagram of wavelet transform. DETAILED DESCRIPTION
[0060] The present invention will be further described below in conjunction with the accompanying drawings:
[0061] like Figures 1 to 3 As shown, a robot wavelet transform vision enhancement method in a low-light environment is characterized by comprising the following steps:
[0062] Step S1, the acquired low-light image is used as an input image model input, and the brightness is changed by HSV conversion to obtain several roughly enhanced images with different brightness;
[0063] Step S2, the input image and the roughly enhanced image are spliced by channel and then passed through a convolution layer to generate an illumination map. At the same time, the input image and the roughly enhanced image are also input into a noise estimation and denoising module to generate a noise feature map, and are also input into a detail recovery module to extract high-frequency information;
[0064] Step S3, multiplying the input image and the illumination map element by element to generate a preliminary brightness enhanced image;
[0065] Step S4, the preliminary brightness enhanced image is further processed through the convolution layer to extract the feature map, and enters the multi-layer wavelet transform module for wavelet transform, the high-frequency information extracted by the detail recovery module is integrated into the wavelet transform, and the noise estimation and denoising module denoises the feature map after wavelet transform under the guidance of the noise feature map;
[0066] Step S5: The denoised feature map is reassembled into a large-size image by integrating image information layer by layer through up-convolution inverse wavelet transform in the multi-layer wavelet transform module. The reorganized image is denoised by the noise estimation and denoising module, and finally passes through the convolution layer to obtain the final enhanced image.
[0067] like Figure 1 As shown, the entire model can be divided into two stages, including the preprocessing stage and the enhancement stage.
[0068] In the preprocessing stage, the low-light image is first used as the input of the Input Image model, and several coarse enhanced images are obtained through HSV conversion. The coarse enhanced image provides a preliminary brightness improvement, so that the image has a certain lighting improvement effect before enhancement. Then, the coarse enhanced image and the low-light image are spliced by channel and then passed through the convolution layer to generate a light map Light Map, which represents the lighting distribution in different areas of the image. The light map will be used to guide the brightness improvement in the enhancement stage. After that, the input image is multiplied element by element with the light map to generate a brightness enhancement image Lit-upImage, which is used to improve the overall brightness of the image, especially the brightness of the dark area, and provide a more balanced brightness basis for the enhancement stage. On the other hand, the low-light image and the coarse enhanced image are input to the noise estimation and denoising module NEM and the detail recovery module DRM. NEM receives the coarse enhanced image and the original input image, analyzes the difference between them, and estimates the noise distribution in the image. By identifying the noise area, the model can suppress noise in a targeted manner in subsequent processing. DRM extracts high-frequency details and texture information from the input image and the roughly enhanced image. This module ensures that the edges and details of the image are preserved during the subsequent enhancement process.
[0069] In the image enhancement stage, the initial brightness enhanced image is further processed through the convolution layer to extract richer illumination and feature information, and then enters the next multi-layer wavelet transform module, which includes an encoder, a bottleneck layer, and a decoder. In the encoder, the initial feature map passes through multiple wavelet blocks in turn. Each wavelet block is subjected to feature extraction and downsampling through wavelet transform to extract feature information of different scales. At the same time, the high-frequency information extracted by the detail recovery module through multiple SCMs is integrated into the wavelet transform through the sum operation; the feature map further integrates deep features through the bottom bottleneck layer, and the processed feature map is passed to the decoder with multiple upsampling convolution layers; in the decoder, the inverse wavelet transform operation in the upsampling convolution layer will gradually integrate the image information layer by layer, restore the image resolution, and finally generate the final enhanced image with a 3*3 convolution layer. In the wavelet transform in the encoder and the inverse wavelet transform in the decoder, NEM is used for denoising after the operation to obtain better image effects.
[0070] The step of generating noise features by the noise estimation and denoising module in step S2 specifically includes:
[0071] S21, NEM module accepts two images as input: the input image and the roughly enhanced image. In the first step, the NEM module accepts two images as input: the original image , roughly enhance the image , where the roughly enhanced image is obtained by converting the original image into the HSV color space and enhancing the brightness of its V channel.
[0072] S22, calculate the pixel difference between the original low-light image and the roughly enhanced image, convert the original image and the roughly enhanced image into grayscale images, and calculate the element-by-element absolute difference. The second step is to extract noise information. Under low-light conditions, the difference between the original image and the enhanced image mainly comes from noise. The original image and the roughly enhanced image are converted into grayscale images and the element-by-element absolute difference is calculated:
[0073]
[0074] in is the original image grayscale image, is the grayscale image of the enhanced image, is the noise difference map, which represents the noise difference between the roughly enhanced image and the original image.
[0075] S23. After estimating the noise difference map, it is necessary to establish a mapping relationship between noise distribution at different brightness levels. A coordinate graph can be constructed with the x-axis representing the brightness values of different pixels in the original image and the y-axis representing the noise value at the corresponding brightness level (obtained through the noise difference map) to count the noise values at different brightness levels, so as to construct a mapping relationship function between brightness and noise. ,in Indicates brightness, represents the noise intensity at this brightness, and the formula can be expressed as:
[0076]
[0077] in, Indicates that at a specific brightness Through this mapping relationship, the noise level at different brightness can be predicted.
[0078] Generally speaking, when using HSV to increase the brightness of low-light images, the noise value in low-light scenes will be amplified. The higher the brightness, the higher the noise value. The data distribution is nonlinear, and the linear regression model is obviously not suitable. Although the random forest model can fit the data nonlinearly, the curve fluctuates too much, and it is easy to lose detail information while removing noise. Therefore, this application chooses the gradient boosting model as the noise estimation function. It can better capture the changing trend of noise in both low-brightness and high-brightness areas, thereby more accurately predicting noise.
[0079] S24, according to the fitted mapping function , a noise signature can be generated , used to represent the noise distribution at different brightness levels. This noise feature map can be used as a guide for subsequent denoising and enhancement processing. The relationship function is:
[0080]
[0081] in, represents the final noise feature map, which provides an estimate of the noise at each pixel intensity.
[0082] like Figure 2 As shown, in order to compensate for the loss of image edge information caused by wavelet transform in the illumination enhancement stage, the final image details are affected. The present invention designs a detail recovery module DRM: the module receives the coarse enhanced image and the original image as input, splices them into feature maps by channel, and then inputs them into SCM. SCM is composed of N spatial attention SAs and channel cross attention CCAs. There is a convolution layer after the last channel cross attention mechanism. The last convolution layer is used to adjust the size and dimension of the feature map to facilitate splicing with the feature maps of each level of the encoder. The explanation of SA and CCA of SCM is as follows:
[0083] Spatial Attention SA: Input Feature Map , respectively, the maximum pooling and average pooling are used for the feature map, the pooled feature map is concatenated by channel and then convolution is performed, and finally the attention score map is calculated by the sigmoid function :
[0084]
[0085]
[0086]
[0087] in" " represents the concatenation in the channel dimension, Represents the sigmoid activation function.
[0088] Depend on And the attention score map Multiply to get the attention feature map :
[0089]
[0090] in Represents element-wise multiplication.
[0091] The overall spatial attention (SA) can be expressed as:
[0092]
[0093] The spatial attention mechanism captures global spatial information by performing maximum pooling and average pooling operations on the input feature map, and fuses this information through convolution operations to generate an enhanced spatial feature map. The spatial attention mechanism can highlight important areas in the image, suppress unimportant areas, improve the spatial feature expression ability of the feature map, and help capture detailed information such as edges and textures.
[0094] Channel Cross Attention CCA: There are two feature map inputs, which are the features after SA Figure 1 (Feature_1) and features that have not undergone SA Figure 2 (Feature_2), for the input feature Figure 1 and Features Figure 2 , perform maximum pooling and average pooling, perform a fully connected layer (FC) operation on the pooled feature map, and use the convolution layer in the fully connected layer to reduce the dimension of the pooled features, activate ReLU, and increase the dimension. Generate two feature tensors and add them as the input of the weights module and calculate the cross attention. The formula for calculating the sum of feature tensors is:
[0095]
[0096]
[0097] In the weights module, we first increase the dimensions of the two features to facilitate element-wise multiplication. Each element of the resulting matrix represents the feature Figure 1 A feature and characteristic of Figure 2 The interaction strength of a feature in , and then sum the last dimension of the interaction matrix, at this time each element represents the feature Figure 1 A feature and characteristic of Figure 2 The sum of the interaction strengths of all features in can be regarded as "interactive attention". This process can be expressed as:
[0098]
[0099] The weights module interactively calculates the information features of two different input sources. This module can effectively combine two sets of feature maps to enhance the model's ability to capture key information.
[0100] The channel cross attention mechanism captures global channel information by performing maximum pooling and average pooling operations on each channel of the input feature map, and fuses this information through a fully connected layer and a weighted sum operation to generate an enhanced channel feature map. The channel cross attention mechanism can highlight important channels in the image, suppress unimportant channels, improve the channel feature expression ability of the feature map, and help capture global features such as color and contour.
[0101] By alternating the use of two attention mechanisms in the information supplementation module, the model can capture key information at different scales. Spatial attention can process local details, while channel attention can process global information. This combination enables the model to obtain effective features at different scales, extract and fuse important information in feature maps more effectively, dynamically adjust the importance of different features during feature fusion, and enhance the overall feature expression ability of the network.
[0102] like Figure 3 As shown, the wavelet transform is performed in the wavelet convolution block in the encoder, and the wavelet transform on the image is a two-dimensional discrete wavelet transform DWT, using the Haar wavelet transform. Assume that the size of the input image is , in and are the height and width of the image, is the number of channels (usually 3, indicating RGB channels). Through wavelet transform, the image is decomposed into four sub-bands, the size of each sub-band is halved, and the number of channels becomes four times the original. Wavelet transform DWT:
[0103]
[0104] in, It is almost self-contained. is the horizontal detail subband, is the vertical detail subband, are diagonal detail subbands. The size of each subband is , after wavelet transformation, the total size of the combined result is: .
[0105] Then the noise estimation and denoising module NEM performs denoising operations on the four sub-bands, which is defined as , then the denoising process can be expressed as:
[0106] NEM(DWT(Image))=( )
[0107] in, is the denoised subband. The size is kept as This is an operation after a wavelet transform layer.
[0108] After three wavelet transform layers, the input is sent to the bottleneck layer, which consists of a dimensionality reduction convolution layer, a nonlinear activation function ReLU, and a dimensionality increase convolution layer. Its input is the output result of the wavelet transform layer. The feature map with a high number of channels is reduced in dimension through the dimensionality reduction convolution layer. The bottleneck layer effectively reduces the computational complexity and the number of parameters. The ReLU activation function introduces nonlinear transformation and improves the feature expression ability. Finally, the dimensionality increase convolution layer is used to restore the reduced dimensionality feature map to a high number of channels to maintain the richness of the features. The output of the bottleneck layer is used as the input of the decoder to gradually reconstruct the high-definition image. The process can be expressed as:
[0109]
[0110] in is the input feature map, and are the weights and biases of the dimension-reducing convolutional layer, and are the weights and biases of the up-dimensional convolutional layer, is the activation function.
[0111] In the upsampling convolution layer upconv of the decoder, it is an inverse small transform, such as Figure 3 The inverse wavelet transform in the lower part recombines the four sub-bands into a large-size image, and then the enhanced image is obtained through denoising. The upsampling convolution layer upconv can be expressed as:
[0112] Reconstructed Image= NEM(IDWT( )
[0113] Where IDWT stands for inverse wavelet transform.
[0114] After three upsampling convolution layers upconv, the image size is gradually restored, and finally the final enhanced image Output is obtained after the convolution layer.
[0115] In general, the present invention is a neural network model based on wavelet transform. Through deep learning technology training, the original low-light image is transformed through HSV transform, and a roughly enhanced image of different brightness is generated by changing the value of the brightness value. On the one hand, the original low-light image and the roughly enhanced image are used as inputs of the noise estimation module and the detail supplement module. On the other hand, after channel splicing, the illuminated features and lighting mapping are obtained through a large-kernel convolution layer. The illuminated feature map is input into the wavelet transform layer for processing to obtain the final enhanced image, which can output the low-light image as a normal-light image, so that the colors and details in the dark are more in line with people's visual effects.
Claims
1. A robot wavelet transform vision enhancement method in a low-light environment, characterized in that: The following steps are involved: Step S1, the acquired low-light image is used as an input image model input, and the brightness is changed by HSV conversion to obtain several roughly enhanced images with different brightness; Step S2, the input image and the roughly enhanced image are spliced by channel and then passed through a convolution layer to generate an illumination map. At the same time, the input image and the roughly enhanced image are also input into a noise estimation and denoising module to generate a noise feature map, and are also input into a detail recovery module to extract high-frequency information; The step of generating noise features by the noise estimation and denoising module in step S2 comprises: S21, NEM module accepts two images, input image and coarse enhanced image, as input; S22. Calculate the pixel difference between the original low-light image and the roughly enhanced image, convert the original image and the roughly enhanced image into grayscale images, and calculate the element-by-element absolute difference: ; in is the original image grayscale image, is the grayscale image of the enhanced image, is the noise difference map, which represents the noise difference between the roughly enhanced image and the original image; S23. After estimating the noise difference map, a mapping relationship of noise distribution is established at different brightness levels. The mapping relationship function between brightness and noise is: ; in, Indicates brightness, Indicates that at a specific brightness The expected noise intensity under Indicates the noise intensity at this brightness; S24, according to the fitted mapping function , generate a noise feature map for representing the noise distribution at different brightness levels, the relationship function is: ; in, represents the final noise feature map, which provides an estimate of the noise at each pixel brightness; Step S3, multiplying the input image and the illumination map element by element to generate a preliminary brightness enhanced image; Step S4, the preliminary brightness enhanced image is further processed through the convolution layer to extract the feature map, and enters the multi-layer wavelet transform module for wavelet transform, the high-frequency information extracted by the detail recovery module is integrated into the wavelet transform, and the noise estimation and denoising module denoises the feature map after wavelet transform under the guidance of the noise feature map; Step S5: The denoised feature map is reassembled into a large-size image by integrating image information layer by layer through up-convolution inverse wavelet transform in the multi-layer wavelet transform module. The reorganized image is denoised by the noise estimation and denoising module, and finally passes through the convolution layer to obtain the final enhanced image.
2. The method for enhancing robot vision through wavelet transform in a low-light environment according to claim 1, characterized in that: The multi-layer wavelet transform module includes an encoder, a bottleneck layer and a decoder. The encoder has multiple wavelet blocks for feature maps to pass through in sequence, and each wavelet block performs feature extraction and downsampling through wavelet transform to extract feature information of different scales. The bottleneck layer further integrates deep-level features and passes the processed feature maps into the decoder. The decoder has multiple upsampling convolutional layers, and the inverse wavelet transform in the upsampling convolutional layer integrates image information layer by layer with the feature maps output by the bottleneck layer.
3. The method for enhancing robot vision through wavelet transform in a low-light environment according to claim 1, characterized in that: In step S2, the detail recovery module receives the input image and the coarse enhanced image, splices them into a feature map by channel, and then inputs it into the SCM. The SCM includes N spatial attention mechanisms and N channel cross-attention mechanisms. After the last channel cross-attention mechanism, a convolution layer for adjusting the size and dimension of the feature map is provided.
4. The method for enhancing robot vision through wavelet transform in a low-light environment according to claim 3, characterized in that: Using the spatial attention mechanism to process the feature map, including: input feature map , respectively, the maximum pooling and average pooling are used for the feature map, the pooled feature map is concatenated by channel and then convolution is performed, and finally the attention score map is calculated by the sigmoid function : ; ; ; in" " represents the concatenation in the channel dimension, Represents the sigmoid activation function; The overall spatial attention is: ; in Represents element-wise multiplication.
5. The method for enhancing robot vision through wavelet transform in a low-light environment according to claim 3, characterized in that: The channel cross attention mechanism is used to process the feature map, including: inputting the feature map processed by the spatial attention mechanism and the feature map not processed by the spatial attention mechanism, performing maximum pooling and average pooling, performing a fully connected layer operation on the pooled feature map, and performing dimension reduction, ReLU activation and dimension increase on the pooled features through the convolution layer in the fully connected layer, generating two feature tensors for addition, and calculating the formula for the sum of the feature tensors is: ; ; The two feature tensors are added together as the input to the weights module and the cross attention is calculated: 。 6. The wavelet transform vision enhancement method for robots in low light environments according to claim 2, characterized in that: After wavelet transform, the image is decomposed into four sub-bands, wavelet transform DWT: ; in It is almost self-contained. is the horizontal detail subband, is the vertical detail subband, It is a diagonal detail subband; The noise estimation and denoising module performs denoising operations on the four sub-bands, defined as NEM ( ), the denoising process can be expressed as: NOT(DWT(Image))=( ); in is the denoised subband.
7. The method for enhancing robot vision through wavelet transform in a low-light environment according to claim 6, characterized in that: The bottleneck layer is composed of a dimensionality reduction convolution layer that can reduce the dimension of a feature map with a high number of channels, a nonlinear activation function ReLU for improving the feature expression capability, and a dimensionality increase convolution layer that can restore the feature map reduced by the dimensionality reduction convolution layer to a feature map with a high number of channels. The process is: ; in is the input feature map, and are the weights and biases of the dimension-reducing convolutional layer, and are the weights and biases of the up-dimensional convolutional layer, is the activation function.
8. The method for enhancing robot vision through wavelet transform in a low-light environment according to claim 6, characterized in that: The inverse wavelet transform in the upsampling convolution layer recombines the four sub-bands, and the image size is gradually restored after three upsampling convolution layers. The upsampling convolution layer is: Reconstructed Image= NEM(IDWT ( )); Where IDWT stands for inverse wavelet transform.
9. The method for enhancing robot vision through wavelet transform in a low-light environment according to claim 1, characterized in that: The roughly enhanced image in step S1 is obtained by converting the original image into the HSV color space and enhancing the brightness of its V channel.
Citation Information
Patent Citations
Maritime image enhancement method in low-illumination environment
CN111489303A
Low-illumination image enhancement method based on curve wavelet attention and Fourier
CN118822908A