Low-light image enhancement method and system based on space-frequency domain characteristic resolution self-adjustment
The low-light image enhancement method, which uses an encoder-decoder architecture and a ResNet-Transformer hybrid module, solves the problems of noise sensitivity and high computational complexity in existing technologies, and achieves efficient enhancement and detail restoration of low-light images. It is suitable for scenarios such as underground coal mines and smartphones.
Patent Information
- Application Number
- CN202511182809.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing low-light image enhancement technologies are sensitive to noise, lack cross-scene generalization capabilities, have high computational complexity, and have limited detail recovery capabilities, making it difficult to perform well in practical applications.
An encoder-decoder architecture is adopted, combined with a hybrid module of ResNet and Transformer, to perform convolution operations on the upper and lower half features of the image respectively. The feature resolution is self-adjusted through the spatial enhancement module and the frequency enhancement module. The multi-head self-attention mechanism and fast Fourier transform are used to optimize the loss functions in the spatial domain and frequency domain to achieve multi-scale feature learning and recovery of the image.
It significantly improves image quality in low-light environments, improves computational efficiency, and enhances the color fidelity and structural integrity of images. It is suitable for security sign recognition in extreme low-light environments and low-light photography for smartphones, while reducing hardware costs.
Smart Images

Figure CN120672637A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and image processing, and in particular relates to a low-light image enhancement method and system based on self-adjustment of spatial-frequency domain feature resolution. Background Art
[0002] Low-light image enhancement, a key research area in computer vision and image processing, addresses image quality degradation in low-light conditions, including noise, detail loss, and color distortion. This technology not only significantly improves human visual perception but also provides high-quality input data for subsequent computer vision tasks such as object detection and face recognition. However, achieving high-quality image enhancement in low-light environments remains a significant challenge due to complex lighting conditions and sensor noise.
[0003] Traditional image enhancement methods primarily employ techniques such as histogram equalization and gamma correction based on image statistics, or decomposition methods based on Retinex theory. These methods achieve enhancement by adjusting the brightness distribution or assuming that images can be decomposed into illumination and reflectance components. However, their core assumptions often ignore the degradation factors present in real low-light images, resulting in enhanced results prone to problems such as noise amplification and color distortion, and limited detail recovery, particularly in complex scenes. With the development of deep learning technology, enhancement methods based on convolutional neural networks and Transformers have made significant progress. These methods utilize end-to-end network architectures, either learning illumination and reflectance components in conjunction with Retinex theory, or directly mapping low-light images to normal-light images. However, existing deep learning methods still have significant shortcomings: high sensitivity to noise and artifacts, insufficient cross-scene generalization, computational complexity that restricts real-time applications, and a need to balance degradation modeling with detail recovery. These limitations severely impact the performance of low-light enhancement techniques in practical applications, necessitating the development of more effective solutions.
[0004] The invention patent, "Neural Network for Raw Low-Light Image Enhancement" (publication number CN115004220B, publication date 20240820), proposes using a U-net convolutional neural network to downsample and upsample image restoration, directly generating full-color RGB images from the monochrome pixels of the image sensor's Bayer array and achieving low-light enhancement. However, this method suffers from inductive bias, making it difficult to effectively restore the illumination of certain images. Its high computational complexity makes it difficult to meet demanding real-time requirements. Furthermore, its generalization capabilities for sensor noise patterns require further verification.
[0005] Invention patent: A low-light image enhancement method based on dual-domain perception of frequency and spatial domains (publication number CN118674628A, publication date 20240920) proposes the use of U-net, a low-light image enhancement method based on dual-domain perception of frequency and spatial domains. However, the U-net method requires multiple calls to the dual-domain restoration module, which greatly increases the computing power requirements. In addition, the spatial enhancement module cannot restore complex scenes using only a simple convolutional network.
[0006] The invention patent, "A Transformer-Based Method for Underground Low-Light Image Enhancement" (publication number CN116152117A, publication date 20230721), provides a Transformer-based method for underground low-light image enhancement. This method utilizes a dual-branch MobileViT module to predict a multiplication map and an addition map, respectively, and combines it with a Cross Attention module to generate a 3×3 color matrix and parameters. Ultimately, the input image and the predicted results are fused to achieve brightness enhancement, color preservation, and detail preservation. While this method has some effectiveness in underground low-light scenes, it typically requires downsampling the image when calculating self-attention, which can result in the loss of important information. Furthermore, existing methods primarily focus on the spatial domain and neglect the use of spectral information in the frequency domain.
[0007] Invention patent: A low-light image enhancement method and device based on a Fourier transform Transformer network (publication number CN117808710A, publication date 20250321) provides a low-light image enhancement method based on a Fourier transform Transformer network. By feeding the frequency of the low-light image into the Transformer network for training, this method relies on jump connections between multiple Transformer layers, making the network structure and parameters very complex. It also only considers the frequency characteristics of the image and ignores the spatial characteristics of the image. Therefore, a method that can simultaneously utilize spatial and frequency domain information and effectively capture key areas is urgently needed to further improve the performance and robustness of low-light image enhancement. Summary of the Invention
[0008] To overcome the problems existing in the related art, the embodiments disclosed in the present invention provide a low-light image enhancement method and system based on self-adjustment of spatial-frequency domain feature resolution. This method can adaptively fuse spatial domain and frequency domain features at different resolution levels, significantly improving image quality in low-light environments.
[0009] The technical solution is as follows: a low-light image enhancement method based on self-adjustment of spatial-frequency domain feature resolution, comprising the following steps: S1 adopts an encoder-decoder architecture and uses a hybrid module of ResNet and Transformer to perform convolution operations on the upper and lower features of the image respectively, encodes the input dark and light images, and completes multi-scale feature learning; S2, flattens the multi-scale features output by the encoder and applies the multi-head self-attention mechanism of the spatial enhancement module to obtain the spatial enhancement features of the image; S3: Input the spatial enhancement features into the convolution layer of the frequency enhancement module to extract high-frequency features. The extracted high-frequency features and spatial enhancement features are used to calculate the image frequency and feature frequency through fast Fourier transform to complete the extraction and enhancement of edge information. S4, features of the output through the decoder Decode, restore image resolution, restore low-light images to normal images, and optimize the network by combining spatial and frequency domain losses.
[0010] Step S1 specifically includes: S101, feed the input image into a 3×3 convolutional network to extract image features F; S102, divide the image feature F into and Two parts, Input to a network consisting of 4 residual blocks, each of which contains two 3×3 convolutional layers and ReLU activation function, The convolution with a step size of 2 is used to perform 1 / 2 downsampling, which is then input to the Transformer module to capture long-range dependencies. The resolution is then restored by upsampling with bilinear interpolation, and the features are output. and Splicing along the channel dimension to obtain fusion features ; S103, the fusion features Put it into ResBlock to get 1 / 2 feature scale ,Will Put in ResBlock to get 1 / 4 of the feature scale .
[0011] Step S2 specifically includes: S201, the obtained characteristic scale , , , , flattened into a 2D feature matrix through a linear neural network, and the number of channels expanded through a linear layer; S202, apply multi-head self-attention to the features obtained in step S201 Mechanism, through residual connection and layer normalization Keep the original information, the expression is: ; Where, For normalization, is the feature processed by the attention mechanism, is the feature after linear mapping, It is processed by the multi-head self-attention mechanism. is the feature after linear mapping of dimension 1, is the feature after linear mapping of dimension 2, It is the feature after linear mapping of dimension 3; S203, input the features into the feedforward network , contains two linear layers with ReLU activation function and output through residual connection; S204, reshape the features back to the original spatial dimensions , and obtain spatial enhancement features , the expression is: ; Where, is the linear mapping layer, is the rectified linear unit activation function, are the height, width, and number of feature channels of the original feature map, respectively.
[0012] Step S3 specifically includes: S301, enhance features in output space Apply cascaded 3×3 and 1×1 depthwise separable convolutions to generate high-frequency features ; S302, high frequency features and spatial enhancement features The dual-domain enhanced features are output by fast Fourier transform addition. .
[0013] Step S4 specifically includes: S401, strengthen the dual domain feature The upsampling is achieved by bilinear interpolation and concatenated with the encoder features of the corresponding scale. After concatenation, the features are processed by ResBlock, which contains two 3×33×3 convolutional layers and residual connections. The number of output channels remains consistent with the encoder scale. The image resolution is gradually restored through ResBlock, and the enhanced image is finally output. ; S402, constraining spatial features through L1 absolute value loss function; S403 , performing fast Fourier transform on the original image and the predicted image, and constraining the frequency features through the L1 absolute value loss function.
[0014] In step S402, the spatial features are constrained by the L1 absolute value loss function, which is expressed as: ; Where, is the spatial feature loss, is the absolute value operation, To enhance the image pixel matrix, is the original image pixel matrix.
[0015] In step S403, the original image and the predicted image are subjected to fast Fourier transform, and the frequency feature is constrained by the L1 absolute value loss function, which is expressed as: ; Where, is the frequency characteristic loss, is the Fast Fourier Transform.
[0016] Another object of the present invention is to provide a low-light image enhancement system based on self-adjustment of spatial-frequency domain feature resolution, which implements the low-light image enhancement method based on self-adjustment of spatial-frequency domain feature resolution. The system comprises: The feature resolution self-adjuster includes an encoder and a decoder. The encoder uses a hybrid module of ResNet and Transformer to perform convolution operations on the upper and lower features of the image, encodes the input dark and light images, and completes multi-scale feature learning; the decoder performs convolution operations on the output features. Decode, restore image resolution, restore low-light images to normal images, and optimize the network by combining spatial and frequency domain losses; The spatial enhancement module flattens the input feature map into a two-dimensional matrix, expands the number of channels through a linear layer, applies a multi-head self-attention mechanism combined with residual connections, refines features through layer normalization and a feedforward network, reshapes the processed features back to the original spatial dimension, and repairs them through a frequency enhancement module; The frequency enhancement module applies deep convolution to the spatially enhanced features to generate high-frequency features, and generates the final output through element-wise multiplication and residual connection.
[0017] Furthermore, the system is applied in coal and rock scenes to process noise judgment in complex interference scenes of underground dust obstruction and restore high-frequency edge information of safety signs.
[0018] Furthermore, the system is carried on a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can realize the functions of the above-mentioned low-light image enhancement system based on self-adjustment of spatial-frequency domain feature resolution.
[0019] In combination with all the above technical solutions, the beneficial effects of the present invention are as follows: First, compared with the existing technology that directly uses Transformer or convolutional layer to extract features, the present invention designs a feature resolution self-adjuster to replace simple feature splicing, and dynamically screens multi-scale features through channel-space dual attention. The feature resolution self-adjuster fuses Resblock and Transformer blocks together and extracts features through splitting and downsampling operations, which not only provides the ability of long-distance pixel association but also maintains low complexity to simplify calculations. The amount of calculation is greatly reduced without losing accuracy. For the spatial enhancement module, compared with directly repairing the image from feature convolution,
[0020] Second, the present invention further flattens the multi-scale feature map into a two-dimensional matrix, and performs image restoration from features at different scales through a linear layer and a feedforward neural network, further maintaining the integrity of the receptive field. In addition, after extracting important spatial features, the present invention performs frequency restoration on important areas, effectively solving the common problems of over-enhancement of dark areas and loss of details in bright areas, and performs better in terms of color fidelity and structural integrity in low-light scenes.
[0021] Third, as a low-light image enhancement technology based on self-adjustment of spatial-frequency domain feature resolution, this invention's creative value is reflected in multiple important dimensions. In terms of technological transformation and commercial value, this solution has been successfully applied to practical scenarios such as underground coal mine safety monitoring and smartphone low-light photography, significantly improving safety sign recognition accuracy and image signal-to-noise ratio performance in extreme low-light environments. According to third-party market analysis data, this technology reduces hardware costs by optimizing algorithm efficiency.
[0022] Fourth, this invention achieves adaptive adjustment of spatial-frequency domain feature resolution for the first time, overcoming the inherent shortcomings of traditional methods in spatial and frequency domain processing. Through innovative network architecture design, it achieves breakthroughs in key performance indicators. Its objective evaluation metrics and real-time processing speed on standard test datasets surpass those of existing mainstream methods, demonstrating significant technical advantages.
[0023] Fifth, this invention solves the long-standing technical challenge of "overexposure in dark areas and distortion in bright areas" in low-light image enhancement. Addressing the limitations of traditional Retinex theory in processing complex scenes, this method achieves high-quality restoration of image details under non-uniform lighting conditions through the coordinated optimization of spatial attention and frequency-domain convolution, achieving significant performance improvements on professional test datasets.
[0024] Sixth, to address the real-time processing limitations of the Transformer architecture, we optimized computational efficiency through a feature splitting strategy. We also innovatively simplified the frequency domain processing flow, combining lightweight depthwise convolution with precise frequency domain loss constraints to significantly reduce computational resource consumption while maintaining enhanced results. These technological innovations provide new technical routes and practical solutions for the development of low-light image processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure; Figure 1 This is a training flow chart of a low-light image enhancement method based on self-adjustment of spatial-frequency domain feature resolution provided by an embodiment of the present invention; Figure 2 This is a flow chart of a low-light image enhancement method based on self-adjustment of spatial-frequency domain feature resolution provided by an embodiment of the present invention; Figure 3 1 is a network structure diagram of a feature resolution self-adjusting encoder provided by an embodiment of the present invention; Figure 4 This is a network structure diagram of a spatial enhancement module provided by an embodiment of the present invention; Figure 5 This is a network structure diagram of a frequency enhancement module provided by an embodiment of the present invention; Figure 6 This is a network structure diagram of a feature resolution self-adjusting decoder provided by an embodiment of the present invention; Figure 7 This is a low-light enhanced image of an indoor building provided by an embodiment of the present invention; wherein (a) is the input image, and (b) is the enhanced image; Figure 8 This is a low-light enhanced image of a table tennis table provided by an embodiment of the present invention; wherein (a) is the input image, and (b) is the enhanced image; Figure 9 This is a low-light enhanced image of a swimming pool provided by an embodiment of the present invention; wherein (a) is the input image, and (b) is the enhanced image. DETAILED DESCRIPTION
[0026] To make the above-mentioned objects, features, and advantages of the present invention more readily apparent, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. The following description sets forth numerous specific details to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art may make similar modifications without departing from the scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0027] The innovation of this invention lies in: it proposes a self-adjusting mechanism for spatial-frequency domain feature resolution, achieves multi-scale feature adaptive fusion by fusing a hybrid encoder of ResNet and Transformer, and significantly improves low-light image enhancement by combining dual-domain collaborative optimization of spatial-domain attention enhancement and frequency-domain deep convolution. Compared with existing technologies, its breakthroughs are reflected in: the first feature splitting downsampling strategy, which divides the input features into spatial / semantic branches for parallel processing. While retaining high-resolution spatial information in 50% of the channels, the other half of the channels capture long-range dependencies through downsampling Transformers, balancing computational efficiency and feature integrity; the design of a cascaded spatial-frequency enhancement module. The spatial enhancement module locates degraded areas through Transformer multi-head attention, and the frequency enhancement module uses lightweight deep separable convolution to enhance edge high-frequency information, avoiding the computational overhead of traditional FFT frequency domain transformation.
[0028] In Example 1, the present invention discloses a low-light image enhancement method based on spatial-frequency domain feature resolution self-adjustment. By jointly utilizing spatial and frequency domain information and introducing a feature resolution self-adjuster, the enhancement effect of low-light images is significantly improved. This method solves the problem that traditional methods rely solely on spatial domain brightness adjustment and cannot effectively handle noise, artifacts, and chromatic aberration. Figure 1 , the embodiment of the present invention provides a training process for a low-light image enhancement method based on self-adjustment of spatial-frequency domain feature resolution.
[0029] For example, Figure 2 As shown, the embodiment of the present invention discloses a low-light image enhancement method based on self-adjustment of spatial-frequency domain feature resolution. It proposes a feature splitting and splicing process, a network architecture based on an encoder-decoder framework for self-adjustment of spatial-frequency domain feature resolution, and a complete processing flow from low-light image input to normal-light image output. Specifically, it includes the following steps: S1 adopts an encoder-decoder architecture and uses a hybrid module of ResNet and Transformer to perform convolution operations on the upper and lower features of the image respectively, encodes the input dark and light images, and completes multi-scale feature learning; like Figure 3 This is a network structure diagram of a feature resolution self-adjusting encoder provided by an embodiment of the present invention. The specific process is as follows: S101, feed the input image into a 3×3 convolutional network to extract image features F; S102, divide the image feature F into and Two parts, Input to a network consisting of 4 residual blocks, each of which contains two 3×3 convolutional layers and ReLU activation function, The convolution with a step size of 2 is used to perform 1 / 2 downsampling, which is then input to the Transformer module to capture long-range dependencies. The resolution is then restored by upsampling with bilinear interpolation, and the features are output. and Splicing along the channel dimension to obtain fusion features ; S103, the fusion features Put it into ResBlock to get 1 / 2 feature scale ,Will Put in ResBlock to get 1 / 4 of the feature scale .
[0030] S2, flattens the multi-scale features output by the encoder and applies the multi-head self-attention mechanism of the spatial enhancement module to obtain the spatial enhancement features of the image; Figure 4 This is a network structure diagram of the spatial enhancement module provided by an embodiment of the present invention. The specific process is as follows: S201, the obtained characteristic scale , , , , flattened into a 2D feature matrix through a linear neural network, and the number of channels expanded through a linear layer; S202, apply multi-head self-attention to the features obtained in step S201 Mechanism, through residual connection and layer normalization Keep the original information, the expression is: ; Where, For normalization, is the feature processed by the attention mechanism, is the feature after linear mapping, It is processed by the multi-head self-attention mechanism. is the feature after linear mapping of dimension 1, is the feature after linear mapping of dimension 2, It is the feature after linear mapping of dimension 3; S203, input the features into the feedforward network , contains two linear layers with ReLU activation function and output through residual connection; S204, reshape the features back to the original spatial dimensions , and obtain spatial enhancement features , the expression is: ; Where, is the linear mapping layer, is the rectified linear unit activation function, are the height, width, and number of feature channels of the original feature map, respectively.
[0031] S3, the spatial enhancement features are input into the convolution layer of the frequency enhancement module to extract high-frequency features. The extracted high-frequency features and spatial enhancement features are used to calculate the image frequency and feature frequency through fast Fourier transform to complete the extraction and enhancement of edge information; The two features are respectively subjected to fast Fourier transform (FFT) to image (size ), the formula is as follows: ; Where, is the pixel value in the spatial domain, is a complex value in the frequency domain, are the image height and width respectively, is the spatial coordinate, is the frequency index, Is an imaginary unit.
[0032] Figure 5 This is a network structure diagram of a frequency enhancement module provided by an embodiment of the present invention, specifically including: S301, enhance features in output space Apply cascaded 3×3 and 1×1 depthwise separable convolutions to generate high-frequency features ; S302, high frequency features and spatial enhancement features The dual-domain enhanced features are output by fast Fourier transform addition. .
[0033] Dual-domain enhancement features Represents the comprehensive image features after restoration, consisting of high-frequency features and spatial enhancement features Transformation and addition are obtained to enhance the features of the dual domains The restored image can be obtained by decoding.
[0034] S4, features of the output through the decoder Decode, restore image resolution, restore low-light images to normal images, and optimize the network by combining spatial and frequency domain losses.
[0035] Figure 6 This is a network structure diagram of a feature resolution self-adjusting decoder provided by an embodiment of the present invention; that is, a decoder including two ResNets is used to decode the extracted features and restore the low-light image to a normal image; S401, strengthen the dual domain feature The upsampling is achieved by bilinear interpolation and concatenated with the encoder features of the corresponding scale. After concatenation, the features are processed by ResBlock, which contains two 3×33×3 convolutional layers and residual connections. The number of output channels remains consistent with the encoder scale. The image resolution is gradually restored through ResBlock, and the enhanced image is finally output. ; S402, constrain the spatial features through the L1 absolute value loss function; the expression is: ; Where, is the spatial feature loss, is the absolute value operation, To enhance the image pixel matrix, is the original image pixel matrix.
[0036] S403, perform fast Fourier transform on the original image and the predicted image, and constrain the frequency features by the L1 absolute value loss function, which is expressed as: ; Where, is the frequency characteristic loss, is the Fast Fourier Transform.
[0037] As can be seen from the above embodiments, the present invention provides a low-light image enhancement method based on self-adjustment of spatial-frequency domain feature resolution. This method combines spatial and frequency domain information to restore low-light images to normal images through self-adjustment of feature resolution. Specifically, this method focuses on key areas of low-light images and fully understands the differences between low-light and normal areas in both the spatial and frequency domains. The spatial enhancement module takes features of the image at different scales as input and repairs the spatial regions using linear layers and a feedforward neural network. Subsequently, the frequency enhancement model enhances high-frequency signals or complex regions. To reduce computational overhead, the present invention uses the Transformer only for half of the features. ResBlocks are used for the remaining half of the features that are not spatially relevant, significantly reducing computational effort without sacrificing accuracy. This method addresses the problem that traditional methods, which rely solely on spatial domain brightness adjustment, cannot effectively handle noise, artifacts, and chromatic aberration. It overcomes the limitations of existing deep learning methods that struggle to fully restore image illumination and detail due to inductive bias or high computational complexity. It also addresses the drawback of existing methods that ignore frequency domain information, resulting in suboptimal enhancement results in complex scenes. This method is suitable for image enhancement in various low-light environments, such as nighttime surveillance, low-light photography, and medical image processing. It effectively restores image brightness, contrast, and detail, suppresses noise and artifacts, and improves image visual quality and the performance of subsequent computer vision tasks. By combining spatial and frequency enhancement modules, and introducing a feature resolution self-adjuster, this method achieves efficient enhancement of key areas in low-light images, significantly improving both the enhancement effect and robustness.
[0038] In Example 2, the low-light image enhancement system based on spatial-frequency domain feature resolution self-adjustment provided by the embodiment of the present invention proposes a Transformer-based spatial enhancement module, a frequency enhancement module through deep convolution and a sequential combination thereof, and a feature resolution self-adjuster; Figure 2-Figure 6 As shown; The feature resolution self-adjuster includes an encoder and a decoder. The encoder uses a hybrid module of ResNet and Transformer to perform convolution operations on the upper and lower features of the image, encodes the input dark and light images, and completes multi-scale feature learning; the decoder performs convolution operations on the output features. Decode, restore image resolution, restore low-light images to normal images, and optimize the network by combining spatial and frequency domain losses; The spatial enhancement module flattens the input feature map into a two-dimensional matrix, expands the number of channels through a linear layer, applies a multi-head self-attention mechanism combined with residual connections, further refines the features through layer normalization and a feed-forward network, and finally reshapes the processed features back to the original spatial dimensions. This is then further repaired by the frequency enhancement module. The frequency enhancement module applies deep convolution to the spatially enhanced features to generate high-frequency features, and then generates the final output through element-wise multiplication and residual connection.
[0039] Specifically, the feature resolution self-adjuster: the present invention splits the high-resolution features and introduces a multi-scale receptive field to fuse and reuse multi-scale and multi-level features, while improving computational efficiency. Specifically: the entire encoder and decoder network includes three scales. The three scales correspond to different resolution features. The encoder is composed of Tranformer and ResBlock. The feature resolution self-adjuster divides the input features into two equal parts along the channel dimension, and uses Transformer for the important features in half of the space. ResBlock is used for the features in the other half that do not care about the spatial position, and then the results are spliced to generate the final output. This is because Transformer provides the ability to associate long-distance pixels, but its computational complexity is extremely high. By combining the advantages of the two networks, the present invention enhances the receptive field of view and feature extraction capabilities of the overall network.
[0040] Spatial Enhancement Module: Utilizing a Transformer-based multi-head self-attention mechanism and a multi-layer feedforward network, the module enhances degraded regions in the spatial domain, providing location information for subsequent frequency domain enhancement. Specifically, the module flattens the input feature map into a two-dimensional matrix, expands the number of channels through linear layers, applies a multi-head self-attention mechanism combined with residual connections to maintain information integrity, further refines features through layer normalization and a feedforward network, and ultimately reshapes the processed features back to their original spatial dimensions. This is then further restored by the frequency enhancement module.
[0041] Frequency Enhancement Module: This module processes the output of the spatial enhancement module, extracting high-frequency signals from the image through deep convolution. This method highlights edges or complex areas in low-light images, further optimizing the enhancement effect. Specifically, deep convolution is first applied to the spatially enhanced features to generate high-frequency features. The final output is then generated through element-wise multiplication and residual connections.
[0042] It can be seen from the above embodiments that the present invention makes full use of spatial and frequency domain information, optimizes the entire enhancement process through a multi-input and multi-output scale strategy, and combines the loss functions of the spatial and frequency domains to generate high-quality normal illumination images.
[0043] The present invention realizes deep collaborative optimization of spatial and frequency domain information in low-light image processing through an innovative low-light image enhancement method with self-adjustment of spatial-frequency domain feature resolution. Combined with the feature resolution self-adjuster, it integrates the advantages of Transformer and ResBlock, increases the receptive field of the network while taking into account the computational complexity, and can dynamically capture the long-range spatial correlation of the degraded area, accurately locate the dark area and noise interference characteristics, and enhance the overall illumination restoration ability of the image. In the spatial enhancement module, a spatial enhancement module is constructed based on the improved Transformer architecture, and the receptive field is expanded and the context understanding ability is enhanced through multi-scale two-dimensional combination and feedforward network. In the frequency enhancement module, the frequency enhancement module constructed by deep separable convolution effectively extracts high-frequency detail information, strengthens the image edge and texture features on the basis of retaining the low-frequency illumination distribution, and significantly alleviates the contradiction between detail blurring and noise amplification in traditional methods.
[0044] In coal and rock scenarios, this invention not only addresses noise misjudgments in complex interference scenarios like underground dust obscuration, but also restores high-frequency edge information of safety signs. Its modular, lightweight architecture ensures dual-domain enhancement performance while optimizing computational efficiency. Its cross-scenario adaptability is demonstrated by robust handling of both natural low-light environments and specialized industrial scenarios (such as underground strong light interference and equipment vibration blur). Through multi-scale and multi-level feature fusion and reuse, it achieves a balanced approach to illumination restoration, detail preservation, and noise suppression, providing a high-quality image foundation for subsequent computer vision tasks.
[0045] To further illustrate the effects of the embodiments of the present invention, the following experiment was conducted. The evaluation metrics selected in the experiment were PSNR and SSIM, which can effectively evaluate the lighting restoration effect of an image. Higher PSNR and SSIM values indicate that the restored image is closer to the true value.
[0046] Table 1 Related evaluation indicators
[0047] Analysis reveals that traditional methods perform the weakest in terms of PSNR and SSIM, primarily due to their reliance on global adjustments and inability to adapt to the complex local degradation characteristics of low-light images. U-Net-based methods offer significant improvements over traditional methods, demonstrating the advantages of deep learning in feature learning and nonlinear mapping. However, their PSNR remains limited to below 20, indicating that single spatial domain processing is unable to fully restore image quality. The Transformer-based method achieves a PSNR of 23.649, demonstrating the effectiveness of the attention mechanism in capturing long-range dependencies. However, its SSIM is slightly lower than that of the U-Net solution, reflecting room for improvement in the Transformer's ability to preserve structural similarity.
[0048] The dual-domain enhancement method of the present invention achieved the best results in all indicators, with a PSNR of 27.727 and an SSIM of 0.888, which is significantly better than other comparison methods. This advantage mainly comes from three aspects: first, the spatial-frequency dual-domain collaborative enhancement mechanism achieves the best balance between illumination restoration and detail preservation; second, multi-scale feature learning effectively handles degradation patterns of different scales; finally, the improved Transformer architecture maintains a strong feature expression capability while reducing computational complexity. It is particularly noteworthy that the outstanding performance of the present invention in the SSIM indicator (an improvement of 7.5% over the second-best method) proves that the restored image is closer to the real lighting conditions in terms of structural similarity, which is particularly important for subsequent visual analysis tasks.
[0049] The stitching results of related images in the test set are shown as follows Figure 7-Figure 9 As shown in the figure, from the experimental results, it can be found that the finely carved patterns and paintings on the building facade are almost invisible under low light conditions, while the dual-domain enhancement mechanism of the present invention not only accurately restores the geometric features of the building structure, but also restores the details of the murals very well.
[0050] Under extreme side lighting conditions, the metal components at the edge of the table tennis table produce complex specular reflections and interlaced shadows. Other methods either overexpose these areas or lose detail. However, through precise control of the frequency domain enhancement module, this invention successfully restores the three-dimensionality and material properties of mechanical structures such as the axle while preserving detail in highlight areas.
[0051] At long distances from the subject, traditional methods are completely unable to distinguish the small buoys at the far end of the swimming pool. However, this invention accurately restores the outlines of these tiny objects through multi-scale feature fusion and long-range attention mechanism.
[0052] Case Study: Low-light image enhancement technology faces numerous unique challenges in the specialized application scenario of underground coal mines. The underground environment not only suffers from severe light deficiency but also from numerous interfering factors such as dust, mist, and equipment vibration, creating extremely complex imaging conditions. Traditional image enhancement methods are often helpless in this environment. Histogram equalization can over-enhance dust particles, creating snow-like noise; methods based on Retinex theory struggle to handle the localized glare caused by direct illumination from mining lamps; and conventional deep learning enhancement networks often misclassify dust particles as image details, resulting in severe image distortion.
[0053] The shortcomings of existing methods in coal mine scenarios are mainly reflected in three dimensions: First, in spatial domain processing, most methods cannot distinguish between the actual equipment outline and dust interference, and the enhanced images often mistake floating coal dust for equipment components; second, in terms of dynamic range control, traditional methods have difficulty simultaneously processing the highlight areas of miners' headlamps and the shadow areas deep in the tunnel; most importantly, existing methods generally lack the ability to model specific noise patterns underground, resulting in important safety warning signs either being submerged in noise or being incorrectly smoothed during the enhancement process.
[0054] In contrast, the present invention, specifically designed for the specific needs of underground coal mines, demonstrates significant advantages. The present invention's spatial enhancement module, leveraging prior knowledge of the underground scene, accurately identifies structural features that need to be preserved (such as hydraulic supports and conveyor belts), while identifying floating dust as interference that needs to be suppressed. The frequency enhancement module specifically enhances the edge features of text on safety signs, ensuring that key warnings such as "No Trespassing" are clearly legible. Particularly noteworthy is the present invention's handling of the area directly illuminated by mining lamps—preserving the brightness of the light source itself while clearly displaying the status of the equipment behind it through the strong light. This feature plays a crucial role in predicting underground equipment failures.
[0055] The above description is only a preferred specific implementation method of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A low-light image enhancement method based on self-adjustment of spatial-frequency domain feature resolution, characterized in that: The method comprises the following steps: S1 adopts an encoder-decoder architecture and uses a hybrid module of ResNet and Transformer to perform convolution operations on the upper and lower features of the image respectively, encodes the input dark and light images, and completes multi-scale feature learning; S2, flattens the multi-scale features output by the encoder and applies the multi-head self-attention mechanism of the spatial enhancement module to obtain the spatial enhancement features of the image; S3, the spatial enhancement features are input into the convolution layer of the frequency enhancement module to extract high-frequency features. The extracted high-frequency features and spatial enhancement features are used to calculate the image frequency and feature frequency through fast Fourier transform to complete the extraction and enhancement of edge information; S4, features of the output through the decoder Decode, restore image resolution, restore low-light images to normal images, and optimize the network by combining spatial and frequency domain losses.
2. The low-light image enhancement method based on spatial-frequency domain feature resolution self-adjustment according to claim 1, characterized in that: Step S1 specifically includes: S101, feed the input image into a 3×3 convolutional network to extract image features F; S102, divide the image feature F into and Two parts, Input to a network consisting of 4 residual blocks, each of which contains two 3×3 convolutional layers and ReLU activation function, The convolution with a step size of 2 is used to perform 1 / 2 downsampling, which is then input to the Transformer module to capture long-range dependencies. The resolution is then restored by upsampling with bilinear interpolation, and the features are output. and Splicing along the channel dimension to obtain fusion features ; S103, the fusion features Put it into ResBlock to get 1 / 2 feature scale ,Will Put in ResBlock to get 1 / 4 of the feature scale .
3. The low-light image enhancement method based on spatial-frequency domain feature resolution self-adjustment according to claim 2, characterized in that: Step S2 specifically includes: S201, the obtained characteristic scale , , , , flattened into a 2D feature matrix through a linear neural network, and the number of channels expanded through a linear layer; S202, apply multi-head self-attention to the features obtained in step S201 Mechanism, through residual connection and layer normalization Keep the original information, the expression is: ; Where, For normalization, is the feature processed by the attention mechanism, is the feature after linear mapping, It is processed by the multi-head self-attention mechanism. is the feature after linear mapping of dimension 1, is the feature after linear mapping of dimension 2, It is the feature after linear mapping of dimension 3; S203, input the features into the feedforward network , contains two linear layers with ReLU activation function and output through residual connection; S204, reshape the features back to the original spatial dimensions , and obtain spatial enhancement features , the expression is: ; Where, is the linear mapping layer, is the rectified linear unit activation function, are the height, width, and number of feature channels of the original feature map, respectively.
4. The low-light image enhancement method based on spatial-frequency domain feature resolution self-adjustment according to claim 3, characterized in that: Step S3 specifically includes: S301, enhance features in output space Apply cascaded 3×3 and 1×1 depthwise separable convolutions to generate high-frequency features ; S302, high frequency features and spatial enhancement features The dual-domain enhanced features are output by fast Fourier transform addition. .
5. The low-light image enhancement method based on spatial-frequency domain feature resolution self-adjustment according to claim 4, characterized in that: Step S4 specifically includes: S401, strengthen the dual domain feature The upsampling is achieved by bilinear interpolation and concatenated with the encoder features of the corresponding scale. After concatenation, the features are processed by ResBlock, which contains two 3×33×3 convolutional layers and residual connections. The number of output channels remains consistent with the encoder scale. The image resolution is gradually restored through ResBlock, and the enhanced image is finally output. ; S402, constraining spatial features through L1 absolute value loss function; S403 , performing fast Fourier transform on the original image and the predicted image, and constraining the frequency features through the L1 absolute value loss function.
6. The low-light image enhancement method based on spatial-frequency domain feature resolution self-adjustment according to claim 5, characterized in that: In step S402, the spatial features are constrained by the L1 absolute value loss function, which is expressed as: ; Where, is the spatial feature loss, is the absolute value operation, To enhance the image pixel matrix, is the original image pixel matrix.
7. The low-light image enhancement method based on spatial-frequency domain feature resolution self-adjustment according to claim 5, characterized in that: In step S403, the original image and the predicted image are subjected to fast Fourier transform, and the frequency feature is constrained by the L1 absolute value loss function, which is expressed as: ; Where, is the frequency characteristic loss, is the Fast Fourier Transform.
8. A low-light image enhancement system based on self-adjustment of spatial-frequency domain feature resolution, characterized in that: The system implements the low-light image enhancement method based on self-adjustment of spatial-frequency domain feature resolution according to any one of claims 1 to 7, and the system comprises: The feature resolution self-adjuster includes an encoder and a decoder. The encoder uses a hybrid module of ResNet and Transformer to perform convolution operations on the upper and lower features of the image, encodes the input dark and light images, and completes multi-scale feature learning; the decoder performs convolution operations on the output features. Decode, restore image resolution, restore low-light images to normal images, and optimize the network by combining spatial and frequency domain losses; The spatial enhancement module flattens the input feature map into a two-dimensional matrix, expands the number of channels through a linear layer, applies a multi-head self-attention mechanism combined with residual connections, refines features through layer normalization and a feedforward network, reshapes the processed features back to the original spatial dimension, and repairs them through a frequency enhancement module; The frequency enhancement module applies deep convolution to the spatially enhanced features to generate high-frequency features, and generates the final output through element-wise multiplication and residual connection.
9. The low-light image enhancement system based on spatial-frequency domain feature resolution self-adjustment according to claim 8, characterized in that: The system is used in coal and rock scenarios to process noise judgment in complex interference scenarios of underground dust obstruction and to restore high-frequency edge information of safety signs.
10. The low-light image enhancement system based on spatial-frequency domain feature resolution self-adjustment according to claim 8, characterized in that: The system is mounted on a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can realize the functions of the low-light image enhancement system based on self-adjustment of spatial-frequency domain feature resolution.
Citation Information
Patent Citations
Night defogging method and device based on double-domain feature learning and cross-dimensional feature optimization
CN117541663A
CNN-Transform-based photon counting image high-efficiency joint denoising super-division method and CNN-Transform-based photon counting image high-efficiency joint denoising super-division system
CN117557467A
Seismic data reconstruction method based on spatial domain and frequency domain fusion architecture
CN118033732A
Low-light image enhancement method based on frequency domain and spatial domain perception
CN118674628A
Mine time shift resistivity monitoring data screening method, storage medium and software
CN119128564A
Cited By
Frequency domain segmentation collaborative gradient driven inverse lithography method, system and equipment and medium
CN121879048A
Low-illumination image enhancement method and device based on wavelet frequency domain residual quotient learning
CN121921183A