Low-light Image Enhancement Method and Device, Medium, and Equipment Based on Fourier Transform
By jointly optimizing the amplitude and phase components of the Fourier frequency domain, using a multi-layer attention mechanism and infrared image enhancement strategy, the problems of brightness distortion and structural information loss in the existing low-light image enhancement methods are solved, and a more natural and stable image enhancement effect is achieved.
Patent Information
- Application Number
- CN202510664959.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing low-light image enhancement methods based on Fourier transform usually focus on amplifying the amplitude component and directly copying the phase component, resulting in image brightness distortion and loss of key structural information, making it difficult to achieve a stable and reliable enhancement effect.
By jointly optimizing the amplitude and phase components of the Fourier frequency domain, a multi-layer attention mechanism is used to generate a brightness attention map, an infrared image is introduced to extract the second phase feature, and spatial and texture feature enhancement is performed, and the image is finally reconstructed through the inverse Fourier transform.
It significantly improves the robustness of low-light image enhancement, realizes targeted enhancement of underexposed areas, avoids degradation of correct exposure areas, and improves the accuracy of dynamic range control and the stability of structural information.
Smart Images

Figure CN120182154B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image enhancement technology, and in particular, to a low-light image enhancement method and device, medium, and equipment based on Fourier transform. Background Art
[0002] In many fields such as autonomous driving, medical imaging, and security monitoring, image quality is the key to ensuring the reliable operation of the system. However, images taken in low-illumination environments generally have problems such as low visibility, narrow dynamic range, poor signal-to-noise ratio, and color channel imbalance. For example, in the autonomous driving scenario, low-light images are difficult to clearly present key information such as roads and pedestrians, threatening driving safety; in the field of medical imaging, low-light images interfere with doctors' accurate judgment of lesions and affect disease diagnosis and treatment. Non-linear illumination will further exacerbate image degradation, making it difficult for traditional RGB-space-based image enhancement methods to effectively separate the illumination and reflection components, further increasing the complexity of low-light image enhancement and highlighting the urgency of researching low-light image enhancement technology.
[0003] Low-light image enhancement technology is an important research direction in the field of computer vision, aiming to solve problems such as reduced visibility, uneven exposure, loss of details, and color distortion caused by environmental and hardware limitations. Traditional non-learning methods have gradually been replaced by neural network model-based methods due to their dependence on artificial priors. The latter significantly improves the enhancement effect of low-light images through data-driven end-to-end training. However, existing methods still struggle to balance performance and computational efficiency: high-performance models often come with high computational costs, while lightweight models may sacrifice image quality. This indicates that a stable and reliable image enhancement solution has not yet been formed in this field, and there is an urgent need to explore new enhancement paradigms.
[0004] In recent years, the method of image enhancement based on Fourier transform has provided a cost-effective solution for low-light image enhancement. After an image is Fourier-transformed, the decomposed amplitude and phase components respectively dominate the illumination distribution and structural features of the image. This decoupling property enables frequency-domain methods to independently control the illumination and details of the image. However, existing Fourier-transform-based image enhancement methods usually focus on amplifying the amplitude component and directly copying the phase component. Such methods usually cause brightness distortion of the image and lose key structural information, thereby directly affecting the final image enhancement effect. Summary of the Invention
[0005] In view of this, the present application provides a low-light image enhancement method and device, a storage medium, and a computer device based on Fourier transform. By jointly optimizing the amplitude component and the phase component in the Fourier frequency domain, the robustness of low-light image enhancement is significantly improved. This dual-component collaborative optimization strategy overcomes the overexposure or distortion problems caused by single amplitude component adjustment and achieves a more natural enhancement effect. Through the brightness attention map, amplitude enhancement features are generated, achieving targeted enhancement of underexposed areas while avoiding degradation of correctly exposed areas, and significantly improving the accuracy of dynamic range control. An infrared image is introduced to extract the second phase feature, which is fused with the first phase feature. The illumination invariance of the infrared image strengthens the expression of contour information and ensures the stability of structural enhancement, greatly improving the enhancement effect of the finally obtained enhanced image.
[0006] According to one aspect of the present application, there is provided a low-light image enhancement method based on Fourier transform, including:
[0007] Performing Fourier transform processing on the low-light image to be enhanced to convert the low-light image from a spatial domain image to a frequency domain representation, and separating the first amplitude feature and the first phase feature corresponding to the low-light image according to the Fourier transform result of the frequency domain representation;
[0008] Based on the low-light image, determining the brightness attention map corresponding to the low-light image through a multi-layer attention mechanism, determining the second amplitude feature according to the brightness attention map, and generating amplitude enhancement features according to the first amplitude feature and the second amplitude feature;
[0009] Performing infrared conversion on the low-light image to obtain the infrared image corresponding to the low-light image, determining the second phase feature according to the infrared image, and generating phase enhancement features according to the first phase feature and the second phase feature;
[0010] According to the amplitude enhancement features and the phase enhancement features, obtaining a preliminary enhanced image through inverse Fourier transform, and performing downsampling processing on the preliminary enhanced image to generate a feature image after downsampling processing;
[0011] Performing spatial feature enhancement on the feature image after downsampling processing to obtain a spatial feature enhancement result, and performing texture feature enhancement on the feature image after downsampling processing to obtain a texture feature enhancement result. Fusing the spatial feature enhancement result and the texture feature enhancement result, and obtaining the enhanced image corresponding to the low-light image according to the fusion processing result.
[0012] According to another aspect of the present application, there is provided a low-light image enhancement device based on Fourier transform, including:
[0013] A Fourier transform module is used to perform Fourier transform processing on the low-light image to be enhanced, so as to convert the low-light image from a spatial-domain image into a frequency-domain representation, and separate the first amplitude feature and the first phase feature corresponding to the low-light image according to the Fourier transform result of the frequency-domain representation;
[0014] An amplitude feature enhancement module is used to determine the brightness attention map corresponding to the low-light image based on the low-light image through a multi-layer attention mechanism, determine the second amplitude feature according to the brightness attention map, and generate an amplitude enhancement feature according to the first amplitude feature and the second amplitude feature;
[0015] A phase feature enhancement module is used to perform infrared conversion on the low-light image to obtain an infrared image corresponding to the low-light image, determine the second phase feature according to the infrared image, and generate a phase enhancement feature according to the first phase feature and the second phase feature;
[0016] A downsampling module is used to obtain a preliminary enhanced image through inverse Fourier transform according to the amplitude enhancement feature and the phase enhancement feature, and perform downsampling processing on the preliminary enhanced image to generate a feature image after downsampling processing;
[0017] An image enhancement module is used to perform spatial feature enhancement on the feature image after downsampling processing to obtain a spatial feature enhancement result, and perform texture feature enhancement on the feature image after downsampling processing to obtain a texture feature enhancement result, fuse the spatial feature enhancement result and the texture feature enhancement result, and obtain an enhanced image corresponding to the low-light image according to the fusion processing result.
[0018] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned low-light image enhancement method based on Fourier transform is implemented.
[0019] According to still another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the program, the above-mentioned low-light image enhancement method based on Fourier transform is implemented.
[0020] With the above technical solution, a low-light image enhancement method, device, storage medium, and computer device based on Fourier transform provided by the present application. First, perform Fourier transform processing on the low-light image to be enhanced, map the low-light image from the spatial domain to the frequency domain, and obtain a complex spectrum. Then, the first amplitude feature of the low-light image can be extracted according to the amplitude spectrum in the spectrum, and the first phase feature of the low-light image can be extracted according to the phase spectrum in the spectrum. Further, based on the low-light image, a brightness attention map corresponding to the low-light image is generated through a multi-layer attention mechanism. Perform Fourier transform processing on the brightness attention map, and according to the spectrum after Fourier transform processing, obtain the amplitude spectrum corresponding to the brightness attention map. Then, the second amplitude feature of the brightness attention map can be extracted according to the amplitude spectrum. Based on the first amplitude feature and the second amplitude feature, amplitude feature enhancement is performed, and finally, an amplitude-enhanced feature can be obtained. Similarly, perform infrared conversion on the low-light image to obtain an infrared image. Then, perform Fourier transform processing on the infrared image, and according to the spectrum after Fourier transform processing, obtain the phase spectrum corresponding to the infrared image. Then, the second phase feature of the infrared image can be extracted according to the phase spectrum. Based on the first phase feature and the second phase feature, phase feature enhancement is performed, and finally, a phase-enhanced feature can be obtained. Subsequently, the obtained amplitude-enhanced feature and phase-enhanced feature are fused, and the inverse Fourier transform is performed on the fusion result to restore a preliminary enhanced image. After obtaining the preliminary enhanced image, the preliminary enhanced image can be input into an encoder for downsampling processing, and then a feature image after downsampling processing can be obtained. Subsequently, using a multi-scale convolution method, spatial feature enhancement is performed on the feature image after downsampling processing to obtain a spatial feature enhancement result. Similarly, using a Fourier convolution method, texture feature enhancement is performed on the feature image after downsampling processing to obtain a texture feature enhancement result. The texture feature enhancement result obtained after Fourier convolution is fused with the spatial feature enhancement result obtained after multi-scale convolution. Finally, based on the fusion result, a final enhanced image with enhanced brightness and optimized details is obtained. By jointly optimizing the amplitude component and phase component in the Fourier frequency domain in the embodiments of the present application, the robustness of low-light image enhancement is significantly improved. This two-component collaborative optimization strategy overcomes the overexposure or distortion problems caused by single amplitude component adjustment and achieves a more natural enhancement effect. By generating an amplitude-enhanced feature through a brightness attention map, targeted enhancement of underexposed areas is achieved, while avoiding the degradation of correctly exposed areas, significantly improving the accuracy of dynamic range control. The infrared image is introduced to extract the second phase feature and fuse it with the first phase feature. The illumination invariance of the infrared image strengthens the expression of contour information and ensures the stability of structure enhancement, greatly improving the enhancement effect of the finally obtained enhanced image.
[0021] The above description is only an overview of the technical solution of the present application. In order to better understand the technical means of the present application, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specific embodiments of the present application are given. Brief Description of the Drawings
[0022] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0023] Figure 1 A schematic flowchart of a low-light image enhancement method based on Fourier transform provided by an embodiment of the present application is shown;
[0024] Figure 2 A schematic flowchart of the generation process of a feature image provided by an embodiment of the present application is shown;
[0025] Figure 3 A schematic flowchart of the generation process of a first amplitude feature and a first phase feature provided by an embodiment of the present application is shown;
[0026] Figure 4 A schematic flowchart of the generation process of an amplitude enhancement feature provided by an embodiment of the present application is shown;
[0027] Figure 5 A schematic structural diagram of a deep learning model provided by an embodiment of the present application is shown;
[0028] Figure 6 A schematic flowchart of the generation process of a phase enhancement feature provided by an embodiment of the present application is shown;
[0029] Figure 7 A schematic flowchart of the generation process of a spatial feature enhancement result provided by an embodiment of the present application is shown;
[0030] Figure 8 A schematic structural diagram of a sampling processing unit provided by an embodiment of the present application is shown;
[0031] Figure 9 A schematic flowchart of the generation process of a texture feature enhancement result provided by an embodiment of the present application is shown;
[0032] Figure 10 A schematic structural diagram of a spectrum transformation network provided by an embodiment of the present application is shown;
[0033] Figure 11 A schematic diagram of an enhanced image of a low-light image provided by an embodiment of the present application is shown;
[0034] Figure 12 The structural schematic diagram of a low-light image enhancement device provided by an embodiment of the present application based on Fourier transform is shown;
[0035] Figure 13 The device structure schematic diagram of a computer device provided by an embodiment of the present application is shown. Detailed implementation manners
[0036] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0037] In this embodiment, a low-light image enhancement method based on Fourier transform is provided. As Figure 1 shown, the method includes:
[0038] Step 101: Perform Fourier transform processing on the low-light image to be enhanced, so as to convert the low-light image from a spatial-domain image to a frequency-domain representation, and separate the first amplitude feature and the first phase feature corresponding to the low-light image according to the Fourier transform result of the frequency-domain representation.
[0039] Step 102: Based on the low-light image, determine the brightness attention map corresponding to the low-light image through a multi-layer attention mechanism, determine the second amplitude feature according to the brightness attention map, and generate an amplitude enhancement feature according to the first amplitude feature and the second amplitude feature.
[0040] Step 103: Perform infrared conversion on the low-light image to obtain the infrared image corresponding to the low-light image, determine the second phase feature according to the infrared image, and generate a phase enhancement feature according to the first phase feature and the second phase feature.
[0041] Step 104: According to the amplitude enhancement feature and the phase enhancement feature, obtain a preliminary enhanced image through inverse Fourier transform, and perform downsampling processing on the preliminary enhanced image to generate a feature image after downsampling processing.
[0042] Step 105: Perform spatial feature enhancement on the feature image after downsampling processing to obtain a spatial feature enhancement result, and perform texture feature enhancement on the feature image after downsampling processing to obtain a texture feature enhancement result. Perform fusion processing on the spatial feature enhancement result and the texture feature enhancement result, and obtain the enhanced image corresponding to the low-light image according to the fusion processing result.
[0043] A low-light image enhancement method based on Fourier transform provided by an embodiment of the present application converts a low-light image from the spatial domain to the frequency domain, separates the amplitude feature (reflecting the illumination intensity) and the phase feature (reflecting the structural texture) of the image, enhances the amplitude and phase respectively in a targeted manner, and finally reconstructs an enhanced image with sufficient brightness and clear details.
[0044] First, perform Fourier transform processing on the low-light image to be enhanced, map the low-light image from the spatial domain to the frequency domain, and obtain a complex spectrum. The amplitude spectrum in the spectrum represents the energy distribution (i.e., brightness information) of the low-light image; the phase spectrum in the spectrum represents the structural information (such as edges, textures) of the low-light image. Then, the first amplitude feature of the low-light image can be extracted according to the amplitude spectrum, and the first phase feature of the low-light image can be extracted according to the phase spectrum.
[0045] Furthermore, in order to enhance the amplitude feature, based on the low-light image, through a multi-layer attention mechanism, multi-scale feature extraction is performed on the low-light image, and important brightness feature enhancement is carried out, and finally a brightness attention map corresponding to the low-light image is generated. Perform Fourier transform processing on the brightness attention map, and according to the spectrum after Fourier transform processing, obtain the amplitude spectrum corresponding to the brightness attention map. Then, the second amplitude feature of the brightness attention map can be extracted according to the amplitude spectrum. Based on the first amplitude feature and the second amplitude feature, amplitude feature enhancement is carried out, and finally the amplitude enhancement feature can be obtained.
[0046] Similarly, in order to enhance the phase feature, infrared conversion can be performed on the low-light image to obtain an infrared image. In a specific embodiment, the input low-light image can be converted into a corresponding infrared image through a pre-trained infrared conversion model. This conversion process reduces the interference of factors such as illumination conditions and color distortion in the visible light band, and significantly enhances key structural information such as object edges and textures. Then, perform Fourier transform processing on the infrared image, and according to the spectrum after Fourier transform processing, obtain the phase spectrum corresponding to the infrared image. Then, the second phase feature of the infrared image can be extracted according to the phase spectrum. Based on the first phase feature and the second phase feature, phase feature enhancement is carried out, and finally the phase enhancement feature can be obtained.
[0047] Subsequently, as Figure 2 shown, fuse the obtained amplitude enhancement feature and phase enhancement feature, and perform inverse Fourier transform on the fusion result to restore the preliminary enhanced image. Specifically, through inverse Fourier transform, the amplitude enhancement feature and the phase enhancement feature are reconstructed to the spatial domain, and its mathematical expression is:
[0048] ;
[0049] where,Y 1 represents the preliminary enhanced image. After obtaining the preliminary enhanced image, the preliminary enhanced image (represented by a feature vector) can be input into the encoder for downsampling processing, and then the feature image after downsampling processing (represented by a feature vector) can be obtained.
[0050] Although the Fourier space conveys global information, it lacks the enhancement of spatial details. Therefore, a process of spatial feature enhancement (multi-scale convolution) and texture feature enhancement (Fourier convolution) is designed. Specifically, by using the multi-scale convolution method, spatial feature enhancement is performed on the feature image after downsampling processing, and the spatial feature enhancement result can be obtained. Similarly, by using the Fourier convolution method, texture feature enhancement is performed on the feature image after downsampling processing, and the texture feature enhancement result can be obtained. To optimize the enhancement effect of low-light images, the texture feature enhancement result obtained after Fourier convolution is fused with the spatial feature enhancement result obtained after multi-scale convolution, and this process can be achieved by adding the two feature enhancement results element by element. The fused processing result obtained after fusion contains both fine texture information and global structural features, forming a feature representation with rich levels. Finally, based on this fusion result, the final enhanced image with enhanced brightness and optimized details is obtained.
[0051] By applying the technical solution of this embodiment, first, perform Fourier transform processing on the low-light image to be enhanced, map the low-light image from the spatial domain to the frequency domain, and obtain a complex spectrum. Then, the first amplitude feature of the low-light image can be extracted according to the amplitude spectrum in the spectrum, and the first phase feature of the low-light image can be extracted according to the phase spectrum in the spectrum. Further, based on the low-light image, a brightness attention map corresponding to the low-light image is generated through a multi-layer attention mechanism. Perform Fourier transform processing on the brightness attention map, and obtain the amplitude spectrum corresponding to the brightness attention map according to the spectrum after the Fourier transform processing. Then, the second amplitude feature of the brightness attention map can be extracted according to the amplitude spectrum. Based on the first amplitude feature and the second amplitude feature, amplitude feature enhancement is performed, and finally, an amplitude-enhanced feature can be obtained. Similarly, the low-light image is subjected to infrared conversion to obtain an infrared image. Then, perform Fourier transform processing on the infrared image, and obtain the phase spectrum corresponding to the infrared image according to the spectrum after the Fourier transform processing. Then, the second phase feature of the infrared image can be extracted according to the phase spectrum. Based on the first phase feature and the second phase feature, phase feature enhancement is performed, and finally, a phase-enhanced feature can be obtained. Subsequently, the obtained amplitude-enhanced feature and phase-enhanced feature are fused, and the inverse Fourier transform is performed on the fusion result to restore a preliminary enhanced image. After obtaining the preliminary enhanced image, the preliminary enhanced image can be input into an encoder for downsampling processing, and then a feature image after the downsampling processing can be obtained. Subsequently, using a multi-scale convolution method, spatial feature enhancement is performed on the feature image after the downsampling processing, and a spatial feature enhancement result can be obtained. Similarly, using a Fourier convolution method, texture feature enhancement is performed on the feature image after the downsampling processing, and a texture feature enhancement result can be obtained. The texture feature enhancement result obtained after Fourier convolution is fused with the spatial feature enhancement result obtained after multi-scale convolution. Finally, based on the fusion result, a final enhanced image with enhanced brightness and optimized details is obtained. By jointly optimizing the amplitude component and the phase component in the Fourier frequency domain in the embodiment of the present application, the robustness of low-light image enhancement is significantly improved. This two-component collaborative optimization strategy overcomes the overexposure or distortion problems caused by single amplitude component adjustment and achieves a more natural enhancement effect; by generating an amplitude-enhanced feature through a brightness attention map, targeted enhancement of underexposed areas is achieved, while avoiding the degradation of correctly exposed areas, significantly improving the accuracy of dynamic range control; introducing an infrared image to extract the second phase feature and fusing it with the first phase feature strengthens the expression of contour information through the illumination invariance of the infrared image, ensuring the stability of structural enhancement, and greatly improving the enhancement effect of the finally obtained enhanced image.
[0052] In an embodiment of the present application, optionally, step 101 includes: performing Fourier transform processing on the low-light image to be enhanced to obtain a Fourier transform result, and based on the Fourier transform result, extracting a first amplitude component and a first phase component corresponding to the low-light image; respectively performing non-linear feature extraction on the first amplitude component and the first phase component to obtain a first amplitude feature corresponding to the first amplitude component and a first phase feature corresponding to the first phase component.
[0053] In this embodiment, in the process of extracting frequency domain information through Fourier transform, first, a fast Fourier transform (FFT) is performed on the input low-light image to convert the spatial domain image into a frequency domain representation, thereby separating the first amplitude component and the first phase component. It should be particularly noted that in the Fourier frequency domain representation, the brightness features of the image are mainly concentrated in the first amplitude component, while the spatial structure features are mainly carried by the first phase component. Given that the low-light image enhancement task needs to focus on optimizing the first amplitude component, but if only the amplitude is amplified while keeping the original phase unchanged, it will inevitably cause detail distortion in the high-light area. Therefore, the embodiment of the present application adopts a synchronous optimization strategy, that is, the influence of the phase component is considered jointly during the adjustment of the amplitude component. To maximize the preservation of the integrity of the frequency domain features and avoid feature loss caused by the local receptive field of traditional convolution operations, after extracting the first amplitude component and the first phase component, non-linear feature extraction can be performed on the first amplitude component and the first phase component respectively, and then a first amplitude feature corresponding to the first amplitude component and a first phase feature corresponding to the first phase component can be obtained.
[0054] In a specific embodiment, as Figure 3 shown, the non-linear feature extraction can be achieved through the following steps: Two 1×1 convolutional kernels are introduced in each branch in cooperation with a parametric rectified linear unit (PReLU) to achieve non-linear feature extraction of the first amplitude component and the first phase component, and finally the optimized first amplitude feature and the first phase feature are output. This can effectively address the signal attenuation problem existing in the process of frequency domain feature extraction by traditional spatial domain convolution operations, laying an accurate frequency domain foundation for subsequent enhancement processing. Among them, the mathematical representation of the aforementioned Fourier transform can be realized according to the Fourier transform operator F, which maps the spatial domain image (where H and W respectively represent the height and width dimensions of the low-light image) to the frequency domain representation . F can be expressed as:
[0055] ;
[0056] where , is the spatial coordinate corresponding to the low-light image, , corresponding frequency-domain coordinates, is the imaginary unit, and its inverse transform is expressed as , mainly realizing the reconstruction from the frequency domain to the spatial domain. In addition, the complex components in the frequency domain can be respectively expressed as a first amplitude component with a clear physical meaning and a first phase component , which are respectively expressed as:
[0057] ;
[0058] ;
[0059] wherein 、 respectively refer to the real and imaginary components of
[0060] In an embodiment of the present application, optionally, step 102 includes: converting the low-light image from the RGB color space to the YCbCr color space, and determining the target brightness component corresponding to the low-light image based on the spatial conversion result; inputting the target brightness component into a deep learning model, encoding the target brightness component into a first extracted feature through the convolutional neural network of the deep learning model, and encoding the target brightness component into a second extracted feature through the self-attention mechanism network including a multi-layer attention mechanism of the deep learning model, obtaining the brightness attention map through a preset decoder of the deep learning model based on the first extracted feature and the second extracted feature; performing Fourier transform processing on the brightness attention map, extracting the second amplitude component corresponding to the brightness attention map, and performing non-linear feature extraction on the second amplitude component to obtain the second amplitude feature; performing a normalization processing operation on the second amplitude feature to obtain a first normalized processing feature, multiplying the first normalized processing feature by the first amplitude feature, and adding the product result to the first amplitude feature to obtain an amplitude enhancement feature.
[0061] In this embodiment, during the amplitude enhancement process, a deep learning-based method is mainly used to generate a brightness attention map to adaptively control the enhancement intensity of amplitude components in different regions. As Figure 4 shown, first, the RGB image can be converted to the YCbCr color space, the luminance (Y) component and the chrominance (Cb, Cr) components are separated, and the luminance component therein is used as the target brightness component. Then, a deep learning model is used to generate a brightness attention map corresponding to the target brightness component, because this component has higher noise robustness compared to other color channels. Specifically, as Figure 5As shown in the figure, the deep learning model (i.e., the HAFormer model) can include three parts, namely a convolutional neural network (CNN branch), a self-attention mechanism network with a multi-layer attention mechanism (Transformer branch), and a preset decoder (Decoder). Among them, the CNN branch and the Transformer branch are two parallel branches. Through the CNN branch, local features of the target luminance component (such as Figure 5 the luminance component in Y in ) are extracted to obtain the first extracted feature; the Transformer branch is used to capture global dependencies to obtain the second extracted feature. Finally, the first extracted feature and the second extracted feature are input into the preset decoder, and through the fusion mechanism in the preset decoder, the first extracted feature and the second extracted feature can be fused, that is, local and global features are fused, and a luminance attention map is generated through the preset decoder. Thus, the cooperation between the local receptive field and the global context features is realized, significantly improving the feature representation ability. Here, the high-brightness area in the generated luminance attention map corresponds to the dark area that needs to be enhanced in the low-light image, and vice versa. In the embodiment of the present application, the CNN-Transformer dual-branch encoder in the deep learning model fuses global and local features to accurately identify low-light areas and suppress the risk of overexposure.
[0062] Subsequently, the generated luminance attention map is subjected to Fourier transform to obtain the second amplitude component in the frequency domain, and nonlinear transformation is performed through two 1×1 convolutional layers in cooperation with a parametric rectified linear unit (PReLU) to output the second amplitude feature. After the second amplitude feature is normalized by the sigmoid function, the first normalized feature is obtained. Then, the first normalized feature and the first amplitude feature are multiplied element by element, and the product result is added to the first amplitude feature to generate an amplitude enhancement feature. The mathematical expression of this process is:
[0063] ;
[0064] where is the amplitude enhancement feature, is the first amplitude feature, is the second amplitude feature.
[0065] In the embodiment of the present application, through the amplitude frequency domain attention weighting strategy, the differential enhancement of the amplitude components in different regions of the image is realized. While maintaining the spectral characteristics of the normal exposure region, the amplitude components of the underexposed region are specifically enhanced, thus effectively avoiding the over-saturation phenomenon caused by traditional global enhancement methods.
[0066] In an embodiment of the present application, optionally, in step 103, "determining a second phase feature according to the infrared image, and generating a phase enhancement feature according to the first phase feature and the second phase feature" includes: after performing Fourier transform processing on the infrared image, extracting a second phase component corresponding to the infrared image, and performing non-linear feature extraction on the second phase component to obtain the second phase feature; respectively performing dimensionality reduction processing on the first phase feature and the second phase feature to obtain a first dimensionality reduction feature corresponding to the first phase feature and a second dimensionality reduction feature corresponding to the second phase feature; adding the first dimensionality reduction feature and the second dimensionality reduction feature to obtain a target dimensionality reduction feature, and performing normalization processing on the target dimensionality reduction feature to obtain a second normalized processing feature; multiplying the second normalized processing feature by the first phase feature, and adding the product result to the first phase feature to obtain a phase enhancement feature.
[0067] In this embodiment, during the infrared enhancement process, a cross-modal infrared image guidance mechanism is mainly introduced to optimize the structural information representation of the Fourier frequency domain phase component. Since infrared imaging is highly robust to illumination changes and can effectively retain the geometric contour features of the scene, infrared images are used for phase feature enhancement. As Figure 6 shown, first, the low-light image can be converted into a corresponding infrared image. This conversion process reduces the interference of factors such as illumination conditions and color distortion in the visible light band, and significantly enhances key structural information such as object edges and textures. Given that the phase component dominates the structural features of the image in the Fourier transform domain, a fast Fourier transform is performed on the infrared image to extract its phase spectrum to obtain the second phase component. Subsequently, two 1×1 convolutional layers are used in conjunction with a parametric rectified linear unit (PReLU) for feature refinement, and the enhanced second phase feature is output.
[0068] Next, to achieve cross-modal information fusion, the second phase feature of the infrared image and the first phase feature are co-optimized. Specifically, the cross-modal attention weights are calculated through matrix flattening flatten operation (i.e., dimensionality reduction processing) and softmax normalization processing operation, as shown in the following formula:
[0069] ;
[0070] where, The weight matrix (i.e., the second normalized processing feature) is used to regulate the structural enhancement intensity of the phase component, is the first phase feature, is the second phase feature, is the first dimensionality reduction feature, is the second dimensionality reduction feature. Finally, a residual learning strategy is adopted to generate the optimized phase enhancement feature , and its calculation process is expressed as:
[0071] .
[0072] In an embodiment of the present application, optionally, the "performing spatial feature enhancement on the downsampled feature image to obtain a spatially enhanced feature result" in step 105 includes: successively processing the downsampled feature image through a first downsampling processing unit, a second downsampling processing unit, and a third downsampling processing unit; inputting the second downsampled feature output by the second downsampling processing unit and the third downsampled feature output by the third downsampling processing unit into a first upsampling unit to obtain a first upsampled feature; inputting the first downsampled feature output by the first downsampling processing unit and the first upsampled feature into a second upsampling unit to obtain a second upsampled feature; performing convolution processing on the second upsampled feature respectively based on three convolutional layers with different convolutional kernel parameters to obtain a first convolution result, a second convolution result, and a third convolution result, multiplying the first convolution result, the second convolution result, and the third convolution result, and adding the product result after convolution processing to the second upsampled feature to obtain a spatially enhanced feature result.
[0073] In this embodiment, as Figure 7 shown, in the multi-scale convolution process (i.e., the spatial feature enhancement process), a pyramid-style feature extraction structure is constructed by operating on the downsampled feature image through a three-level downsampling processing unit and a two-level upsampling processing unit. Each sampling processing unit (including the upsampling processing unit and the downsampling processing unit) mainly uses the Conv2Former Block operation, and the feature consistency between the sampling processing units is maintained through skip connections. Among them, the structure of the Conv2Former Block is as Figure 8 shown. Specifically, the first downsampling processing unit, the second downsampling processing unit, and the third downsampling processing unit are connected in sequence. The downsampled feature image is first input into the first downsampling processing unit to obtain a first downsampled feature. Then, the first downsampled feature is input into the second downsampling processing unit to obtain a second downsampled feature, and the second downsampled feature is input into the third downsampling processing unit to obtain a third downsampled feature. Further, the second downsampled feature and the third downsampled feature are input into the first upsampling unit to obtain a first upsampled feature, and the first downsampled feature and the first upsampled feature are input into the second upsampling unit to obtain a second upsampled feature.
[0074] Subsequently, the second upsampled feature is simultaneously input into three independent convolutional layers with different convolutional kernel parameters (specifically, it can be a 1×1 convolutional layer, etc., which can be determined according to requirements), and the first convolutional result A1, the second convolutional result A2, and the third convolutional result A3 are respectively generated. These three convolutional results functionally correspond to the query vector, key vector, and value vector in the self-attention mechanism.
[0075] After that, the calculation of the attention weights is realized by measuring the similarity between the query vector (A1) and the key vector (A2) to generate a correlation matrix (AF), mainly obtained by element-wise multiplication of the query vector (A1) and the key vector (A2). Subsequently, the correlation matrix (AF) is element-wise multiplied by the value vector (A3), and the product result is convolved to obtain a global representation. The obtained global representation is element-wise added to the second upsampled feature to obtain a spatial feature enhancement result, realizing the collaborative optimization of global context information and local detail features.
[0076] In the embodiment of the present application, optionally, the "performing texture feature enhancement on the downsampled feature image to obtain a texture feature enhancement result" in step 105 includes: inputting the downsampled feature image into the first convolutional network, the second convolutional network, the third convolutional network, and the spectral transformation network respectively to obtain the fourth convolutional result, the fifth convolutional result, the sixth convolutional result, and the spectral transformation result; summing the fourth convolutional result and the sixth convolutional result, performing normalization processing on the summation result, and performing non-linear feature extraction on the normalized result to obtain the first feature extraction result; summing the fifth convolutional result and the spectral transformation result, performing normalization processing on the summation result, and performing non-linear feature extraction on the normalized result to obtain the second feature extraction result; fusing the first feature extraction result and the second feature extraction result to obtain the texture feature enhancement result.
[0077] In this embodiment, in the Fourier convolution process (i.e., the texture feature enhancement process), a fast Fourier convolution architecture is adopted to enhance the recovery ability of periodic textures in high-resolution images. This architecture adopts a dual-branch parallel processing mechanism: the local branch (including the first convolutional network and the second convolutional network) extracts local features in the spatial domain through conventional convolution operations, while the spectral transformation network in the global branch (including the third convolutional network and the spectral transformation network) transforms the input downsampled feature image to the frequency domain through fast Fourier transform, performs convolution operations in the frequency domain, and then restores it to the spatial domain through inverse Fourier transform to obtain the spectral transformation result. The output features of the two branches are fused by channel concatenation, which not only retains the ability of traditional convolution to capture local details but also obtains a global receptive field through frequency domain operations.
[0078] Specifically, asFigure 9 As shown, the feature image after downsampling processing is input into the first convolutional network to obtain a fourth convolutional result; the feature image after downsampling processing is input into the second convolutional network to obtain a fifth convolutional result; the feature image after downsampling processing is input into the third convolutional network to obtain a sixth convolutional result; the feature image after downsampling processing is input into the spectrum transformation network to obtain a spectrum transformation result. Subsequently, the fourth convolutional result and the sixth convolutional result are summed, and the sum result is normalized, and non-linear feature extraction is performed on the normalized result to obtain a first feature extraction result; the fifth convolutional result and the spectrum transformation result are summed, and the sum result is normalized, and non-linear feature extraction is performed on the normalized result to obtain a second feature extraction result. Finally, the first feature extraction result and the second feature extraction result are fused to obtain a texture feature enhancement result. In a specific embodiment, the first convolutional network, the second convolutional network, and the third convolutional network may be 3×3 convolutional networks; the normalization processing operation may be implemented by selecting a module with a normalization function and can be determined according to specific requirements; the non-linear feature extraction is implemented by ReLU (Rectified Linear Unit).
[0079] In the embodiment of the present application, optionally, inputting the feature image after downsampling processing into the spectrum transformation network to obtain a spectrum transformation result includes: inputting the feature image after downsampling processing into the first joint processing unit in the spectrum transformation network to obtain a first joint processing result, where the first joint processing unit includes a first convolutional sub-unit, a first normalization sub-unit, and a first non-linear transformation sub-unit connected in sequence; performing a Fourier transform on the first joint processing result, and inputting the Fourier transform result into the second joint processing unit to obtain a second joint processing result, where the second joint processing unit includes a second convolutional sub-unit, a second normalization sub-unit, and a second non-linear transformation sub-unit connected in sequence; performing an inverse Fourier transform on the second joint processing result, adding the inverse Fourier transform result to the first joint processing result to obtain an addition result, and performing a convolutional processing on the addition result to obtain a spectrum transformation result.
[0080] In this embodiment, the internal structure of the spectrum transformation network is as Figure 10As shown in the figure. Among them, the first joint processing unit may include a first convolutional subunit, a first normalization subunit, and a first non-linear transformation subunit connected in sequence, and the second joint processing unit may include a second convolutional subunit, a second normalization subunit, and a second non-linear transformation subunit connected in sequence. Among them, the first convolutional subunit and the second convolutional subunit may specifically be convolutional blocks, which can be determined according to specific requirements; the first normalization subunit and the second normalization subunit may specifically be modules with normalization functions, which can be determined according to specific requirements; the first non-linear transformation subunit and the second non-linear transformation subunit may be ReLU (Rectified Linear Unit, rectified linear unit). In addition, the convolutional layer for performing convolutional processing on the addition result may be a 1×1 convolutional layer.
[0081] In one embodiment, after obtaining the spatially enhanced feature result and the texture enhanced feature result, the spatially enhanced feature result and the texture enhanced feature result can be fused, and this process is achieved by element-wise addition of the two enhanced feature results. The fused features contain both fine texture information and global structural features, forming a feature representation with rich levels. Finally, the fused features can be input into the decoder network to output the final enhanced image with enhanced brightness and optimized details, effectively solving the limitations of traditional methods in texture preservation and global consistency, and achieving high-quality enhancement of complex scenes. As Figure 11 shown, a schematic diagram of the final enhanced image of a low-light image is given.
[0082] Furthermore, as Figure 1 a specific implementation of the method, an embodiment of the present application provides a low-light image enhancement device based on Fourier transform, as Figure 12 shown, the device includes:
[0083] A Fourier transform module, configured to perform Fourier transform processing on the low-light image to be enhanced, so as to convert the low-light image from a spatial domain image to a frequency domain representation, and separate the first amplitude feature and the first phase feature corresponding to the low-light image according to the Fourier transform result of the frequency domain representation;
[0084] An amplitude feature enhancement module, configured to determine a brightness attention map corresponding to the low-light image based on the low-light image through a multi-layer attention mechanism, determine a second amplitude feature according to the brightness attention map, and generate an amplitude enhancement feature according to the first amplitude feature and the second amplitude feature;
[0085] A phase feature enhancement module, configured to perform infrared conversion on the low-light image to obtain an infrared image corresponding to the low-light image, determine a second phase feature according to the infrared image, and generate a phase enhancement feature according to the first phase feature and the second phase feature;
[0086] A downsampling module, configured to obtain a preliminary enhanced image through inverse Fourier transform according to the amplitude enhanced feature and the phase enhanced feature, and perform downsampling processing on the preliminary enhanced image to generate a feature image after downsampling processing;
[0087] An image enhancement module, configured to perform spatial feature enhancement on the feature image after downsampling processing to obtain a spatial feature enhancement result, and perform texture feature enhancement on the feature image after downsampling processing to obtain a texture feature enhancement result, fuse the spatial feature enhancement result and the texture feature enhancement result, and obtain an enhanced image corresponding to the low-light image according to the fusion processing result.
[0088] Optionally, the Fourier transform module is configured to:
[0089] Perform Fourier transform processing on the low-light image to be enhanced to obtain a Fourier transform result, and based on the Fourier transform result, extract a first amplitude component and a first phase component corresponding to the low-light image;
[0090] Perform non-linear feature extraction on the first amplitude component and the first phase component respectively to obtain a first amplitude feature corresponding to the first amplitude component and a first phase feature corresponding to the first phase component.
[0091] Optionally, the amplitude feature enhancement module is configured to:
[0092] Convert the low-light image from the RGB color space to the YCbCr color space, and determine a target luminance component corresponding to the low-light image based on the spatial conversion result;
[0093] Input the target luminance component into a deep learning model, encode the target luminance component into a first extracted feature through the convolutional neural network of the deep learning model, and encode the target luminance component into a second extracted feature through the self-attention mechanism network including a multi-layer attention mechanism of the deep learning model. Based on the first extracted feature and the second extracted feature, obtain the luminance attention map through a preset decoder of the deep learning model;
[0094] After performing Fourier transform processing on the luminance attention map, extract a second amplitude component corresponding to the luminance attention map, and perform non-linear feature extraction on the second amplitude component to obtain the second amplitude feature;
[0095] Perform a normalization processing operation on the second amplitude feature to obtain a first normalized processing feature, multiply the first normalized processing feature by the first amplitude feature, and add the product result to the first amplitude feature to obtain an amplitude enhanced feature.
[0096] Optionally, the phase feature enhancement module is configured to:
[0097] After performing Fourier transform processing on the infrared image, extract the second phase component corresponding to the infrared image, and perform non-linear feature extraction on the second phase component to obtain the second phase feature;
[0098] Perform dimensionality reduction processing on the first phase feature and the second phase feature respectively to obtain the first dimensionality reduction feature corresponding to the first phase feature and the second dimensionality reduction feature corresponding to the second phase feature;
[0099] Add the first dimensionality reduction feature and the second dimensionality reduction feature to obtain a target dimensionality reduction feature, and perform normalization processing on the target dimensionality reduction feature to obtain a second normalized processing feature;
[0100] Multiply the second normalized processing feature by the first phase feature, and add the product result to the first phase feature to obtain a phase enhancement feature.
[0101] Optionally, the image enhancement module is configured to:
[0102] Process the feature image after downsampling through a first downsampling processing unit, a second downsampling processing unit, and a third downsampling processing unit in sequence;
[0103] Input the second downsampling feature output by the second downsampling processing unit and the third downsampling feature output by the third downsampling processing unit into a first upsampling unit to obtain a first upsampling feature;
[0104] Input the first downsampling feature output by the first downsampling processing unit and the first upsampling feature into a second upsampling unit to obtain a second upsampling feature;
[0105] Based on three convolutional layers with different convolutional kernel parameters, perform convolutional processing on the second upsampling feature respectively to obtain a first convolutional result, a second convolutional result, and a third convolutional result. Multiply the first convolutional result, the second convolutional result, and the third convolutional result, and add the product result after convolutional processing to the second upsampling feature to obtain a spatial feature enhancement result.
[0106] Optionally, the image enhancement module is further configured to:
[0107] Input the feature image after downsampling into a first convolutional network, a second convolutional network, a third convolutional network, and a spectral transformation network respectively to obtain a fourth convolutional result, a fifth convolutional result, a sixth convolutional result, and a spectral transformation result;
[0108] Sum the fourth convolution result and the sixth convolution result, perform normalization processing on the sum result, and perform non-linear feature extraction on the normalized result to obtain a first feature extraction result;
[0109] Sum the fifth convolution result and the spectrum transformation result, perform normalization processing on the sum result, and perform non-linear feature extraction on the normalized result to obtain a second feature extraction result;
[0110] Fuse the first feature extraction result and the second feature extraction result to obtain a texture feature enhancement result.
[0111] Optionally, the image enhancement module is further configured to:
[0112] Input the downsampled feature image into a first joint processing unit in the spectrum transformation network to obtain a first joint processing result, where the first joint processing unit includes a first convolution sub-unit, a first normalization sub-unit, and a first non-linear transformation sub-unit connected in sequence;
[0113] Perform a Fourier transform on the first joint processing result, and input the Fourier transform result into a second joint processing unit to obtain a second joint processing result, where the second joint processing unit includes a second convolution sub-unit, a second normalization sub-unit, and a second non-linear transformation sub-unit connected in sequence;
[0114] Perform an inverse Fourier transform on the second joint processing result, add the inverse Fourier transform result to the first joint processing result to obtain an addition result, and perform convolution processing on the addition result to obtain a spectrum transformation result.
[0115] It should be noted that for other corresponding descriptions of each functional unit involved in the low-light image enhancement device based on Fourier transform provided in the embodiments of the present application, reference can be made to Figures 1 to 11 the corresponding description in the method, which will not be elaborated here.
[0116] The embodiments of the present application further provide a computer device, which may specifically be a personal computer, a server, a network device, etc., such as Figure 13As shown, the computer device includes a bus, a processor, a memory, and a communication interface, and may further include an input / output interface and a display device. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps in the method embodiments are implemented.
[0117] Those skilled in the art can understand that Figure 13 the structure shown in [figure reference] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component arrangement.
[0118] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium may be non-volatile or volatile, and stores a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0119] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0120] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties.
[0121] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0122] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0123] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A low-light image enhancement method based on Fourier transform, characterized in that Including: Performing Fourier transform processing on the low-light image to be enhanced to convert the low-light image from a spatial-domain image to a frequency-domain representation, and separating a first amplitude feature and a first phase feature corresponding to the low-light image according to the Fourier transform result of the frequency-domain representation; Based on the low-light image, determining a brightness attention map corresponding to the low-light image through a multi-layer attention mechanism, determining a second amplitude feature according to the brightness attention map, and generating an amplitude enhancement feature according to the first amplitude feature and the second amplitude feature; Performing infrared conversion on the low-light image to obtain an infrared image corresponding to the low-light image, determining a second phase feature according to the infrared image, and generating a phase enhancement feature according to the first phase feature and the second phase feature; According to the amplitude enhancement feature and the phase enhancement feature, performing inverse Fourier transform to obtain a preliminary enhanced image, and performing downsampling processing on the preliminary enhanced image to generate a feature image after downsampling processing; Performing spatial feature enhancement on the feature image after downsampling processing to obtain a spatial feature enhancement result, and performing texture feature enhancement on the feature image after downsampling processing to obtain a texture feature enhancement result, fusing the spatial feature enhancement result and the texture feature enhancement result, and obtaining an enhanced image corresponding to the low-light image according to the fusion processing result.
2. The method according to claim 1, wherein The performing Fourier transform processing on the low-light image to be enhanced to convert the low-light image from a spatial-domain image to a frequency-domain representation, and separating a first amplitude feature and a first phase feature corresponding to the low-light image according to the Fourier transform result of the frequency-domain representation, includes: Performing Fourier transform processing on the low-light image to be enhanced to obtain a Fourier transform result, and extracting a first amplitude component and a first phase component corresponding to the low-light image based on the Fourier transform result; Performing non-linear feature extraction on the first amplitude component and the first phase component respectively to obtain a first amplitude feature corresponding to the first amplitude component and a first phase feature corresponding to the first phase component.
3. The method according to claim 1, characterized in that, The based on the low-light image, determining a brightness attention map corresponding to the low-light image through a multi-layer attention mechanism, determining a second amplitude feature according to the brightness attention map, and generating an amplitude enhancement feature according to the first amplitude feature and the second amplitude feature, includes: Converting the low-light image from the RGB color space to the YCbCr color space, and determining a target brightness component corresponding to the low-light image based on the spatial conversion result; Inputting the target brightness component into a deep learning model, encoding the target brightness component into a first extracted feature through the convolutional neural network of the deep learning model, and encoding the target brightness component into a second extracted feature through the self-attention mechanism network including a multi-layer attention mechanism of the deep learning model, and obtaining the brightness attention map through a preset decoder of the deep learning model based on the first extracted feature and the second extracted feature; After performing Fourier transform processing on the luminance attention map, extract the second amplitude component corresponding to the luminance attention map, and perform non-linear feature extraction on the second amplitude component to obtain the second amplitude feature; Perform a normalization operation on the second amplitude feature to obtain a first normalized feature, multiply the first normalized feature by the first amplitude feature, and add the product result to the first amplitude feature to obtain an amplitude enhancement feature.
4. The method according to claim 1, wherein Said determining a second phase feature according to the infrared image, and generating a phase enhancement feature according to the first phase feature and the second phase feature, including: After performing Fourier transform processing on the infrared image, extract the second phase component corresponding to the infrared image, and perform non-linear feature extraction on the second phase component to obtain the second phase feature; Perform dimensionality reduction processing on the first phase feature and the second phase feature respectively to obtain a first dimensionality reduction feature corresponding to the first phase feature and a second dimensionality reduction feature corresponding to the second phase feature; Add the first dimensionality reduction feature and the second dimensionality reduction feature to obtain a target dimensionality reduction feature, and perform normalization processing on the target dimensionality reduction feature to obtain a second normalized feature; Multiply the second normalized feature by the first phase feature, and add the product result to the first phase feature to obtain a phase enhancement feature.
5. The method according to claim 1, wherein Said performing spatial feature enhancement on the downsampled feature image to obtain a spatial feature enhancement result, including: Successively process the downsampled feature image through a first downsampling processing unit, a second downsampling processing unit, and a third downsampling processing unit; Input the second downsampling feature output by the second downsampling processing unit and the third downsampling feature output by the third downsampling processing unit into a first upsampling unit to obtain a first upsampling feature; Input the first downsampling feature output by the first downsampling processing unit and the first upsampling feature into a second upsampling unit to obtain a second upsampling feature; Based on three convolutional layers with different convolution kernel parameters, perform convolution processing on the second upsampling feature respectively to obtain a first convolution result, a second convolution result, and a third convolution result, multiply the first convolution result, the second convolution result, and the third convolution result, and add the product result to the second upsampling feature after convolution processing to obtain a spatial feature enhancement result.
6. The method according to claim 1, wherein Said performing texture feature enhancement on the downsampled feature image to obtain a texture feature enhancement result, including: Input the downsampled feature image into a first convolutional network, a second convolutional network, a third convolutional network, and a spectral transformation network respectively to obtain a fourth convolution result, a fifth convolution result, a sixth convolution result, and a spectral transformation result; Sum the fourth convolution result and the sixth convolution result, perform normalization processing on the sum result, and perform non-linear feature extraction on the normalization processing result to obtain a first feature extraction result; Sum the fifth convolution result and the spectrum transformation result, perform normalization processing on the sum result, and perform non-linear feature extraction on the normalized result to obtain a second feature extraction result; Fuse the first feature extraction result and the second feature extraction result to obtain a texture feature enhancement result.
7. The method according to claim 6, wherein Input the downsampled feature image into a spectrum transformation network to obtain a spectrum transformation result, including: Input the downsampled feature image into a first joint processing unit in the spectrum transformation network to obtain a first joint processing result, where the first joint processing unit includes a first convolution sub-unit, a first normalization sub-unit, and a first non-linear transformation sub-unit connected in sequence; Perform a Fourier transform on the first joint processing result, and input the Fourier transform result into a second joint processing unit to obtain a second joint processing result, where the second joint processing unit includes a second convolution sub-unit, a second normalization sub-unit, and a second non-linear transformation sub-unit connected in sequence; Perform an inverse Fourier transform on the second joint processing result, add the inverse Fourier transform result to the first joint processing result to obtain an addition result, and perform convolution processing on the addition result to obtain a spectrum transformation result.
8. A low-light image enhancement device based on Fourier transform, characterized in that, Including: A Fourier transform module for performing Fourier transform processing on the low-light image to be enhanced, so as to convert the low-light image from a spatial domain image to a frequency domain representation, and separate the first amplitude feature and the first phase feature corresponding to the low-light image according to the Fourier transform result of the frequency domain representation; An amplitude feature enhancement module for determining a brightness attention map corresponding to the low-light image based on the low-light image through a multi-layer attention mechanism, determining a second amplitude feature according to the brightness attention map, and generating an amplitude enhancement feature according to the first amplitude feature and the second amplitude feature; A phase feature enhancement module for performing infrared conversion on the low-light image to obtain an infrared image corresponding to the low-light image, determining a second phase feature according to the infrared image, and generating a phase enhancement feature according to the first phase feature and the second phase feature; A downsampling module for obtaining a preliminary enhanced image through an inverse Fourier transform according to the amplitude enhancement feature and the phase enhancement feature, and performing downsampling processing on the preliminary enhanced image to generate a downsampled feature image; An image enhancement module for performing spatial feature enhancement on the downsampled feature image to obtain a spatial feature enhancement result, and performing texture feature enhancement on the downsampled feature image to obtain a texture feature enhancement result, fusing the spatial feature enhancement result and the texture feature enhancement result, and obtaining an enhanced image corresponding to the low-light image according to the fusion processing result.
9. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Low-illumination image enhancement method based on curve wavelet attention and Fourier
CN118822908A
Visible light and infrared image fusion method based on space-frequency domain characteristics
CN119963958A