Mama low-illumination image enhancement method combining light color decoupling and cross guidance frequency-space attention and computer readable storage medium
The Mamba low-light image enhancement method, which employs light-color decoupling and cross-guided frequency-space attention, solves the problems of local over-enhancement, under-enhancement, and color shift in existing low-light image enhancement methods in complex low-light scenes, thereby improving the naturalness, detail restoration, and structure preservation of the image.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INST OF TECH AT WEIHAI
- Filing Date
- 2026-04-20
- Publication Date
- 2026-05-19
AI Technical Summary
Existing low-light image enhancement methods struggle to effectively distinguish between brightness and color changes in real, complex low-light scenes, leading to local over-enhancement, under-enhancement, or color shift. Furthermore, traditional uniform enhancement methods lack specificity and are difficult to adapt to local lighting conditions, affecting the naturalness, detail fidelity, and structural consistency of the image.
A Mamba low-light image enhancement method with light-color decoupling and cross-guided frequency-space attention is adopted. By acquiring the hue, saturation and luminance components of the input image, an illumination estimation input is generated. Under the guidance of illumination perception features, a dual-branch collaborative enhancement framework in the frequency domain and spatial domain is constructed. The frequency branch is used to restore texture and high-frequency details, and the spatial branch is used to restore edge structure. An illumination perception-guided cross-domain fusion module is designed to adaptively coordinate the fusion intensity of frequency features and spatial features.
It significantly improves the problems of insufficient brightness, texture obliteration and structural blur in low-light images, enhances the naturalness of the enhancement results, the ability to restore details and the effect of structure preservation, and adapts to complex lighting scenes.
Smart Images

Figure CN122066599A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of digital image processing and computer vision technology, and more specifically, to a Mamba low-light image enhancement method that combines light-color decoupling and cross-guided frequency-space attention, and a computer-readable storage medium. Background Technology
[0002] Low-light image enhancement is a crucial problem in digital image processing and computer vision, with wide applications in scenarios such as nighttime surveillance, autonomous driving, intelligent security, medical imaging, remote sensing, and mobile terminal photography. Due to the combined effects of insufficient lighting, uneven local illumination, and sensor noise in the imaging environment, low-light images typically suffer from overall low brightness, loss of detail in dark areas, blurred edge structures, color distortion, and amplified noise, severely impacting the accuracy of subsequent image analysis and visual perception tasks.
[0003] Most existing low-light image enhancement methods directly boost the brightness of the input image in the RGB space or perform end-to-end mapping, such as the Chinese invention patent with publication number CN116703783A based on CNN. A low-light image enhancement method based on Transformer hybrid modeling and bilateral interaction utilizes a normal illumination image as a label image to train a CNN-Transformer-based low-light image interactive enhancement network. An output mapping module converts the enhanced feature maps into the final image. Chinese invention patent CN114663300A, "Low-Light Image Enhancement Method, System, and Related Devices Based on DCE," takes image data as input and saves output parameters to an iteration parameter module. Each parameter in the iteration parameter module determines a curve shape, and the pixels in the RGB channels of the image data adjust their original RGB pixel values according to the curve shape. The curve iteration module takes the iteration parameters in the iteration parameter module as input and calculates a higher-order brightness enhancement curve (LE) through iteration and pixelation. The higher-order brightness enhancement curve LE is then used to perform pixel-level adjustments to the variation range of the input low-light image.
[0004] While the above solutions can improve image visibility to some extent, they still have significant shortcomings in real-world, complex, low-light scenes. First, relying solely on RGB space modeling makes it difficult to effectively distinguish between brightness and color changes, easily leading to local over-enhancement, under-enhancement, or color shifts after enhancement. Second, low-light degradation not only suppresses high-frequency texture information in the image but also disrupts the consistency of spatial structure, making it difficult for single spatial or frequency domain methods to simultaneously achieve detail restoration and structural preservation. Third, the degree of degradation and lighting conditions often vary significantly across different regions, and traditional uniform enhancement methods lack specificity, making it difficult to adaptively adjust enhancement strategies based on local lighting conditions.
[0005] Therefore, there is an urgent need for a low-light image enhancement method that can explicitly model lighting priors and combine frequency domain and spatial domain information for collaborative restoration, so as to improve the naturalness, detail fidelity and structural consistency of the enhancement results. Summary of the Invention
[0006] To address the aforementioned problems, this application employs a Mamba low-light image enhancement method that combines light-color decoupling with cross-guided frequency-space attention, comprising the following steps: The input image is acquired and color space mapping is performed to obtain the hue component, saturation component, and lightness component; Generate illumination estimation input based on the input image, saturation component, and luminance component; The illumination estimation input and the hue component are input into the illumination estimation module to obtain the initial brightened image and illumination perception features; The initial brightened image is input into the frequency domain enhancement branch and the spatial domain enhancement branch respectively to extract multi-level frequency domain features and spatial domain features; The frequency domain features and spatial domain features at each level are input into a cross-domain enhancement network composed of multi-level frequency-space fusion units. Under the guidance of illumination perception features, the frequency domain features and spatial domain features are fused step by step to output the final enhanced image.
[0007] Optionally, each stage of the frequency domain enhancement branch includes: The initial brightened image is mapped to the frequency domain and decomposed into amplitude and phase components to obtain amplitude and phase features; Spatial attention and channel attention are extracted from amplitude features and phase features respectively, and spatial attention weights and channel attention weights corresponding to amplitude features and phase features are obtained. The updated amplitude and phase features are obtained by cross-guided enhancement based on each attention weight; The updated amplitude and phase features are input into the inverse Fourier transform unit to obtain the frequency reconstruction features. Spatial feature scanning mapping and residual mapping are then performed on the frequency reconstruction features to obtain the frequency enhancement features.
[0008] Optionally, cross-boot enhancements include: By fusing the spatial attention weights and channel attention weights of amplitude and phase features, a comprehensive guided response for the amplitude component is obtained. Integrated guidance response of phase components , represented as: ; In the formula, and These represent the spatial attention weights and channel attention weights corresponding to the amplitude features, respectively. and These represent the spatial attention weights and channel attention weights corresponding to the phase features, respectively. Indicates a fusion operation; The guiding information of the phase component is used to correct the amplitude component, and the guiding information of the amplitude component is used to correct the phase component. Furthermore, channel rearrangement enhances the cross-channel information interaction between the two types of features, resulting in an updated amplitude feature. and phase characteristics , represented as: ; In the formula, and These represent the amplitude update function and the phase update function under cross-guided conditions, respectively.
[0009] Optionally, the spatial feature scanning mapping and residual mapping of the frequency reconstruction features include: Multi-mode Spatial Feature Scanning Unit (MPSFS) to perform path rearrangement and local feature propagation on the frequency reconstruction features to obtain scan-enhanced features. : ; in, This represents a multi-modal spatial feature scan mapping.
[0010] Subsequently, the scan-enhanced features and input features are residually fused and then passed through a normalization unit and a feedforward mapping unit to obtain the final frequency-enhanced features. The process is represented as follows: ; ; in, Presentation layer normalization operation, This represents the feedforward mapping unit.
[0011] Optionally, each level of the airspace enhancement branch includes: The initial brightened image is normalized to obtain the normalized input features; Multi-scale spatial enhancement features are extracted from normalized input features using pyramid pooling attention; Multi-mode spatial scanning is performed on multi-scale spatial enhancement features to obtain scan enhancement features; Features are enhanced in the output space through residual connections and feedforward mapping.
[0012] Optionally, multi-scale spatial augmentation features extracted from the normalized input features via pyramid pooling attention include: Normalized input features Input a pyramid pooled attention unit and perform pooling operations on it at multiple scales. Let the th... The pooling operation corresponding to each scale is: Then we have: ; in, Indicates the first Pooling ratio, Indicates the number of pooling branches. Indicates the first Pooling features at various scales.
[0013] Pooling features at multiple scales are used as key-value pair features, i.e.: ; At the same time, the normalized input features As a query feature and query features Key-value pair features Spatial relationship modeling is performed using multi-head self-attention units to obtain multi-scale spatial augmentation features. ,Right now: ; in, This represents a multi-head self-attention mapping.
[0014] Optionally, multimodal spatial scanning includes adding multi-scale spatial enhancement features. The multi-mode spatial feature scanning unit (MPSFS) is input, and its spatial reconstruction and propagation along different paths are performed to obtain enhanced scanning features. ,Right now: ; In the formula, This represents a multi-modal spatial feature scan mapping.
[0015] The multi-mode spatial feature scanning unit includes one or a combination of horizontal scanning, vertical scanning, zigzag scanning and interlaced scanning.
[0016] Optionally, output spatial enhancement features via residual connections and feedforward mapping include scanning enhancement features. Input features Perform residual connections to obtain intermediate features: ; in, This indicates an element-wise summation operation.
[0017] Then the intermediate features The updated features are sequentially input into the normalization unit and the feedforward mapping unit, and then output as the final spatial augmentation features through residual connections. ,Right now: ; in, This represents the feedforward mapping unit.
[0018] Optionally, the stepwise fusion of frequency domain features and spatial domain features guided by illumination perception features includes: The cross-domain enhanced network includes The cascaded stages, the first The frequency domain characteristics, spatial domain characteristics, decoding characteristics, and fusion output of each stage are denoted as follows: , , and Then we have: ; in, Indicates the first Level frequency sensing enhancement mapping, Indicates the first Level space enhancement mapping.
[0019] For the deepest level Its fused output is represented as: ; in, Indicates the first Frequency-space fusion mapping.
[0020] For the remaining levels First, fuse the features of the next level. After being mapped to the current scale by the decoding module, the result is obtained. ; Then, the current level frequency domain features Current level airspace characteristics With decoding features The pieces are stitched together and combined with lighting perception features. Input the current-level frequency-space fusion module to obtain ; in, Indicates the first Level decoding mapping, This indicates a channel-based splicing operation.
[0021] Finally, the first-stage fusion output is mapped to the enhanced image Output, i.e.: ; in, This indicates the output mapping function.
[0022] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the Mamba low-light image enhancement method combining optical-color decoupling and cross-guided frequency-space attention as described above.
[0023] The beneficial effects of the Mamba low-light image enhancement method combining light-color decoupling and cross-guided frequency-space attention provided in this application are as follows: First, an illumination estimation stage is constructed using RGB images and brightness and saturation information in HSV space to generate initial brightening results and illumination perception features, thereby explicitly representing the distribution of dark areas, brightness gradients, and local illumination unevenness patterns in low-light images. Second, guided by illumination perception features, a dual-branch collaborative enhancement framework in the frequency and spatial domains is constructed. The frequency branch restores texture and high-frequency details, while the spatial branch restores edge structures and regional consistency. Furthermore, an illumination perception-guided cross-domain fusion module is designed to adaptively coordinate the fusion intensity of frequency and spatial features according to the illumination state of different regions, thereby improving the adaptability of the enhancement results to complex low-light scenes. In addition, in the frequency domain feature modeling, a cross-guided attention mechanism of amplitude and phase is introduced to enable texture energy recovery and structural localization information recovery to proceed in tandem, further enhancing the image detail and contour preservation capabilities.
[0024] The method proposed in this application can effectively improve problems such as insufficient brightness, texture obscuration, structural blurring, and color distortion in low-light images. It has the advantages of enhancing the naturalness of the results, strong detail recovery ability, good structure preservation effect, and strong adaptability to complex lighting scenes. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0026] Figure 1 This is a schematic diagram of the overall architecture of the Mamba low-light image enhancement method combining light-color decoupling and cross-guided frequency-space attention provided in the embodiments of this application; Figure 2 yes Figure 1 A magnified schematic diagram of the HSV-based illumination estimation stage architecture on the left. Figure 3 yes Figure 1 Enlarged schematic diagram of the cross-domain enhancement stage architecture on the right; Figure 4 This is a schematic diagram of the frequency domain enhancement branch architecture provided in the embodiments of this application; Figure 5This is a schematic diagram of the spatial enhancement tributary architecture provided in the embodiments of this application; Figure 6 These are low-light enhancement comparison images of book text details provided in the embodiments of this application, wherein (a) is the image before enhancement and (b) is the image after enhancement; Figure 7 These are comparison images of mid-field low-light enhancement in large-space venues provided in the embodiments of this application, wherein (a) is the image before enhancement and (b) is the image after enhancement; Figure 8 These are comparison images of enhanced low-light performance in distant views of large venues provided in this application embodiment, where (a) is the image before enhancement and (b) is the image after enhancement. Detailed Implementation
[0027] To make the technical problems, technical solutions, and beneficial effects to be solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this application.
[0028] Example 1 like Figures 1-3 As shown, this application provides a Mamba low-light image enhancement method that combines light and color decoupling with cross-guided frequency-space attention, including the following steps: acquiring the input image and performing color space mapping to obtain hue component, saturation component and luminance component; Generate illumination estimation input based on the input image, saturation component, and luminance component; The illumination estimation input and the hue component are input into the illumination estimation module to obtain the initial brightened image and illumination perception features; The initial brightened image is input into the frequency domain enhancement branch and the spatial domain enhancement branch respectively to extract multi-level frequency domain features and spatial domain features; The frequency domain features and spatial domain features at each level are input into a cross-domain enhancement network composed of multi-level frequency-space fusion units. Under the guidance of illumination perception features, the frequency domain features and spatial domain features are fused step by step to output the final enhanced image.
[0029] Figure 1 In the diagram, the left side represents the HSV-based illumination estimation stage, and the right side represents the cross-domain enhancement stage. RGB represents the features of the input image in the RGB space, and HSV represents the features of the input image in the HSV space. The C inside the white circle represents the feature stitching operation.
[0030] In this embodiment, it specifically includes: S1. Acquire the low-light input image and perform color space mapping. Get the input low-light image (i.e., the input image), where and These represent the height and width of the image, respectively. The input image is mapped from RGB space to HSV space to obtain the hue components. saturation component and brightness component , represented as: ; in, This indicates a color space conversion operation from RGB to HSV.
[0031] This step decouples the brightness and color information in the original image, allowing for the explicit extraction of prior information related to the lighting state in subsequent processing.
[0032] S2. Construct an HSV auxiliary representation and generate illumination estimation input. Input image Channel fusion is performed with the luminance-related components in the HSV space to construct a joint input representation for the illumination estimation stage. The RGB image is then combined with the saturation components. and brightness component By splicing along the channel dimension, we obtain: ; in, , This indicates a splicing operation based on channel dimensions.
[0033] By introducing a saturation component and brightness component The joint input representation not only preserves the semantic and texture information in the original image, but also explicitly enhances the network's ability to perceive brightness distribution and color stability.
[0034] S3. Generate initial lighting results and lighting perception features through the lighting estimation module.
[0035] Combined input representation hue components Input lighting estimation module This yields the initial brightening result (initial brightened image). and lighting perception features ,Right now: ; in, , representing the initial brightened image, , indicating the characteristics of lighting perception, Indicates the number of feature channels.
[0036] The lighting estimation module includes: (1) Adopt Convolution performs channel fusion on the joint input representation; (2) Depth-separable convolution is used to model the local illumination relationship between regions of different brightness; (3) Output the initial brightening results through convolution mapping. and lighting perception features .
[0037] In this application, This can be abstractly represented as the input image undergoing a brightness mapping. The enhanced result under the action is: ; in, This is a three-channel lighting representation. This indicates element-wise multiplication.
[0038] S4. Recover texture detail information through frequency domain branching, such as Figure 4 As shown.
[0039] Figure 4 middle, Indicates amplitude characteristics, Indicates phase characteristics, This represents the updated amplitude feature. This represents the updated phase characteristics. This indicates an element-wise summation operation.
[0040] S401. Map the input features to the frequency domain and decompose them into amplitude and phase components.
[0041] Let the input features be First, perform a Fourier transform on the input features to obtain their spectral representation: ; in, Indicates Fourier transform, Represents the amplitude component of the spectrum. Represents the phase component of the spectrum. It represents the imaginary unit.
[0042] In this application, amplitude component It primarily reflects the energy intensity of an image at different frequencies and is closely related to image contrast, texture intensity, and high-frequency details; phase component It primarily reflects the geometric location information of structures in an image and is closely related to edges, contours, and local structures. Decomposing the spectrum into amplitude and phase components is beneficial for modeling the texture recovery and structure localization capabilities of the image separately.
[0043] With amplitude components As amplitude feature , with phase component As a phase feature .
[0044] S402. Extract spatial attention and channel attention for amplitude and phase components, respectively.
[0045] The amplitude feature and phase characteristics Input the Cross-Guided Amplitude-Phase Attention (CGAPA) module separately. For any input feature... First, construct its spatial attention branch and channel attention branch.
[0046] The spatial attention branch extracts spatial statistics using global max pooling and global average pooling, concatenates them along the channel dimension, and then obtains the spatial attention weights through convolution. This process can be represented as follows: ; in, Indicates global max pooling. Indicates global average pooling. Indicates the kernel size as Convolution operation, This represents the spatial attention weights.
[0047] The channel attention branch extracts channel descriptors using global average pooling and through two layers. Convolution is used for mapping to obtain channel attention weights, represented as: ; in, This represents the activation function. This represents the channel attention weight.
[0048] S403. Cross-guided enhancement of amplitude and phase components.
[0049] The spatial attention weights and channel attention weights are fused to obtain a combined guided response of amplitude and phase components. Preferably, the fusion method uses element-wise summation, expressed as: ; in, and These represent the spatial attention weights and channel attention weights corresponding to the amplitude features, respectively. and These represent the spatial attention weights and channel attention weights corresponding to the phase features, respectively. This indicates a fusion operation.
[0050] Furthermore, to enhance the complementary modeling capability between amplitude and phase, the CGAPA module employs a cross-guided approach to update amplitude and phase features. Specifically, the guiding information of the phase component is used to correct the amplitude component, and the guiding information of the amplitude component is used to correct the phase component. Channel rearrangement operations are then used to enhance cross-channel information interaction between the two types of features, resulting in updated amplitude features. and phase characteristics The process can be abstractly represented as follows: ; in, and These represent the amplitude update function and the phase update function under cross-guided conditions, respectively.
[0051] Through this step, the amplitude component no longer focuses solely on the frequency energy itself, but achieves more accurate texture restoration with the assistance of phase structure information; the phase component can also maintain a more stable structural positioning under the guidance of the amplitude energy distribution, thereby enhancing the accuracy of frequency domain detail restoration.
[0052] S404. Reconstructing the frequency domain enhancement result through inverse Fourier transform.
[0053] Updated amplitude features and phase characteristics Inputting the inverse Fourier transform unit yields the frequency reconstruction features: ; in, Indicates the inverse Fourier transform. This indicates the frequency reconstruction characteristics after amplitude-phase joint enhancement.
[0054] S405. Perform spatial feature scanning and residual mapping on the frequency reconstruction features to obtain frequency enhancement features. To further enhance local texture modeling capabilities, a multi-mode spatial feature scanning unit (MPSFS) is introduced based on the frequency reconstruction features after inverse Fourier transform. This unit performs path rearrangement and local feature propagation on the frequency reconstruction features to obtain enhanced scanning features. : ; in, This represents a multi-modal spatial feature scan mapping.
[0055] Subsequently, the scan enhancement features and input features are residually fused, and then passed through a normalization unit and a feedforward mapping unit to obtain the final frequency enhancement features. The process is represented as follows: ; ; in, Presentation layer normalization operation, This represents the feedforward mapping unit.
[0056] S5. Recover structural and contextual information through spatial domain branching, such as Figure 5 As shown.
[0057] Figure 5 In the diagram, GAP1~GAP4 represent pooling layers of different scales, and P1~P4 represent pooling features of different scales. This represents the query feature, K represents the number of pooling branches, and V represents the luminance component. This indicates an element-wise summation operation.
[0058] S501. Normalize the input features.
[0059] Let the input features be First, the input features are normalized using a normalization unit to obtain: ; in, Presentation layer normalization operation.
[0060] This step helps improve the numerical stability of subsequent multi-scale pooling and attention computation processes.
[0061] S502. Extract multi-scale contextual features through pyramid pooling attention.
[0062] Normalized input features Input a pyramid pooled attention unit and perform pooling operations at multiple scales to obtain contextual representations under different receptive fields. Let the... The pooling operation corresponding to each scale is: Then we have: ; in, Indicates the first Pooling ratio, Indicates the number of pooling branches. Indicates the first Pooling features at various scales.
[0063] Preferably, pooling features at multiple scales are used as key-value pair features, i.e.: ; At the same time, the normalized input features As a query feature and the query features With the key-value pair features Spatial relationship modeling is performed using multi-head self-attention units to obtain multi-scale spatial augmentation features. ,Right now: ; in, This represents a multi-head self-attention mapping.
[0064] In this way, the input features can simultaneously perceive contextual information at multiple scales, thereby enhancing the ability to jointly model global structure and local details.
[0065] S503. Perform multi-mode spatial scanning on multi-scale spatial enhancement features.
[0066] The multi-scale spatial enhancement features obtained in step S502 The multi-mode spatial feature scanning unit (MPSFS) is input, and its spatial reconstruction and propagation along different paths are performed to obtain enhanced scanning features. ,Right now: ; in, This represents a multi-modal spatial feature scan mapping.
[0067] In a preferred embodiment, the multi-mode spatial feature scanning unit includes at least one of horizontal scanning, vertical scanning, zigzag scanning, and staggered scanning, and more preferably, all four scanning methods are used. Through this multi-mode scanning, the propagation capability of spatial features in different directions can be enhanced, alleviating the problem of insufficient recovery of complex structural information by single-path modeling.
[0068] S504. Enhance features in the output space through residual connections and feedforward mapping.
[0069] The scan enhancement features obtained in step S503 Input features Perform residual connections to obtain intermediate features: ; in, This indicates an element-wise summation operation.
[0070] Then the intermediate features The updated features are sequentially input into the normalization unit and the feedforward mapping unit, and then output as the final spatial augmentation features through residual connections. ,Right now: ; in, This represents the feedforward mapping unit.
[0071] S6. Input the initially enhanced image into a cross-domain enhancement network composed of multi-level frequency-spatial fusion units. Guided by illumination perception features, the network fuses frequency domain features and spatial domain features step by step, and outputs the final enhanced image.
[0072] The preliminary enhanced image (initial brightened image) obtained in step S3. The frequency domain enhancement branch and the spatial domain enhancement branch are input separately to extract frequency domain features and spatial domain features at each level; simultaneously, the illumination sensing features are... Each frequency-space fusion module is input to guide the fusion process of frequency domain features and spatial domain features at each level.
[0073] The cross-domain enhanced network includes The cascaded stages, the first The frequency domain characteristics, spatial domain characteristics, decoding characteristics, and fusion output of each stage are denoted as follows: , , and Then we have: ; in, Indicates the first Level frequency sensing enhancement mapping, Indicates the first Level space enhancement mapping.
[0074] For the deepest level Its fused output is represented as: ; in, Indicates the first Frequency-space fusion mapping.
[0075] For the remaining levels First, fuse the features of the next level. After being mapped to the current scale by the decoding module, the result is obtained. ; Then, the current level frequency domain features Current level airspace characteristics With the decoding features The pieces are stitched together and combined with lighting perception features. Input the current-level frequency-space fusion module to obtain ; in, Indicates the first Level decoding mapping, This indicates a channel-based splicing operation.
[0076] Finally, the first-stage fusion output is mapped to the enhanced image Output, i.e.: ; in, This indicates the output mapping function.
[0077] Through the aforementioned multi-level, progressive cross-domain enhancement process, it is possible to enhance the lighting perception features. Guided by this, the frequency domain detail information and spatial domain structure information are gradually fused and restored, resulting in an enhanced image output with higher brightness, richer texture, and more complete structure.
[0078] like Figure 6 As shown in (a), the text in the input image is blurry and has low detail recognition, such as... Figure 6 As shown in (b), the text on the book in the enhanced image is fully recognizable without any blurring.
[0079] like Figure 7 As shown in (a), the input image is generally dark, with the ground, walls, and basketball hoop structure almost invisible. Figure 7 As shown in (b), the enhanced image shows that the overall brightness of the venue is improved, there are no local dark areas, and the spatial structure is intact.
[0080] like Figure 8 As shown in (a), the input image has a dark basket and a dark ground, as... Figure 8 As shown in (b), the enhanced image displays clear texture details of the basketball hoop, backboard, and netting. Figure 6 and Figure 8 The black border in the image is an overlay to conceal the logo and is not generated by the solution in this application.
[0081] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A Mamba low-light image enhancement method combining light-color decoupling and cross-guided frequency-space attention, characterized in that, Includes the following steps: The input image is acquired and color space mapping is performed to obtain the hue component, saturation component, and lightness component; Generate illumination estimation input based on the input image, saturation component, and luminance component; The illumination estimation input and the hue component are input into the illumination estimation module to obtain the initial brightened image and illumination perception features; The initial brightened image is input into the frequency domain enhancement branch and the spatial domain enhancement branch respectively to extract multi-level frequency domain features and spatial domain features; The frequency domain features and spatial domain features at each level are input into a cross-domain enhancement network composed of multi-level frequency-space fusion units. Under the guidance of illumination perception features, the frequency domain features and spatial domain features are fused step by step to output the final enhanced image.
2. The Mamba low-light image enhancement method combining light-color decoupling and cross-guided frequency-space attention as described in claim 1, characterized in that: Each stage of the frequency domain enhancement branch includes: The initial brightened image is mapped to the frequency domain and decomposed into amplitude and phase components to obtain amplitude and phase features; Spatial attention and channel attention are extracted from the amplitude feature and phase feature respectively to obtain the spatial attention weight and channel attention weight corresponding to the amplitude feature, and the spatial attention weight and channel attention weight corresponding to the phase feature; The updated amplitude and phase features are obtained by cross-guided enhancement based on each attention weight; The updated amplitude and phase features are input into the inverse Fourier transform unit to obtain the frequency reconstruction features. Spatial feature scanning mapping and residual mapping are then performed on the frequency reconstruction features to obtain the frequency enhancement features.
3. The Mamba low-light image enhancement method combining optical-color decoupling and cross-guided frequency-space attention as described in claim 2, characterized in that: The cross-guided enhancement includes: By fusing the spatial attention weights and channel attention weights of amplitude and phase features, a comprehensive guided response for the amplitude component is obtained. Integrated guidance response of phase components , is represented as: ; In the formula, and These represent the spatial attention weights and channel attention weights corresponding to the amplitude features, respectively. and These represent the spatial attention weights and channel attention weights corresponding to the phase features, respectively. Indicates a fusion operation; The guiding information of the phase component is used to correct the amplitude component, and the guiding information of the amplitude component is used to correct the phase component. Furthermore, channel rearrangement enhances the cross-channel information interaction between the two types of features, resulting in an updated amplitude feature. and phase characteristics , is represented as: ; In the formula, and These represent the amplitude update function and the phase update function under cross-guided conditions, respectively.
4. The Mamba low-light image enhancement method combining optical-color decoupling and cross-guided frequency-space attention as described in claim 2, characterized in that: The spatial feature scanning mapping and residual mapping of the frequency reconstruction features include: a multi-mode spatial feature scanning unit (MPSFS) that performs path rearrangement and local feature propagation on the frequency reconstruction features to obtain scan-enhanced features. : ; in, Represents a multi-modal spatial feature scan mapping. Indicates frequency reconstruction characteristics; Subsequently, the scan enhancement features Input features Residual fusion is performed, and the data are then passed through a normalization unit and a feedforward mapping unit to obtain the final frequency enhancement features. The process is represented as follows: ; ; in, Indicates residual fusion characteristics, Presentation layer normalization operation, This represents the feedforward mapping unit.
5. The Mamba low-light image enhancement method combining optical-color decoupling and cross-guided frequency-space attention as described in claim 1, characterized in that: Each stage of the spatial enhancement branch includes: The initial brightened image is normalized to obtain the normalized input features; Multi-scale spatial enhancement features are extracted from normalized input features using pyramid pooling attention; Multi-mode spatial scanning is performed on multi-scale spatial enhancement features to obtain scan enhancement features; Features are enhanced in the output space through residual connections and feedforward mapping.
6. The Mamba low-light image enhancement method combining optical-color decoupling and cross-guided frequency-space attention as described in claim 5, characterized in that: The multi-scale spatial enhancement features extracted through pyramid pooling attention to normalize the input features include: Normalized input features Input a pyramid pooled attention unit and perform pooling operations on it at multiple scales. Let the th... The pooling operation corresponding to each scale is: Then we have: ; in, Indicates the first Pooling ratio, Indicates the number of pooling branches. Indicates the first Pooling features at various scales; Pooling features at multiple scales are used as key-value pair features, i.e.: ; At the same time, using the normalized input features As a query feature and the query features With the key-value pair features Spatial relationship modeling is performed using multi-head self-attention units to obtain multi-scale spatial augmentation features. ,Right now: ; in, This represents a multi-head self-attention mapping.
7. The Mamba low-light image enhancement method combining optical-color decoupling and cross-guided frequency-space attention as described in claim 5, characterized in that: The multi-mode spatial scanning includes multi-scale spatial enhancement features. The multi-mode spatial feature scanning unit (MPSFS) is input, and its spatial reconstruction and propagation along different paths are performed to obtain enhanced scanning features. ,Right now: ; In the formula, Represents a multi-modal spatial feature scan mapping; The multi-mode spatial feature scanning unit includes one or a combination of horizontal scanning, vertical scanning, zigzag scanning, and interlaced scanning.
8. The Mamba low-light image enhancement method combining optical-color decoupling and cross-guided frequency-space attention as described in claim 5, characterized in that: The method of outputting spatial enhancement features through residual connection and feedforward mapping includes scanning enhancement features. Input features Perform residual connections to obtain intermediate features: ; in, This represents an element-wise summation operation; Then the intermediate features The updated features are sequentially input into the normalization unit and the feedforward mapping unit, and then output as the final spatial augmentation features through residual connections. ,Right now: ; in, This represents the feedforward mapping unit.
9. The Mamba low-light image enhancement method combining optical-color decoupling and cross-guided frequency-space attention as described in claim 1, characterized in that: The stepwise fusion of frequency domain features and spatial domain features guided by illumination perception features includes: The cross-domain enhanced network includes The cascaded stages, the first The frequency domain characteristics, spatial domain characteristics, decoding characteristics, and fusion output of each stage are denoted as follows: , , and Then we have: ; in, Indicates the first Level frequency sensing enhancement mapping, Indicates the first Level space augmentation mapping; For the deepest level Its fused output is represented as: ; in, Indicates the first Frequency-space fusion mapping; For the remaining levels First, fuse the features of the next level. After being mapped to the current scale by the decoding module, we get: ; Then, the current level frequency domain features Current level airspace characteristics With the decoding features The pieces are stitched together and combined with lighting perception features. Input the current-level frequency-space fusion module to obtain: ; in, Indicates the first Level decoding mapping, This indicates a channel-based splicing operation; Finally, the first-stage fusion output is mapped to the enhanced image Output, i.e.: ; in, This indicates the output mapping function.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the Mamba low-light image enhancement method as described in any one of claims 1-9, which combines light-color decoupling with cross-guided frequency-space attention.