A low-light image enhancement method, device and electronic equipment

CN122841218APending Publication Date: 2026-09-29CHONGQING LANDIAN AUTOMOBILE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610933222.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]相关技术中,在进行暗光图像增强时,缺乏对图像局部内容的自适应能力,容易产生过度增强或产生细节丢失,难以有效处理复杂的非均匀光照退化

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841218A_ABST
    Figure CN122841218A_ABST
Patent Text Reader

Abstract

The application discloses a dark light image enhancement method and device and electronic equipment, and relates to the technical field of image processing, and is used for realizing enhancement processing of a dark light image, and comprises the following steps: acquiring a to-be-processed image through a vehicle-mounted sensor, converting the to-be-processed image from a three-primary-color color space to a target color space, and obtaining a first channel component and a second channel component; extracting a first feature map of the first channel component, determining a first weight corresponding to the first feature map, extracting a second feature map of the second channel component, and determining a second weight corresponding to the second feature map; performing feature fusion processing on the first feature map and the second feature map through the first weight and the second weight, and obtaining a fused feature map; and finally, outputting a target image according to the fused feature map. Through the above method, the dark light image can be enhanced in real time, a large amount of computing resources is not required, and the dark light image enhancement effect is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus and electronic device for enhancing low-light images. Background Technology

[0002] Low-light image enhancement is a technique used to improve image quality, making the enhanced image more suitable for human or machine vision processing. In low-light environments, due to insufficient light, images obtained by imaging tools often suffer from poor global visibility, low contrast, color distortion, and local degradation. This significantly impacts image usability and visual quality, affecting not only human visual perception but also the performance of downstream advanced vision tasks such as object detection and semantic segmentation.

[0003] Therefore, in order to improve the quality of low-light images, low-light image enhancement processing has become a current research hotspot. Low-light image enhancement processing can improve image brightness, suppress noise interference, and correct color deviation, so that the processed image has a better visual effect.

[0004] In related technologies, low-light image enhancement lacks adaptability to local image content, easily leading to over-enhancement or loss of detail, and struggles to effectively handle complex non-uniform illumination degradation. Furthermore, these technologies often struggle to balance computational complexity with enhancement effectiveness; some methods with high enhancement rates tend to have high computational complexity, resulting in significant overhead for high-resolution images and failing to meet real-time requirements. For example, when vehicles travel through tunnels or at night, the ambient light intensity drops sharply, resulting in low-quality images unsuitable for subsequent tasks such as target recognition and environmental detection. Therefore, image enhancement is necessary to obtain high-quality images. However, current image enhancement methods often suffer from high computational overhead, failing to meet real-time requirements. Thus, in scenarios with high real-time data requirements, such as intelligent driving or autonomous driving, the inability to perform image enhancement in real-time with low computational overhead can negatively impact driving safety. Summary of the Invention

[0005] This application provides a low-light image enhancement method, apparatus, and electronic device to enhance low-light images with low computational overhead, thereby improving image quality.

[0006] In a first aspect, this application provides a low-light image enhancement method, the method comprising: The image to be processed is acquired by the vehicle-mounted sensor and converted from the three primary color space to the target color space to obtain the first channel component and the second channel component. Extract the first feature map of the first channel component and determine the first weight corresponding to the first feature map; extract the second feature map of the second channel component and determine the second weight corresponding to the second feature map. By using the first weight and the second weight, feature fusion is performed on the first feature map and the second feature map to obtain a fused feature map; Based on the fused feature map, the target image is output; wherein, the target image is the image obtained by performing feature enhancement processing on the fused feature map using a target gating network determined based on the fused feature map, and then converting it to the three primary color space.

[0007] The above method significantly improves the computational efficiency of low-light image enhancement processing, meeting the real-time image processing requirements of vehicle terminals during operation. It also significantly enhances the effect of low-light image enhancement with low computational overhead, ensuring that images acquired by onboard sensors have significantly improved image quality after enhancement. This allows for subsequent use in downstream tasks such as target recognition and environmental detection, reducing misidentification due to loss of detail. Furthermore, image enhancement avoids the need to replace high-resolution cameras, fully utilizing the hardware's performance potential and reducing hardware costs.

[0008] In one optional implementation, extracting a first feature map of the first channel component and determining a first weight corresponding to the first feature map, and extracting a second feature map of the second channel component and determining a second weight corresponding to the second feature map, includes: Extract the first scale feature map corresponding to the first channel component according to the first feature scale and determine the first scale weight corresponding to the first scale feature map; and extract the second scale feature map corresponding to the second channel component according to the first feature scale and determine the second scale weight corresponding to the second scale feature map. By using the first scale weight and the second scale weight, feature fusion is performed on the first scale feature map and the second scale feature map to obtain the first scale fused feature map. The first feature map corresponding to the first channel component is extracted from the first scale fused feature map according to the second feature scale, and the first weight corresponding to the first feature map is determined. The second feature map corresponding to the second channel component is extracted from the first scale fused feature map according to the second feature scale, and the second weight corresponding to the second feature map is determined.

[0009] By using the above method, when performing feature extraction and feature fusion on bi-branch paths, only one weight is used for each feature map, which significantly reduces computational complexity and improves feature representation ability.

[0010] In one optional implementation, a target-gated feedforward network is used to perform feature enhancement processing on the fused feature map, including: Determine the feature scale corresponding to the fused feature map; The target gated feedforward network is selected based on the correspondence between the feature scale and the preset gated feedforward network type; A target-gated feedforward network is used to perform feature enhancement processing on the fused feature map to obtain the target fused feature map.

[0011] Using the above method, different gated feedforward networks can be adaptively selected for feature enhancement processing for fused feature maps with different feature complexities, enabling feature enhancement of fused feature maps with different feature scales.

[0012] In one optional implementation, feature fusion is performed on the first-scale feature map and the second-scale feature map using a first-scale weight and a second-scale weight to obtain a first-scale fused feature map, including: A first intermediate feature map is obtained based on a first-scale feature map and a first-scale weight, and a second intermediate feature map is obtained based on a second-scale feature map and a second-scale weight. The first and second intermediate feature maps are fused to obtain the third intermediate feature map. A target-gated feedforward network based on the third intermediate feature map is used to perform feature enhancement processing on the third intermediate feature map to obtain the fourth intermediate feature map. The fourth intermediate feature map is downsampled according to the preset K feature scales to obtain K downsampled feature maps; The first-scale fused feature map is obtained by applying channel attention weighted modulation to the K downsampled feature maps.

[0013] In one optional implementation, a target-gated feedforward network is used to perform feature enhancement processing on the fused feature map to obtain the target fused feature map, including: When the feature scale of the fused feature map is greater than a preset threshold, the fused feature map is subjected to Fourier transform to obtain frequency domain features. The amplitude component of the fused feature map is enhanced based on the frequency domain features to obtain the target amplitude component. Convert the target amplitude components into complex form and perform an inverse Fourier transform to obtain the target fusion feature map; or... When the feature scale of the fused feature map is less than a preset threshold, the fused feature map is modulated according to the third weight, and the modulated fused feature map is modulated by scaling parameters and / or translation parameters to obtain the target fused feature map.

[0014] In one optional implementation, the K downsampled feature maps are subjected to channel attention weighted modulation to obtain a first-scale fused feature map, including: Based on a preset number of channels, convolution operations are performed on the K downsampled feature maps to obtain K channel attention feature maps; The first-scale fused feature map is obtained by performing a weighted linear combination of the K channel attention feature maps.

[0015] By using the above method, features of different receptive fields can be extracted through multi-scale dilated convolution, which can enhance the network's adaptability to dark light regions of different scales.

[0016] Secondly, this application provides a low-light image enhancement device, the device comprising: The conversion module is used to acquire the image to be processed through the vehicle-mounted sensor and convert the image to be processed from the three primary color space to the target color space to obtain the first channel component and the second channel component. The processing module is used to extract a first feature map of the first channel component and determine a first weight corresponding to the first feature map, and to extract a second feature map of the first channel component and determine a second weight corresponding to the second feature map; The fusion module is used to fuse the first feature map and the second feature map using a first weight and a second weight to obtain a fused feature map. An enhancement module is used to output a target image based on the fused feature map; wherein the target image is an image obtained by performing feature enhancement processing on the fused feature map using a target gating network determined based on the fused feature map and converting it to a three-primary-color space.

[0017] In an optional implementation, when extracting the first feature map of the first channel component and determining the first weight corresponding to the first feature map, and extracting the second feature map of the second channel component and determining the second weight corresponding to the second feature map, the processing module is specifically used for: Extract the first scale feature map corresponding to the first channel component according to the first feature scale and determine the first scale weight corresponding to the first scale feature map; and extract the second scale feature map corresponding to the second channel component according to the first feature scale and determine the second scale weight corresponding to the second scale feature map. By using the first scale weight and the second scale weight, feature fusion is performed on the first scale feature map and the second scale feature map to obtain the first scale fused feature map. The first feature map corresponding to the first channel component is extracted from the first scale fused feature map according to the second feature scale, and the first weight corresponding to the first feature map is determined. The second feature map corresponding to the second channel component is extracted from the first scale fused feature map according to the second feature scale, and the second weight corresponding to the second feature map is determined.

[0018] In one optional implementation, when the target gated feedforward network performs feature enhancement processing on the fused feature map to obtain the target fused feature map, the enhancement module is specifically used for: Determine the feature scale corresponding to the fused feature map; The target gated feedforward network is selected based on the correspondence between the feature scale and the preset gated feedforward network type; A target-gated feedforward network is used to perform feature enhancement processing on the fused feature map to obtain the target fused feature map.

[0019] In an optional implementation, when fusing the first-scale feature map and the second-scale feature map using a first-scale weight and a second-scale weight to obtain a first-scale fused feature map, the processing module is specifically used for: A first intermediate feature map is obtained based on a first-scale feature map and a first-scale weight, and a second intermediate feature map is obtained based on a second-scale feature map and a second-scale weight. The first and second intermediate feature maps are fused to obtain the third intermediate feature map. A target-gated feedforward network based on the third intermediate feature map is used to perform feature enhancement processing on the third intermediate feature map to obtain the fourth intermediate feature map. The fourth intermediate feature map is downsampled according to the preset K feature scales to obtain K downsampled feature maps; The first-scale fused feature map is obtained by applying channel attention weighted modulation to the K downsampled feature maps.

[0020] In one optional implementation, when using a target-gated feedforward network to perform feature enhancement processing on the fused feature map to obtain the target fused feature map, the enhancement module is specifically used for: When the feature scale of the fused feature map is greater than a preset threshold, the fused feature map is subjected to Fourier transform to obtain frequency domain features. The amplitude component of the fused feature map is enhanced based on the frequency domain features to obtain the target amplitude component. Convert the target amplitude components into complex form and perform an inverse Fourier transform to obtain the target fusion feature map; or... When the feature scale of the fused feature map is less than a preset threshold, the fused feature map is modulated according to the third weight, and the modulated fused feature map is modulated by scaling parameters and / or translation parameters to obtain the target fused feature map.

[0021] In one optional implementation, when obtaining the first-scale fused feature map after passing K downsampled feature maps through channel attention weighted modulation, the processing module is specifically used for: Based on a preset number of channels, convolution operations are performed on the K downsampled feature maps to obtain K channel attention feature maps; The first-scale fused feature map is obtained by performing a weighted linear combination of the K channel attention feature maps.

[0022] Thirdly, this application provides an electronic device including a processor and a memory, wherein the memory stores program code that, when executed by the processor, causes the processor to perform the steps of the low-light image enhancement method described in the first aspect.

[0023] Fourthly, this application provides a computer-readable storage medium including program code that, when executed on an electronic device, causes the electronic device to perform the steps of the low-light image enhancement method described in the first aspect.

[0024] Fifthly, this application provides a computer program product that, when invoked by a computer, causes the computer to perform the steps of the low-light image enhancement method as described in the first aspect.

[0025] Furthermore, other features and advantages of this application will be set forth in the following description and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A schematic diagram of a suitable system architecture is provided for an embodiment of this application; Figure 2 A schematic diagram of the architecture of an in-vehicle edge computing platform provided in an embodiment of this application; Figure 3 A schematic diagram illustrating the implementation process of a low-light image enhancement method provided in this application embodiment; Figure 4 A schematic diagram of the structure of an encoder module provided in the application embodiment; Figure 5 A schematic diagram of a gated cross-attention module provided in an embodiment of this application; Figure 6 A schematic diagram of a gated cross-fusion module provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a decoder module provided in an embodiment of this application; Figure 8A schematic diagram of a depth-separable sensing module provided in an embodiment of this application; Figure 9 A schematic diagram of a gated feedforward network provided in an embodiment of this application; Figure 10 A schematic diagram of the structure of a FreMLP module in a gated feedforward network provided in an embodiment of this application; Figure 11 A schematic diagram of another gated feedforward network provided in the application embodiment; Figure 12 A schematic diagram of a dual-path processing architecture is provided in the application embodiment; Figure 13 This application provides a schematic diagram of the structure of a low-light image enhancement device according to an embodiment; Figure 14 This application provides a schematic diagram of the structure of an electronic device. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0028] It should be noted that in the description of this application, "multiple" is understood as "at least two". "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. A connected to B can represent: A and B directly connected, or A and B connected through C. Furthermore, in the description of this application, terms such as "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order.

[0029] Furthermore, the data collection, dissemination, and use in the technical solution of this application all comply with the requirements of relevant national laws and regulations.

[0030] The design concept of the embodiments of this application is briefly introduced below: Low-light image enhancement is a technique used to improve image quality, making the enhanced image more suitable for human or machine vision processing. In low-light environments, due to insufficient light, images obtained by imaging tools often suffer from poor global visibility, low contrast, color distortion, and local degradation. This significantly impacts image usability and visual quality, affecting not only human visual perception but also the performance of downstream advanced vision tasks such as object detection and semantic segmentation.

[0031] Therefore, in order to improve the quality of low-light images, low-light image enhancement processing has become a current research hotspot. Low-light image enhancement processing can improve image brightness, suppress noise interference, and correct color deviation, so that the processed image has a better visual effect.

[0032] In related technologies, low-light image enhancement lacks adaptability to local image content, easily leading to over-enhancement or loss of detail, and struggles to effectively handle complex non-uniform illumination degradation. Furthermore, these technologies often struggle to balance computational complexity with enhancement effectiveness; some methods with high enhancement results tend to have high computational complexity, resulting in significant overhead for high-resolution images and failing to meet real-time requirements. For example, when vehicles travel through tunnels or at night, the ambient light intensity drops sharply, resulting in low-quality images unsuitable for subsequent tasks such as target recognition and environmental detection. Therefore, image enhancement is necessary to obtain high-quality images. However, current image enhancement methods often suffer from high computational overhead, failing to meet real-time requirements. Thus, in scenarios with high real-time data requirements, such as intelligent driving or autonomous driving, the inability to perform image enhancement in real-time with low computational overhead can negatively impact driving safety.

[0033] In view of this, this application provides a low-light image enhancement method, which includes: first, acquiring an image to be processed through an on-board sensor and converting the image to be processed from a three-primary-color space to a target color space to obtain a first channel component and a second channel component; then, extracting a first feature map of the first channel component and determining a first weight corresponding to the first feature map, and extracting a second feature map of the second channel component and determining a second weight corresponding to the second feature map; second, performing feature enhancement processing on the first feature map and the second feature map using the first weight and the second weight to obtain a target fusion feature map; further, using a target-gated feedforward network determined based on the fusion feature map to perform feature enhancement processing on the fusion feature map to obtain a target fusion feature map; finally, converting the target fusion feature map from the target color space to a three-primary-color space to obtain a target image. This method can significantly improve the computational efficiency of low-light image enhancement processing, meet the real-time requirements of image processing, and significantly improve the effect of low-light image enhancement with low computational overhead compared to traditional cross-attention mechanisms.

[0034] It should be noted that the acquisition, storage, use, and processing of data in this application embodiment all comply with the relevant provisions of national laws and regulations.

[0035] See Figure 1 The diagram shown illustrates a system architecture according to an embodiment of this application. This system architecture includes a target terminal 101 and a server 102. The target terminal 101 and the server 102 can interact via a communication network. The communication network can employ wireless communication or wired communication methods.

[0036] For example, the target terminal 101 can access the network and communicate with the server 102 through cellular mobile communication technology, wherein the cellular mobile communication technology includes, for example, 5th generation mobile networks (5G) technology.

[0037] Optionally, the target terminal 101 can access the network and communicate with the server 102 via short-range wireless communication, wherein the short-range wireless communication method includes, for example, Wireless Fidelity (Wi-Fi) technology.

[0038] This application embodiment does not impose any limitation on the number of communication devices involved in the above system architecture. For example, there may be more target terminals, or no target terminals, or other network devices may be included, such as... Figure 1 As shown, only the target terminal 101 and server 102 are described as examples. The following is a brief introduction to each of the above devices and their respective functions.

[0039] The target terminal 101 is a device that can provide voice and / or data connectivity to a user, and may be a device that supports limited and / or wireless connectivity.

[0040] For example, the target terminal 101 includes, but is not limited to: mobile phones, tablets, laptops, handheld computers, mobile internet devices (MID), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminal devices in industrial control, wireless terminal devices in autonomous driving, wireless terminal devices in smart grids, wireless terminal devices in transportation safety, wireless terminal devices in smart cities, or wireless terminal devices in smart homes, etc.

[0041] Furthermore, a relevant client can be installed on the target terminal 101. This client can be software, such as an application (APP), browser, short video software, or a network element, mini-program, etc. In this embodiment, the target terminal 101 can use the aforementioned client related to low-light image enhancement to send the acquired image to be processed to the server 102 for subsequent low-light image enhancement methods.

[0042] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0043] It is worth noting that, in the embodiments of this application, the methods in these embodiments can be executed by an electronic device, which can be the target terminal 101 or the server 102. That is, the method can be executed by the terminal device 101 or the server 102 alone, or by the target terminal device 101 and the server 102 together. Furthermore, the executing entities for each method can be the same or different, and this embodiment of the application does not impose any restrictions on this.

[0044] For example, when the target terminal 101 executes the low-light image enhancement method provided in this application alone, the target terminal 101 can directly acquire the image to be processed and convert the image to be processed from the three primary color space to the target color space to obtain the first channel component and the second channel component; then, it extracts the first feature map of the first channel component and determines the first weight corresponding to the first feature map, and extracts the second feature map of the second channel component and determines the second weight corresponding to the second feature map; next, it performs feature fusion on the first feature map and the second feature map through the first weight and the second weight to obtain a fused feature map; further, it uses a target gated feedforward network determined based on the fused feature map to perform feature enhancement processing on the fused feature map to obtain a target fused feature map; finally, it converts the target fused feature map from the target color space to the three primary color space to obtain the target image.

[0045] For example, when the target terminal 101 and the server 102 jointly execute the low-light image enhancement method of this application, the target terminal 101 can acquire the image to be processed in real time and send it to the server 102. The server 102, based on the acquired image to be processed, converts the image to be processed from the three primary color space to the target color space to obtain the first channel component and the second channel component. Then, it extracts the first feature map of the first channel component and determines the first weight corresponding to the first feature map, and extracts the second feature map of the second channel component and determines the second weight corresponding to the second feature map. Next, it performs feature fusion on the first feature map and the second feature map through the first weight and the second weight to obtain a fused feature map. Further, it uses a target gated feedforward network determined based on the fused feature map to perform feature enhancement processing on the fused feature map to obtain a target fused feature map. Finally, it converts the target fused feature map from the target color space to the three primary color space to obtain the target image.

[0046] See Figure 2 As shown, the low-light enhancement method described in this application can be applied to an on-vehicle edge computing platform. Low-light images are acquired through on-vehicle sensors, and then the acquired low-light images are input to the on-vehicle edge computing platform. The low-light enhancement processing described in this application is performed on the acquired low-light images in the on-vehicle edge computing platform to obtain enhanced low-light images. The enhanced low-light images can then be displayed through an on-vehicle display device. The low-light images can also be used for subsequent downstream advanced vision tasks such as target detection and semantic segmentation.

[0047] The low-light image enhancement method provided by the exemplary embodiments of this application will now be described with reference to the accompanying drawings.

[0048] See Figure 3 The diagram shown illustrates the implementation flow of a low-light image enhancement method provided in this application. The specific implementation flow of this method is as follows: S1: Acquire the image to be processed through the vehicle-mounted sensor, and convert the image to be processed from the three primary color space to the target color space to obtain the first channel component and the second channel component.

[0049] In this embodiment, when a vehicle is driving through a tunnel or at night, the ambient light intensity drops drastically. If the vehicle's intelligent driving or autonomous driving functions are activated, the image data collected by the onboard sensors is often of low quality, exhibiting problems such as poor global visibility, low contrast, color distortion, and localized degradation, making it unsuitable for direct processing. Therefore, when a vehicle travels to an area with low light intensity, the image data collected by the onboard sensors needs to undergo feature enhancement processing to improve image quality and facilitate subsequent processing for different tasks. For example, vehicles are typically equipped with different types of cameras that can collect images of the area around the vehicle, such as front-view cameras, rear-view cameras, and surround-view cameras. These cameras can capture low-light images (i.e., images to be processed are obtained through onboard sensors) in low-light environments, facilitating subsequent low-light image enhancement processing.

[0050] Furthermore, after acquiring the image to be processed, in order to decouple the brightness and color information in the low-light image enhancement task and eliminate the inherent red discontinuity noise and black plane noise in the image to be processed in the three primary color space (in this embodiment, the three primary color space refers to the RGB three primary color space), it is necessary to convert the image to be processed from the three primary color space to the target color space to obtain the first channel component and the second channel component. In this embodiment, the first channel component can be the chromaticity component, that is, the channel component corresponding to the three HVI channels, and the second channel component can be the illuminance component, that is, the channel component corresponding to a single I channel. By extracting the HVI channel component and the I channel component respectively, the decoupling of brightness and color information is achieved. In the following description, HVI channel component is used to refer to the first channel component and I channel component is used to refer to the second channel component.

[0051] It should be noted that, in the embodiments of this application, the target color space can be a horizontal / vertical-intensity color space (HVI).

[0052] Specifically, in the embodiments of this application, when converting the image to be processed from the three primary color space to the HVI color space, the hue H of the image to be processed in the hue-saturation-value (HSV) color space can first be calculated. HSVSaturation (S) and brightness (V). Among them, hue (H) HSV Saturation (S) and brightness (V) can be calculated using the following formulas: (1) (2) hue_index; (3) Where max represents the maximum value, min represents the minimum value, R represents the red channel, G represents the green channel, and B represents the blue channel. Indicates hue index, To prevent small constants from being divided by zero.

[0053] Furthermore, the aforementioned hue index The relationship between the three channels and the maximum brightness value V can be determined as follows: (4) Then, the color sensitivity factor can be calculated using the learnable density parameter k, thereby determining the color sensitivity factor and hue H. HSV The three channels corresponding to the HVI color space are calculated from the saturation (S) and luminance (V). The specific calculation formulas for the color sensitivity factor, H, V, and I channels are as follows: (5) (6) (7) I=V; Where color_sensitive represents the color sensitivity factor, H HSV S represents the hue value, S represents the saturation value, V represents the brightness value, and k represents the learnable density parameter.

[0054] It is worth noting that in this embodiment, the learnable density parameter k is a scalar network parameter, automatically optimized through end-to-end training with a fixed initial value. The range of k is implicitly constrained by the optimizer and learning rate policy, without explicit hard boundary restrictions. Furthermore, the role of k is to modulate... The power of the function adaptively controls the steepness of the color sensitivity factor curve as a function of brightness. When the k value is small, the color sensitivity changes smoothly between dark and bright areas; when the k value is large, the color sensitivity changes significantly between dark and bright areas. This allows for adaptive adjustment of the color information preservation and enhancement strategy based on the data distribution. It should also be noted that after enhancing the low-light image, the same k value is used when inversely converting the enhanced low-light image from the HVI color space to the three primary color space, as used when converting from the three primary color space to the HVI color space, to ensure consistency between the forward and reverse color space transformations.

[0055] The above method converts the image to be processed from the three primary color space to the HVI color space, so as to decouple the brightness information and color information, eliminate the red discontinuity noise and black plane noise present in the three primary color space, thereby improving the robustness and accuracy of subsequent feature extraction.

[0056] S2: Extract the first feature map of the first channel component and determine the first weight corresponding to the first feature map, and extract the second feature map of the second channel component and determine the second weight corresponding to the second feature map.

[0057] In this embodiment, a gated cross-attention mechanism is used based on a dual-path encoder to extract features from the acquired HVI channel components and I channel components respectively, thereby obtaining a first feature map corresponding to the HVI channel components and a second feature map corresponding to the I channel components.

[0058] Specifically, after obtaining the HVI channel components and I channel components, the HVI channel components and I channel components can be input into the reference respectively. Figure 4 In the encoder module shown, further details can be found in the documentation. Figure 5 As shown, the encoder module includes multiple gated cross-attention modules (EGCAs), each connected in series. The EGCAs use a gated cross-attention mechanism to extract the first-scale feature map corresponding to the HVI channel component and the second-scale feature map corresponding to the I channel component according to the first feature scale.

[0059] For example, the first feature scale can be 1, 1 / 2, 1 / 4, or 1 / 8, and this application does not impose any restrictions on it in the embodiments.

[0060] Furthermore, after extracting the first-scale feature map corresponding to the HVI channel components and the second-scale feature map corresponding to the I channel components, it is necessary to determine the first-scale weights corresponding to the first-scale feature map and the second-scale weights corresponding to the second-scale feature map using the GCF module in EGCA. Further, see... Figure 6As shown, the GCF module can perform feature fusion on the first-scale feature map and the second-scale feature map using the first-scale weight and the second-scale weight to obtain the first-scale fused feature map.

[0061] In this embodiment, the HVI channel components are input into the gated cross-attention module, and a 3×3 convolutional layer is used to extract features from the HVI channel components to obtain the first scale feature map corresponding to the HVI channel components. The I channel components are also extracted to obtain the second scale feature map corresponding to the I channel components.

[0062] For example, in this embodiment of the application, a gated cross-attention mechanism is used to extract features from the HVI channel components and the I channel components, as expressed by the following expression: (8) (9) Where, x local Represents the first-scale feature map, y local The second-scale feature map is represented by x, which represents the HVI channel component, and y, which represents the I channel component.

[0063] Furthermore, in EGCA, the first-scale feature map also corresponds to a first-scale weight, and the second-scale feature map also corresponds to a second-scale weight. The first-scale feature map is modulated using the first-scale weight to obtain a first intermediate feature map, and the second-scale feature map is modulated using the second-scale weight to obtain a second intermediate feature map. In this embodiment, the first-scale feature map and the second-scale feature map are feature maps obtained from feature extraction using the same feature scale from the HVI channel component and the I channel component, respectively. Furthermore, the first-scale weight and the second-scale weight are also weights obtained based on the same feature scale for both paths.

[0064] In this embodiment, the first scale weight and the second scale weight are represented by the following expressions: (10) (11) in, Indicates the first-scale weight. Indicates the second-scale weight. , This represents the weights of a 1×1 convolutional layer with a compression ratio of 4. This represents the Sigmoid activation function.

[0065] It is worth noting that, in this embodiment, only one corresponding first-scale weight is used when modulating the first-scale feature map, and only one corresponding second-scale weight is used when modulating the second-scale feature map. Thus, the computational complexity is O(N), where N represents the height of the feature map. Wide. Compared to traditional cross-attention mechanisms, this approach maps two different inputs to Query(Q), Key(K), and Value(V) matrices through learnable linear projections, then determines the attention weights A based on the Q and K matrices, and finally multiplies the attention weights A with the V matrix to obtain the cross-attention output. The computational complexity of the traditional cross-attention mechanism is O(N). 2 The spatial resolution of the input is quadratic, while the computational complexity of the gated cross-attention mechanism used in this embodiment is only O(N), which significantly reduces the computational complexity of feature fusion.

[0066] Furthermore, such as Figure 5 As shown, feature fusion is performed on the first intermediate feature map and the second intermediate feature map to obtain the third intermediate feature map, which is represented by the following expression: (12) in, This represents the first-scale feature map. Indicates the first-scale weight. This represents the second-scale feature map. This represents the weight of the second scale.

[0067] Also see Figure 7 As shown, the decoder module also includes multiple gated cross-attention modules (EGCAs). Each EGCA is connected in series and upsamples the HVI and I channel components using different feature scales to obtain feature maps corresponding to the HVI and I channel components, respectively. The feature weights for each HVI and I channel component are then determined, thereby achieving cross-fusion of the HVI and I channel components. It should be noted that the EGCA in the decoder module works on the same principle as the EGCA in the encoder module; the only difference is that the upsampled feature maps are processed in the decoder.

[0068] Then, based on the difference in feature scale of the third feature map, different types of gated feedforward networks can be used to perform feature enhancement processing on the fused feature map, thereby obtaining the fourth feature map.

[0069] Furthermore, in this embodiment, a depth-separable scale-aware module is used to process the fourth intermediate feature map, and multi-scale feature fusion is achieved through multi-scale feature extraction and adaptive channel attention.

[0070] Specifically, the fourth intermediate feature map can be downsampled according to a preset K feature scale to obtain K downsampled feature maps. In this way, k downsampled feature maps can be obtained for each path in both the HVI and I channel components.

[0071] For example, in this embodiment, the fourth intermediate feature map can be downsampled using the original feature scale, half the feature scale, and quarter the feature scale, respectively, to obtain downsampled feature maps at the original feature scale, half the feature scale, and quarter the feature scale. It should be noted that this embodiment does not impose specific limitations on the feature scale; a suitable feature scale can be selected according to requirements to obtain the corresponding number of downsampled feature maps.

[0072] Optionally, in this embodiment, the depth-separable scale-aware module uses a multi-scale dilated convolution and dense connection strategy with a dilation rate of [1, 2, 1] or [1, 2, 3, 2, 1] to downsample the fourth intermediate feature map, obtaining K downsampled feature maps. The specific expression is: (13) (14) (15) (16) in, The void ratio is d. i Depth-separable convolution, This represents pointwise convolution.

[0073] Then, the fourth intermediate feature map is downsampled according to K feature scales. The obtained K downsampled feature maps are then subjected to channel attention weighted modulation to obtain the first scale fused feature map.

[0074] In one alternative implementation, see [reference] Figure 8 As shown, multiple depth-separable sensing modules are used to perform convolution operations on K downsampled feature maps based on a preset number of channels, thereby obtaining K channel attention feature maps.

[0075] Secondly, a weighted linear combination of the K channel attention feature maps is performed to obtain the first-scale fused feature map.

[0076] Specifically, when performing convolution operations on the K downsampled feature maps, each channel of the input downsampled feature map is convolved independently according to a preset number of channels. For example, if the input has C channels, then C convolution kernels are used, with each kernel operating on one channel. Thus, the number of output channels is the same as the number of input channels, and there is no information exchange between channels. Finally, based on the weights learned through channel attention, the K channel feature maps are weighted and linearly combined to obtain the first-scale fused feature map.

[0077] In this embodiment, the first-scale fused feature map is obtained through channel attention weighted modulation, which can be represented by the following expression: (17) (18) (19) (20) in, The weights are learned through channel attention, and Down and Up represent downsampling and upsampling operations, respectively.

[0078] After obtaining the first-scale fused feature map, the first feature map corresponding to the HVI channel component can be extracted from the first-scale fused feature map according to the second feature scale, and the first weight corresponding to the first feature map can be determined. The second feature map corresponding to the I channel component can be extracted from the first-scale fused feature map according to the second feature scale, and the second weight corresponding to the second feature map can be determined.

[0079] It should be noted that the processes described above for extracting the first and second feature maps and determining the first and second weights are the same as those for extracting the first and second scale feature maps and determining the first and second scale weights. In other words, obtaining the first scale fused feature map and subsequently extracting the first and second feature maps and determining the first and second weights are cascaded processes of EGCA. The first scale fused feature map is input into a second EGCA to extract the first and second feature maps and determine the first and second weights. Therefore, the specific steps for extracting the first and second feature maps and determining the first and second weights will not be detailed here; the detailed process is as described above.

[0080] S3: Using the first weight and the second weight, perform feature fusion on the first feature map and the second feature map to obtain a fused feature map.

[0081] In this embodiment of the application, after obtaining the first weight, the second weight, the first feature map, and the second feature map, the first feature map and the second feature map are weighted and fused. That is, the first feature map is modulated using the first weight, and the second feature map is modulated using the second weight. The modulated results of the two are added together to obtain the fused feature map.

[0082] It should be noted that the process of obtaining the fusion feature map described above is the same as the method of obtaining the fusion feature map in the aforementioned formula (12), and will not be repeated here.

[0083] S4: Output the target image based on the fused feature map.

[0084] It should be noted that, in the embodiments of this application, the target image is an image obtained by performing feature enhancement processing on the fused feature map using a target-gated feedforward network determined based on the fused feature map, and then converting it to the three primary color space.

[0085] Specifically, in the embodiments of this application, different types of gated feedforward networks can be used to perform feature enhancement processing on the fused feature maps according to the difference in feature scale of the fused feature maps. Different types of gated feedforward networks can perform different feature enhancement processing for different feature complexities of the fused feature maps.

[0086] In one optional implementation, firstly, the feature scale corresponding to the fused feature map is determined, for example, the feature scale of the fused feature map is 1, 1 / 2, 1 / 4, 1 / 8, etc. Then, a target gated feedforward network is selected according to the correspondence between the feature scale and the preset gated feedforward network type.

[0087] It should be noted that the correspondence between the feature scale and the preset gated feedforward network type is learned in advance through multiple iterations of training based on sample data. In specific applications, the target gated feedforward network can be directly selected based on the feature scale to perform feature enhancement processing on the fused feature map.

[0088] In one alternative implementation, when the fused feature map enters the shallow structure of the encoder, the corresponding features in the fused feature map are relatively complex. A traditional IEL-gated feedforward network can be used to enhance the fused feature map. First, based on the preset FFN expansion factor and the number of input channels, the number of expanded hidden channels is calculated. Then, a 1×1 convolution projects the number of input feature channels onto 2h (i.e., twice the number of input channels) to obtain intermediate features. Next, a 3×3 depthwise separable convolution is applied to the intermediate features, dividing the output result into two branches of the same shape along the channel dimension. Each branch is then subjected to a 3×3 depthwise separable convolution. The convolution result is passed through the Tanh activation function, and the activated features are added to the original input of that branch using residuals to obtain the enhanced intermediate features. Finally, the two enhanced intermediate features are multiplied element-wise to achieve interactive fusion of the two-branch information. A 1×1 convolution projects the number of channels back to the original number of input channels, thus obtaining the enhanced target fused feature map.

[0089] Specifically, the above process can be represented by the following expression: , ;(twenty one) , ;(twenty two) , ;(twenty three) ;(twenty four) in, This represents the preset FFN expansion factor, C is the number of input channels, DWconv3×3 represents a 3×3 depthwise separable convolution, and Chunk ( ,2) indicates that it is divided into two parts along the channel dimension. This indicates the sum of the residuals. This indicates element-wise multiplication.

[0090] Traditional IEL gated feedforward networks can enhance features such as brightness in the shallow structure of the encoder, better adapting to more complex features in the shallow structure while maintaining computational efficiency.

[0091] For further details, please refer to [link / reference]. Figure 9 , Figure 10As shown, when the feature scale of the fused feature map is greater than a preset threshold, for example, when the feature scale of the fused feature map is greater than 1 / 4, it means that the fused feature map enters the middle layer structure of the encoder. At this time, the feature complexity of the fused feature map is moderate. The FreMLP module in the target gated feedforward network can be used to perform Fourier transform on the fused feature map to obtain frequency domain features. Then, based on the frequency domain features, the amplitude component of the fused feature map is enhanced while keeping the phase unchanged to obtain the target amplitude component. Then, the target amplitude component is converted into a complex form and subjected to inverse Fourier transform to obtain the target fused feature map.

[0092] Optionally, the above process can be represented by the following expression: X = FFT(x); (25) (26) (27) (28) (29) Where FFT and IFFT represent Fast Fourier Transform and Inverse Fast Fourier Transform based on DFT matrix multiplication, respectively, M represents the amplitude spectrum of the fused feature map, and P represents the phase spectrum of the fused feature map.

[0093] Alternatively, when the feature scale of the fused feature map is less than a preset threshold, a third weight is used to modulate the fused feature map, and then the modulated fused feature map is modulated by scaling parameters and / or translation parameters to obtain the target fused feature map.

[0094] Specifically, see Figure 11 As shown, the fused feature map is modulated by a third weight. Then, the modulated fused feature map is linearly transformed by a 1×1 convolution, expanding the channel dimension to twice the original number of channels. Next, the expanded fused feature map is passed through a 3×3 depthwise separable convolution to modulate the scaling and / or translation parameters of the modulated fused feature map without changing the number of channels. Finally, a 1×1 convolution is passed to map the number of channels back to the original input dimension, resulting in the final feature transformation result, the target fused feature map.

[0095] Optionally, the above process can be represented by the following expression: , (30) (31) in, This represents a gating operation that divides the fused feature map into two equal parts along the channel dimension and then multiplies them element by element, where C represents the number of channels.

[0096] In this embodiment of the application, the enhanced target image can be obtained by converting the enhanced target fusion feature map from the target color space to the three primary color space.

[0097] Specifically, the inverse transformation process of the target fusion feature map first calculates the color sensitivity factor based on the learnable density parameter k and the brightness I, and then calculates the hue H based on the color sensitivity factor. HSV Saturation (S). Among these, color sensitivity factor and hue (H) HSV The saturation S is calculated using the following formulas: (32) (33) (34) in, The color sensitivity factor is represented by k, the learnable density parameter is represented by I, and the brightness value is represented by H in the HSV color space. HSV This represents the hue value in the HSV color space. S represents the saturation value in the HSV color space.

[0098] Finally, the enhanced target image is obtained by using the standard HSV color space to three primary color space conversion formula.

[0099] In summary, see the following: Figure 12 As shown, this application employs a dual-path processing architecture. The image to be processed is converted from the three primary color space to the target color space, obtaining HVI channel components and I channel components. These HVI and I channel components are used as inputs, and the encoder and decoder utilize multi-layer EGCA to achieve cross-fusion of the HVI and I channel components, thereby enhancing the low-light image. The processed HVI and I channel components are then output from their respective paths, undergo color space transformation, and are inversely converted from the target color space back to the three primary color space, resulting in the final enhanced image. This architecture enhances low-light images by structuring color and brightness information, making the enhancement process more targeted towards brightness information. The multi-scale feature extraction method can extract features from different receptive fields, enhancing adaptability to low-light regions at different scales. Furthermore, this application significantly reduces computational load and meets the real-time requirements of low-light image processing in various scenarios.

[0100] Furthermore, based on the same technical concept, embodiments of this application provide a low-light image enhancement device, which is used to implement the above-described method flow of the embodiments of this application. See also... Figure 13As shown, the device includes: a conversion module 1301, a processing module 1302, a fusion module 1303, an enhancement module 1304, and an inverse conversion module 1305, wherein, The conversion module 1301 is used to acquire the image to be processed through the vehicle-mounted sensor and convert the image to be processed from the three primary color space to the target color space to obtain the first channel component and the second channel component. The processing module 1302 is used to extract the first feature map of the first channel component and determine the first weight corresponding to the first feature map, and to extract the second feature map of the second channel component and determine the second weight corresponding to the second feature map; The fusion module 1303 is used to perform feature fusion on the first feature map and the second feature map using the first weight and the second weight to obtain a fused feature map; The enhancement module 1304 is used to output a target image based on the fused feature map; wherein the target image is an image obtained by performing feature enhancement processing on the fused feature map using a target gating network determined based on the fused feature map and converting it to a three-primary-color space.

[0101] In an optional implementation, when extracting the first feature map of the first channel component and determining the first weight corresponding to the first feature map, and extracting the second feature map of the second channel component and determining the second weight corresponding to the second feature map, the processing module 1202 is specifically used for: Extract the first scale feature map corresponding to the first channel component according to the first feature scale and determine the first scale weight corresponding to the first scale feature map; and extract the second scale feature map corresponding to the second channel component according to the first feature scale and determine the second scale weight corresponding to the second scale feature map. By using the first scale weight and the second scale weight, feature fusion is performed on the first scale feature map and the second scale feature map to obtain the first scale fused feature map. The first feature map corresponding to the first channel component is extracted from the first scale fused feature map according to the second feature scale, and the first weight corresponding to the first feature map is determined. The second feature map corresponding to the second channel component is extracted from the first scale fused feature map according to the second feature scale, and the second weight corresponding to the second feature map is determined.

[0102] In one optional implementation, the fused feature map is enhanced in a target-gated feedforward network, wherein the enhancement module 1304 is specifically used for: Determine the feature scale corresponding to the fused feature map; The target gated feedforward network is selected based on the correspondence between the feature scale and the preset gated feedforward network type; A target-gated feedforward network is used to perform feature enhancement processing on the fused feature map to obtain the target fused feature map.

[0103] In an optional implementation, when performing feature fusion on the first-scale feature map and the second-scale feature map using a first-scale weight and a second-scale weight to obtain a first-scale fused feature map, the processing module 1302 is specifically used for: A first intermediate feature map is obtained based on a first-scale feature map and a first-scale weight, and a second intermediate feature map is obtained based on a second-scale feature map and a second-scale weight. The first and second intermediate feature maps are fused to obtain the third intermediate feature map. A target-gated feedforward network based on the third intermediate feature map is used to perform feature enhancement processing on the third intermediate feature map to obtain the fourth intermediate feature map. The fourth intermediate feature map is downsampled according to the preset K feature scales to obtain K downsampled feature maps; The first-scale fused feature map is obtained by applying channel attention weighted modulation to the K downsampled feature maps.

[0104] In an optional implementation, when using a target-gated feedforward network to perform feature enhancement processing on the fused feature map to obtain the target fused feature map, the enhancement module 1304 is specifically used for: When the feature scale of the fused feature map is greater than a preset threshold, the fused feature map is subjected to Fourier transform to obtain frequency domain features. The amplitude component of the fused feature map is enhanced based on the frequency domain features to obtain the target amplitude component. Convert the target amplitude components into complex form and perform an inverse Fourier transform to obtain the target fusion feature map; or... When the feature scale of the fused feature map is less than a preset threshold, the fused feature map is modulated according to the third weight, and the modulated fused feature map is modulated by scaling parameters and / or translation parameters to obtain the target fused feature map.

[0105] In an optional implementation, when obtaining a first-scale fused feature map after passing K downsampled feature maps through channel attention weighted modulation, the processing module 1302 is specifically used for: Based on a preset number of channels, convolution operations are performed on the K downsampled feature maps to obtain K channel attention feature maps; The first-scale fused feature map is obtained by performing a weighted linear combination of the K channel attention feature maps.

[0106] Based on the same technical concept, embodiments of this application also provide an electronic device that can implement the low-light image enhancement method flow provided in the above embodiments of this application. In one embodiment, the electronic device may be a server, a terminal device, or other electronic device. See also... Figure 14 As shown, the electronic device may include: At least one processor 1401 and a memory 1402 connected to at least one processor 1401. In this embodiment, the specific connection medium between the processor 1401 and the memory 1402 is not limited. Figure 14 The example shown is the connection between processor 1401 and memory 1402 via bus 1400. Bus 1400 is... Figure 14 The connections between other components are shown in thick lines only and are not intended to be limiting. The Bus 1400 can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 14 The term is represented by a single thick line, but this does not imply that there is only one bus or one type of bus. Alternatively, the processor 1401 can also be called a controller; there is no restriction on the name.

[0107] In this embodiment, memory 1402 stores instructions executable by at least one processor 1401. By executing the instructions stored in memory 1402, at least one processor 1401 can execute a low-light image enhancement method described above. Processor 1401 can implement... Figure 13 The functions of each module in the device shown.

[0108] The processor 1401 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 1402 and calling data stored in memory 1402, the processor can perform various functions and process data, thereby monitoring the device as a whole.

[0109] In one possible design, processor 1401 may include one or more processing units. Processor 1301 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 1401. In some embodiments, processor 1401 and memory 1402 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.

[0110] Processor 1401 can be a general-purpose processor, such as a CPU, digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the low-light image enhancement method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0111] Memory 1402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 1402 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. Memory 1402 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 1402 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0112] By designing and programming the processor 1401, the code corresponding to the low-light image enhancement method described in the foregoing embodiments can be embedded into the chip, thereby enabling the chip to execute the code during operation. Figure 3 The illustrated embodiment presents the steps of a low-light image enhancement method. How to design and program the processor 1401 is a technique well-known to those skilled in the art and will not be described further here.

[0113] Based on the same inventive concept, embodiments of this application also provide a storage medium storing computer instructions that, when executed on a computer, cause the computer to perform a low-light image enhancement method described above.

[0114] In some possible implementations, this application also provides a low-light image enhancement method that can also be implemented as a program product including program code that, when the program product is run on a device, causes the control device to perform the steps in a low-light image enhancement method according to various exemplary embodiments of this application as described above.

[0115] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0116] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0117] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0118] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable low-light image enhancement device to produce a server, such that the instructions, which execute via the processor of the computer or other programmable low-light image enhancement device, generate instructions for implementing the process... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0119] Program code for performing the operations of this application can be written using any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0120] These computer program instructions can also be loaded onto a computer or other programmable low-light image enhancement device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0121] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for enhancing low-light images, characterized in that, The method includes: The image to be processed is acquired by the vehicle-mounted sensor, and the image to be processed is converted from the three primary color space to the target color space to obtain the first channel component and the second channel component. Extract a first feature map of the first channel component and determine a first weight corresponding to the first feature map; extract a second feature map of the second channel component and determine a second weight corresponding to the second feature map. By using a first weight and a second weight, feature fusion is performed on the first feature map and the second feature map to obtain a fused feature map; Based on the fused feature map, a target image is output; wherein the target image is an image obtained by performing feature enhancement processing on the fused feature map using a target gating network determined based on the fused feature map, and then converting it to a three-primary-color space.

2. The method as described in claim 1, characterized in that, The steps of extracting the first feature map of the first channel component and determining the first weight corresponding to the first feature map, and extracting the second feature map of the second channel component and determining the second weight corresponding to the second feature map, include: Extract the first scale feature map corresponding to the first channel component according to the first feature scale and determine the first scale weight corresponding to the first scale feature map; and extract the second scale feature map corresponding to the second channel component according to the first feature scale and determine the second scale weight corresponding to the second scale feature map. By using the first scale weight and the second scale weight, feature fusion is performed on the first scale feature map and the second scale feature map to obtain the first scale fused feature map; The first feature map corresponding to the first channel component is extracted from the first scale fused feature map according to the second feature scale, and the first weight corresponding to the first feature map is determined. The second feature map corresponding to the second channel component is extracted from the first scale fused feature map according to the second feature scale, and the second weight corresponding to the second feature map is determined.

3. The method as described in claim 1, characterized in that, The step of using a target-gated feedforward network determined based on the fused feature map to perform feature enhancement processing on the fused feature map includes: Determine the feature scale corresponding to the fused feature map; The target gated feedforward network is selected based on the correspondence between the feature scale and the preset gated feedforward network type; The target-gated feedforward network is used to perform feature enhancement processing on the fused feature map to obtain the target fused feature map.

4. The method as described in claim 2, characterized in that, The step of fusing features from the first-scale feature map and the second-scale feature map using the first-scale weight and the second-scale weight to obtain the first-scale fused feature map includes: A first intermediate feature map is obtained based on the first scale feature map and the first scale weight, and a second intermediate feature map is obtained based on the second scale feature map and the second scale weight; The first intermediate feature map and the second intermediate feature map are fused to obtain a third intermediate feature map; A target-gated feedforward network determined based on the third intermediate feature map is used to perform feature enhancement processing on the third intermediate feature map to obtain a fourth intermediate feature map. The fourth intermediate feature map is downsampled according to K preset feature scales to obtain K downsampled feature maps; The first scale fused feature map is obtained by applying channel attention weighted modulation to the K downsampled feature maps.

5. The method as described in claim 3, characterized in that, The step of using the target-gated feedforward network to perform feature enhancement processing on the fused feature map to obtain the target fused feature map includes: When the feature scale of the fused feature map is greater than a preset threshold, a Fourier transform is performed on the fused feature map to obtain frequency domain features. Based on the frequency domain features, the amplitude component of the fused feature map is enhanced to obtain the target amplitude component. The target amplitude component is converted into a complex form and then subjected to an inverse Fourier transform to obtain the target fusion feature map; or, When the feature scale of the fused feature map is less than a preset threshold, the fused feature map is modulated according to the third weight, and the modulated fused feature map is modulated by scaling parameters and / or translation parameters to obtain the target fused feature map.

6. The method as described in claim 4, characterized in that, The step of obtaining the first scale fused feature map by performing channel attention weighted modulation on the K downsampled feature maps includes: Based on a preset number of channels, convolution operations are performed on the K downsampled feature maps to obtain K channel attention feature maps; The first scale fusion feature map is obtained by performing a weighted linear combination of the K channel attention feature maps.

7. A low-light image enhancement device, characterized in that, The device includes: The conversion module is used to acquire the image to be processed through the vehicle-mounted sensor and convert the image to be processed from the three primary color space to the target color space to obtain the first channel component and the second channel component. The processing module is used to extract a first feature map of the first channel component and determine a first weight corresponding to the first feature map, and to extract a second feature map of the second channel component and determine a second weight corresponding to the second feature map; The fusion module is used to perform feature fusion on the first feature map and the second feature map using a first weight and a second weight to obtain a fused feature map; An enhancement module is used to output a target image based on the fused feature map; wherein the target image is an image obtained by performing feature enhancement processing on the fused feature map using a target gating network determined based on the fused feature map and converting it to a three-primary-color space.

8. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes: computer program code, which, when run on a computer, causes the computer to perform the method described in any one of claims 1-6.