Infrared dark light enhancement and model training method, device, equipment, medium and product

By extracting and processing low-light infrared images and generating multi-channel images using basic blocks, the problem of low infrared dark light enhancement efficiency is solved and high-quality infrared image processing is achieved.

CN120219192APending Publication Date: 2025-06-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510193106.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently achieve infrared dark light enhancement, resulting in poor quality of low-light infrared images.

Method used

By performing feature extraction of single-channel low-light infrared images, the initial feature map is processed using basic blocks (including attention networks and feedforward networks), multi-channel images are generated, and normal light infrared images are obtained by combining low-light infrared images.

Benefits of technology

It realizes efficient processing of infrared dark light enhancement, improving the quality and visual effects of infrared images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219192A_ABST
    Figure CN120219192A_ABST
Patent Text Reader

Abstract

The invention provides an infrared dark light enhancement and model training method and device, equipment, a medium and a product, and relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like. The infrared dark light enhancement method comprises the following steps: carrying out feature extraction on a single-channel low-light infrared image to obtain an initial feature map; processing the initial feature map based on a basic block to obtain a target feature map; acquiring a multi-channel image based on the target feature map; and obtaining a normal light infrared image based on the low light infrared image and the multi-channel image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to technical fields such as computer vision, deep learning, and large models. In particular, it relates to an infrared low-light enhancement and model training method, apparatus, device, medium, and product. Background Art

[0002] Infrared low-light enhancement is to enhance low-light infrared images to obtain normal-light infrared images. How to efficiently achieve infrared low-light enhancement is a problem that needs to be solved. Summary of the Invention

[0003] The present disclosure provides an infrared low-light enhancement method, apparatus, device, medium, and product.

[0004] According to one aspect of the present disclosure, there is provided an infrared low-light enhancement method, including: extracting features from a single-channel low-light infrared image to obtain an initial feature map; processing the initial feature map based on a basic block to obtain a target feature map; obtaining a multi-channel image based on the target feature map; and obtaining a normal-light infrared image based on the low-light infrared image and the multi-channel image.

[0005] According to another aspect of the present disclosure, there is provided an infrared low-light enhancement model training method, including: using a teacher model to process a multi-channel first sample image to obtain a multi-channel first output image; using a student model to process a single-channel second sample image to obtain a multi-channel second output image; the second sample image is obtained based on the first sample image; constructing a total loss function based on the first output image, the second output image, and the normal-light image corresponding to the first sample image; and adjusting the model parameters of the student model based on the total loss function to obtain an infrared low-light enhancement model.

[0006] According to another aspect of the present disclosure, there is provided an infrared low-light enhancement apparatus, including: an extraction module for extracting features from a single-channel low-light infrared image to obtain an initial feature map; a processing module for processing the initial feature map using a basic block to obtain a target feature map; an acquisition module for obtaining a multi-channel image according to the target feature map; and an enhancement module for obtaining a normal-light infrared image according to the low-light infrared image and the multi-channel image.

[0007] According to another aspect of the present disclosure, there is provided an infrared low-light enhancement model training device, including: a first processing module, configured to process a multi-channel first sample image by using a teacher model to obtain a multi-channel first output image; a second processing module, configured to process a single-channel second sample image by using a student model to obtain a multi-channel second output image; the second sample image is obtained based on the first sample image; a construction module, configured to construct a total loss function according to the first output image, the second output image, and the normal light image corresponding to the first sample image; an adjustment module, configured to adjust the model parameters of the student model according to the total loss function to obtain an infrared low-light enhancement model.

[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of the above aspects.

[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of the above aspects.

[0010] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, which implements the method according to any one of the above aspects when executed by a processor.

[0011] According to the embodiments of the present disclosure, infrared low-light enhancement can be efficiently achieved.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0014] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;

[0015] Figure 2 is a schematic diagram of an implementation system for implementing the embodiments of the present disclosure;

[0016] Figure 3 is a schematic diagram of an infrared low-light enhancement model provided according to the embodiments of the present disclosure;

[0017] Figure 4 is a schematic diagram of a basic block provided according to an embodiment of the present disclosure;

[0018] Figure 5 is a schematic diagram of an SGAU provided according to an embodiment of the present disclosure;

[0019] Figure 6 is a schematic diagram according to a second embodiment of the present disclosure;

[0020] Figure 7 is a schematic diagram according to a third embodiment of the present disclosure;

[0021] Figure 8 is a schematic diagram according to a fourth embodiment of the present disclosure;

[0022] Figure 9 is a comparison schematic diagram of the infrared enhancement effects between the method provided according to an embodiment of the present disclosure and other methods;

[0023] Figure 10 is a schematic diagram according to a fifth embodiment of the present disclosure;

[0024] Figure 11 is a schematic diagram according to a sixth embodiment of the present disclosure;

[0025] Figure 12 is a schematic diagram of an electronic device for implementing the infrared low-light enhancement method or the infrared low-light enhancement model training method according to an embodiment of the present disclosure. Detailed implementation manners

[0026] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0027] To efficiently implement infrared low-light enhancement, the present disclosure provides the following embodiments.

[0028] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure. This embodiment provides an infrared low-light enhancement method. As Figure 1 shown, the method includes:

[0029] 101. Extract features from a single-channel low-light infrared image to obtain an initial feature map.

[0030] 102. Process the initial feature map to obtain a target feature map.

[0031] 103. Obtain a multi-channel image based on the target feature map.

[0032] 104. Based on the low-light infrared image and the multi-channel image, obtain a normal-light infrared image.

[0033] Wherein, the basic block includes: an attention network and a feed-forward network, and the attention network is a residual network composed of an attention layer and a convolutional layer, and the feed-forward network is a residual network composed of a gated unit.

[0034] In the infrared low-light enhancement scenario, the image to be processed is a low-light infrared image, and the low-light infrared image is a single-channel image, and a normal-light infrared image after enhancement needs to be obtained.

[0035] The initial feature map refers to the feature map obtained after feature extraction of the low-light infrared image. Specifically, a convolutional network can be used to process the low-light infrared image to obtain the initial feature map.

[0036] The target feature map refers to the feature map obtained after processing the initial feature map.

[0037] The multi-channel image is an image generated based on the target feature map, and this image is multi-channel, for example, a three-channel image.

[0038] After obtaining the multi-channel image, based on the low-light infrared image and the multi-channel image, obtain a normal-light infrared image.

[0039] In this embodiment, obtaining the target feature map based on the initial feature map can improve the feature representation ability. Furthermore, obtaining the multi-channel image based on the target feature map and obtaining the normal-light infrared image based on the low-light infrared image and the multi-channel image can efficiently achieve infrared low-light enhancement.

[0040] In some embodiments, the initial feature map can be processed based on the basic block.

[0041] The basic block is the basic unit for processing the initial feature map, and the specific number can be set according to actual needs.

[0042] The basic block includes: an attention network and a feed-forward network (FFN).

[0043] Both the attention network and the FFN are residual networks, and the attention network is composed of an attention layer and a convolutional layer, and the FFN is composed of a gated unit.

[0044] The Residual Network (ResNet) is a network architecture in deep learning. It generally consists of two parts, namely the shortcut path and the residual path. The shortcut path usually does not contain non-linear transformations; the residual path usually contains non-linear transformations such as multiple convolutional layers. The output of the residual path is added to the output of the shortcut path to obtain the final output. Expressed by the formula:

[0045] H(x) = F(x) + A(x)

[0046] Where x is the input of the residual network, H(x) is the output of the residual network, F(x) is the output of the residual path, A(x) is the output of the shortcut path. If the shortcut path is a direct path, then A(x) = x.

[0047] In this embodiment, the attention network is a residual network composed of an attention layer and a convolutional layer. Among them, the residual path includes an attention layer and a convolutional layer, and the shortcut path is a direct path.

[0048] FFN is a residual network built based on gated units. Among them, the residual path includes gated units, and the shortcut path is a direct path.

[0049] In this embodiment, the attention network is a residual network composed of an attention layer and a convolutional layer, which can introduce the attention mechanism of the Transformer network into the Convolutional Neural Networks (CNNs). In this way, it can improve the network's ability to extract feature maps and enhance the infrared low-light enhancement effect. FFN is a residual network composed of gated units, which can simplify the network structure, improve the representation ability while being lightweight, and enhance the infrared low-light enhancement effect. Therefore, infrared low-light enhancement can be efficiently and lightly realized.

[0050] To better understand the present disclosure, the application scenarios related to the present disclosure are described as follows:

[0051] Figure 2 It is a schematic diagram of the implementation system for implementing the embodiments of the present disclosure.

[0052] As Figure 2 shown, an application (APP) for infrared low-light enhancement can be installed on the user terminal 201. Taking the example that the user can upload a low-light infrared image through this APP for infrared low-light enhancement locally on the user terminal, after the user terminal obtains the low-light infrared image, it can use a pre-deployed infrared low-light enhancement model to process the low-light infrared image to obtain a normal-light infrared image after low-light enhancement. After obtaining the normal-light infrared image, it can be displayed to the user through the APP.

[0053] Taking the local deployment of the infrared low-light enhancement model on the user terminal as an example, it can be understood that the infrared low-light enhancement model can also be deployed on the server. The user terminal interacts with the server, sends the low-light infrared image to the server, and the server performs enhancement processing on the low-light infrared image based on the infrared low-light enhancement model to obtain a normal-light infrared image, which is then returned to the user terminal for display.

[0054] Figure 3 It is a schematic diagram of the infrared low-light enhancement model provided according to an embodiment of the present disclosure.

[0055] As Figure 3 shown, the infrared low-light enhancement model includes: an input-end convolutional block 301, a basic block 302 (only one is marked for simplicity), an output-end convolutional block 303, a format conversion module 304, and a single-channel extraction module 305.

[0056] The input-end convolutional block 301 is used to extract features from the low-light infrared image to obtain an initial feature map. Among them, the low-light infrared image is a single-channel image, and its size is represented by H*W*1, where H and W are the height and width respectively.

[0057] The basic block 302 is used to process the initial feature map to obtain a target feature map.

[0058] The output-end convolutional block 303 is used to process the target feature map to obtain a multi-channel image. Taking three channels as an example, the size of the multi-channel image is represented by H*W*3.

[0059] The format conversion module 304 is used to convert the initial format image into a target format image. In this embodiment, taking the conversion from the RGB format to the YCbCr format as an example, where both the RGB format and the YCbCr format are color standards. The RGB format includes three channels: red (R), green (G), and blue (B), and the YCbCr format includes three channels: luminance (Y), blue chrominance (Cb), and red chrominance (Cr).

[0060] The single-channel extraction module 305 is used to extract a preset channel image from the target format image. In this embodiment, taking the extraction of the Y channel image of the YCbCr format image as an example, the Y channel image is used as the normal-light infrared image after infrared low-light enhancement. The size of the normal-light infrared image is the same as that of the low-light infrared image, represented by H*W*1.

[0061] In terms of specific composition, the input-end convolutional block and the output-end convolutional block can adopt two-dimensional 3x3 convolution (Conv2d3x3).

[0062] Functionally, the basic block can be divided into an encoding function and a decoding function. Therefore, the basic block can be divided into an encoding block and a decoding block. Correspondingly, based on the encoding block, the initial feature map can be encoded to obtain an intermediate feature map; based on the decoding block, the intermediate feature map can be decoded to obtain the target feature map.

[0063] For example, in Figure 3 the left part, that is, the basic block corresponding to the downsampling module can be called an encoding block. In Figure 3 the right part of, that is, the basic block corresponding to the upsampling module and the convolutional block at the output end can be called a decoding block. The encoding block is used to encode the initial feature map to obtain an intermediate feature map; the decoding block is used to decode the intermediate feature map to obtain the target feature map.

[0064] In this way, through the encoding and decoding process, the initial feature map can be processed to obtain the target feature map, improving the representation effect of the target feature map and the processing efficiency.

[0065] To improve the model effect, feature maps with different resolutions can be obtained and the feature maps with different resolutions can be fused. Therefore, as Figure 3 shown, the infrared low-light enhancement model also includes a downsampling module and an upsampling module. Through the up and down sampling modules, feature maps with multiple resolutions can be obtained.

[0066] The number of basic blocks and up and down sampling modules can be set according to actual needs. These modules can form a U-shaped network (Unet) structure.

[0067] In this embodiment, to ensure real-time performance, both the downsampling module and the upsampling module are 3, and the number of basic blocks under each resolution feature map is 1. In addition, the network width can be selected as 8, and the upsampling and downsampling coefficients are both 2, which can reduce the redundant parameters of the network and thus ensure the real-time performance of the model.

[0068] Assume that the initial resolution (the resolution of the initial feature map) is represented by 1, and the up and down sampling coefficients are both 2. Then, as Figure 3 shown, feature maps with resolutions of 1, 1 / 2, 1 / 4, and 1 / 8 can be obtained respectively.

[0069] After obtaining feature maps with different resolutions, the feature maps with different resolutions can be fused. As Figure 3 shown, for the feature maps from the encoding block and the feature maps from the upsampling module with the same resolution, they are fused, such as element-wise addition, to obtain a fused feature map, and then the fused feature map is decoded using the decoding block corresponding to the resolution.

[0070] The downsampling module can be implemented by a 3x3 convolution and a Rectified Linear Unit (ReLU) function. ReLU is a commonly used activation function in artificial neural networks, usually referring to non-linear functions represented by the ramp function and its variants.

[0071] The upsampling module can be implemented by a 1x1 convolution and a resize operator. The resize operator adopts nearest neighbor interpolation to improve the real-time performance of the model. Nearest neighbor interpolation means that for a pixel to be calculated, the value of the other pixel closest to it is used as the value of the pixel to be calculated.

[0072] Figure 4 It is a schematic diagram of the basic block provided according to an embodiment of the present disclosure.

[0073] As Figure 4 shown, the basic block includes: an attention network and an FFN.

[0074] The input feature map of the basic block is first processed by the attention network, and the feature map output by the attention network is input into the FFN. After being processed by the FFN, the output feature map of the basic block is obtained.

[0075] The attention network is a residual network, and its residual path includes: a convolutional layer and an attention layer; its shortcut path is a direct path, represented by a straight line.

[0076] As Figure 4 shown, the attention layer can specifically adopt a channel attention layer.

[0077] The channel attention layer processes the feature map based on the channel attention mechanism.

[0078] The channel attention mechanism assigns a weight to each feature map channel, and this weight reflects the importance of the channel for the final task. The input feature map is multiplied by the corresponding weight to obtain a weighted feature map. In this way, important channel feature maps are amplified, while unimportant channel feature maps are suppressed.

[0079] As Figure 4 shown, the convolutional layer can specifically include: a common convolutional layer and a depthwise separable convolutional layer (Depthwiseconv2d). The common convolutional layer is represented by Conv2d 3x3, and the depthwise separable convolutional layer is represented by DWConv2d 3x3.

[0080] The ordinary convolutional layer performs convolutional operations on each channel of the input feature map separately, and then accumulates the results to obtain the output feature map. The depthwise separable convolution divides this process into two steps: depthwise convolution and pointwise convolution. Depthwise convolution performs convolution operations on each channel of the input using an independent convolutional kernel respectively; pointwise convolution uses a 1x1 convolutional kernel to perform convolution operations on the output of the depthwise convolution.

[0081] Compared with ordinary convolution, depthwise separable convolution can reduce the amount of computation and the number of model parameters, thereby improving the training and inference speed of the model.

[0082] The activation function uses the LReLU function. The LReLU (Leaky ReLU) function is a variant of the rectified linear unit (ReLU) function, which can improve the non-linear representation ability of the network compared with the ReLU function.

[0083] After the input feature map of the attention network is processed by the convolutional layer and the channel attention layer, the output feature map of the residual path is obtained. After adding the output feature map of the residual path to the input feature map of the attention network, the output feature map of the attention network is obtained, and this output feature map is used as the input feature map of the FFN and input into the FFN.

[0084] The FFN is a residual network based on gated units.

[0085] As Figure 4 shown, the shortcut path of the FFN is a direct path, represented by a straight line; the residual path includes a convolutional layer and a gated unit.

[0086] The convolutional layer includes: a depthwise separable convolutional layer (DWConv2d 3x3) and an ordinary convolutional layer (Conv2d 1x1).

[0087] The gated unit is specifically a simplified gated attention unit (SGAU).

[0088] Figure 5 is a schematic diagram of the SGAU provided according to an embodiment of the present disclosure.

[0089] As Figure 5 shown, the size of the input feature map of the SGAU is H*W*C (the height is H, the width is W, and the number of channels is C), which is divided into two feature maps along the channel direction. One feature map passes through the Swish activation function to obtain the activated feature map, and then is multiplied element-wise with the other feature map to obtain the output feature map of the SGAU, with a size of H*W*C / 2.

[0090] The Swish activation function is a non-linear activation function that can improve the performance of neural networks, thereby enhancing the efficiency of machine learning.

[0091] SGAU aims to achieve lightweight by splitting the feature map along the channel direction, and at the same time realizes the gating mechanism through the Swish activation function, thereby dynamically adjusting the influence of different feature layers, helping the network to better select useful information, and essentially being an attention mechanism to enhance the network's representation ability.

[0092] Combined with the above application scenarios, the present disclosure also provides the following embodiments.

[0093] Figure 6 FIG. is a schematic diagram according to the second embodiment of the present disclosure. This embodiment provides an infrared low-light enhancement method. Taking a multi-channel image as a three-channel RGB image as an example, the method includes:

[0094] 601. Extract features from a single-channel low-light infrared image to obtain an initial feature map.

[0095] For example, referring to Figure 3 , a convolutional block at the input end can be used to perform convolutional processing on the low-light infrared image to obtain an initial feature map. The low-light infrared image is single-channel and its size can be expressed as H*W*1.

[0096] 602. Process the initial feature map based on the basic block to obtain a target feature map.

[0097] Wherein, the basic block includes: an attention network and a feed-forward network, and the attention network is a residual network based on an attention layer and a convolutional layer, and the feed-forward network is a residual network based on a gated unit.

[0098] Specifically, based on the initial feature map, an input feature map of the attention network can be obtained; the attention network is used to process the input feature map of the attention network to obtain an output feature map of the attention network, which is used as the input feature map of the feed-forward network; the feed-forward network is used to process the input feature map of the feed-forward network to obtain an output feature map of the feed-forward network; based on the output feature map of the feed-forward network, the target feature map is obtained.

[0099] For example, when there is one basic block, the initial feature map is used as the input feature map of the attention network, the output feature map of the attention network is used as the input feature map of the feed-forward network, and the output feature map of the feed-forward network is used as the target feature map. Or,

[0100] When there are multiple basic blocks, the initial feature map is used as the input feature map of the attention network in the first basic block; inside each basic block, the output feature map of the attention network is used as the input feature map of the feed-forward network; the output feature map of the feed-forward network in non-last basic blocks is used as the input feature map of the attention network in its adjacent next basic block, or, in the scenario of multi-resolution fusion, after downsampling or upsampling the output feature map of the feed-forward network in non-last basic blocks, it is used as the input feature map of the attention network in its adjacent next basic block; and so on, the output feature map of the feed-forward network in the last basic block is used as the target feature map.

[0101] Reference Figure 4 , inside a single basic block, the attention network and the feed-forward network (FFN) are connected in series. The input feature map of the attention network is obtained based on the initial feature map. The output feature map of the attention network is used as the input feature map of the FFN, and the target feature map is obtained based on the output feature map of the FFN.

[0102] In this embodiment, the target feature map is obtained by using the attention network and the feed-forward network connected in series, which can improve the representation ability of the target feature map, and further improve the infrared low-light enhancement effect.

[0103] Furthermore, for the attention network: the attention layer and the convolutional layer can be used to process the input feature map of the attention network to obtain the residual feature map of the attention network; the input feature map of the attention network and the residual feature map of the attention network are added together to obtain the output feature map of the attention network.

[0104] For example, a residual path can be composed of a convolutional layer and an attention layer. The output feature map of the residual path can be called the residual feature map. After adding the input feature map of the attention network and the residual feature map, the output feature map of the attention network is obtained.

[0105] Specifically, as Figure 4 shown, the attention layer can be a channel attention layer; the convolutional layer includes: a depthwise separable convolutional layer.

[0106] The channel attention mechanism assigns a weight to each feature map channel, and this weight reflects the importance of the channel for the final task. The input feature map is multiplied by the corresponding weight to obtain a weighted feature map. In this way, important channel feature maps are amplified, while unimportant channel feature maps are suppressed.

[0107] The ordinary convolutional layer performs convolutional operations on each channel of the input feature map separately, and then accumulates the results to obtain the output feature map. The depthwise separable convolutional layer divides this process into two steps: depthwise convolution and pointwise convolution. Depthwise convolution performs convolutional operations on each channel of the input using an independent convolutional kernel respectively; pointwise convolution uses a 1x1 convolutional kernel to perform convolutional operations on the output of the depthwise convolution.

[0108] Compared with ordinary convolution, depthwise separable convolution can reduce the amount of computation and the number of model parameters, thereby improving the training and inference speed of the model.

[0109] In this embodiment, obtaining the output feature map of the attention network based on the input feature map and the residual feature map of the attention network can simply and efficiently obtain the output feature map of the attention network, and then simply and efficiently obtain the target feature map.

[0110] In addition, by using a channel attention layer to focus on the important channels of the feature map, the feature map representation ability is improved, and thus the infrared low-light enhancement effect is improved. Through the depthwise separable convolutional layer, the amount of computation can be reduced and the processing efficiency can be improved.

[0111] For the FFN: Using the gating unit to process the input feature map of the feed-forward network to obtain the residual feature map of the feed-forward network; adding the input feature map of the feed-forward network and the residual feature map of the feed-forward network to obtain the output feature map of the feed-forward network.

[0112] For example, referring to Figure 4 , a residual path can be composed of a gating unit and a convolutional layer. The output feature map of the residual path can be called the residual feature map. After adding the input feature map of the feed-forward network and the residual feature map, the output feature map of the feed-forward network is obtained.

[0113] In this embodiment, obtaining the output feature map of the feed-forward network based on the input feature map and the residual feature map of the feed-forward network can simply and efficiently obtain the output feature map of the feed-forward network, and then simply and efficiently obtain the target feature map.

[0114] Further, for the gating unit, the input feature map of the gating unit can be split into a first sub-feature map and a second sub-feature map; using a preset activation function to perform activation processing on the first sub-feature map to obtain a weighted feature map; weighting the second sub-feature map based on the weighted feature map to obtain the output feature map of the gating unit.

[0115] For example, referring to Figure 5, the first sub-feature map and the second sub-feature map are obtained by splitting by channel, and both have a size of H*W*C / 2. The feature map obtained after processing the first sub-feature map through a preset activation function (such as the Swish activation function) is called the weighted feature map. This weighted feature map is weighted with the second sub-feature map, such as element-wise multiplication, to obtain the output feature map of the gating unit.

[0116] In this embodiment, splitting the feature map can simplify the operation and achieve lightweight. Through the processing of the activation function, the gating mechanism can be realized, the feature representation ability of the feature map can be improved, and thus the infrared low-light enhancement effect can be enhanced.

[0117] 603. Obtain a three-channel RGB image based on the target feature map.

[0118] For example, refer to Figure 3 , use the output-end convolutional block to process the target feature map to obtain a three-channel RGB image, and its size is expressed as H*W*3.

[0119] 604. Add the low-light infrared image and the three-channel RGB image to obtain an initial format image.

[0120] Among them, when adding, the low-light infrared image can be filled into a three-channel image, so that the initial format image with a size of H*W*3 is obtained after addition.

[0121] 605. Convert the initial format image to a target format image.

[0122] 606. Extract the preset channel image of the target format image as the normal-light infrared image.

[0123] For example, refer to Figure 3 , the initial format image is an RGB image, and the target format image is a YCbCr image. After converting the RGB image to a YCbCr image, extract the Y channel image as the normal-light infrared image.

[0124] In this embodiment, through format conversion and single-channel extraction, the normal-light infrared image can be obtained efficiently and accurately.

[0125] Figure 7 is a schematic diagram according to the third embodiment of the present disclosure. This embodiment provides an infrared low-light enhancement method. Taking the fusion of multi-resolution feature maps as an example, this method includes:

[0126] 701. Use the input-end convolutional block in the infrared low-light enhancement model to perform feature extraction on the low-light infrared image to obtain an initial feature map.

[0127] For example, refer to Figure 3, input a low-light infrared image with a size of H*W*1 into the input convolutional block, and the output is the initial feature map.

[0128] 702. Use multiple encoding blocks in the infrared low-light enhancement model to encode the initial feature map to obtain intermediate feature maps with multiple resolutions.

[0129] For example, referring to Figure 3 , the basic block corresponding to downsampling can be called an encoding block. Through multiple encoding blocks and their corresponding downsampling processes, intermediate feature maps with multiple resolutions can be obtained. For example, the multiple resolutions are represented by 1, 1 / 2, 1 / 4, and 1 / 8 respectively.

[0130] 703. For the current resolution among the multiple resolutions, fuse the intermediate feature map of the current resolution and the upsampled feature map to obtain a fused feature map; the upsampled feature map is obtained by upsampling the decoded feature map of the previous resolution.

[0131] For example, referring to Figure 3 , taking addition as an example for fusion, assuming the current resolution is 1 / 2, then add the intermediate feature map of resolution 1 / 2 and the upsampled feature map to obtain the fused feature map of resolution 1 / 2.

[0132] 704. Use the decoding block corresponding to the current resolution in the infrared low-light enhancement model to decode the fused feature map to obtain the decoded feature map of the current resolution.

[0133] For example, referring to Figure 3 , the basic block corresponding to upsampling can be called a decoding block. Taking the current resolution of 1 / 2 as an example, then use the decoding block corresponding to resolution 1 / 2 to decode the fused feature map of resolution 1 / 2 to obtain the decoded feature map of resolution 1 / 2.

[0134] This decoded feature map is input into the upsampling module, and the output is the upsampled feature map. For example, after upsampling the decoded feature map of resolution 1 / 2, a feature map of resolution 1 is obtained.

[0135] 705. Use the decoded feature map when the current resolution is the initial resolution of the initial feature map as the target feature map.

[0136] For example, referring to Figure 3 , after being processed by the rightmost basic block, a decoded feature map of resolution 1 is obtained. Since the initial resolution is represented by 1, then use this decoded feature map, that is, the output feature map of the rightmost basic block as the target feature map.

[0137] In this embodiment, by fusing feature maps with different resolutions, the feature representation ability can be improved, thereby enhancing the infrared low-light enhancement effect.

[0138] 706. Use the output convolutional block in the infrared low-light enhancement model to perform convolution on the target feature map to obtain a multi-channel image.

[0139] For example, referring to Figure 3 , taking three channels as an example, input the target feature map into the output convolutional block, and the output of this convolutional block is a multi-channel image with a size of H*W*3.

[0140] 707. Add the low-light infrared image and the multi-channel image to obtain an initial format image.

[0141] Among them, when adding, the low-light infrared image can be filled into a three-channel image, so that the initial format image with a size of H*W*3 is obtained after addition.

[0142] 708. Convert the initial format image into a target format image.

[0143] 709. Extract the preset channel image of the target format image as the normal-light infrared image.

[0144] For example, referring to Figure 3 , if the initial format image is an RGB image and the target format image is a YCbCr image, after converting the RGB image into a YCbCr image, extract the Y channel image as the normal-light infrared image.

[0145] In this embodiment, through format conversion and single-channel extraction, the normal-light infrared image can be obtained efficiently and accurately.

[0146] Figure 8 is a schematic diagram according to the fourth embodiment of the present disclosure. This embodiment provides a method for training an infrared low-light enhancement model, and the method includes:

[0147] 801. Use a teacher model to process a multi-channel first sample image to obtain a multi-channel first output image.

[0148] 802. Use a student model to process a single-channel second sample image to obtain a multi-channel second output image; the second sample image is obtained based on the first sample image.

[0149] 803. Based on the first output image, the second output image, and the normal-light image corresponding to the first sample image, construct a total loss function.

[0150] 804. Adjust the model parameters of the student model based on the total loss function to obtain an infrared low-light enhancement model.

[0151] Specifically, in the existing sample set, low-light images and their corresponding normal-light images can be stored correspondingly. Taking the multi-channel image as an RGB image as an example, the low-light RGB images and their corresponding normal-light RGB images can be stored correspondingly in the sample set.

[0152] After that, the low-light RGB image can be used as the first sample image, and the normal-light RGB image can be used as the normal-light image corresponding to the first sample image.

[0153] After obtaining the first sample image (low-light RGB image), it can be converted to the YCbCr format, which can be called the low-light YCbCr image, and the Y-channel image of the low-light YCbCr image (which can be called the low-light Y-channel image) can be extracted as the second sample image.

[0154] Both the input and output of the teacher model are multi-channel images. For example, if the input is a three-channel low-light RGB image, its output can be called the first output image. The first output image is a three-channel image, which is the image obtained by the teacher model performing low-light processing on the low-light RGB image.

[0155] The input of the student model is a single-channel image, such as the above-mentioned low-light Y-channel image, and its output can be called the second output image. The second output image is a three-channel image, which is the image obtained by the student model performing low-light processing on the low-light Y-channel image.

[0156] The structures of the teacher model and the student model are the same, such as Figure 3 the model structure shown.

[0157] After obtaining the above-mentioned first output image, second output image, and normal-light image, a total loss function can be constructed based on this.

[0158] Specifically, a first sub-loss function can be constructed based on the first output image and the second output image, and a second sub-loss function can be constructed based on the second output image and the normal-light image; then the total loss function can be constructed based on the first sub-loss function and the second sub-loss function. For example, the total loss function is obtained after weighted summation.

[0159] Both the first sub-loss function and the second sub-loss function can be L1 loss functions. The L1 loss function calculates the average of the absolute differences between the predicted value and the true value, and is used to measure the accuracy of model prediction.

[0160] It is expressed by the formula as:

[0161] loss = α·L1(I s ,I t)+(1-α)·L1(I s ,I truth )

[0162] Among them, loss is the total loss function;

[0163] L1() is the L1 loss function;

[0164] α is a preset weighting coefficient;

[0165] I t is the first output image, I s is the second output image, I truth is the normal light image.

[0166] After obtaining the total loss function, the model parameters of the student model can be adjusted by means of backpropagation or the like until a preset end condition is reached, such as reaching a preset number of iterations or model convergence. The student model when the end condition is reached is used as the final infrared low-light enhancement model. In the inference stage, this infrared low-light enhancement model can be used for the infrared low-light enhancement process.

[0167] In this way, the student model can be trained by means of knowledge distillation. Specifically, the RGB information of the normal light image and the RGB information of the output image of the teacher model are used to supervise the learning of the student model.

[0168] In this embodiment, a loss function is constructed based on the first output image output by the teacher model, the second output image output by the student model, and the normal light image, and then the model parameters are adjusted, so as to realize supervised distillation training, improve the performance of the infrared low-light enhancement model, and further improve the visual effect and multi-exposure adaptation ability.

[0169] In some embodiments, the student model includes a basic block, and the basic block includes: an attention network and a feedforward network, and the attention network is a residual network composed of an attention layer and a convolutional layer, and the feedforward network is a residual network composed of a gated unit.

[0170] The attention network is a residual network composed of an attention layer and a convolutional layer, which can introduce the attention mechanism of the Transformer network into the Convolutional Neural Networks (CNN), so as to improve the network's ability to extract feature maps and improve the effect of the infrared low-light enhancement model. The FFN is a residual network composed of a gated unit, which can simplify the network structure, improve the representation ability while being lightweight, and improve the effect of the infrared low-light enhancement model. For this reason, an efficient and lightweight infrared low-light enhancement model can be obtained, so that infrared low-light enhancement can be realized efficiently and lightweight.

[0171] In some embodiments, the student model is used to process the single-channel second sample image to obtain a multi-channel second output image, including:

[0172] Feature extraction is performed on the second sample image to obtain an initial feature map;

[0173] Based on the basic block, the initial feature map is processed to obtain a target feature map;

[0174] The second output image is obtained based on the target feature map.

[0175] Among them, the process of the student model obtaining the second output image is similar to the inference process. For example, referring to Figure 3 , the input convolutional block of the student model is used to perform feature extraction on the second sample image to obtain an initial feature map; the basic block of the student model is used to process the initial feature map to obtain a target feature map; the output convolutional block of the student model is used to perform convolution on the target feature map to obtain a multi-channel image; the second sample image and the multi-channel image are added to obtain the second output image.

[0176] In this way, a second output image enhanced by the student model can be obtained. Furthermore, a total loss function is constructed based on the second output image, and the model parameters of the student model are adjusted based on the total loss function, so as to obtain an efficient and lightweight infrared low-light enhancement model.

[0177] In some embodiments, the processing of the initial feature map based on the basic block to obtain a target feature map includes:

[0178] Based on the initial feature map, the input feature map of the attention network is obtained;

[0179] The attention network is used to process the input feature map of the attention network to obtain the output feature map of the attention network, which is used as the input feature map of the feed-forward network;

[0180] The feed-forward network is used to process the input feature map of the feed-forward network to obtain the output feature map of the feed-forward network;

[0181] Based on the output feature map of the feed-forward network, the target feature map is obtained.

[0182] For example, inside a single basic block, the attention network and the feed-forward network (FFN) are connected in series. The input feature map of the attention network is obtained based on the initial feature map, the output feature map of the attention network is used as the input feature map of the FFN, and the target feature map is obtained based on the output feature map of the FFN.

[0183] In this embodiment, an attention network and a feed-forward network connected in series are used to obtain a target feature map, which can improve the representation ability of the target feature map and thus enhance the infrared low-light enhancement effect.

[0184] In some embodiments, using the attention network to process the input feature map of the attention network to obtain the output feature map of the attention network includes:

[0185] Using the attention layer and the convolutional layer to process the input feature map of the attention network to obtain the residual feature map of the attention network;

[0186] Adding the input feature map of the attention network and the residual feature map of the attention network to obtain the output feature map of the attention network.

[0187] For example, a residual path can be composed of a convolutional layer and an attention layer, and the output feature map of the residual path can be called the residual feature map. After adding the input feature map of the attention network and the residual feature map, the output feature map of the attention network is obtained.

[0188] In this embodiment, obtaining the output feature map of the attention network based on the input feature map and the residual feature map of the attention network can simply and efficiently obtain the output feature map of the attention network, and thus simply and efficiently obtain the target feature map.

[0189] In some embodiments, using the feed-forward network to process the input feature map of the feed-forward network to obtain the output feature map of the feed-forward network includes:

[0190] Using the gated unit to process the input feature map of the feed-forward network to obtain the residual feature map of the feed-forward network;

[0191] Adding the input feature map of the feed-forward network and the residual feature map of the feed-forward network to obtain the output feature map of the feed-forward network.

[0192] For example, referring to Figure 4 , a residual path can be composed of a gated unit and a convolutional layer, and the output feature map of the residual path can be called the residual feature map. After adding the input feature map of the feed-forward network and the residual feature map, the output feature map of the feed-forward network is obtained.

[0193] In this embodiment, obtaining the output feature map of the feed-forward network based on the input feature map and the residual feature map of the feed-forward network can simply and efficiently obtain the output feature map of the feed-forward network, and thus simply and efficiently obtain the target feature map.

[0194] In some embodiments, using the gating unit to process the input feature map of the feed-forward network to obtain the residual feature map of the feed-forward network includes:

[0195] Splitting the input feature map of the gating unit into a first sub-feature map and a second sub-feature map;

[0196] Using a preset activation function to activate the first sub-feature map to obtain a weighted feature map;

[0197] Weighting the second sub-feature map based on the weighted feature map to obtain the output feature map of the gating unit.

[0198] For example, referring to Figure 5 , the first sub-feature map and the second sub-feature map are obtained by splitting by channel, both with a size of H*W*C / 2. The feature map obtained after processing the first sub-feature map with a preset activation function (such as the Swish activation function) is called the weighted feature map. This weighted feature map is weighted with the second sub-feature map, such as element-wise multiplication, to obtain the output feature map of the gating unit.

[0199] In this embodiment, splitting the feature map can simplify the operation and achieve lightweight. Through the processing of the activation function, a gating mechanism can be realized, improving the feature map representation ability, and further enhancing the infrared low-light enhancement effect.

[0200] In some embodiments, the basic block includes: an encoding block and a decoding block;

[0201] Processing the initial feature map based on the basic block to obtain the target feature map includes:

[0202] Encoding the initial feature map based on the encoding block to obtain an intermediate feature map;

[0203] Decoding the intermediate feature map based on the decoding block to obtain the target feature map.

[0204] For example, in Figure 3 the left part, that is, the basic block corresponding to the downsampling module can be called the encoding block. In Figure 3 the right part, that is, the basic block corresponding to the upsampling module and the convolutional block at the output end can be called the decoding block. The encoding block is used to encode the initial feature map to obtain an intermediate feature map; the decoding block is used to decode the intermediate feature map to obtain the target feature map.

[0205] In this way, through the encoding and decoding process, the initial feature map can be processed to obtain the target feature map, improving the representation effect of the target feature map and the processing efficiency.

[0206] In some embodiments, there are multiple encoding blocks and multiple decoding blocks, and each encoding block and each decoding block respectively correspond to one of multiple resolutions;

[0207] There are multiple intermediate feature maps, and each intermediate feature map respectively corresponds to one resolution;

[0208] Decoding the intermediate feature map based on the decoding block to obtain the target feature map includes:

[0209] For the current resolution among the multiple resolutions, fusing the intermediate feature map of the current resolution and the upsampled feature map to obtain a fused feature map; the upsampled feature map is obtained by upsampling the decoded feature map of the previous resolution;

[0210] Using the decoding block corresponding to the current resolution to decode the fused feature map to obtain the decoded feature map of the current resolution;

[0211] Taking the decoded feature map when the current resolution is the initial resolution of the initial feature map as the target feature map.

[0212] In this embodiment, by fusing feature maps of different resolutions, the feature representation ability can be improved, thereby enhancing the infrared low-light enhancement effect.

[0213] Through the above overall infrared low-light enhancement process, lightweight and efficient infrared low-light enhancement can be achieved, improving the real-time performance of the infrared low-light enhancement process and being applicable to user-side processing scenarios.

[0214] After experiments, infrared low-light enhancement verification is carried out on publicly available datasets such as LOLv1, LOLv2_real, and LOLv2_syn, and compared with other methods (Method a, Method b). The method (Method c) provided by the present disclosure embodiment has smaller parameter quantity and flops, indicating a reduction in computational complexity; PSNR and SSIM are greatly improved compared to other algorithms, indicating an improvement in the infrared low-light enhancement effect. Among them, LOLv1, LOLv2_real, and LOLv2_syn are publicly available datasets, flops is the number of floating-point operations per second, PSNR is the peak signal-to-noise ratio, and SSIM is the structural similarity index.

[0215] The specific description is as follows:

[0216] Method a: The network input is a single-channel low-light infrared image, and the output is a single-channel;

[0217] Method b: The network input is a single-channel low-light infrared image, and the output is three channels. Only the normal light RGB image is used for supervised training.

[0218] Method c: The network input is a single-channel low-light infrared image, and the output is three channels. The normal light RGB image and the inference results of the RGB three-channel low-light teacher model are used for supervised distillation training.

[0219] The psnr / ssim metrics of the three methods on three public datasets are shown in Table 1 below. The flops calculation defines the input resolution as 256x256.

[0220] Table 1

[0221]

[0222] It can be seen that the number of parameters and flops of the baseline network provided by the present disclosure embodiment are only 0.160M and 0.66G, which are more lightweight and efficient than other CNN networks. Method c provided by the present disclosure embodiment obtains the best metric results on three public datasets through supervised distillation training with RGB information added at the output end.

[0223] Figure 9 It is a comparison schematic diagram of the infrared enhancement effects of the method provided by the present disclosure embodiment and other methods.

[0224] To further compare the visual image effects of the three methods, the inference results of the actual images of low light and overexposure are as Figure 9 shown. The original image is a low-light infrared image, and the images enhanced by the above three methods are represented by Method a to Method c respectively. It can be seen that the single-channel input image information of Method a is limited, resulting in overfitting of network training, making the overall image turn white and overexposed to improve the brightness index; comparing Method b and Method c, it is reflected that the model of non-distillation training (Method b) is still a bit white, and the model of distillation training (Method c) has stronger ability to handle multi-exposure. Through supervised distillation training with RGB information, the student model better inherits the relevant characteristics of the RGB three-channel teacher model.

[0225] Figure 10 It is a schematic diagram according to the fifth embodiment of the present disclosure. The present embodiment provides an infrared low-light enhancement device, and the device 1000 includes: an extraction module 1001, a processing module 1002, an acquisition module 1003, and an enhancement module 1004.

[0226] The extraction module 1001 is used to extract features from a single-channel low-light infrared image to obtain an initial feature map; the processing module 1002 is used to process the initial feature map with a basic block to obtain a target feature map; the acquisition module 1003 is used to obtain a multi-channel image according to the target feature map; the enhancement module 1004 is used to obtain a normal-light infrared image according to the low-light infrared image and the multi-channel image.

[0227] In this embodiment, obtaining a target feature map based on the initial feature map can improve the feature representation ability. Furthermore, obtaining a multi-channel image based on the target feature map and obtaining a normal-light infrared image based on the low-light infrared image and the multi-channel image can efficiently achieve infrared low-light enhancement.

[0228] In some embodiments, the initial feature map can be processed based on a basic block.

[0229] In this embodiment, the attention network is a residual network composed of an attention layer and a convolutional layer, which can introduce the attention mechanism of the Transformer network into the Convolutional Neural Networks (CNNs), thereby improving the network's ability to extract feature maps and enhancing the infrared low-light enhancement effect. The FFN is a residual network composed of gated units, which can simplify the network structure, improve the representation ability while being lightweight, and enhance the infrared low-light enhancement effect. Therefore, infrared low-light enhancement can be efficiently and lightly achieved.

[0230] In some embodiments, the initial feature map is processed with a basic block, and the basic block includes: an attention network and a feed-forward network. The attention network is a residual network composed of an attention layer and a convolutional layer, and the feed-forward network is a residual network composed of gated units; the processing module 1002 is further used to:

[0231] Based on the initial feature map, obtain the input feature map of the attention network;

[0232] Use the attention network to process the input feature map of the attention network to obtain the output feature map of the attention network and use it as the input feature map of the feed-forward network;

[0233] Use the feed-forward network to process the input feature map of the feed-forward network to obtain the output feature map of the feed-forward network;

[0234] Based on the output feature map of the feed-forward network, obtain the target feature map.

[0235] In this embodiment, an attention network and a feed-forward network connected in series are used to obtain a target feature map, which can improve the representation ability of the target feature map and thus enhance the infrared low-light enhancement effect.

[0236] In some embodiments, the processing module 1002 is further configured to:

[0237] Use the attention layer and the convolutional layer to process the input feature map of the attention network to obtain the residual feature map of the attention network;

[0238] Add the input feature map of the attention network and the residual feature map of the attention network to obtain the output feature map of the attention network.

[0239] In this embodiment, obtaining the output feature map of the attention network based on the input feature map and the residual feature map of the attention network can simply and efficiently obtain the output feature map of the attention network, and thus simply and efficiently obtain the target feature map.

[0240] In some embodiments, the processing module 1002 is further configured to:

[0241] Use the gated unit to process the input feature map of the feed-forward network to obtain the residual feature map of the feed-forward network;

[0242] Add the input feature map of the feed-forward network and the residual feature map of the feed-forward network to obtain the output feature map of the feed-forward network.

[0243] In this embodiment, obtaining the output feature map of the feed-forward network based on the input feature map and the residual feature map of the feed-forward network can simply and efficiently obtain the output feature map of the feed-forward network, and thus simply and efficiently obtain the target feature map.

[0244] In some embodiments, the processing module 1002 is further configured to:

[0245] Split the input feature map of the gated unit into a first sub-feature map and a second sub-feature map;

[0246] Use a preset activation function to activate the first sub-feature map to obtain a weighted feature map;

[0247] Weight the second sub-feature map based on the weighted feature map to obtain the output feature map of the gated unit.

[0248] In this embodiment, feature map splitting can simplify the operation and achieve lightweight. Through activation function processing, a gate mechanism can be implemented to improve the feature map representation ability, and thus enhance the infrared low-light enhancement effect.

[0249] In some embodiments, the initial feature map is processed using a basic block, and the basic block includes: an encoding block and a decoding block;

[0250] The processing module 1002 is further configured to:

[0251] Based on the encoding block, encode the initial feature map to obtain an intermediate feature map;

[0252] Based on the decoding block, decode the intermediate feature map to obtain the target feature map.

[0253] In this embodiment, through the encoding and decoding process, the initial feature map can be processed to obtain the target feature map, improving the representation effect of the target feature map and the processing efficiency.

[0254] In some embodiments, there are multiple encoding blocks and multiple decoding blocks, and each encoding block and each decoding block respectively correspond to one of multiple resolutions;

[0255] There are multiple intermediate feature maps, and each intermediate feature map respectively corresponds to one resolution;

[0256] The processing module 1002 is further configured to:

[0257] For the current resolution among the multiple resolutions, fuse the intermediate feature map of the current resolution and the upsampled feature map to obtain a fused feature map; the upsampled feature map is obtained by upsampling the decoded feature map of the previous resolution;

[0258] Use the decoding block corresponding to the current resolution to decode the fused feature map to obtain the decoded feature map of the current resolution;

[0259] Take the decoded feature map when the current resolution is the initial resolution of the initial feature map as the target feature map.

[0260] In this embodiment, by fusing feature maps of different resolutions, the feature representation ability can be improved, thereby enhancing the infrared low-light enhancement effect.

[0261] In some embodiments, the enhancement module 1004 is further configured to:

[0262] Add the low-light infrared image and the multi-channel image to obtain an initial format image;

[0263] Convert the initial format image to a target format image;

[0264] Extract the preset channel image of the target format image as the normal-light infrared image.

[0265] In this embodiment, through format conversion and single-channel extraction, a normal light infrared image can be obtained efficiently and accurately.

[0266] Figure 11 FIG. 4 is a schematic diagram according to the sixth embodiment of the present disclosure. This embodiment provides an infrared low-light enhancement model training device, and the device 1100 includes: a first processing module 1101, a second processing module 1102, a construction module 1103, and an adjustment module 1104.

[0267] The first processing module 1101 is configured to process a multi-channel first sample image by using a teacher model to obtain a multi-channel first output image; the second processing module 1102 is configured to process a single-channel second sample image by using a student model to obtain a multi-channel second output image; the second sample image is obtained based on the first sample image; the construction module 1103 is configured to construct a total loss function according to the first output image, the second output image, and the normal light image corresponding to the first sample image; the adjustment module 1104 is configured to adjust the model parameters of the student model according to the total loss function to obtain an infrared low-light enhancement model.

[0268] In this embodiment, a loss function is constructed based on the first output image output by the teacher model, the second output image output by the student model, and the normal light image, and then the model parameters are adjusted, so as to implement supervised distillation training, improve the performance of the infrared low-light enhancement model, and further improve the visual effect and multi-exposure adaptation ability.

[0269] In some embodiments, the student model includes a basic block, and the basic block includes: an attention network and a feed-forward network, and the attention network is a residual network based on an attention layer and a convolutional layer, and the feed-forward network is a residual network based on a gated unit;

[0270] The second processing module 1102 is further configured to:

[0271] Extract features from the second sample image to obtain an initial feature map;

[0272] Process the initial feature map based on the basic block to obtain a target feature map;

[0273] Obtain the second output image based on the target feature map.

[0274] In this embodiment, a second output image enhanced by the student model can be obtained, and then a total loss function is constructed based on the second output image, and the model parameters of the student model are adjusted based on the total loss function, so as to obtain an efficient and lightweight infrared low-light enhancement model.

[0275] In addition, the attention network is a residual network composed of an attention layer and a convolutional layer, which can introduce the attention mechanism of the Transformer network into the Convolutional Neural Networks (CNN), thereby improving the network's ability to extract feature maps and enhancing the effect of the infrared low-light enhancement model. The FFN is a residual network composed of gated units, which can simplify the network structure, improve the representation ability while reducing the weight, and enhance the effect of the infrared low-light enhancement model. Therefore, an efficient and lightweight infrared low-light enhancement model can be obtained, enabling efficient and lightweight implementation of infrared low-light enhancement.

[0276] In some embodiments, the second processing module 1102 is further configured to:

[0277] Based on the initial feature map, obtain the input feature map of the attention network;

[0278] Use the attention network to process the input feature map of the attention network to obtain the output feature map of the attention network, and use it as the input feature map of the feed-forward network;

[0279] Use the feed-forward network to process the input feature map of the feed-forward network to obtain the output feature map of the feed-forward network;

[0280] Based on the output feature map of the feed-forward network, obtain the target feature map.

[0281] In this embodiment, using a series-connected attention network and feed-forward network to obtain the target feature map can improve the representation ability of the target feature map, and further enhance the infrared low-light enhancement effect.

[0282] In some embodiments, the second processing module 1102 is further configured to:

[0283] Use the attention layer and the convolutional layer to process the input feature map of the attention network to obtain the residual feature map of the attention network;

[0284] Add the input feature map of the attention network and the residual feature map of the attention network to obtain the output feature map of the attention network.

[0285] In this embodiment, obtaining the output feature map of the attention network based on the input feature map and the residual feature map of the attention network can simply and efficiently obtain the output feature map of the attention network, and further simply and efficiently obtain the target feature map.

[0286] In some embodiments, the second processing module 1102 is further configured to:

[0287] Using the gating unit, process the input feature map of the feed-forward network to obtain the residual feature map of the feed-forward network;

[0288] Add the input feature map of the feed-forward network and the residual feature map of the feed-forward network to obtain the output feature map of the feed-forward network.

[0289] In this embodiment, obtaining the output feature map of the feed-forward network based on the input feature map and the residual feature map of the feed-forward network can simply and efficiently obtain the output feature map of the feed-forward network, and thus simply and efficiently obtain the target feature map.

[0290] In some embodiments, the second processing module 1102 is further configured to:

[0291] Split the input feature map of the gating unit into a first sub-feature map and a second sub-feature map;

[0292] Use a preset activation function to activate the first sub-feature map to obtain a weighted feature map;

[0293] Weight the second sub-feature map based on the weighted feature map to obtain the output feature map of the gating unit.

[0294] In this embodiment, feature map splitting can simplify the operation and achieve lightweight, and through activation function processing, a gating mechanism can be implemented to improve the feature map representation ability, thereby enhancing the infrared low-light enhancement effect.

[0295] In some embodiments, the basic block includes: an encoding block and a decoding block;

[0296] The second processing module 1102 is further configured to:

[0297] Based on the encoding block, encode the initial feature map to obtain an intermediate feature map;

[0298] Based on the decoding block, decode the intermediate feature map to obtain the target feature map.

[0299] In this way, through the encoding and decoding process, the initial feature map can be processed to obtain the target feature map, improving the representation effect of the target feature map and the processing efficiency.

[0300] In some embodiments, there are multiple encoding blocks and multiple decoding blocks, and each encoding block and each decoding block respectively correspond to one of multiple resolutions;

[0301] There are multiple intermediate feature maps, and each intermediate feature map respectively corresponds to one resolution;

[0302] The second processing module 1102 is further configured to:

[0303] For the current resolution among the multiple resolutions, fuse the intermediate feature map and the upsampled feature map of the current resolution to obtain a fused feature map; the upsampled feature map is obtained by upsampling the decoded feature map of the previous resolution.

[0304] Use the decoding block corresponding to the current resolution to decode the fused feature map to obtain the decoded feature map of the current resolution.

[0305] Take the decoded feature map when the current resolution is the initial resolution of the initial feature map as the target feature map.

[0306] In this embodiment, by fusing feature maps of different resolutions, the feature representation ability can be improved, thereby enhancing the infrared low-light enhancement effect.

[0307] It can be understood that in the embodiments of the present disclosure, the same or similar content in different embodiments can be referred to each other.

[0308] It can be understood that the "first", "second", etc. in the embodiments of the present disclosure are only used for distinction and do not indicate the level of importance, the sequence of time, etc.

[0309] It can be understood that if there is no special limitation description for the sequence of steps involved in the process, it means that the timing relationship between these steps is not limited.

[0310] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0311] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0312] Figure 12 FIG. shows a schematic block diagram of an exemplary electronic device 1200 that can be used to implement the embodiments of the present disclosure. The electronic device 1200 is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smartphone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0313] As Figure 12As shown, the electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 1202 or computer programs loaded from a storage unit 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the electronic device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0314] Multiple components in the electronic device 1200 are connected to the I / O interface 1205, including: an input unit 1206, such as a keyboard, a mouse, etc.; an output unit 1207, such as various types of displays, speakers, etc.; a storage unit 1208, such as a magnetic disk, an optical disc, etc.; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 allows the electronic device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0315] The computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1201 executes the various methods and processes described above, such as an infrared low-light enhancement method or an infrared low-light enhancement model training method. For example, in some embodiments, the infrared low-light enhancement method or the infrared low-light enhancement model training method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1209. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the infrared low-light enhancement method or the infrared low-light enhancement model training method described above can be executed. Alternatively, in other embodiments, the computing unit 1201 can be configured to execute the infrared low-light enhancement method or the infrared low-light enhancement model training method in any other appropriate manner (e.g., by means of firmware).

[0316] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0317] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable task processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0318] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0319] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0320] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0321] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0322] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0323] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. An infrared dark light enhancement method, comprising: Perform feature extraction on a single-channel low-light infrared image to obtain an initial feature map; Processing the initial feature map to obtain a target feature map; Acquire a multi-channel image based on the target feature map; A normal-light infrared image is acquired based on the low-light infrared image and the multi-channel image.

2. The method according to claim 1, wherein: The initial feature map is processed by a basic block, wherein the basic block includes an attention network and a feedforward network, wherein the attention network is a residual network composed of an attention layer and a convolutional layer, and the feedforward network is a residual network composed of a gated unit; The processing of the initial feature map to obtain a target feature map includes: Based on the initial feature map, obtaining an input feature map of the attention network; Using the attention network, processing the input feature map of the attention network to obtain the output feature map of the attention network as the input feature map of the feedforward network; Using the feedforward network, processing the input feature map of the feedforward network to obtain the output feature map of the feedforward network; Based on the output feature map of the feedforward network, the target feature map is obtained.

3. The method according to claim 2, wherein: The using the attention network to process the input feature map of the attention network to obtain the output feature map of the attention network includes: Using the attention layer and the convolution layer, the input feature map of the attention network is processed to obtain a residual feature map of the attention network; The input feature map of the attention network and the residual feature map of the attention network are added to obtain the output feature map of the attention network.

4. The method according to claim 2, wherein: The adopting the feedforward network to process the input feature map of the feedforward network to obtain the output feature map of the feedforward network includes: Using the gating unit, processing the input feature map of the feedforward network to obtain a residual feature map of the feedforward network; An input feature map of the feedforward network and a residual feature map of the feedforward network are added to obtain an output feature map of the feedforward network.

5. The method according to claim 4, wherein: The adopting the gating unit to process the input feature map of the feedforward network to obtain the residual feature map of the feedforward network includes: Splitting the input feature map of the gate control unit into a first sub-feature map and a second sub-feature map; Using a preset activation function, activating the first sub-feature map to obtain a weighted feature map; The second sub-feature map is weighted based on the weighted feature map to obtain an output feature map of the gating unit.

6. The method according to claim 1, wherein: The initial feature map is processed using a basic block, and the basic block includes: an encoding block and a decoding block; The processing of the initial feature map to obtain a target feature map includes: Based on the encoding block, encoding the initial feature map to obtain an intermediate feature map; Based on the decoding block, the intermediate feature map is decoded to obtain the target feature map.

7. The method according to claim 6, wherein: There are multiple encoding blocks and multiple decoding blocks, and each encoding block and each decoding block corresponds to one of multiple resolutions; There are multiple intermediate feature maps, and each intermediate feature map corresponds to a resolution; The decoding of the intermediate feature map based on the decoding block to obtain the target feature map includes: For a current resolution among the multiple resolutions, fusing the intermediate feature map of the current resolution and the upsampled feature map to obtain a fused feature map; the upsampled feature map is obtained by upsampling the decoded feature map of the previous resolution; Using a decoding block corresponding to the current resolution, decoding the fused feature map to obtain a decoded feature map of the current resolution; The decoded feature map when the current resolution is the initial resolution of the initial feature map is used as the target feature map.

8. The method according to claim 1, wherein: The acquiring of a normal-light infrared image based on the low-light infrared image and the multi-channel image comprises: Adding the low-light infrared image and the multi-channel image to obtain an initial format image; Converting the initial format image into a target format image; A preset channel image of the target format image is extracted as the normal light infrared image.

9. A method for training an infrared dark light enhancement model, comprising: Using the teacher model, processing the multi-channel first sample image to obtain the multi-channel first output image; Using the student model, the single-channel second sample image is processed to obtain a multi-channel second output image; The second sample image is obtained based on the first sample image; constructing a total loss function based on the first output image, the second output image, and a normal light image corresponding to the first sample image; Based on the total loss function, the model parameters of the student model are adjusted to obtain an infrared dark light enhancement model.

10. An infrared dark light enhancement device, comprising: An extraction module, used for extracting features from a single-channel low-light infrared image to obtain an initial feature map; A processing module, configured to process the initial feature map using a basic block to obtain a target feature map; An acquisition module, used for acquiring a multi-channel image according to the target feature map; an enhancement module, configured to obtain a normal-light infrared image based on the low-light infrared image and the multi-channel image; Among them, the basic block includes: an attention network and a feedforward network, and the attention network is a residual network composed of an attention layer and a convolutional layer, and the feedforward network is a residual network composed of a gating unit.

11. An infrared dark light enhanced model training device, comprising: A first processing module is used to process the multi-channel first sample image using a teacher model to obtain a multi-channel first output image; A second processing module is used to process the single-channel second sample image using a student model to obtain a multi-channel second output image; The second sample image is obtained based on the first sample image; A construction module, configured to construct a total loss function according to the first output image, the second output image, and a normal light image corresponding to the first sample image; An adjustment module is used to adjust the model parameters of the student model according to the total loss function to obtain an infrared dark light enhancement model.

12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.

14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.