Low-light scene target detection method with infrared sensing and system thereof
By using infrared prediction branching and feature fusion technology in low-light environments, the problem of reduced target detection accuracy in traditional methods is solved, high-precision target detection is achieved, and cost and computing resource requirements are reduced.
Patent Information
- Application Number
- CN202510156789.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
AI Technical Summary
In low-light environments, traditional RGB image processing methods lead to a decrease in the accuracy of object detection, infrared images lack resolution and clarity during long-distance object detection, and multimodal fusion technology has problems with high computing resources and costs.
A low-light scene object detection method with infrared perception is adopted, infrared prediction branches are predicted, multi-scale features are extracted in combination with backbone networks, and feature fusion is used to fusion using complementary fusion filter modules and hybrid feature pyramid structures, and finally mapped as the target category, size and position through the 2D detection head.
Achieve high-precision object detection under low light conditions, improving detection performance while controlling costs and simplifying processes.
Smart Images

Figure CN120088607A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a low-light scene target detection method and system with infrared perception; in particular, it relates to the field of target detection under low-light conditions. Background Art
[0002] Object Detection, as one of the core tasks in the field of computer vision, has been widely studied and applied in many fields such as intelligent monitoring, autonomous driving, security, and military in recent years. Object Detection aims to identify objects in an image through a computer vision system and accurately mark their positions. With the rapid development of deep learning technology, methods based on convolutional neural networks (CNNs) have become the mainstream.
[0003] However, in low-light environments, the visibility of targets in traditional RGB image processing methods significantly decreases under low-light conditions, resulting in the loss of image information, which in turn affects the accuracy of target detection. In such cases, RGB images cannot provide sufficient details and clarity, increasing the difficulty of target recognition and resulting in poor detection effects.
[0004] Infrared imaging technology has unique advantages in low-light environments and can identify targets with large temperature differences. Therefore, it has more advantages than RGB images in the case of insufficient light. However, the performance of infrared images also has certain limitations. When detecting distant targets, the resolution and clarity of infrared images are usually inferior to those of visible light images. In addition, the quality of infrared images is affected by temperature differences. Especially when the temperatures of objects and the background are close, it becomes more difficult to distinguish between the target and the background, resulting in blurred image contours and further reducing the detection accuracy.
[0005] Although multi-modal fusion methods improve the detection accuracy by combining the advantages of RGB and infrared images, this technology also faces many challenges. Multi-modal fusion requires higher computing resources and complex processing processes. At the same time, the cost of fusion devices and technologies is relatively high, which limits its popularization in practical applications. Therefore, although multi-modal technology has good performance potential, how to control costs and simplify processes while maintaining high performance remains an urgent problem to be solved. Summary of the Invention
[0006] The present invention aims at the defects existing in the prior art and provides a low-light scene target detection method and system with infrared perception to solve the problem of insufficient detection accuracy in the prior art under low-light conditions.
[0007] To achieve the above object, the present invention adopts the following technical solutions. A low-light scene target detection method with infrared perception includes the steps:
[0008] S0. Obtain the visible light image of the scene using a visible light camera, and read its RGB features;
[0009] S1. Use the infrared prediction branch to predict the infrared image according to the RGB features;
[0010] S2. Use the backbone network to extract the multi-scale features of the infrared image and the visible light image;
[0011] S3. Use the complementary fusion filtering module to fuse the thermal radiation features and visual features obtained in S2;
[0012] S4. Adopt a hybrid feature pyramid structure to further fuse the results in S3 to enhance the multi-scale fusion of global and local features;
[0013] S5. Map the fused features obtained in S4 to the size, category, and position of the 2D object through the 2D detection head.
[0014] Furthermore, step S0 includes: obtaining the visible light image through the visible light camera, reading its RGB features, and converting the RGB features into a tensor format suitable for processing by the deep learning model.
[0015] Furthermore, step S1 includes:
[0016] S1.1. Input the RGB feature tensor, and the shape of the tensor is 512×512×3; where 512×512 represents the spatial resolution of the image, and 3 represents the three color channels of RGB;
[0017] S1.2. Perform local feature extraction and downsampling operations through the convolutional module in the encoder to generate feature maps of different scales;
[0018] S1.3. Generate an infrared image with a shape of (1, 3, 512, 512) through the decoder module and multi-scale feature extraction and upsampling operations.
[0019] Furthermore, in step S2, the backbone network performs feature extraction on the infrared image and the RGB image respectively, where the shapes of the infrared image and the RGB image are both (1, 3, 512, 512); generating multi-scale features, specifically including:
[0020] The first layer of features: with a shape of (1, 32, 128, 128);
[0021] The second layer of features: with a shape of (1, 64, 64, 64);
[0022] The third layer of features: with a shape of (1, 128, 32, 32);
[0023] Fourth - layer feature: The shape is (1, 256, 16, 16).
[0024] Furthermore, step S3 includes:
[0025] S3.1: Using visual features to guide thermal radiation features, including performing translational transformation, summation, generating weights, and weighted fusion on visual features;
[0026] S3.2: Using thermal radiation features to guide visual features, including performing translational transformation, summation, generating weights, and weighted fusion on thermal radiation features;
[0027] S3.3: Balancing the two guided features to obtain the final multi - scale fusion features;
[0028] Among them, the thermal radiation features and visual features respectively include three - scale features; the three - scale thermal radiation features have shapes of (1, 64, 64, 64), (1, 128, 32, 32), and (1, 256, 16, 16); the three - scale visual features have shapes of (1, 64, 64, 64), (1, 128, 32, 32), and (1, 256, 16, 16).
[0029] Furthermore, step S4 specifically includes:
[0030] S4.1: Extracting local features from the complementary - fused features F1(1, 64, 64, 64), F2(1, 128, 32, 32), and F3(1, 256, 16, 16) to generate local features Out1, Out2, and Out3;
[0031] S4.2: Performing global feature processing on the complementary - fused features F1, F2, and F3 to generate a global feature T4; mapping the global feature through a self - attention mechanism and MLP to generate a feature Out4;
[0032] S4.3: Fusing Out4 with Out1, Out2, and Out3 through learnable weights to obtain the final multi - scale fusion features.
[0033] Furthermore, S4.1 includes:
[0034] Inputting three feature maps F1(1, 64, 64, 64), F2(1, 128, 32, 32), and F3(1, 256, 16, 16) with different scales;
[0035] Upsampling F3 to the same spatial resolution as F2, concatenating it with F2, and further processing through a convolutional layer to obtain a feature T1;
[0036] Next, upsample T1 to the same spatial resolution as F1, concatenate it with F1, and further process it through a convolutional layer to obtain the feature Out1;
[0037] Downsample Out1 to the same spatial resolution as T1 and concatenate it with T1, and further process it through a convolutional layer to obtain the feature Out2;
[0038] Downsample Out2 to the same spatial resolution as F3 and concatenate it with F3, and further process it through a convolutional layer to obtain the feature Out3;
[0039] S4.2 includes:
[0040] Use 1*1 convolution on F1, F2, and F3 to adjust the number of channels to 256, and obtain T1(1, 256, 64, 64), T2(1, 256, 32, 32), and T3(1, 256, 16, 16) respectively;
[0041] Perform Flatten operation on T1, T2, and T3, and convert them into one-dimensional tensors T1(1, 256, 4096), T2(1, 256, 1024), T3(1, 256, 256);
[0042] Concatenate the one-dimensional tensors into a tensor T4(1, 256, 5376) according to the third dimension;
[0043] Use the self-attention mechanism module to calculate the attention weights of T4, and then further map the features through an MLP (Multi-Layer Perceptron) to obtain the feature Out4(1, 256, 5376);
[0044] Split Out4 into three tensors according to the channels: Out5(1, 256, 64, 64), Out6(1, 256, 32, 32), Out7(1, 256, 16, 16);
[0045] S4.3 includes:
[0046] Through the learned adjustable weights, perform weighted fusion of the local features (Out1, Out2, Out3) and the corresponding global features (Out5, Out6, Out7), so as to obtain an enhanced multi-scale feature representation.
[0047] Furthermore, in step S5, the 2D detection head maps the fused features into the category, size, and position of the object through multiple convolutional layers, and outputs a vector at each position, including 4 parameters of the predicted object bounding box, object confidence, and category probability distribution.
[0048] A low-light scene target detection system with infrared perception includes:
[0049] A visible light camera, which is used to acquire visible light images and read their RGB features;
[0050] An infrared prediction branch, which is used to predict infrared images based on the RGB features;
[0051] A backbone network, which is used to extract multi-scale features of the infrared images and visible light images;
[0052] A complementary fusion filtering module, which is used to fuse thermal radiation features and visual features;
[0053] A hybrid feature pyramid structure, which is used to enhance the multi-scale fusion of global and local features;
[0054] A 2D detection head, which is used to map the fused features to the size, category, and location of 2D objects.
[0055] Furthermore, the infrared prediction branch includes:
[0056] A decoder module is used in the decoder, with normalization operations, a GeLU activation function, and a GRN structure. The decoder module extracts image features at different levels through multi-scale convolution; the Squeeze-and-Excitation (SE) attention mechanism is also introduced to enable the decoder to focus on the parts of the image that need to be emphasized;
[0057] The complementary fusion filtering module includes:
[0058] A visual feature guiding module, which is used to guide thermal radiation features using visual features;
[0059] A thermal radiation feature guiding module, which is used to guide visual features using thermal radiation features;
[0060] A feature balancing module, which is used to balance the two guided features;
[0061] The hybrid feature pyramid structure includes:
[0062] A local feature extraction module, which is used to extract local features;
[0063] A global feature extraction module, which is used to extract global features;
[0064] A feature fusion module, which is used to fuse local features and global features.
[0065] The 2D detection head includes:
[0066] A convolutional layer, which is used to map the fused features to the category, size, and location of an object;
[0067] An output module, which is used to output the parameters of the object bounding box, the object confidence, and the class probability distribution.
[0068] Advantages of the present invention compared with the prior art.
[0069] The present invention proposes a low-light scene target detection method and system with infrared perception. The system and method have infrared perception function and can achieve high-precision target detection under low-light conditions. In particular, through the infrared prediction branch, the model can perceive the thermal radiation characteristics and effectively fuse them with the visual characteristics, thereby improving the detection performance in low-light environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The present invention will be further described below in conjunction with the drawings and specific embodiments. The protection scope of the present invention is not limited to the description of the following content.
[0071] Figure 1 It is the overall structure diagram of the low-light scene target detection method and system with infrared perception;
[0072] Figure 2 It is the network structure diagram of the infrared prediction branch;
[0073] Figure 3 It is the structural schematic diagram of the complementary fusion module;
[0074] Figure 4 It is the structural schematic diagram of the hybrid feature pyramid;
[0075] Figure 5 It is the prediction result demonstration diagram.
[0076] Figure 6 It is the overall network architecture diagram of the target detection system. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0078] As referred to below, the low-light scene is a scene where the light intensity is between 0.1 lux and 50 lux.
[0079] The 2D detection head refers to an anchor-free decoupled detection head. Anchor-free means no fixed anchor box, and decoupled means separate prediction of classification and regression.
[0080] A low-light scene target detection method and system with infrared perception includes:
[0081] S0. Obtain a visible light image through a visible light camera and read its RGB features; specifically:
[0082] The visible light camera is used to obtain visible light images. Modern cameras are usually equipped with RGB (Red, Green, Blue) sensors, which can collect image data within these specific wavelength bands. The visible light camera captures the visible light images in the scene. The visible light camera can record the RGB (Red, Green, Blue) three-channel information of the scene, and this information contains rich color and texture features. Subsequently, the system reads the RGB features of the image and converts them into a tensor format suitable for processing by the deep learning model. RGB features are important inputs in the object detection task and can provide key information such as the color, shape, and texture of the object for the model.
[0083] S1. Predict the infrared image, i.e., the thermal radiation image, through the infrared prediction branch. Specifically:
[0084] 1.1 Input the RGB feature tensor. First, the system reads the original visible light image and converts it into an RGB feature tensor. The shape of this tensor is 512×512×3, where 512×512 represents the spatial resolution of the image, and 3 represents the three RGB color channels. This RGB feature tensor serves as the input to the infrared prediction branch for subsequent multi-scale feature extraction and infrared image generation.
[0085] 1.2 Multi-scale feature extraction. The infrared prediction branch encodes the input RGB feature tensor into the multi-scale space by gradually downsampling it. This process generates a series of feature maps of different scales, specifically including feature tensors of the following shapes: (1, 64, 256, 256), (1, 128, 128, 128), (1, 256, 64, 64), (1, 512, 32, 32), (1, 512, 16, 16), (1, 512, 8, 8), (1, 512, 4, 4), (1, 512, 2, 2). These multi-scale features capture different levels of information of the image from global to local and provide a rich feature representation for subsequent infrared image generation.
[0086] 1.3 Feature mapping and infrared image generation. The encoder is responsible for mapping these multi-scale features back to the high-resolution infrared space. Through a series of upsampling and feature fusion operations, an infrared image with a shape of (1, 3, 512, 512) is finally generated. Although infrared images are usually single-channel, for computational convenience, it is designed to have 3 channels here, and the values of each channel are the same. The purpose of this design is to keep the number of channels consistent with that of the RGB image for subsequent calculation and processing.
[0087] 1.4 Supervised learning in the training phase. In the training phase, the infrared prediction branch generates an infrared image based on the input RGB features and is supervised and trained using the real infrared image. Specifically, the system calculates the loss between the generated infrared image and the real infrared image and updates the model parameters through the backpropagation algorithm. Through continuous iterative training, the infrared prediction branch gradually learns the prior knowledge of the infrared image, enabling it to more accurately predict the corresponding infrared image from the RGB image.
[0088] 1.5 Generation of infrared images in the prediction phase. In the prediction phase, the system directly generates the corresponding infrared image based on the input RGB image and the trained model parameters. At this time, the model already has the mapping ability from the RGB image to the infrared image and can generate high-quality infrared prediction results without relying on real infrared images. This process can be applied to various practical scenarios, such as night vision imaging and target detection.
[0089] S2. Use the backbone network to extract multi-scale features of the thermal radiation image and the visual image; specifically:
[0090] 1. Input the infrared image and the RGB image.
[0091] In step S2, the system inputs the infrared image and the RGB image simultaneously. The infrared image is (1, 3, 512, 512), and the RGB image has a shape of (1, 3, 512, 512). These two images respectively represent thermal radiation information and visible light information and are the basis for subsequent multi-scale feature extraction.
[0092] 2. Feature extraction by the backbone network.
[0093] The infrared image and the RGB image are respectively subjected to feature extraction by two backbone networks. The role of the backbone network is to extract multi-level feature representations from the input images. For each image, the backbone network gradually downsamples to generate feature maps of different scales.
[0094] 3. Generation of multi-scale features.
[0095] After being processed by the backbone network, the infrared image and the RGB image respectively generate three types of multi-scale features, and the specific shapes are as follows:
[0096] The first layer of features: The shape is (1, 32, 128, 128), representing the most hierarchical semantic information, with a relatively high spatial resolution (128×128) and 32 channels.
[0097] The second layer of features: The shape is (1, 64, 64, 64), representing lower-level semantic information, with a relatively high spatial resolution (64×64) and 64 channels.
[0098] The second - layer features: with a shape of (1, 128, 32, 32), representing medium - level semantic information, the spatial resolution is reduced to 32×32, and the number of channels is increased to 128.
[0099] The third - layer features: with a shape of (1, 256, 16, 16), representing higher - level semantic information, the spatial resolution is further reduced to 16×16, and the number of channels is increased to 256.
[0100] These multi - scale features capture different levels of information in the image: shallow features contain more details and edge information, while deep features contain more semantic and global information. For reducing the computational cost, only the last three layers of features are retained for subsequent calculations.
[0101] S3. Use a complementary fusion filtering module to fuse the thermal radiation features and visual features obtained in step S2.
[0102] Input features.
[0103] The input of the complementary fusion filtering module is the six multi - scale features output in step S2, including:
[0104] Infrared features: Three - scale thermal radiation features with shapes of (1, 64, 64, 64), (1, 128, 32, 32), and (1, 256, 16, 16) respectively.
[0105] Visual features: Three - scale RGB features with shapes of (1, 64, 64, 64), (1, 128, 32, 32), and (1, 256, 16, 16) respectively.
[0106] Guiding of thermal radiation features (using visual features to guide thermal radiation features).
[0107] 1. Perform translation transformation on visual features. For visual features of each scale (e.g., (1, 64, 64, 64)), the system performs nine different - direction translation transformations on it. The translation directions include up, down, left, right, upper - left, upper - right, lower - left, lower - right, and no translation. Each translation operation generates a feature map of the same size as the original feature, and a total of 9 feature maps are obtained.
[0108] 2. Sum the translated features. Sum these 9 translated feature maps element - by - element to obtain a comprehensive feature map. This feature map fuses the local information of visual features in different directions and can better capture the spatial context.
[0109] 3. Change the translation step size and repeat the summation.
[0110] Next, the system will change the translation step sizes, set to 2 and 3 respectively, and repeat the above translation and summation operations. In this way, for the visual features at each scale, three sets of feature maps after translation and summation (step sizes of 1, 2, and 3) will be obtained. These feature maps capture local information at different scales respectively.
[0111] 4. Generate weights and perform weighted fusion. Use the thermal radiation features corresponding to the scale to generate three weights (through convolutional or fully connected layers), corresponding to the three sets of translation and summation features with step sizes of 1, 2, and 3 respectively. These weights are used to balance the contributions of different displacement features. Multiply the weights element-wise with the corresponding translation and summation features to obtain the weighted visual features.
[0112] 5. Multiply element-wise with the thermal radiation features. Multiply the weighted visual features element-wise with the thermal radiation features corresponding to the scale to obtain the thermal radiation features guided by the visual features. The purpose of this step is to utilize the local information of the visual features to enhance the expressive ability of the thermal radiation features.
[0113] Guiding of visual features (using thermal radiation features to guide visual features).
[0114] For the guiding of visual features, the steps are completely symmetric to those of the above guiding of thermal radiation features:
[0115] 1. Perform translation transformations on the thermal radiation features in nine different directions to generate 9 feature maps.
[0116] 2. Sum the translated feature maps to obtain the comprehensive feature map.
[0117] 3. Change the translation step size (step sizes of 1, 2, and 3), and repeat the translation and summation operations.
[0118] 4. Use the visual features to generate three weights and perform weighted fusion on the features after translation and summation.
[0119] 5. Multiply the weighted thermal radiation features element-wise with the visual features corresponding to the scale to obtain the visual features guided by the thermal radiation features.
[0120] Balance the features guided by both.
[0121] After completing the guiding of the above two modalities, the system will obtain two sets of guided features:
[0122] Thermal radiation features guided by visual features.
[0123] Visual features guided by thermal radiation features.
[0124] To balance the contributions of these two features, the system introduces two learnable weights (obtained through training). These two weights are respectively used to weight the two guided features and add them together to obtain the final multi-scale fusion feature.
[0125] Output the fused multi-scale features: After being processed by the complementary fusion filtering module, the system outputs the fused multi-scale features. These features contain both thermal radiation information and visual information, and enhance the feature expression ability through mutual guidance.
[0126] Specifically, the visual features are guided by the thermal radiation features to help the visual features capture the position information of the target more accurately. Conversely, the thermal radiation features are guided by the visual features so that the thermal radiation features can better capture the contour information of the target.
[0127] S4: Further fuse the results in step S3 using a hybrid feature pyramid structure to enhance the multi-scale fusion of global and local features; specifically:
[0128] Specifically, after complementary fusion, three features are obtained: F1(1, 64, 64, 64), F2(1, 128, 32, 32), and F3(1, 256, 16, 16).
[0129] For local feature extraction, first upsample F3 and concatenate it with F2, and further extract the feature T1 through convolution. Then upsample T1 and concatenate it with F1, and further extract the feature Out1 through convolution. Downsample Out1 and concatenate it with T1, and further extract the feature Out2 through convolution. Downsample Out2 and concatenate it with F3, and further extract the feature Out3 through convolution.
[0130] For global features, first use 1*1 convolution to change F1, F2, and F3 to 256-channel operations to obtain T1(1, 256, 64, 64), T2(1, 256, 32, 32), and T3(1, 256, 16, 16). Then perform the Flatten operation to get T1(1, 256, 4096), T2(1, 256, 1024), and T3(1, 256, 256). Concatenate them according to the third dimension to get T4(1, 256, 5376). Calculate its attention through the self-attention mechanism module, and then further map the features through the MLP to obtain the feature Out4.
[0131] Finally, Out4 is split into Out5(1, 256, 64, 64), Out6(1, 256, 32, 32), and Out7(1, 256, 16, 16) according to channels and fused with Out1(1, 256, 64, 64), Out2(1, 256, 32, 32), and Out3(1, 256, 16, 16) through learnable weights.
[0132] S5. Use a 2D detection head to map the fused features obtained in step S4 into the size and category of 2D objects.
[0133] The 2D detection head maps the features output by S4 into the category, size, and position of the object through multiple convolutional layers. The input feature map contains multi-scale context information. After convolutional processing, a vector is output at each position, including 4 parameters (such as the center coordinates and width and height) of the predicted object bounding box, the object confidence, and the category probability distribution.
[0134] Example 1. Overall architecture.
[0135] Object Detection Network with Infrared Sensing Capabilities(InSCnet): Designed an object detection network InSCnet with infrared sensing capabilities. Its overall architecture is as Figure 1 、 Figure 6 shown, consisting of five parts, including: a backbone network, an infrared prediction branch (IPB), an inductive guidance and fusion module (IGCF), a hybrid feature pyramid module (HyFP), and a 2D detection head. The present invention uses the improved C2FDark53 in yolov8. Given the RGB feature H in ×W in ×3, the output feature of the backbone network is where 1 / 8, 1 / 16, and 1 / 32 respectively represent In the present invention, the input RGB image is mapped into a single-channel visible light image through the infrared prediction branch (IPB). Similarly, the backbone network maps the infrared image H in ×W in ×1 into features The complementary fusion filtering module (COFF) is used to perform complementary fusion on visual and thermal radiation features. The hybrid feature pyramid module (HyFP) performs global and local feature fusion on the features at three scales through a feature pyramid and a TransBlock. Finally, the size and category of the 2D target are mapped through the detection head.
[0136] Example 2. Infrared prediction branch (IPB): As Figure 2As shown, it is the IPB infrared prediction branch network designed by the present invention. Its encoder part adopts the ConvNeXt V2 network, which inherits the advantages of traditional CNNs and draws on the advantages of Transformers. Compared with standard convolutional networks, ConvNeXt V2 is designed to provide stronger expressive power and higher computational efficiency when dealing with complex image tasks by optimizing the structure of convolutional layers, kernel size, and network depth.
[0137] The decoder first undergoes normalization operations, the GeLU activation function, and the GRN structure. The module uses depthwise convolution, which decomposes the standard convolution operation into per-channel convolution and 1x1 convolution, significantly reducing the computational complexity and the number of parameters while retaining rich local feature information. This design not only reduces the consumption of computing resources but also improves the efficiency of the model when dealing with large-scale data. Secondly, multiple convolution kernel sizes (5x5, 3x3 convolutions) are used in the module. By extracting image features at different levels through multi-scale convolutions, it can capture fine-grained local information and global structural information, helping to better reconstruct the details of the image. The Squeeze-and-Excitation (SE) attention mechanism is introduced, which can enhance the response to important features and suppress the interference of irrelevant features by adaptively adjusting the weights between channels. It effectively improves the sensitivity of the model to key regions, enabling the decoder to focus more on the parts of the image that need to be emphasized.
[0138] Example 3: Effectively fusing thermal radiation features and visual features is also an important task. The present invention refers to the idea of D4LCN, as shown in Figure 3 (a). Traditional 2D convolution kernels cannot effectively perform feature fusion or effectively infer the spatial relationship between foreground and background pixels. Referring to the idea of depth convolution, thermal radiation and its local features are used to assign a weight to each pixel point of visual information. Specifically, the input visual feature Visual∈R C×H×W and the thermal radiation feature IR∈R C×H×W where C is the number of channels, and W and H are the width and height of the features respectively. A translation transformation is performed on the depth features within each channel to increase the local thermal radiation features. Due to the large intra-class and inter-class scale differences in RGB images, different dilation rates are also added to the translation transformation. And weights are generated for each dilation rate through the visual features, enabling the network to adaptively learn the effect of different receptive fields to solve the scale sensitivity problem. At the same time, to increase the information flow between channels, D4LCN performs a shift and sum operation on the visual features in the channel dimension. The formula for this part is:
[0139] IGF out = IR Local (x,y)·Visual(x,y)
[0140] Among them, k ∈ (1, 2, 3), i, j ∈ (-k, 0, k), and IR represents the thermal radiation feature.
[0141] The feature shift operation in the original method did not effectively promote the information exchange and flow between different channels. Moreover, only one feature was simply used to guide another feature, but this way failed to fully exploit the complementarity and fusion potential of the two features. Therefore, to solve this problem and enhance the feature fusion ability of the model, the present invention designs an improved complementary fusion strategy, as shown in Figure 3 (b) of. The present invention first fuses the visual feature and the thermal radiation feature through a guiding fusion mechanism. Specifically, the thermal radiation feature is used to guide the visual feature, thereby helping the visual feature to more accurately capture the position information of the target. Conversely, the visual feature is used to guide the thermal radiation feature, so that the thermal radiation feature can better capture the contour information of the target. The rich edge and detail information in the visual feature can effectively enhance the spatial resolution of the thermal radiation feature. Through this two-way guiding feature fusion method, the present invention can make full use of the respective advantages of the two modalities of thermal radiation and vision. The formula for this part is:
[0142]
[0143] IGF out = IRLoc al (x, y)·Vis(x, y)
[0144]
[0145] VGF out = Vis Local (x, y)·IR(x, y)
[0146] COF out = Mean(VGF out , IGF out )
[0147] Among them, IR and Vis represent the thermal radiation and visual features respectively, IR Local , Vis Local represent the local features of thermal radiation and vision respectively, IGF out , VGF out represent the features after the fusion of thermal radiation and visual guidance. COF out is the final output.
[0148] Embodiment 4, Hybrid Feature Pyramid Module (HyFP): Considering the differences in feature modeling between convolutional neural networks (CNNs) and Transformer-based encoder-decoder architectures, the Transformer architecture benefits from the self-attention mechanism and has excellent long-range dependence modeling capabilities, enabling effective capture of global information. However, while focusing on global features, it may overlook local details, which can limit its performance in some tasks. Convolutional neural networks (CNNs), on the other hand, excel in local feature extraction, can effectively capture local spatial information, and have strong translational invariance, giving them an advantage in dealing with detailed features.
[0149] To make up for the respective deficiencies of these two architectures, the present invention adopts a fusion strategy of CNN and Transformer in network design. As Figure 4 shown, it is a hybrid feature pyramid structure, and its core idea is to combine the hierarchical structure of the feature pyramid with the attention mechanism to achieve comprehensive feature extraction of images from coarse to fine and from local to global.
[0150] Specifically, this structure first outputs three different scales of features from the backbone network through the feature pyramid. These are fused and extracted. Each layer of the pyramid represents feature information of different scales. This feature information not only includes the local details of the image but also covers the overall structure of the image. At the same time, the attention mechanism is introduced into the feature pyramid in parallel. The attention mechanism can automatically capture the important regions in the image and perform weighted processing on them, thereby highlighting key features and suppressing irrelevant information. By introducing the attention mechanism, the hybrid feature pyramid structure can more accurately capture the key information in the image, improving the accuracy and robustness of feature extraction.
[0151] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "schematic embodiments", "preferred embodiments", "specific implementation manners", or "preferred implementation manners", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0152] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; thus, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A low-light scene target detection method with infrared perception, characterized in that: Includes steps: S0, using a visible light camera to obtain a visible light image of the scene and read its RGB features; S1, using the infrared prediction branch to predict the infrared image according to the RGB features; S2. Extracting multi-scale features of the infrared image and the visible light image using a backbone network; S3, using a complementary fusion filter module to fuse the thermal radiation features and visual features obtained in S2; S4, using a hybrid feature pyramid structure to further fuse the results in S3 to enhance the multi-scale fusion of global and local features; S5. Map the fused features obtained in S4 to the size, category and position of the 2D object through the 2D detection head.
2. The method according to claim 1, characterized in that Step S0 includes: A visible light image is acquired through a visible light camera, and its RGB features are read, and the RGB features are converted into a tensor format suitable for processing by a deep learning model.
3. The method according to claim 2, characterized in that Step S1 includes: S1.
1. Input RGB feature tensor, the shape of which is 512×512×3; 512×512 represents the spatial resolution of the image, and 3 represents the three RGB color channels; S1.2, local feature extraction and downsampling operations are performed through the convolution module in the encoder to generate feature maps of different scales; S1.
3. Generate an infrared image with a shape of (1, 3, 512, 512) through a decoder module and multi-scale feature extraction and upsampling operations.
4. The method according to claim 1, characterized in that In step S2, the backbone network extracts features from the infrared image and the RGB image respectively, wherein the shapes of the infrared image and the RGB image are both (1, 3, 512, 512); and generates multi-scale features, specifically including: The first layer features: shape is (1, 32, 128, 128); The second layer features: shape is (1, 64, 64, 64); The third layer features: shape is (1, 128, 32, 32); The fourth layer features: shape is (1, 256, 16, 16).
5. The method according to claim 1, characterized in that Step S3 includes: S3.
1. Using visual features to guide thermal radiation features, including translation transformation, summation, weight generation and weighted fusion of visual features; S3.2, using thermal radiation features to guide visual features, including translation transformation, summation, weight generation and weighted fusion of thermal radiation features; S3.3, balance the two guided features to obtain the final multi-scale fusion features; Among them, the thermal radiation features and visual features include features of three scales respectively; the shapes of the thermal radiation features of the three scales are (1, 64, 64, 64), (1, 128, 32, 32), and (1, 256, 16, 16); the shapes of the visual features of the three scales are (1, 64, 64, 64), (1, 128, 32, 32), and (1, 256, 16, 16).
6. The method according to claim 1, characterized in that Step S4 specifically includes: S4.
1. Extract local features from the complementary fused features F1 (1, 64, 64, 64), F2 (1, 128, 32, 32), and F3 (1, 256, 16, 16) to generate local features Out1, Out2, and Out3. S4.2, perform global feature processing on the complementary fused features F1, F2, and F3 to generate global feature T4; map the global features through the self-attention mechanism and MLP to generate feature Out4; S4.
3. Fuse Out4 with Out1, Out2, and Out3 through learnable weights to obtain the final multi-scale fusion features.
7. The method according to claim 6, characterized in that S4.1 includes: Input three feature maps of different scales: F1 (1, 64, 64, 64), F2 (1, 128, 32, 32) and F3 (1, 256, 16, 16); F3 is upsampled to the same spatial resolution as F2, concatenated with F2, and further processed through a convolutional layer to obtain feature T1; Then T1 is upsampled to the same spatial resolution as F1, concatenated with F1, and further processed through the convolution layer to obtain feature Out1; Downsample Out1 to the same spatial resolution as T1 and concatenate it with T1, and further process it through the convolution layer to obtain feature Out2; Downsample Out2 to the same spatial resolution as F3 and concatenate it with F3, and further process it through the convolution layer to obtain feature Out3; S4.2 includes: Use 1*1 convolution to adjust the number of channels to 256 for F1, F2, and F3, and get T1(1, 256, 64, 64), T2(1, 256, 32, 32), and T3(1, 256, 16, 16) respectively; Perform Flatten operation on T1, T2, and T3 and convert them into one-dimensional tensors T1(1, 256, 4096), T2(1, 256, 1024), and T3(1, 256, 256); Concatenate the one-dimensional tensors into a tensor T4(1, 256, 5376) according to the third dimension; Use the self-attention mechanism module to calculate the attention weight of T4, and then further map the features through MLP to obtain feature Out4 (1, 256, 5376); Split Out4 into three tensors according to channels: Out5(1, 256, 64, 64), Out6(1, 256, 32, 32), Out7(1, 256, 16, 16); S4.3 includes: Through the learned adjustable weights, the local features (Out1, Out2, Out3) are weightedly fused with the corresponding global features (Out5, Out6, Out7) to obtain the enhanced multi-scale feature representation.
8. The method according to claim 1, characterized in that In step S5, the 2D detection head maps the fused features into the category, size and position of the object through multiple convolutional layers, and outputs a vector for each position, which includes the four parameters of the predicted object bounding box, the object confidence and the category probability distribution.
9. A low-light scene target detection system with infrared sensing, characterized in that: include: A visible light camera, used to acquire visible light images and read their RGB features; An infrared prediction branch, used for predicting an infrared image based on the RGB features; A backbone network, used for extracting multi-scale features of the infrared image and the visible light image; Complementary fusion filtering module, used to fuse thermal radiation features and visual features; Hybrid feature pyramid structure to enhance the multi-scale fusion of global and local features; The 2D detection head is used to map the fused features into the size, category, and position of 2D objects.
10. The system according to claim 9, characterized in that The infrared prediction branch includes: The decoder module is used in the decoder, with normalization operation, GeLU activation function and GRN structure. The decoder module extracts image features at different levels through multi-scale convolution. It also introduces the Squeeze-and-Excitation attention mechanism, which enables the decoder to focus on the parts of the image that need to be focused on. The complementary fusion filtering module comprises: A visual feature guiding module, used for guiding thermal radiation features using visual features; A thermal radiation feature guiding module, used to guide visual features using thermal radiation features; Feature balancing module, used to balance the two guided features; The hybrid feature pyramid structure includes: A local feature extraction module, used to extract local features; A global feature extraction module, used to extract global features; The feature fusion module is used to fuse local features with global features. The 2D detection head comprises: Convolutional layer, used to map the fused features to the category, size and position of the object; The output module is used to output the parameters of the object bounding box, object confidence and category probability distribution.
Citation Information
Cited By
Photovoltaic hot spot detection method and system based on visible light and thermal infrared image fusion
CN120495302A
Unmanned aerial vehicle small target detection method and device based on YOLOv12 and infrared radiation characteristics
CN121074736A