A Real-time Pedestrian Detection Method for Thermal Infrared Images Based on Weak Salient Maps

Through the real-time thermal infrared image pedestrian detection method based on weakly significant graphs, the lightweight LFFD network is used to generate and combine weakly significant graphs, which solves the shortcomings of traditional methods in small temperature differences and real-time detection, and achieves higher detection accuracy and real-time performance.

CN114170617BActive Publication Date: 2025-06-10WUHAN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010952388.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-11
Publication Date
2025-06-10
Estimated Expiration
2040-09-11

AI Technical Summary

Technical Problem

Traditional thermal infrared pedestrian detection methods have poor detection effects when there is insufficient ambient light or small temperature difference, and the complexity of deep convolutional neural networks makes real-time detection difficult to achieve.

Method used

Using a real-time thermal infrared image pedestrian detection method based on weakly significant images, two-level improvements are made through the lightweight single-object detection network LFFD, weakly significant images are generated and combined with thermal infrared images, and sent to the target detection network to improve detection accuracy, and the detection stability is improved through the fusion of detection results.

Benefits of technology

The accuracy of pedestrian detection in thermal infrared images is improved, and the ability to work in real time with limited hardware resources is achieved, especially in the case of small temperature difference during the day, the detection effect is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170617B_ABST
    Figure CN114170617B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time thermal infrared image pedestrian detection method based on a weak saliency map in the field of pedestrian detection. The method includes the following steps: 1) sending the thermal infrared image into the SD-LFFD network for prediction to generate a preliminary pedestrian detection result and a weak saliency map of the pedestrian area; 2) combining the weak saliency map with the original input thermal infrared image to "light up" the pedestrian area; 3) sending the combined image into the SF-LFFD network for pedestrian detection again to generate a new pedestrian detection result; 4) fusing the pedestrian detection results generated by the above two improved LFFD networks, namely SD-LFFD and SF-LFFD, to obtain the final pedestrian detection result. The present invention can effectively improve the problem of poor pedestrian detection effect from thermal infrared images during the day when the temperature difference between the human body and the background is small, can effectively improve the accuracy of pedestrian detection in thermal infrared images, and can also achieve real-time operation under the condition of limited hardware resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the research field of pedestrian detection, and particularly relates to a real-time thermal infrared image pedestrian detection method based on a weak saliency map. Background Art

[0002] Automatic pedestrian detection technology is widely applied to computer vision tasks such as vehicle-mounted safety systems and video surveillance systems. Pedestrian detection algorithms based on visible light images have poor effects in cases of insufficient or uneven ambient light, while pedestrian detection algorithms based on thermal infrared images are less affected by light conditions due to their thermal radiation imaging principle and are very suitable for all-weather pedestrian detection. Traditional thermal infrared pedestrian detection algorithms are mainly implemented by extracting artificial features and combining classifiers, such as methods that combine artificial features such as Histogram of Oriented Gradient (HOG) and Histogram of Local Intensity Differences (HLID) with an SVM classifier. Such traditional pedestrian detection methods often have disadvantages such as weak robustness and low accuracy due to their dependence on feature design. With the development of deep learning, using deep convolutional neural networks to solve the pedestrian detection problem has become the current mainstream method. Deep convolutional neural networks can automatically learn more reliable and more expressive image features, making pedestrian detection methods have stronger generalization ability and higher detection accuracy. However, for thermal infrared images, the imaging of human body targets is not obvious enough during the day with small temperature differences, which will lead to a deterioration in the detection effect. Ghose used thermal infrared images as the input of a deep convolutional neural network (reference: Ghose D, Desai S M, Bhattacharya S, et al. Pedestrian Detection in Thermal Images using Saliency Maps[C] / / Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. 2019:0-0.), and combined strong saliency map detection to alleviate the problem that the pedestrian area cannot be highlighted from the background area when the temperature difference is not large. However, when there is a missed detection in the saliency map, the pedestrian area is regarded as the background, thus affecting the final detection result. This method uses a complex saliency detection network and performs strong saliency map detection with pixel-level annotation as the label of the saliency map, which takes a long time. In addition, the detection network of this method uses a relatively complex Faster R-CNN network, and real-time detection in practical applications requires relying on expensive hardware costs. Summary of the Invention

[0003] In order to overcome the above-mentioned technical deficiencies, the present invention provides a real-time thermal infrared image pedestrian detection method based on a weak saliency map. This method uses a weakly labeled method to train a saliency detection network to generate a weak saliency map, which is then combined with the thermal infrared image and fed into a target detection network to improve the pedestrian detection accuracy. At the same time, the detection results generated by the saliency detection network and the target detection network are fused to alleviate the impact of missed detections in the saliency map generated by the saliency detection network on the detection results. The present invention can improve the accuracy of pedestrian detection in thermal infrared images, and the entire method can also achieve real-time operation under limited hardware resources.

[0004] In order to achieve the above object, the technical solution proposed in the present invention is: a real-time thermal infrared image pedestrian detection method based on a weak saliency map. In this technical solution, in order to enable the entire algorithm to achieve real-time operation under limited hardware resources, the lightweight single-object detection network LFFD is used as the basic network, which consists of two improved LFFD networks, namely SD-LFFD and SF-LFFD.

[0005] The pedestrian detection method mainly includes the following steps:

[0006] Step 1: Feed the thermal infrared image into SD-LFFD for prediction to generate preliminary pedestrian detection results and a weak saliency map of the pedestrian area.

[0007] Step 2: Combine the weak saliency map with the original input thermal infrared image to "light up" the pedestrian area.

[0008] Step 3: Feed the combined image into the SF-LFFD network for pedestrian detection again to generate new detection results.

[0009] Step 4: Fuse the pedestrian detection results generated by the two-level LFFD network to obtain the final pedestrian detection result.

[0010] In step 1 of the above technical solution, SD-LFFD (Saliency Detection-LFFD) adds a target saliency detection function compared with LFFD. The SD-LFFD network mainly consists of two parts: (1) The target detection part, which has the same structure as LFFD, is mainly used to generate target position information, category information, and confidence; (2) The target saliency detection part, which is modified based on LFFD, is mainly used to generate a weak saliency map to roughly enhance the pedestrian area in the thermal infrared image. Further, the original thermal infrared image is sent into the SD-LFFD network to generate a preliminary pedestrian detection result, and at the same time, a weak saliency map of the pedestrian area is also generated. The target saliency detection part inserts convolutional layers and upsampling layers at the four output branches C11, C14, C17, and C20 in the network structure of the original LFFD, connects the obtained feature maps in the channel dimension, changes the channels through a 1×1 convolutional layer, and then outputs through the sigmod activation function. Finally, the output feature map is scaled by bilinear interpolation to obtain the final saliency map. The loss function of SD-LFFD is:

[0011]

[0012] Among them, i represents the i-th output branch, j represents the j-th pixel point, S represents the area of the current output branch S = w×h, and the first term is the classification loss function L c , and the cross-entropy loss function is used. When the j-th pixel point of the i-th output branch falls into the true box, then c ij = 1, otherwise c ij = 0; the second term is the regression loss function L r , and the L2 loss function is used. t ij represents the relative displacement between the coordinate position corresponding to the receptive field of the current pixel point and the coordinate position of the true box; the third term is the loss function L s of the saliency detection part, and the cross-entropy loss function is used. k represents the k-th pixel point, p represents the label of the saliency map, p k = 1 in the pedestrian area, and p k = 0 in the background area.

[0013] In step 2 of the above technical solution, the original thermal infrared image input to the LFFD network is in RGB format, but the pixel values of the 3 channels are the same, and it is actually a grayscale image. To keep the number of input channels of the LFFD network unchanged, therefore, in this step, two of the channels are taken and combined with the weak saliency map generated by SD-LFFD to form a new three-channel image.

[0014] In step 3 of the above technical solution, SF-LFFD (Saliency Fusion-LFFD) is an LFFD network that further detects by fusing the above weak saliency map information. Its input is the weak saliency map and the original thermal infrared image, and the output is the target position information, category information, and confidence. The loss function of SF-LFFD is:

[0015]

[0016] The difference between the loss function of the SF-LFFD network and that of the SD-LFFD is that the loss function L of the saliency detection part is removed s , while the classification loss function L c and the regression loss function L r are retained.

[0017] In step 4 of the above technical solution, the pedestrian detection results generated by the two-stage improved LFFD networks, namely SD-LFFD and SF-LFFD, are fused to achieve the complementarity of the two methods to obtain more accurate pedestrian detection results. The formula for fusing the two-stage networks is:

[0018]

[0019] where the confidence and position information generated by SD-LFFD are respectively expressed as C SD-LFFD and B SD-LFFD , the confidence and position information generated by SF-LFFD are respectively expressed as C SF-LFFD and B SF-LFFD , and the confidence and position information of the final output of the pedestrian category are C out and B out . Further, in the above formula, take In this technical solution, using the two-stage improved LFFD network for pedestrian detection is equivalent to deepening the original LFFD network structure, enhancing the network's information processing and feature expression capabilities.

[0020] Compared with the prior art, the present invention has the following advantages:

[0021] (1) Designed a weak saliency detection network structure for pedestrians, which can generate a weak saliency map of pedestrians while generating pedestrian detection results; and combines the detected weak saliency map with the original thermal infrared image, enabling the deep network to pay more attention to the potential areas of pedestrians, thereby improving the problem of poor pedestrian detection effect in thermal infrared images during the day when the temperature difference between the human body and the background is small;

[0022] (2) Fusing the pedestrian detection results generated by the two-stage improved LFFD network can achieve the complementarity of the detection results of the two-stage network, which helps to improve the overall detection accuracy of the algorithm. Since the lightweight object detection network LFFD is used as the basic network, even if the pedestrian detection is completed by fusing the two-stage LFFD network, real-time detection can still be achieved under limited hardware resources. Description of the Drawings

[0023] Figure 1 is the flowchart of the real-time thermal infrared image pedestrian detection method based on weak saliency map of the present invention

[0024] Figure 2 is the structural diagram of the LFFD pedestrian detection network

[0025] Figure 3 is the structural diagram of the saliency detection part in SD-LFFD Detailed Embodiment

[0026] The technical solution of the present invention will be further described below with reference to the drawings.

[0027] See Figure 1 , a real-time thermal infrared image pedestrian detection method described in the present invention includes the following steps:

[0028] Step 1: Send the thermal infrared image into SD-LFFD for prediction to generate preliminary pedestrian detection results and a weak saliency map of the pedestrian area.

[0029] Step 2: Combine the weak saliency map with the original input thermal infrared image to "light up" the pedestrian area.

[0030] Step 3: Send the combined image into the SF-LFFD network for pedestrian detection again to generate new detection results.

[0031] Step 4: Fuse the pedestrian detection results generated by the two-stage LFFD ( Figure 2 is the structural diagram of the LFFD network based on pedestrian detection) network to obtain the final pedestrian detection result.

[0032] In step 1 of the above technical solution, SD-LFFD (Saliency Detection-LFFD) adds a target saliency detection function compared with LFFD. The SD-LFFD network mainly consists of two parts: (1) The target detection part, which has the same structure as LFFD, is mainly used to generate target position information, category information, and confidence; (2) The target saliency detection part, which is modified based on LFFD, is mainly used to generate a weak saliency map to roughly enhance the pedestrian area in the thermal infrared image. Further, the original thermal infrared image is sent into the SD-LFFD network to generate a preliminary pedestrian detection result, and at the same time, a weak saliency map of the pedestrian area is also generated. The target saliency detection part inserts a convolutional layer and an upsampling layer at the four output branches C11, C14, C17, and C20 in the network structure of the original LFFD, connects the obtained feature maps in the channel dimension, changes the channels through a 1×1 convolutional layer, and then outputs through a sigmod activation function. Finally, the output feature map is scaled by bilinear interpolation to obtain the final saliency map. The network structure of the SD-LFFD saliency detection part is shown in Figure 3 as follows. The loss function of SD-LFFD is:

[0033]

[0034] where i represents the i-th output branch, j represents the j-th pixel, S represents the area of the current output branch S = w×h, the first term is the classification loss function L c , and the cross-entropy loss function is used. When the j-th pixel of the i-th output branch falls within the ground truth box, then c ij = 1, otherwise c ij = 0; the second term is the regression loss function L r , and the L2 loss function is used. t ij represents the relative displacement between the coordinate position of the receptive field corresponding to the current pixel and the coordinate position of the ground truth box; the third term is the loss function L s of the saliency detection part, and the cross-entropy loss function is used. k represents the k-th pixel, p represents the label of the saliency map, p k = 1 for the pedestrian area, and p k = 0 for the background area.

[0035] In step 2 of the above technical solution, the original thermal infrared image input to the LFFD network is in RGB format, but the pixel values of the three channels are the same, and it is essentially a grayscale image. To keep the number of input channels of the LFFD network unchanged, therefore, in this step, two of the channels are taken and combined with the weak saliency map generated by SD-LFFD to form a new three-channel image.

[0036] In step 3 of the above technical solution, SF-LFFD (Saliency Fusion-LFFD) is an LFFD network for further detection by fusing the above-mentioned weakly salient map information. Its input is the weakly salient map and the original thermal infrared image, and the output is the target position information, category information, and confidence. The loss function of SF-LFFD is:

[0037]

[0038] The difference between the loss function of the SF-LFFD network and that of the SD-LFFD is that the loss function L of the saliency detection part is removed s , while the classification loss function L c and the regression loss function L r are retained.

[0039] In step 4 of the above technical solution, the pedestrian detection results generated by the two-stage improved LFFD networks, namely SD-LFFD and SF-LFFD, are fused to achieve the complementarity of the two methods to obtain more accurate pedestrian detection results. The formula for fusing the two-stage networks is:

[0040]

[0041] Among them, the confidence and position information generated by SD-LFFD are respectively represented as C SD-LFFD and B SD-LFFD , the confidence and position information generated by SF-LFFD are respectively represented as C SF-LFFD and B SF-LFFD , and the confidence and position information of the final output of the pedestrian category are C out and B out . Further, in the above formula, take

[0042] To verify the effectiveness of the technical solution of the present invention, pedestrian detection experiments are carried out on two typical thermal infrared image pedestrian datasets, CVC-09 and CVC-14. By comparing the original LFFD (ORI-LFFD, Original LFFD), SD-LFFD, SF-LFFD, SD-LFFD+SF-LFFD in the technical solution of the present invention, and at the same time, to reflect the superiority of the present invention in the same lightweight network, the method in the technical solution is compared with the lightweight object detection network Tiny-YOLOv3.

[0043] Table 1 Comparison of pedestrian detection AP values (%)

[0044]

[0045] Table 1 lists the AP values of pedestrian detection experiments on different networks, where Day, Night, and Total represent three test scenarios of the daytime, night-time, and the overall dataset in the dataset. Compared with the original LFFD network (i.e., ORI-LFFD), the method proposed in the present invention (i.e., SD-LFFD+SF-LFFD) has improved the overall detection effect by nearly 5% on the CVC-09 dataset and 11% on the CVC-14 dataset. Since the temperature difference between the human body and the environment is smaller during the day than at night, the detection effect during the day is often worse than at night. After using the method proposed in the present invention, the detection accuracy during the day and at night has been improved, and the improvement during the day is more obvious, especially on the CVC-14 dataset, where it has increased by 13%. Therefore, the present invention can, to a certain extent, alleviate the problem of poor detection effect of thermal infrared images during the day. In addition, the AP value of SF-LFFD is better than that of ORI-LFFD and SD-LFFD in different datasets and different test scenarios, which shows that the weak saliency map used in the present invention is helpful for improving the object detection effect of the SF-LFFD network. In addition, in the CVC09 dataset, for different test scenarios of Day, Night, and Total, the AP value of the method proposed in the present invention is higher than that of Tiny-YOLOv3. In the CVC14 dataset, in the test scenario of Day, the method proposed in the present invention is slightly behind, but in the test scenario of Night, the AP value of the method proposed in the present invention is about 10% higher, and in the overall dataset, the method proposed in the present invention performs better. This shows that the present invention has a certain accuracy advantage in the same lightweight object detection network.

[0046] Table 2 Speed comparison of different pedestrian detection methods

[0047] Pedestrian detection network Model size (M) Frame rate (FPS) Inference speed (ms) Tiny-YOLOv3 33.99 18.31 54.61 The method proposed by the present invention 14.45 31.25 32

[0048] After testing, the average frame rate of the method proposed in the present invention is 31FPS, that is, the time to process one frame of image is about 0.032s. Compared with Tiny-YOLOv3, the method proposed in the present invention has a large lead in speed, and at the same time shows that the method in this paper can work in real time under limited hardware resources. This benefits from the use of a simple and easy-to-implement object weak saliency detection algorithm and an improved lightweight LFFD network.

[0049] Although relatively detailed implementation schemes are used as illustrations in this paper, they are not used to limit the present invention. Those skilled in the art of the present invention can make various modifications and refinements to the specific implementations described, but the present invention also intends to include these changes and variations.

[0050] The content not described in detail in this specification belongs to the prior art well known to those skilled in the art.

Claims

1. A real-time pedestrian detection method for thermal infrared images based on a weak saliency map, characterized in that, it includes the following steps: Step 1, sending the thermal infrared image into the SD-LFFD network for prediction to generate a preliminary pedestrian detection result and a weak saliency map of the pedestrian area; Step 2, combining the weak saliency map with the original input thermal infrared image to "light up" the pedestrian area; Step 3, sending the combined image into the SF-LFFD network for pedestrian detection again to generate a new detection result; Step 4, fusing the pedestrian detection results generated by the above two-stage improved LFFD networks, namely SD-LFFD and SF-LFFD, to obtain the final pedestrian detection result; The SD-LFFD network in step 1 is an improved structure of the LFFD network. Compared with the LFFD network, it adds a target saliency detection function and consists of two parts: (1) The target detection part, which has the same structure as the LFFD, is used to generate target position information, category information, and confidence; (2) The target saliency detection part, which is modified based on the LFFD, is used to generate a weak saliency map to roughly enhance the pedestrian area in the thermal infrared image. The target saliency detection part inserts convolutional layers and upsampling layers at the four output branches C11, C14, C17, and C20 in the original LFFD network structure, connects the obtained feature maps in the channel dimension, changes the channels through a 1×1 convolutional layer, then outputs through a sigmod activation function, and finally scales the output feature map using bilinear interpolation to obtain the final saliency map. The loss function of SD-LFFD is where i represents the i-th output branch, j represents the j-th pixel, S represents the area of the current output branch S = w×h, and the first term is the classification loss function L c , and the cross-entropy loss function is used. When the j-th pixel of the i-th output branch falls within the ground truth box, then c ij = 1, otherwise c ij = 0; the second term is the regression loss function L r , and the L2 loss function is used. t ij represents the relative displacement between the coordinate position corresponding to the receptive field of the current pixel and the coordinate position of the ground truth box; the third term is the loss function L s of the saliency detection part, and the cross-entropy loss function is used. k represents the k-th pixel, p represents the label of the saliency map, for the pedestrian area p k = 1, and for the background area p k = 0; The SF-LFFD network in step 3 is another improved structure of the LFFD network. The inputs are the weak saliency map and the original thermal infrared image, and the outputs are the target location information, category information, and confidence. The loss function is The difference in the loss function compared to the SD-LFFD network is that the loss function L of the saliency detection part is removed s , while the classification loss function L c and the regression loss function L r are retained.

2. The real-time pedestrian detection method for thermal infrared images based on a weak saliency map according to claim 1, characterized in that, the implementation process of Step 2 is as follows: The original thermal infrared image input to the LFFD network is in RGB format, but the pixel values of the three channels are the same, and it is essentially a grayscale image. To keep the number of input channels of the LFFD network unchanged, in Step 2, two of the channels are taken and combined with the weak saliency map generated by SD-LFFD to form a new three-channel image.

3. The real-time pedestrian detection method for thermal infrared images based on a weak saliency map according to claim 1, characterized in that, the implementation process of Step 4 is as follows: The pedestrian detection results generated by the two-stage improved LFFD networks, namely the SD-LFFD network and the SF-LFFD network, are fused to achieve the complementarity of the two methods to obtain a more accurate pedestrian detection result. The formula for fusing the two-stage networks is: Among them, the confidence and location information generated by the SD-LFFD network are represented as C SD-LFFD and B SD-LFFD , the confidence and location information generated by the SF-LFFD network are represented as C SF-LFFD and B SF-LFFD , the confidence and location information finally output for the pedestrian category are C out and B out , further, take

Citation Information

Patent Citations

  • Pedestrian detection method based on CNN (Convolutional Neural Network) and semantic segmentation

    CN108399361A

  • Method and Apparatus for Detecting Target Objects in Images

    US20200234072A1