Infrared Image Pedestrian Reflection Detection Method Based on Deep Learning and Image Mask

Through deep learning and image mask methods, a two-stage object detection model is constructed, which solves the detection problem of pedestrian and pedestrian reflection in infrared images, and realizes efficient and accurate pedestrian reflection area detection, which is suitable for real-time detection in complex backgrounds.

CN115830632BActive Publication Date: 2025-07-22NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211482624.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2025-07-22
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

The existing infrared pedestrian image detection methods are difficult to effectively distinguish pedestrian and pedestrian reflections under complex backgrounds, resulting in poor detection accuracy and speed, especially the single-stage object detection network is not effective in infrared images.

Method used

A two-stage object detection model based on deep learning is adopted, and data preprocessing is carried out by constructing a prediction model, and a distortion-free rectangle filling is used using an image mask. Combined with the first and second stage object detection models, the joint area and reflection location of pedestrians and pedestrian reflections are predicted respectively, and the spatial relationship is preserved.

Benefits of technology

It improves the detection accuracy of pedestrian and pedestrian reflections in infrared images, meets real-time detection requirements, and improves the detection effect in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115830632B_ABST
    Figure CN115830632B_ABST
Patent Text Reader

Abstract

The present invention discloses an infrared image pedestrian reflection detection method based on deep learning and image masking, which can efficiently eliminate the interference of pedestrian reflections in infrared images under complex backgrounds, facilitating the further processing of infrared pedestrian images. The method includes the following steps: In the first stage, a deep learning object detection network model is constructed. By utilizing the good feature extraction ability of convolution for complex contexts, the complex background of the infrared image is removed, and a joint region containing only "pedestrian - pedestrian reflection" is extracted. Then, an image mask is used to perform distortion-free rectangular filling on the joint region to restore the original position information and construct a transition image. In the second stage, a lightweight neural network is designed to obtain the position of the pedestrian reflection from the transition image. The method of the present invention can completely and effectively detect the pedestrian reflection region in the thermal infrared image and retain the spatial relationship between the pedestrian and the pedestrian reflection, which is highly practical and has good prospects for popularization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses an infrared image pedestrian reflection detection method based on deep learning and image masking, belonging to the field of computer vision. Background Art

[0002] Due to the specular echo effect, when an infrared pedestrian image is formed, the thermal radiation emitted by the pedestrian is reflected by the specular material and then received by the thermal infrared imager again, and presented in various forms, resulting in interference from pedestrian reflections, as shown in the appendix. Figure 3 as shown.

[0003] Pedestrian reflections have contours and gradient information very similar to those of pedestrians. Coupled with the low contrast of infrared images themselves, pedestrian reflections are easily mistaken for pedestrians during detection. The infrared pedestrian reflection detection technology is to specifically detect the positions of pedestrians and their corresponding reflections for subsequent use.

[0004] Traditional infrared target detection methods, such as max-mean, image segmentation, and partial differential equations, are restricted by a single scene and cannot obtain an adaptive threshold according to the environment, and the effect is not prominent when dividing the background and foreground.

[0005] Deep learning network models have better feature extraction capabilities for complex image backgrounds and are more suitable for detecting infrared images. In the field of deep learning object detection, single-stage object detection algorithms that unify region selection and object detection are often used, and good results can be obtained in visible light images. Since infrared pedestrian images are monotonous in color, the contours of infrared pedestrians and pedestrian reflections are very similar, and the contrast between the two is very low. If a single-stage object detection network is directly applied to detect infrared pedestrians and pedestrian reflections, the effect is not ideal, which is reflected in that the detection speed and detection accuracy cannot meet the requirements simultaneously. Summary of the Invention

[0006] To solve the above problems, the present invention provides an infrared image pedestrian reflection detection method based on deep learning and image masking, which can meet real-time detection and improve the detection accuracy of pedestrians and pedestrian reflections in infrared pedestrian images with complex backgrounds.

[0007] To achieve the above object, the present invention provides an infrared image pedestrian reflection detection method based on deep learning and image masking, including:

[0008] Construct a prediction model:

[0009] Obtain an infrared image and perform data preprocessing to obtain a feature map.

[0010] Input the feature map into the first-stage object detection model to remove the background and predict the joint region positions of multiple "pedestrian - pedestrian reflection", and cache them as coordinates L1.

[0011] Obtain the number of transitional images to be constructed, perform image masking on the infrared image according to the position of the "pedestrian - pedestrian reflection" joint region, achieve distortion - free rectangular filling, and construct transitional images.

[0012] Input the transitional images into the pre - set second - stage object detection model, predict the coordinates of the "pedestrian reflection" in the infrared image, and cache them as coordinate L2.

[0013] Take out the coordinates L1 and L2 of the two stages, mark the "pedestrian reflection" region in the infrared image according to the coordinate positions, and clear the cache;

[0014] Train the prediction model;

[0015] Input the infrared image to be processed into the trained prediction model to obtain the pedestrian reflection.

[0016] Furthermore, the data pre - processing includes:

[0017] Based on neural network convolution, distribute the spatial information of pixels to each channel, make the dimensions of the input image the same, and integrate infrared images of different sizes into a feature map of a specified size.

[0018] Furthermore, input the feature map into the pre - set first - stage object detection model to remove the complex background, predict the positions of multiple "pedestrian - pedestrian reflection" joint regions, and cache the coordinates L1, including:

[0019] Take feature extraction and dimension compression as one step, process the feature map three times. The feature map after the first processing is taken as a group, the feature maps after the first and second processings are taken as a group, and the feature maps after the first and third processings are taken as a group. Use the spatial pyramid pooling network to integrate the features, and output 3 feature layers with the size of W×H×3(x + y+w + h+confidence+nc), representing three detection frames of different sizes, where (x, y, w, h) are the x - axis coordinate, y - axis coordinate, width, and height of the predicted target center; confidence is the confidence of this target; nc is the number of categories to be predicted;

[0020] Pre - set the head threshold conf for non - maximum suppression determination;

[0021] Calculate the ratio of the area overlap of each target with the target with the highest confidence score, IOU. If conf > IOU, determine them as the same target, otherwise discard the detection frame;

[0022] Introduce candidate boxes, calculate the offset positions for the output features, obtain the final prediction result of the "pedestrian - pedestrian reflection" region, and cache the coordinates L1.

[0023] Further, masking the infrared image according to the "pedestrian - pedestrian reflection" joint region position to achieve distortion - free rectangular filling and constructing a transition image, including:

[0024] Obtaining a single - channel grayscale image of the transition image based on the infrared image and the mask kernel:

[0025]

[0026] where, * represents the dot - product operation, which is used to combine the pixel regions outside the target covered by kernel(i); is the image stitching operation; I is the pixel matrix of the infrared image; kernel(i) is the pixel matrix of the mask kernel of the i - th target with dynamic change; i is the sorting number of the current target; k is the total number of targets included in each transition image;

[0027] Traverse P(i), replace the background pixels outside the target with uniform grayscale pixels to complete the construction of the transition image.

[0028] Further, the construction of k includes:

[0029] Calculating k based on the size of the infrared image:

[0030]

[0031] where, a is the standard initial value of the number of targets in a single image; min(.) takes the minimum of the two; || is the OR operation; [] represents rounding; w and h are the width and height of the infrared image.

[0032] Further, obtaining the number of transition images to be constructed, including:

[0033]

[0034] where, T(n) is the number of transition images to be constructed; n is the total number of targets; k is the number of targets included in each transition image; % is the remainder operation.

[0035] Further, the construction of kernel(i) includes:

[0036] Taking out and traversing the joint region positions in L1 to establish a pixel matrix of the dynamically changing mask kernel kernel(i):

[0037]

[0038] where, L1(i represents the pixel of the i - th target, L1(i is fixed; - 1 is the size of the background pixel outside the target.

[0039] Further, the image stitching operation includes:

[0040] Compare each pixel R at the corresponding positions of P(i) and P(i - 1) P(i) and R P(i-1) to obtain the result pixel R:

[0041]

[0042] where -1 is the size of the background pixel outside the target.

[0043] Furthermore, the steps of presetting the second-stage object detection model include:

[0044] Construct a second-stage object detection model by pruning the first-stage object detection model, so that the second-stage object detection model only outputs two feature layers, and add a determination mechanism to the prediction head of the second-stage object detection model to determine whether there is a pedestrian reflection in the joint region.

[0045] Furthermore, predicting the coordinates of "pedestrian reflection" in the infrared image and caching them as coordinates L2 includes:

[0046] Based on the determination mechanism, determine whether there is a pedestrian reflection in the region. If there is a pedestrian reflection, cache it as coordinates L2; otherwise, determine that there is no pedestrian reflection in the joint region and wait for all the transition images to be output.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides an end-to-end infrared image pedestrian and pedestrian reflection detection method, which uses a two-stage deep learning object detection model and can simultaneously detect the positions of pedestrians and pedestrian artifacts while retaining spatial information. It not only improves the accuracy of the traditional single-stage deep learning object detection model in detecting infrared pedestrian images with complex backgrounds but also meets the requirements of real-time detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a schematic flowchart of a specific embodiment of the method of the present invention.

[0049] Figure 2 is a preview diagram of the end-to-end effect of the present invention.

[0050] Figure 3 is a schematic diagram of the formation principle of pedestrian reflection in an infrared pedestrian image.

[0051] Figure 4 is the network structure of the preprocessing module.

[0052] Figure 5 is the structural diagram of the first-stage "pedestrian - pedestrian reflection" joint region detection structure.

[0053] Figure 6 is the structural diagram of the second-stage "pedestrian reflection" region detection network structure. Detailed implementation manners

[0054] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0055] An infrared image pedestrian reflection detection method based on deep learning and image masking. The overall process of this method is as follows:

[0056] I. Constructing a prediction model:

[0057] As shown in Figure 1 and Figure 2 , first, an infrared thermal imager is used to obtain an infrared image as the input image, the image resolution is set to 640×480, and the temperature resolution is 0.1°F. The input image is subjected to data preprocessing, the image dimensions are normalized, and then the first-stage object detection model is used to predict the "pedestrian - pedestrian reflection" area. Then, an image mask is used to perform distortion-free rectangular filling on the joint area. Finally, the preset second-stage object detection model is used to further predict the "pedestrian reflection" area to detect the position of the "pedestrian reflection" in the infrared image and retain the corresponding spatial relationship between the two.

[0058] Step 1: Obtain an infrared pedestrian image and perform data preprocessing.

[0059] In the data preprocessing part designed by the present invention, referring to the network model shown in the attached Figure 4 , specifically, the spatial information of pixels is dispersed into each channel by using neural network convolution to make the dimensions of the input image the same. In this embodiment, the input image with a size of 640×640×3 is preprocessed to obtain a feature map of 160×160×64 after data preprocessing.

[0060] Step 2: Input the feature map after data preprocessing into the first-stage object detection model to predict the joint area positions of multiple "pedestrians - pedestrian reflections", and cache them as coordinates L1.

[0061] Furthermore, the first-stage deep network model is as shown in the attached Figure 5 . Composite convolution is the basis for constructing a deep learning network and is used to extract complex features of infrared pedestrian images. Composite convolution includes ordinary convolution, batch normalization, and activation functions, and the update of parameters is realized by backpropagation. Since the invention involves knowledge related to deep learning, the results of the input image after passing through each module of the deep learning network will be briefly described below.

[0062] The feature map of 160×160×64 after the above data preprocessing is input into the feature extraction module to extract features. Then, a downsampling convolution with a stride of 2 is used to compress the dimensions of the feature map. As shown in the attached Figure 5As shown, the feature extraction module is stacked three times, and at this time, the size of the feature map changes to 20×20×512. The feature maps after the first processing are grouped as one set, the feature maps after the first and second processing are grouped as one set, and the feature maps after the first and third processing are grouped as one set. The spatial pyramid pooling network with a kernel size of [5, 7, 13] is used to integrate the features.

[0063] Furthermore, the construction of the object regression network is based on the traditional feature pyramid network. The feature map with the size of 20×20×512 mentioned above passes through the feature pyramid network and is concatenated with the backbone feature map according to dim = 1 to achieve feature fusion. Finally, three effective feature layers are output, representing three detection frames of different sizes, with the size of W×H×3x + y + w + h + confidence + nc), where (x, y, w, h) are the predicted center coordinates and width and height of the object, confidence is the confidence of this object, and nc is the number of categories to be predicted. In this embodiment, nc = 1.

[0064] Furthermore, set the prediction head threshold conf = 0.5 to determine non-maximum suppression. Calculate the ratio of the area overlap between each object and the object with the highest confidence score, IOU. If conf > IOU, it is determined as the same object, otherwise, discard this detection frame. Introduce candidate boxes [10, 13, 16, 30, 33, 23], [30, 61, 62, 45, 59, 119], [116, 90, 156, 198, 373, 326] to calculate the offset position for the three feature layers, and obtain the final prediction result of the "pedestrian - pedestrian reflection" area, caching the coordinates L1.

[0065] Step 3: Perform image masking on the input image according to the position of the "pedestrian - pedestrian reflection" joint area to achieve distortion-free rectangular filling and restore the coordinates.

[0066] First, calculate the number T(n) of the transition image P according to the total number n of the objects detected in the first stage.

[0067] Furthermore, the number k of objects to be detected in the transition image should not be too large, otherwise it will affect the accuracy of the neural network, and if it is too small, the detection speed will be reduced. Therefore, the value of k should adaptively change according to the size of the input image. According to the best practice of this embodiment, its expression is,

[0068]

[0069] where w and h are the width and height of the input image, a is the standard initial value of the number of objects in a single image. In this embodiment, a = 4, min(.) takes the minimum of the two, || is the OR operation, and [] represents rounding.

[0070] Furthermore, from the above value of k,

[0071]

[0072] In the formula, % represents the modulo operation, and the number T(n) of transitional images P that require an image mask is calculated.

[0073] In the implementation of the image mask in step 3, the combined region positions in L1 are taken out and traversed, and a pixel matrix of a dynamically changing mask kernel kernel(i) is established. kernel(i) is based on the combined region, the pixels within the region remain unchanged, and the background pixels outside the region will be erased and temporarily replaced with -1 for all. Briefly expressed as,

[0074]

[0075] where L1(i) indicates that the pixels of the i-th target remain unchanged, and the background pixel sizes outside the target will all become -1.

[0076] Furthermore, each transitional image P(i) is a single-channel grayscale image, which is obtained from the input image I and the mask kernel. The expression is,

[0077]

[0078] where * represents dot product, and its function is to cover the areas of no interest with -1. For the image stitching operation, specifically, each pixel R at the corresponding positions of P(i) and P(i - 1) is compared P(i) and R P(i-1) , and the resulting pixel R is obtained. The expression is,

[0079]

[0080] Finally, P(i) is traversed, and those with -1 are replaced with unified grayscale pixels, and the construction of the transitional image is completed.

[0081] From the above, each transitional image P(i) contains at most k target numbers, and the overall transitional image P is composed of T(n) P(i) calculated through different mask kernels.

[0082] Step 4: Input the transitional image P into a preset second-stage object detection model to predict the coordinates of "pedestrian reflection" in the infrared image.

[0083] The second-stage deep learning object detection network needs to detect the position of "pedestrian reflection" from the area that only contains the infrared image "pedestrian - pedestrian reflection". It can be considered that the area of the infrared image "pedestrian - pedestrian reflection" does not have a complex background, and a lightweight network can be constructed to complete the detection task, as shown in the appendix Figure 6 as shown.

[0084] The second-stage network model is completed by pruning on the first-stage model. The main differences are that it only outputs two feature layers and a determination mechanism is added to the prediction head.

[0085] Furthermore, the determination mechanism is used to judge whether there is a pedestrian reflection in the joint area. If so, the coordinate L2 is cached until the network output is completed. If not, it indicates that there is no reflection in the joint area, and wait until the entire transition image P is output.

[0086] Step 5: Take out the cached coordinates L1 and L2 of the two stages, mark the "pedestrian reflection" area of the infrared image according to the coordinate positions, and clear the cache.

[0087] II. Training the prediction model:

[0088] 4215 infrared images of pedestrians were collected using an infrared thermal imager. They were divided into a training set, a test set, and a validation set according to a ratio of 7:2:1, in the form of XML files with the format (x0, y0, w, h), where x0, y0, w, and h represent the position of the center point of the target and the width and height respectively.

[0089] The comparison of the model detection accuracy performance is evaluated using the mean average precision mAP with the highest credibility. The calculation of the accuracy is determined by the precision and recall. When the intersection over union IOU of the prediction box and the ground truth box is greater than a certain set value (in this invention, it is set to 0.5), the prediction is considered correct.

[0090] Furthermore, the mean precision p(r) represents the curve formed by the precision and recall. Since there is only one category, here the mean average precision mAP = AP.

[0091] When training the model, to prevent the model oscillation phenomenon caused by overly random initial weights, the models of the two stages are trained for 10 Epochs respectively, and the weights with the largest mAP are taken as the pre-trained weights.

[0092] Furthermore, the infrared image "pedestrian - pedestrian reflection" detection model of the first stage is trained for 300 Epochs, and the infrared image "pedestrian reflection" detection model of the second stage is trained for 100 Epochs to obtain the final prediction model.

[0093] III. Input the infrared graphics to be processed into the trained prediction model to obtain the pedestrian reflection.

[0094] The present invention designs an infrared image pedestrian reflection detection method based on deep learning and image masking. It avoids the problem of low accuracy of single-stage deep learning object detection models in processing infrared pedestrian images with complex backgrounds. The designed models are all lightweight models, meeting real-time detection requirements and achieving good practical results.

[0095] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, improvements, etc. made within the spirit and principle of the present invention are included within the protection scope of the present invention.

Claims

1. An infrared image pedestrian reflection detection method based on deep learning and image masking, characterized in that: Including: Constructing a prediction model: Obtaining an infrared image and performing data preprocessing to obtain a feature map; Inputting the feature map into the first-stage object detection model to remove the background and predict the joint region positions of multiple "pedestrian-pedestrian reflection", caching them as coordinates L1; Obtaining the number of transition images to be constructed, performing image masking on the infrared image according to the joint region positions of "pedestrian-pedestrian reflection", realizing distortion-free rectangular filling, and constructing transition images; Inputting the transition images into the preset second-stage object detection model to predict the coordinates of "pedestrian reflection" in the infrared image, caching them as coordinates L2; Taking out the coordinates L1 and L2 of the two stages, marking the "pedestrian reflection" region of the infrared image according to the coordinate positions, and clearing the cache; Training the prediction model; Inputting the infrared image to be processed into the trained prediction model to obtain the pedestrian reflection.

2. The infrared image pedestrian reflection detection method based on deep learning and image masking according to claim 1, characterized in that: The data preprocessing includes: Based on neural network convolution, distributing the spatial information of pixels to each channel to make the dimensions of the input image the same, and integrating infrared images of different sizes into a feature map of a specified size.

3. The infrared image pedestrian reflection detection method based on deep learning and image masking according to claim 1, characterized in that: Inputting the feature map into the preset first-stage object detection model to remove the complex background and predict the joint region positions of multiple "pedestrian-pedestrian reflection", caching the coordinates L1, including: Taking feature extraction and dimension compression as one step, processing the feature map three times. The feature map after the first processing is a group, the feature maps after the first and second processing are a group, and the feature maps after the first and third processing are a group. Using a spatial pyramid pooling network to integrate the features, and outputting 3 feature layers with a size of W×H×3(x+y+w+h+confidence+nc), representing three detection frames of different sizes, where (x, y, w, h) are the x-axis coordinate, y-axis coordinate, width, and height of the predicted target center; confidence is the confidence of the target; nc is the number of categories to be predicted; Presetting a head threshold conf for determining non-maximum suppression; Calculating the ratio IOU of the area overlap between each target and the target with the highest confidence score. If conf > IOU, it is determined as the same target, otherwise the detection frame is discarded; Introducing a candidate box, calculating the offset position for the output features, obtaining the final prediction result of the "pedestrian-pedestrian reflection" region, and caching the coordinates L1.

4. The infrared image pedestrian reflection detection method based on deep learning and image masking according to claim 1, characterized in that: The performing image masking on the infrared image according to the joint region positions of "pedestrian-pedestrian reflection", realizing distortion-free rectangular filling, and constructing transition images, includes: Obtaining a single-channel grayscale image of the transition image based on the infrared image and the mask kernel: ; Among them, * represents the dot product operation, which is used to combine the pixel regions outside the coverage of kernel(i); is the image stitching operation; I is the pixel matrix of the infrared image; kernel(i) is the mask kernel pixel matrix of the i-th target with dynamic changes; i is the sorting number of the current target; k is the total number of targets included in each transition image; Traversing P(i), replacing the background pixels outside the target with uniform grayscale pixels to complete the construction of the transition image.

5. The infrared image pedestrian reflection detection method based on deep learning and image masking according to claim 4, wherein: The construction of k includes: Calculating k based on the size of the infrared image: ; Where a is the standard initial value of the number of targets in a single image; min(.) takes the minimum of the two; || is an OR operation; [·] represents rounding; w and h are the width and height of the infrared image.

6. The infrared image pedestrian reflection detection method based on deep learning and image masking according to claim 5, wherein: Obtaining the number of transition images to be constructed, including: ; Where T(n) is the number of transition images to be constructed; n is the total number of targets; k is the number of targets included in each transition image; % is the remainder operation.

7. The infrared image pedestrian reflection detection method based on deep learning and image masking according to claim 4, wherein: The image mask includes: Taking out and traversing the joint region positions in L1, and establishing a dynamically changing mask kernel pixel matrix kernel(i). Kernel(i) is based on the joint region, and the pixels within the region remain unchanged, while the background pixels outside the region are erased and replaced with -1: ; Where L1(i) represents the pixels of the i-th target, and L1(i) is fixed; -1 is the size of the background pixels outside the target.

8. The infrared image pedestrian reflection detection method based on deep learning and image masking according to claim 4, wherein: The image stitching operation includes: Compare each pixel R at the corresponding positions of P(i) and P(i - 1) P(i) and R P(i-1) to obtain the resulting pixel R: ; Where -1 is the size of the background pixels outside the target.

9. The infrared image pedestrian reflection detection method based on deep learning and image masking according to claim 1, wherein: The steps of presetting the second-stage object detection model include: Constructing the second-stage object detection model by pruning the first-stage object detection model, so that the second-stage object detection model only outputs two feature layers, and adding a determination mechanism to the prediction head of the second-stage object detection model to determine whether there is a pedestrian reflection in the joint region.

10. The infrared image pedestrian reflection detection method based on deep learning and image masking according to claim 9, wherein: Predicting the coordinates of "pedestrian reflection" in the infrared image and caching them as coordinates L2, including: Judging whether there is a pedestrian reflection in the region based on the determination mechanism. If there is a pedestrian reflection, cache it as coordinates L2; otherwise, determine that there is no pedestrian reflection in the joint region and wait for all the transition images to be output.

Citation Information

Patent Citations

  • Fall detection method and system based on edge calculation

    CN112906548A

  • Inverted image detection method, device and equipment

    CN113269761A