An infrared image vehicle detection method

Through the deep learning Unet network model combined with static and dynamic object detection methods, the problem of poor detection accuracy in infrared image vehicle detection is solved, more efficient vehicle detection and recognition is achieved, and the accuracy and efficiency of infrared monitoring system is improved.

CN115393254BActive Publication Date: 2025-06-10ZHENGZHOU XINDA ADVANCED TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110573859.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-25
Publication Date
2025-06-10
Estimated Expiration
2041-05-25

AI Technical Summary

Technical Problem

The prior art has poor detection accuracy in infrared image vehicle detection, especially when the vehicle edges are blurred, which is easy to lead to missed detection and missed detection.

Method used

The deep learning Unet network model is adopted to combine static and dynamic object detection methods, and infrared image vehicle detection is realized through single-frame image detection and sequence frame image detection. The method includes normalizing the single-frame image in the infrared monitoring video stream, and performing image cropping and repeated detection if no vehicle information is detected; at the same time, performing motion detection method processing on the sequence frame image, and performing vehicle detection after detecting the target.

Benefits of technology

It improves the accuracy of infrared image vehicle detection, can more effectively identify and track vehicles, and enhances the effectiveness of urban management and military surveillance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393254B_ABST
    Figure CN115393254B_ABST
Patent Text Reader

Abstract

The present invention relates to an infrared image vehicle detection method, belonging to the technical field of image processing. In the present invention, a single-frame image is obtained from an infrared monitoring video stream and input into a deep learning network model for vehicle detection; at the same time, a sequence of frame images is obtained from the infrared real-time monitoring video stream for motion detection; if a target is detected, it is input into the deep learning network model for vehicle detection. The present invention combines two methods of static target detection and moving target detection, and adopts the methods of single-frame image detection and sequence-frame image detection, realizing the detection of vehicle targets under infrared imaging conditions and improving the accuracy of infrared image vehicle detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting vehicles in infrared images, belonging to the technical field of image processing. Background Art

[0002] In recent years, infrared monitoring systems have been widely used in monitoring systems due to their characteristics such as concealment, anti-interference, and all-weather operation, and particularly play an important role in the field of military monitoring. Infrared imaging systems can not only detect targets through obstacles such as smoke, dust, and fog, but also have a long operating range. Using infrared imaging to achieve target detection is an important part of the modern optoelectronic technology field and has wide applications in aspects such as border and coastal defense monitoring and urban management.

[0003] Infrared imaging devices are different from visible light imaging devices. There is non-uniformity in the response of photosensitive elements in infrared imaging devices. Therefore, an infrared imaging system cannot simply be considered a linear shift-invariant system. The infrared image obtained through an infrared imaging device is a set composed of a real scene image, imaging noise, and imaging interference. For a generally relatively flat background, the local radiation intensity of an infrared target is much greater than the radiation intensity of background clutter. Therefore, for infrared image target detection, spatial domain filtering methods are mostly used.

[0004] At present, there is an urgent need to implement vehicle detection in infrared images. Obtaining vehicle information in nighttime infrared monitoring, recording the time and direction of vehicle appearance, and identifying and tracking vehicles play a significant role in the handling of traffic accidents in urban management and are of great significance in recording vehicle information that appears at night in military monitoring. In infrared images, the edges of vehicles are blurred and it is not easy to distinguish the contour boundaries. It is very difficult to implement vehicle detection in infrared images using traditional image processing methods, and it is very easy to cause missed detections and false detections. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for detecting vehicles in infrared images to solve the problem of poor detection accuracy in current vehicle detection in infrared images.

[0006] The present invention provides a method for detecting vehicles in infrared images to solve the above technical problems. The method includes the following steps:

[0007] 1) Obtain a single-frame image from an infrared monitoring video stream and perform normalization processing on the single-frame image;

[0008] 2) Use a deep learning Unet network model to detect the single-frame image after normalization processing. The deep learning Unet network model is obtained by training with training samples, and the training samples include infrared images with vehicles and infrared images without vehicles;

[0009] 3) Determine the result of the detection in step 2). If no vehicle information is output, perform an image cropping operation on the single-frame image to obtain a series of cropped images, and input the series of cropped images into the deep learning Unet network model for detection; if vehicle information is output, the vehicle detection in the infrared image is completed.

[0010] 4) Obtain a sequence of frame images from the infrared real-time monitoring video stream and perform normalization processing on the sequence of frame images.

[0011] 5) Use the motion detection method to process the normalized sequence of frame images to obtain the width, height, and upper left corner coordinates of the minimum bounding rectangle of each contour determined by the difference images of any two consecutive frames. Crop the latter frame image in two consecutive frames according to the width, height, and upper left corner coordinates of the minimum bounding rectangle, and input the cropped image into the deep learning Unet network model in step 2) for detection.

[0012] The present invention obtains a single-frame image from an infrared monitoring video stream and inputs it into a deep learning Unet network model for vehicle detection; at the same time, it obtains a sequence of frame images from the infrared real-time monitoring video stream for motion detection; if a target is detected, it is input into the deep learning Unet network model for vehicle detection. The present invention combines two methods of static target detection and dynamic target detection, and adopts the methods of single-frame image detection and sequence-frame image detection to realize the detection of vehicle targets under infrared imaging conditions, and improves the accuracy of vehicle detection in infrared images.

[0013] Further, to solve the problem of poor detection accuracy caused by the small imaging of vehicle targets in infrared images, the deep learning Unet network model includes an encoder, an intermediate layer, and a decoder.

[0014] Further, when the deep learning Unet network model detects a single-frame image, it is used to output the region of interest of the vehicle, calculate the minimum bounding rectangle of the region of interest of the vehicle, and mark the vehicle in the single-frame image according to the coordinates of the minimum bounding rectangle.

[0015] Further, the calculation process of the minimum bounding rectangle is as follows:

[0016] A. Determine the four endpoints of the region of interest of the vehicle.

[0017] B. Construct four tangents of the region of interest of the vehicle through the four endpoints.

[0018] C. If one or two tangents coincide with one side of the region of interest of the vehicle, then calculate the area of the rectangle determined by the four tangents and save it as the current minimum area value, otherwise define the current minimum area value as infinity.

[0019] D. Rotate the line clockwise until one of the tangents coincides with one side of the polygon where the region of interest of the vehicle is located;

[0020] E. Calculate the area of the rotated rectangle and compare it with the current minimum area. If it is less than the current minimum area, update the minimum area and save the rectangle information corresponding to the minimum area;

[0021] F. Repeat steps D and E until the angle by which the line has rotated is greater than 90 degrees;

[0022] G. Take the rectangle corresponding to the final minimum area as the minimum bounding rectangle of the region of interest of the vehicle.

[0023] Furthermore, the process of processing the normalized sequence of frame images by the motion detection method in step 5) is as follows:

[0024] a. Obtain images at any two consecutive moments. Take the image at the previous moment as the background image and perform grayscale processing on the two consecutive frames of images;

[0025] b. Perform a difference operation on the two grayscale images to obtain a difference image, and perform thresholding on the difference image to retain the pixels with a grayscale difference greater than the set threshold to form a threshold image;

[0026] c. Use a convolution operation to perform erosion on the threshold image to obtain an eroded image;

[0027] d. Perform dilation on the eroded image to obtain a dilated image, and perform contour finding on the dilated image;

[0028] e. Traverse each contour to determine the width, height, and top-left coordinates of the minimum bounding rectangle of each contour.

[0029] Furthermore, the normalization in steps 2) and 4) both uses the bilinear interpolation method. Description of the Drawings

[0030] Figure 1 is a flowchart of the infrared image vehicle detection method of the present invention;

[0031] Figure 2-a is a schematic diagram of the semantic segmentation result of the image using the Unet network in an embodiment of the present invention;

[0032] Figure 2-b is a schematic diagram of the color value translation of the semantic segmentation of the image using the Unet network in an embodiment of the present invention;

[0033] Figure 2-c is a schematic diagram of the region of interest obtained after processing the image using the Unet network in an embodiment of the present invention. Detailed Embodiments

[0034] The specific implementation manners of the present invention will be further described below in conjunction with the accompanying drawings.

[0035] Method Embodiment

[0036] To solve the problem of difficult detection of vehicles in infrared images, the present invention combines two methods of static target detection and moving target detection, and adopts the methods of single-frame image detection and sequence-frame image detection. The implementation process of this method is as Figure 1 shown. A single-frame image is obtained from the infrared real-time monitoring video stream and input into the deep learning Unet network model for vehicle detection; if vehicle information is output, the vehicle detection in the infrared image is completed; if no vehicle information is output, the image is randomly cropped and the cropped result is input into the deep learning Unet network model for detection again; at the same time, sequence-frame images are obtained from the infrared real-time monitoring video stream for motion detection; if a target is detected, it is input into the deep learning network model for vehicle detection; if no target is detected, sequence-frame images are continuously obtained. The specific implementation process of this method is as follows.

[0037] 1. Obtain a single-frame infrared image to be measured, and perform normalization processing on the obtained single-frame infrared image.

[0038] A single-frame image is obtained from the infrared monitoring real-time video stream to be measured, that is, the original image, with a size of 480*640. The original image is normalized by the bilinear interpolation method to obtain an image I n ( x , y )), where x ∈[0,511], y ∈[0,511].

[0039] Specifically, the bilinear interpolation calculation formula for the image is:

[0040] I n ( x , y ) = I ( x -1, y -1) × (1 - x )(1 - y ) + I ( x +1, y -1) × ( x (1 - y ) +

[0041] I ( x -1, y( + 1)×(1 - x ) y+ I ( x + 1, y + 1)× xy

[0042] where I ( x - 1, y - 1), I ( x + 1, y - 1), I ( x - 1, y + 1), I ( x + 1, y + 1) are the original image pixel values, I n ( x , y ) are the new pixel values obtained by normalization. If there is no pixel value at this coordinate position in the original image, it is assigned a value of 0.

[0043] 2. Establish a deep learning Unet network model to detect the normalized single-frame infrared image.

[0044] Considering that the infrared lens has a long observation distance and the vehicle target has a small image, in this embodiment, a deep learning Unet network model is used for detection. The image obtained by normalization I n is input into the deep learning semantic segmentation network Unet to obtain the semantic segmentation result image I seg . According to the semantic segmentation result image, color value translation is performed to extract the region of interest of the vehicle. Calculate the minimum bounding rectangle for the region of interest, and the vehicle bounding box information can be obtained. According to the vehicle bounding box information, draw in the image I n and mark the text information "vehicle", thus realizing the vehicle detection of the infrared image.

[0045] The deep learning model (Unet network model) in this embodiment is a trained model, and the training samples used are the normalized images containing vehicles and not containing vehicles, where the images containing vehicles cover various types of vehicles.

[0046] Among them, the Unet network model consists of three parts: an encoder, an intermediate layer, and a decoder. The encoder consists of four convolutional modules, each convolutional module consists of two 3*3 convolutions, and downsamples by 2 times; the number of input channels of the four convolutional modules of the encoder is (3, 64, 128, 256) respectively, and the number of input filters is (64, 128, 256, 512) respectively; the image is input into the Unet network, and a 512-dimensional feature map is obtained after being processed by the encoder. The intermediate layer consists of two convolutional modules, each convolutional module consists of a 1*1 convolution, a batch processing layer, and a relu activation function layer, the number of input channels is (512, 1024) respectively, and the number of input filters is (1024, 1024) respectively; according to the above 512-dimensional feature map, a 1024-dimensional feature map is obtained after being processed by the intermediate layer. The decoder consists of four convolutional modules, each convolutional module consists of two 3*3 convolutions, and upsamples by 2 times, gradually restoring to the size of the input image; the number of input channels of the four convolutional modules is (1024, 512, 256, 128) respectively, and the number of input filters is (512, 256, 128, 64) respectively; according to the above 1024-dimensional feature map, a 64-dimensional feature map is obtained after being processed by the decoder.

[0047] Finally, the above 64-dimensional feature map is input into a 1*1 convolution with 5 channels and 64 filters to obtain a 5-dimensional mask image. Each dimension of the mask image represents the prediction of a label category, and there are 5-dimensional predictions for 5 labels. Then, a 1*1 convolution operation with 1 channel and 5 filters is performed to obtain a 1-dimensional mask image. According to the color value correspondence relationship, the 1-dimensional mask image is mapped into a color image, that is, the semantic segmentation result image I seg 。

[0048] In this embodiment, semantic segmentation labels are defined, 0 represents other; 1 represents vehicle; 2 represents road; 3 represents sky; 4 represents building. Define the color value correspondence relationship of the semantic segmentation result. If it is assumed that in this embodiment, the Unet network is used to detect a certain frame of infrared image, and the semantic segmentation result is as Figure 2-a shown, each square represents a superpixel block. At the junction of classification categories, the superpixel blocks that cannot be clearly judged to belong to a certain category are represented by the "other" category, that is Figure 2-a the superpixel blocks representing "other" in. According to the defined semantic segmentation labels, the color values of the semantic segmentation result output by the Unet network are translated, and the translation result is as Figure 2-b shown. Pay attention to the region of interest of the vehicle, such as Figure 2-b the black frame mark in. All non-"1" superpixel blocks are forced to be "0", and the superpixel blocks with "1" are retained, that is, the region of interest of the vehicle is extracted, such as Figure 2-c shown.

[0049] Since the region of interest of the vehicle output by the above model is generally an irregular polygon, it is necessary to calculate the minimum bounding rectangle of the region of interest of the vehicle to obtain the coordinate values of the four points corresponding to the minimum area of the bounding rectangle; according to the coordinate values of the four points, mark in the image I n and connect them with line segments, that is, draw the bounding box of the vehicle; mark the text "vehicle" in the upper left corner of the bounding box, that is, complete the vehicle detection of the infrared image. The determination process of the minimum bounding rectangle is as follows:

[0050] A. Calculate the four end points of the region of interest of the vehicle, which generally refer to the four extreme points of the upper left, lower left, upper right, and lower right of the polygon where the region of interest of the vehicle is located.

[0051] B. Construct four tangent lines of the region of interest of the vehicle through the four end points.

[0052] C. If one (or two) lines coincide with one side, then calculate the area of the rectangle determined by the four lines and save it as the current minimum value. Otherwise, define the current minimum value as infinity.

[0053] D. Rotate the line clockwise until one of them coincides with one side of the polygon.

[0054] E. Calculate the area of the new rectangle and compare it with the current minimum value. If it is less than the current minimum value, update it and save the rectangle information that determines the minimum value.

[0055] F. Repeat steps D and E until the angle by which the line has rotated is greater than 90 degrees.

[0056] G. Output the coordinate values of the four points corresponding to the minimum area of the bounding rectangle.

[0057] 3. Determine whether there is vehicle information in the single-frame infrared image detected.

[0058] If vehicle information is detected in step 2, the vehicle detection of the current infrared image is completed; if no vehicle information is detected, it may be because the image is too large and the target is too small. Therefore, the present invention performs a random image cropping operation on the image I n and randomly crops out 6 images of size 224*224 I ci , i ∈[0,5], and input the cropped images I ci into the Unet network model for detection in sequence.

[0059] 4. Obtain sequential frame images from the infrared real-time monitoring video stream and perform normalization processing on the sequential frame images.

[0060] Obtain sequential frame images from the infrared real-time monitoring video stream I t , I t-1 , and perform normalization processing using the bilinear interpolation method to obtain an image with a size of 512*512 I nt , I n(t-1) . The formula adopted by the bilinear interpolation method is the same as the formula for normalizing a single-frame infrared image in step 1, which will not be elaborated here

[0061] 5. Detect the normalized sequential frame images using the motion detection method

[0062] The motion detection method is to process two consecutive frames of images. Use the previous frame of image as the background image to obtain the frame difference image of the two frames of images, and process the frame difference image. The specific detection process is as follows

[0063] In this embodiment, t the image at time -1 is used as the background image, and the two frames of images are grayscaled to obtain grayscale images G nt and G n(t-1) . Then perform a difference operation on the two grayscale images to obtain a difference image D nt ; perform thresholding on the difference image D nt , set the threshold to 50, that is, pixels with a gray difference greater than 50 are retained to obtain a threshold image T nt ; perform an erosion operation on the threshold image T nt , set the convolution kernel size to (3,3) to obtain an eroded image E nt . Perform a dilation operation on the eroded image E nt , set the convolution kernel size to (18,18) to obtain a dilated image DL nt ; perform a contour finding operation on the dilated image. Traverse each contour, determine the minimum bounding rectangle of each contour, obtain the width, height and the coordinates of the upper left corner of the minimum bounding rectangle, in the image DL nt I nt ​Draw a regular circumscribed rectangle on it, crop the image according to the width, height and the upper left corner coordinates of the regular circumscribed rectangle, and input the cropped image into the Unet network for vehicle detection. The regular circumscribed rectangle of each contour refers to its minimum circumscribed rectangle, and the method used is the same as the calculation method of the minimum regular circumscribed rectangle in step 2. The process of vehicle detection by the Unet network model is the same as the detection process in step 2, which will not be elaborated here.

[0064] Through the above process, the present invention combines two methods of static target detection and dynamic target detection, and adopts the methods of single-frame image detection and sequence-frame image detection, thereby improving the accuracy of vehicle detection in infrared images.

[0065] System embodiment

[0066] The device proposed in this embodiment includes a processor and a memory. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the method of the above method embodiment. That is to say, the method in the above method embodiment should be understood that the process of the vehicle detection method in infrared images can be realized by computer program instructions. These computer program instructions can be provided to the processor, so that by executing these instructions by the processor, the functions specified by the above method process can be generated.

[0067] The processor referred to in this embodiment refers to a processing device such as a microprocessor MCU or a programmable logic device FPGA, and a GPU can also be used to implement it; the memory referred to in this embodiment includes a physical device for storing information, usually after digitizing the information, it is stored by means of media such as electricity, magnetism or optics. For example: various memories that store information by means of electric energy, such as RAM, ROM, etc.; various memories that store information by means of magnetic energy, such as hard disks, floppy disks, magnetic tapes, magnetic core memories, bubble memories, USB flash drives; various memories that store information by means of optical methods, such as CDs or DVDs. Of course, there are other types of memories, such as quantum memories, graphene memories, etc.

[0068] The device composed of the above memory, processor and computer program is realized by the processor executing corresponding program instructions in the computer. The processor can be equipped with various operating systems, such as the windows operating system, the linux system, android, the iOS system, etc. As another implementation, the device may further include a display for displaying the synchronous detection results for the reference of the staff.

Claims

1. An infrared image vehicle detection method, characterized in that, the detection method comprises the following steps: 1) Obtain a single-frame image from an infrared monitoring video stream and perform normalization processing on the single-frame image; 2) Use the deep learning Unet network model to detect the single-frame image after normalization processing. The deep learning Unet network model is obtained by training with training samples, and the training samples include infrared images with vehicles and infrared images without vehicles; 3) Determine the detection result of step 2). If no vehicle information is output, perform an image cropping operation on the single-frame image to obtain a series of cropped images, and input the series of cropped images into the deep learning Unet network model for detection; if vehicle information is output, the infrared image vehicle detection is completed; 4) Obtain a sequence of frame images in the infrared real-time monitoring video stream and perform normalization processing on the sequence of frame images; 5) Use the motion detection method to process the normalized sequence of frame images to obtain the width, height, and upper left corner coordinates of the minimum circumscribed rectangle of each contour determined by the difference image of any two consecutive frames. Crop the latter frame image in two consecutive frames according to the width, height, and upper left corner coordinates of the minimum circumscribed rectangle, and input the cropped image into the deep learning Unet network model in step 2) for detection.

2. The infrared image vehicle detection method according to claim 1, characterized in that, the deep learning Unet network model includes an encoder, an intermediate layer, and a decoder.

3. The infrared image vehicle detection method according to claim 2, characterized in that, when the deep learning Unet network model detects a single-frame image, it is used to output the region of interest of the vehicle, calculate the minimum circumscribed rectangle of the region of interest of the vehicle, and mark the vehicle in the single-frame image according to the coordinates of the minimum circumscribed rectangle.

4. The infrared image vehicle detection method according to claim 3, characterized in that, the calculation process of the minimum circumscribed rectangle is as follows: A. Determine the four endpoints of the region of interest of the vehicle; B. Construct four tangents of the region of interest of the vehicle through the four endpoints; C. If one or two tangents coincide with one side of the region of interest of the vehicle, then calculate the area of the rectangle determined by the four tangents and save it as the current minimum area value, otherwise define the current minimum area value as infinity; D. Rotate the line clockwise until one of the tangents coincides with one side of the polygon where the region of interest of the vehicle is located; E. Calculate the area of the rotated rectangle and compare it with the current minimum area value. If it is less than the current minimum area value, update the minimum area value and save the rectangle information corresponding to the minimum area value; F. Repeat steps D and E until the angle by which the line has rotated is greater than 90 degrees; G. Take the rectangle corresponding to the final minimum area value as the minimum circumscribed rectangle of the region of interest of the vehicle.

5. The infrared image vehicle detection method according to claim 1, characterized in that, the process of using the motion detection method to process the normalized sequence of frame images in step 5) is as follows: a. Obtain images at any two consecutive moments, use the image at the previous moment as the background image, and perform grayscale processing on the two consecutive frames of images; b. Perform a difference operation on the two grayscale images to obtain a difference image, and perform thresholding on the difference image to retain the pixels with a grayscale difference greater than the set threshold to form a threshold image; c. Use convolution operation to perform erosion operation on the threshold image to obtain an eroded image; d. Perform dilation operation on the eroded image to obtain a dilated image, and perform contour finding operation on the dilated image; e. Traverse each contour to determine the width, height and the upper left corner coordinates of the minimum bounding rectangle of each contour.

6. The infrared image vehicle detection method according to any one of claims 1-5, characterized in that the normalization in the said step 2) and step 4) both adopts the bilinear interpolation method.

Citation Information

Patent Citations

  • Video SAR vehicle target detection method

    CN111798490A

  • Converged network lane line detection method based on attention mechanism and terminal equipment

    CN111950467A