Fast and Robust Detection Method for Infrared Targets in Complex Scenes

By applying grayscale histogram features with spatial position information and weighted kernel function histograms in infrared object detection, combined with Papil distance to measure feature similarity, the problem of poor detection effect of traditional methods in complex scenarios is solved, and higher robustness and real-timeness are achieved.

CN115578667BActive Publication Date: 2025-05-30NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211108273.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-13
Publication Date
2025-05-30
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

Traditional infrared object detection methods have poor detection effects in complex scenarios and cannot effectively deal with clutter interference in the background, making it difficult to achieve robustness and real-timeness.

Method used

The grayscale histogram features with spatial position information are adopted, and the weighted kernel function histogram is extracted and calculated through image enhancement, thresholded segmentation, and the weighted kernel function histogram is extracted and calculated, combining Papist distance to measure feature similarity, improving the robustness and real-timeness of target detection.

Benefits of technology

It improves the robustness and real-time nature of infrared target detection, enhances the ability to express the characteristics of the target, and can position the target more accurately, especially in complex backgrounds and interference scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115578667B_ABST
    Figure CN115578667B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for rapid and robust infrared target detection in complex scenarios. First, the weighted histogram features of the target template are extracted and saved. Then, a saliency region detection is performed using an image segmentation method based on a fixed threshold to quickly extract candidate regions where the target may exist. Finally, the weighted histogram features of the candidate regions are extracted and the matching degree between them and the template features is calculated to determine the position of the target in the current image. The present invention applies the extraction of grayscale histograms with spatial position information to target detection, enhancing the ability to represent the features of the target, improving the detection speed and the robustness of detection, and having high real-time performance. On the one hand, fast and effective features are selected, and the calculation of the weighted histogram only requires one traversal of the region of interest. On the other hand, the design of sub-region division is added, enabling a more accurate position estimation of the target even in scenarios where interference appears and adheres to the target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a fast and robust method for detecting aircraft targets from infrared images. Background Art

[0002] Infrared imaging technology has the characteristics of long working distance, strong anti-interference ability, high measurement accuracy, all-weather operation, etc., and has extremely important applications in the military field. Infrared target detection has always been the core technology of infrared detection systems and is also a hot spot and difficulty in infrared image processing.

[0003] Common traditional methods for target detection include median filters, matched filters, high-pass filters, and some combined forms of filters, etc. Median filters have a simple structure and fast operation speed, but they can only filter out noise and cannot handle clutter interference in the background. Spatial matched filters need to set a filtering template according to the target shape, and the detection performance will be affected when the prior information of the target is unknown. High-pass filters have an obvious effect on filtering slowly changing backgrounds, but they are usually only applicable to the detection of point targets. Gray statistic features are also widely used in the field of image processing due to their good ability to suppress Gaussian noise. Especially the first-order gray histogram feature is useful for better selecting the boundary threshold through the histogram for thresholding when determining the object boundary using the contour line, and is particularly useful for scene segmentation with a strong contrast between the object and the background. In addition, the area and integrated optical density of the object can be obtained while calculating the histogram. However, the first-order gray histogram also has the disadvantages of only reflecting the gray situation of the image, not reflecting the pixel position, one image corresponding to a unique gray histogram, but the same gray histogram may correspond to multiple images.

[0004] In recent years, the technology of artificial neural network (ANN) has developed rapidly. The neural network method needs to give pre-set standard samples to the network for learning, and the detection ability of the network depends on the richness of the training samples. There are also some representative infrared target detection algorithms, such as the method using local entropy, the method based on the facet model, the method based on the SUSAN1 principle for background suppression, etc. These algorithms have played a certain target detection ability in some practical applications, but when the target is in a complex background, the detection effect is not very ideal. In addition, due to the characteristics of large computing amount, poor interpretability, and slow operation speed of the neural network, it is difficult to be deployed and applied on DSP devices.

[0005] In summary, traditional target detection methods usually can only utilize partial gray information due to some constraints and cannot cope with general scenarios, while the neural network method requires a large amount of training data and cannot achieve good results in small-sample scenarios. Summary of the Invention

[0006] Technical Problem to be Solved

[0007] In order to overcome the deficiencies of traditional histograms in infrared target detection, the present invention provides a fast and robust infrared target detection method in complex scenes, applying the gray - level histogram features with spatial position information to the field of target detection to improve the robustness and real - time performance of target detection.

[0008] Technical Solution

[0009] A fast and robust infrared target detection method in complex scenes, characterized by the following steps:

[0010] Step 1: Perform image enhancement on the source image. According to the cumulative frequency histogram, compress the intervals on both sides of the gray level and linearly stretch the middle region to obtain an image with prominent interesting gray - level feature space.

[0011] Step 2: Perform threshold segmentation on the enhanced image to obtain the segmented binary image Q. First, determine a gray - level threshold within the gray - level value range of the image, and then compare the gray level of each pixel in the image with this threshold: pixels greater than or equal to the threshold are foreground pixels, and pixels less than the threshold are background pixels.

[0012] Step 3: Under the framework of tracking - before - detecting, based on the position of the target in the previous frame, set a search window at its position, with a size 9 times that of the target size, search for the foreground connected regions in the window in the binary image to form a set P, and record the position information of each connected region; also include the connected regions on the boundary in the processing.

[0013] Step 4: Quickly extract candidate regions where the target may exist. Each foreground region can be regarded as an object. Traverse the foreground regions in the set P in the binary image Q and record the number of foreground pixels of each object; for objects with the number of foreground pixels less than 25, regard them as noise and eliminate them.

[0014] Step 5: Extract features for each region of interest. If the region of interest is too large, perform sub - region division, and add the divided sub - regions to the set of regions of interest; if it increases in only one direction (length or width), divide it into two sub - regions along the growth direction, and if it increases in both directions, divide it into four sub - regions; the sub - regions are rectangular regions starting from the boundary with an area equal to the size of the target in the previous frame.

[0015] Step 6: Extract features for each object in each set of regions of interest. During the calculation of the gray - level histogram, use the normalized distance of each pixel point as the weight, and count the gray levels corresponding to this pixel point to obtain the gray - level histogram with spatial information.

[0016] Step 7: Feature similarity measurement. The weighted kernel function histograms of the extracted regions of interest are used to measure the similarity with the features of the target in the previous frame according to the Bhattacharyya distance, and the one with the highest score is the target in this frame.

[0017] Step 8: Update parameters. Update each parameter with a certain weight every two frames.

[0018] Step 9: Repeat the above process until all video sequence frames are processed.

[0019] A further technical solution of the present invention: In step 5, the region of interest is too large, specifically, the threshold is set to increase by 20% in both the length and width directions.

[0020] A further technical solution of the present invention: The formula for calculating the Bhattacharyya distance described in step 7:

[0021]

[0022] where p represents the target template vector and q represents the feature vector of the current frame.

[0023] A computer system, comprising: one or more processors, and a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.

[0024] A computer-readable storage medium, storing computer-executable instructions, which are used to implement the above method when executed.

[0025] Beneficial effects

[0026] A fast and robust infrared target detection method in complex scenes provided by the present invention applies the extraction of grayscale histograms with spatial position information to target detection, enhances the ability to represent target features, improves the detection speed and robustness, and has high real-time performance. On the one hand, fast and effective features are selected, and the calculation of the weighted histogram only needs to traverse the region of interest once; on the other hand, the design of sub-region division is added, so that more accurate position estimation of the target can be obtained even in scenes where interference appears and adheres to the target. Description of the drawings

[0027] The drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs represent the same components.

[0028] Figure 1Sub-region division actual scenario diagram; a. The target connected region grows in the width direction; b. The target connected region grows in the height direction; c. The target connected region grows in both directions.

[0029] Figure 2 Flowchart of an infrared target fast and robust detection method in complex scenarios. Specific implementation manners

[0030] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0031] Traditional histograms only count the number of pixels falling into the histogram bins. Weighted histograms further consider the distance between the pixels and the target center, and pixels far from the target center contribute less to the histogram. The idea of the weighted histogram with spatial position information is that when calculating the histogram, a certain weight value is assigned to each point, and the size of the weight value depends on its distance from the center point, and the Epanechnikov Kernal can be used to adjust it. The Epanechnikov Kernal formula of the kernel density function:

[0032]

[0033] c in the above formula (1) d is the volume of the unit sphere in d-dimensional space. If d = 2, c d is the area of the unit circle.

[0034] When statistically analyzing the distribution of color levels, for each pixel, a certain weight value is assigned according to its distance from the center point of the window. When it is closer to the center point, its weight value is larger, and vice versa. This also increases the robustness of object description because when affected by occlusion or complex backgrounds, points closer to the periphery of the object are less reliable.

[0035] Target detection methods can be divided into two categories according to the different orders of using image features, namely the Detect Before Track (DBT) method and the Track Before Detect (TBD) method. In the present invention, the weighted histogram features of the target template are first extracted and saved, and then a significance region detection is performed using an image segmentation method based on a fixed threshold to quickly extract candidate regions where the target may exist. Then, the weighted histogram features of the candidate regions are extracted and the matching degree with the template features is calculated to determine the position of the target in the current image.

[0036] In the solution of the present invention, an object detection method based on a weighted kernel function gray histogram is adopted, and the specific steps are as follows:

[0037] (1) Image enhancement is performed on the source image. According to the cumulative frequency histogram, the intervals on both sides of the gray level are compressed, and the middle region is linearly stretched at the same time, so as to obtain an image with prominent gray features in the region of interest.

[0038] (2) Threshold segmentation is performed on the enhanced image to obtain the binary image Q after segmentation. First, a gray threshold within the gray value range of the image is determined, and then the gray levels of each pixel in the image are compared with this threshold: pixels greater than or equal to this threshold are foreground pixels, and pixels less than this threshold are background pixels.

[0039] (3) Under the framework of tracking first and then detecting, based on the position of the target in the previous frame, a search window is set at its position, and the size is 9 times the size of the target. The foreground connected regions in this window are searched in the binary image to form a set P, and the position information of each connected region is recorded. The connected regions on the boundary are also included in the processing.

[0040] (4) Candidate regions where the target may exist are quickly extracted, and each foreground region can be regarded as an object. The foreground regions in the set P are traversed in the binary image Q, and the number of foreground pixels of each object is recorded. For objects with a very small area (less than 25 pixels), they are regarded as noise and eliminated.

[0041] (5) Feature extraction is performed on each region of interest. If the region of interest is too large (the threshold is set to increase by 20% in the length and width directions), sub-region division is performed, and the divided sub-regions are added to the set of regions of interest. If it is an increase in a single direction (length or width), it is divided into two sub-regions in the growth direction. If it is an increase in both directions, it is divided into four sub-regions. The sub-regions are rectangular regions starting from the boundary and having the size of the target in the previous frame. This is because in the actual scenario, when an aircraft throws out interference and adheres to it, it often tends to one side rather than being at the center of the connected region. See the appendix Figure 1 .

[0042] (6) Feature extraction is performed on each object in each set of regions of interest. In the process of calculating the gray histogram, the normalized distance of each pixel point is used as the weight, and the gray levels corresponding to these pixel points are statistically obtained to get a gray histogram with spatial information.

[0043] (7) Feature similarity measurement. The weighted kernel function histograms of the extracted regions of interest are compared with the features of the target in the previous frame according to the Bhattacharyya distance for similarity measurement, and the one with the highest score is the target of this frame. The Bhattacharyya distance formula is:

[0044]

[0045] Among them, p represents the target template vector, and q represents the current frame feature vector;

[0046] (8) Update the parameters, and update each parameter with a certain weight every two frames.

[0047] (9) Repeat the above process until all video sequence frames are processed.

[0048] A preferred implementation step provided by the present invention is as follows:

[0049] Step 1: Perform image enhancement on the source image. According to the cumulative frequency histogram, compress the intervals on both sides of the gray level and linearly stretch the middle region at the same time to obtain an image with prominent gray feature space of interest.

[0050] (A) Calculate the cumulative frequency histogram. Statistically analyze the gray histogram, and accumulate each gray level to the next level to obtain the cumulative gray histogram. The value of each histogram bar represents the proportion of those less than or equal to this gray level.

[0051] (B) Calculate the high and low thresholds. Through the input parameter x (set to 0.1), find the gray levels L1 where the histogram in the cumulative frequency histogram is equal to 0.1 and the gray level L2 where it is equal to 0.9.

[0052] (C) Highlight the gray space of interest. Set the pixels below L1 to 1, the pixels above L2 to 255, and linearly stretch other pixels according to Equation (3):

[0053]

[0054] Step 2: Perform threshold segmentation on the enhanced image

[0055] (A) First, determine a gray threshold T within the gray value range of the image (T has been optimized according to the threshold segmentation algorithm and a large number of actual infrared videos), and then compare the gray levels of each pixel in the image with this threshold: the pixels greater than or equal to it are foreground pixels, and the pixels less than it are background pixels. Let f(x, y) be the gray value at the point (x, y), and then process the image with Equation (4):

[0056]

[0057] Among them, T is the threshold, and Q(x, y) is the segmented image;

[0058] (B) Record the position information P of each connected domain i (x i ,y i ,w i ,h i ), xi , y i is the upper left corner coordinate of the minimum bounding rectangle, w i , h i represent the width and height. While thresholding each pixel, count and record the number of pixels in each connected component, and exclude connected components with an area less than 20 from consideration.

[0059] Step 3: Search for the region of interest

[0060] (A) Based on the position P of the target in the previous frame 0 (x 0 , y 0 , w 0 , h 0 ), set a search window at the center of the target, with a size 9 times that of the target. The position vector of the search window P s (x s , y s , w s , h s ) is [x 0 - 1.5 * w 0 , y 0 - 1.5 * h 0 , 3w 0 , 3h 0 .

[0061] (B) Traverse the pixel points within the search window in the binary image, calculate the foreground connected components in this window according to the seed filling algorithm, form a set S, and record the position information P of each connected component S i . If a connected component is partially inside and partially outside the window, it is also included in the set S. i

[0062] Step 4: Add sub-region elements to the set of regions of interest as appropriate

[0063] (A) Traverse the set S, and compare the width w i and height h i of the object S i (that is, the w i and h i in the position information P i ) with the width w 0 and height h 0 of the target. If it increases in one direction (height or width), it is divided into two parts in the growth direction. If it increases in both directions, it is divided into four sub-regions. See Appendix Figure 1 .

[0064] (B) When it grows in height in one direction, it is divided into P h1 (x i , y i , wi , h 0 ), and P h2 (x i , y i + h i - h 0 , w i , h 0 ), when growing in one direction, it is divided into P w1 (x i , y i , w 0 , h i ) and P w2 (x i + w i - w 0 , y i , w 0 , h i ), when growing in four directions, it is divided into P i1 (x i , y i , w 0 , h 0 ) and P i2 (x i + w i - w 0 , y i , w 0 , h 0 ), P w2 (x i , y i + h i - h 0 , w 0 , h 0 ) and P w2 (x i + w i - w 0 , y i + h i - h 0 , w 0 , h 0 ) and add them to the set S.

[0065] Step 5: Extract features from the candidate regions of the object

[0066] In the process of calculating the grayscale histogram, the normalized distance of each pixel point is used as the weight, and the grayscale levels corresponding to this pixel point are statistically counted to obtain a grayscale histogram with spatial information.

[0067] (A) Traverse the set S. For the region S i , i = 1, 2,..., n, where n is the number of elements in the set. First, calculate the center point coordinates The maximum distance D from the pixel point to the center point is half of the diagonal length of the minimum circumscribed rectangle, that is

[0068] (B) Traverse all pixels in S i For the pixel point s, its normalized distance is If the gray level of this pixel point is f(x, y), then the feature vector of the weighted histogram

[0069] (C) After traversing all pixel points, perform histogram normalization, accumulate the sum of all gray levels in T, and denote it as T sum , traverse each element in T, and let T[i] = T[i] / T sum .

[0070] Step 6: Feature similarity measurement. Measure the similarity between the weighted kernel function histograms of each region of interest extracted and the features of the target in the previous frame according to the Bhattacharyya distance. The one with the highest score is the target in this frame.

[0071] According to the target template vector T m and the feature vector T of the current frame, calculate the similarity measurement value according to the Bhattacharyya distance formula. The larger this value is, the more similar it indicates.

[0072]

[0073] n is the dimension of the feature vector T, the larger BC(T m , T) is, the more similar the two are.

[0074] Step 7: Record the region S corresponding to the maximum similarity measurement value BC(T m , T) i , as the result of this frame, and update the feature vector with a weight of 0.5 every two seconds:

[0075] T m = 0.5 * T m + 0.5 * T

[0076] Step 8: Repeat the above steps until the processing is completed

[0077] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A method for fast and robust infrared target detection in complex scenes, characterized in that the steps are as follows: Step 1: Perform image enhancement on the source image. According to the cumulative frequency histogram, compress the intervals on both sides of the gray level and linearly stretch the middle region to obtain an image with prominent gray-level features of the region of interest; Step 2: Perform threshold segmentation on the enhanced image to obtain a binary image Q after segmentation. First, determine a gray threshold within the gray-level value range of the image, and then compare the gray levels of each pixel in the image with this threshold: pixels greater than or equal to this threshold are foreground pixels, and pixels less than this threshold are background pixels; Step 3: Under the framework of tracking first and then detecting, based on the position of the target in the previous frame, set a search window at its position, with a size 9 times that of the target size, search for the foreground connected regions in the window in the binary image to form a set P, and record the position information of each connected region; Connected regions on the boundary are also included in the processing; Step 4: Quickly extract candidate regions where the target may exist. Each foreground region can be regarded as an object. Traverse the foreground regions in the set P in the binary image Q and record the number of foreground pixels of each object; For objects with less than 25 foreground pixels, they are regarded as noise and eliminated; Step 5: Extract features for each region of interest. If the region of interest is too large, perform sub-region division and add the divided sub-regions to the set of regions of interest; If it increases in only one direction of length or width, it is divided into two sub-regions in the growth direction. If it increases in both directions, it is divided into four sub-regions; The sub-region is a rectangular region starting from the boundary with an area equal to the size of the target in the previous frame; Step 6: Extract features for each object in each set of regions of interest. During the calculation of the gray histogram, use the normalized distance of each pixel point as the weight and count the gray levels corresponding to this pixel point to obtain a gray histogram with spatial information; Step 7: Feature similarity measurement. Use the weighted kernel function histogram of each extracted region of interest and the features of the target in the previous frame to perform similarity measurement according to the Bhattacharyya distance, and the one with the highest score is the target of this frame; Step 8: Update parameters. Update each parameter with a certain weight every two frames; Step 9: Repeat the above process until all video sequence frames are processed.

2. The method for fast and robust infrared target detection in complex scenes according to claim 1, characterized in that in step 5, the region of interest being too large is specifically that the threshold is set to a 20% increase in the length and width directions.

3. The method for fast and robust infrared target detection in complex scenes according to claim 1, characterized in that the formula for calculating the Bhattacharyya distance in step 7: where p represents the target template vector and q represents the feature vector of the current frame.

4. A computer system, characterized in that it includes: one or more processors, a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to claim 1.

5. A computer-readable storage medium, characterized in that Stores computer-executable instructions that, when executed, are used to implement the method recited in claim 1.

Citation Information

Patent Citations

  • Method for detecting and tracking infrared small target in complex background

    CN102103748A

  • Method for tracking target of interest in infrared image

    CN106485733A