Infrared target detection method based on combination of local contrast and long-short term memory
Through the combination of local contrast and long and short-term memory networks, the problem of high false alarm rate and large calculation amount of infrared small targets in complex backgrounds is solved, and high-precision and robust infrared small target detection and tracking are achieved.
Patent Information
- Application Number
- CN202510678534.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has a high false alarm rate and large calculation amount in infrared small object detection under complex backgrounds, making it difficult to achieve fast and accurate detection.
An infrared object detection method with local contrast combined with long and short-term memory is adopted, and the direction derivative is calculated through the Facet model, a dual local contrast model is constructed, and a long-term convolutional neural network is combined for target tracking and detection, and a short-term convolutional neural network is used to process the occlusion situation.
It significantly reduces the false alarm rate, improves detection accuracy, can accurately locate small targets, adapt to target detection of different backgrounds and sizes, and ensures the robustness and real-timeness of the targets in different backgrounds.
Smart Images

Figure CN120219726A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of small target detection and positioning, and in particular to an infrared target detection method combining local contrast with long short-term memory. Background Art
[0002] Implementing robust infrared small target fast and accurate detection in complex backgrounds is of great significance for infrared search and tracking (IRST) applications. Some small targets are always immersed in complex backgrounds with low signal-to-clutter ratio (SCR), which are usually composed of high-intensity regions, target-like interferences (broken clouds), sharp edges, and high-brightness noise with a few pixel sizes (PNHB), resulting in a large number of false alarms during the target detection process, posing challenges to the fast and accurate detection of infrared small targets.
[0003] HVS-based methods achieve detection by measuring the difference between the target and the local background. On the one hand, the comprehensive consideration of target and background information ensures their high-level detection performance in complex scenes. Among them, Xu Yunkai et al. proposed a small target detection method (Xu Yunkai et al., Infrared Small Target Detection Based on Local Contrast-Weighted Multidirectional Derivative, IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 61, 2023: Art. No. 5000816), which uses a multi-directional derivative model with a penalty factor based on Facet and a dual-local contrast model to reduce the false alarm rate of small target detection. However, this method only targets the small target detection of single-frame images, does not fully utilize the relative relationship of target positions between multiple frames of images, and has a large amount of calculation, resulting in a lack of accurate real-time target tracking ability in practical applications such as target tracking. Therefore, there is an urgent need for an infrared target detection method combining local contrast with long short-term memory to overcome the problems existing in the above technologies. Summary of the Invention
[0004] The purpose of the present invention is to provide an infrared target detection method combining local contrast with long short-term memory to solve the problems of high false alarm rate and large amount of calculation of single-frame images in complex backgrounds existing in the prior art.
[0005] To achieve the above object, the present invention provides an infrared target detection method combining local contrast with long short-term memory, including the following steps: Step 1, input a grayscale image; Step 2, calculate the directional derivative of the grayscale image based on the Facet model; Step 3, calculate the penalty factor using the directional derivative; Step 4: Construct a three - layer double local contrast model; Step 5: Construct an LCWMD model according to the penalty factor and the three - layer double local contrast model to capture the initial position of the target; Step 6: Extract the gray - scale features of the target and the background, and train a long - term convolutional neural network; Step 7: Combine the constructed LCWMD model with the long - term network for long - term infrared target tracking detection, and judge whether the target is occluded. If it is occluded, continue with long - term tracking detection; otherwise, execute the next step; Step 8: Judge again whether the target is detected. If not, detect the local boundary of the occluding object; if the target is detected, execute the next step. Among them, when detecting the local boundary of the occluding object, it is necessary to judge whether the target reappears. If it reappears, perform long - term tracking detection; otherwise, continue with local detection; Step 9: Extract target features to train a short - term convolutional neural network; Step 10: Combine the constructed LCWMD model with the short - term network for short - term infrared target tracking detection, and judge whether the target has left the occlusion area. If not, continue with short - term tracking detection; otherwise, return to Step 7 and execute sequentially until the target is lost.
[0006] Preferably, the process of calculating the directional derivative of the grayscale image based on the Facet model in Step 2 is as follows: S21: Represent the pixels in the neighborhood through a bicubic function as follows: ; ; where and are the row index set and column index set of a 5×5 - sized neighborhood centered at (0,0), is the known discrete orthogonal polynomial basis; represents the fitting coefficient, which is solved by minimizing the distance between the fitted grayscale value and the actual grayscale value as follows: ; where represents the original grayscale value; S22: Based on the orthogonality property of the polynomial, calculate the fitting coefficient with the following calculation expression: ; ; where denotes a fixed convolution kernel; S23. Calculate the directional derivative of the grayscale image, and the expression is as follows: ; ; In the formula, denotes the first-order directional derivative in the direction, denotes the fitting coefficient, denotes the vectorand the included angle between the axis, denotes the second-order directional derivative in the
[0007] Preferably, the process of calculating the penalty factor using the directional derivative in step 3 is as follows: Design the improved multi-directional derivative feature as: ; wherein, is the penalty factor corresponding to each pair of different directions, denotes the second-order directional derivative in the direction, denotes the second-order directional derivative in the direction; ; ; ; wherein, denotes the corresponding maximum value after the derivative maps in two different directions are divided point by point, denotes the absolute value of the difference between the derivative maps in two different directions, denotes a very small offset and respectively denote the results after filtering the derivative maps in two different directions; the reason for filtering is to prevent the target information from shrinking after taking the derivative. Designing four simple filters corresponding to each direction for filtering can expand the overlapping area without destroying the information of the directional derivative map; the reason why the penalty factor can suppress the background is that: generally, small and weak infrared targets have approximate isotropy, while the complex background has weaker isotropy. The designed penalty factor is very strong when the isotropy is weak, so the background can be relatively suppressed; In summary, the expression of the multi-directional derivative with the penalty factor is as follows: .
[0008] Preferably, the process of constructing the three-layer double local contrast model in step 4 is as follows: S41. According to the feature that the target gray intensity remains stable in the center and decays towards the periphery, use the local contrast model with a three-layer sliding window structure to represent the relationship between the target and the background, namely the TC core layer, the TA decay layer, and the LB background layer, and obtain the average gray value relationship. The expression is as follows: ; ; ; Among them, represents the average gray value of the corresponding layer, represents the maximum value among the average gray values of 8 background blocks, represents the number of pixels in the corresponding layer, represents the average gray value of the layer, represents the average gray value of the layer, represents the average gray value of the th background block in the layer; S42. Use the method combining ratio and difference to calculate the local contrast. The calculation expression is as follows: ; ; ; In the formula, represents the local contrast between the TC layer and the TA layer, represents the local contrast between the TC layer and the LB layer, represents the local contrast between the TB layer and the LB layer; S43. Remove background clutter through non-negative constraints. The expression is as follows: ; ; In the formula, represents the difference between the maximum value and the average value of the gray values in the th background block in the layer, represents the standard deviation between, represents calculating the standard deviation, represents the maximum value of the gray values in the th background block in the layer; Obviously, the smoother the background layer, The smaller the value is; S44. Construct the final three-layer dual local contrast model , and the expression is as follows: ; In the formula, is the offset set to 1, and its function is to avoid the denominator from being too small, so as to eliminate the influence of the value being very small in the smooth background area.
[0009] Preferably, in step 5, the LCWMD model is constructed according to the penalty factor and the three-layer dual local contrast model, and the operation of using adaptive threshold segmentation is used to realize detection and capture the initial position of the target. The model and the adaptive threshold The expressions are as follows: ; ; In the formula, and represent the mean and variance of, is a given constant, and the optimal range is 0.6 to 0.8.
[0010] Preferably, in step 6, the gray-scale features of the target and the background are extracted, and the process of training the long-term convolutional neural network is as follows: S61. First, process the first frame of the image. If a suspected target is found, confirm the target position in the corresponding areas of the second and third frames of the image, and generate a border adapted to the target size through the SAM segmentation model; interpolate and scale the target area to a size of 5×5 as a positive example sample (label +1); at the same time, collect background samples at equal intervals from the background around the target, scale them to the same size, and use them as negative example samples (label 0); after mixing the positive and negative samples, form an initial sample space; if the positive example samples are insufficient, use the SMOTE oversampling method to generate additional positive examples; S62. Calculate the size of the convolutional layer; the input channel of the convolutional layer is 1, the output channel is 16, the convolutional kernel size is 3×3, the stride is 1, and the padding is 1. The calculation expression is as follows: ; Among them, represents the feature value at the position on the output channel ; represents the corresponding pixel value in the input image; represents the weight parameter of the convolutional kernel; represents the bias of the S63. Downsample the convolved feature map using a max pooling layer with a pooling kernel size of 2×2 and a stride of 2. The pooling calculation expression is as follows: ; where, represents the value at position in the th channel of the pooled feature map; S64. Use the feature map obtained after convolution and pooling as the input to the first fully connected layer. The calculation expression of the fully connected layer is as follows: ; where, is the input vector; is the weight matrix of the first fully connected layer; is the bias term; is the output of the first layer; S65. Then pass the vector output by the first fully connected layer to the second fully connected layer. This layer outputs a single node for binary classification. The calculation expression is as follows: ; where, is the weight matrix of the second fully connected layer; is the bias term; is the output of the second layer; S66. The network uses the Sigmoid activation function for binary classification on the output . The calculation expression of the Sigmoid function is as follows: ; where, is the final output of the model, with a value range between [0, 1]. It can be understood as the probability that the sample belongs to a certain category. Based on the value of , the category of the sample can be determined. Here, samples with a value greater than 0.5 are judged as suspected positive examples, and the output value represents the confidence that the sample is the target.
[0011] Preferably, in step 7, the constructed LCWMD model is combined with a long-term network for long-term infrared target tracking detection, and it is determined whether the target is occluded. If it is occluded, long-term tracking detection continues; otherwise, the following process for the next step is as follows: First, according to the relationship between consecutive frames, within the local area of the next frame's image, use the trained long-term convolutional neural network for scanning to obtain multiple samples of suspected targets, and select the sample with the highest confidence as the detection result of the neural network. Then, calculate the LCWMD values of all samples with high confidence, and judge whether the target is occluded based on this value. Since the LCWMD value of the target sample will decrease rapidly when the target is occluded, the maximum LCWMD value when the target sample is not occluded is used as the benchmark to calculate the benchmark when detecting the frame , and the expression is as follows: ; Among them, represents the LCWMD value of the -th frame image; when the LCWMD value of the target sample drops below 80% of the benchmark, when is satisfied, it is considered that the target starts to enter the complex background area and is occluded. At this time, the local area is detected and processed according to the detection results.
[0012] Preferably, when detecting the local boundary of the occluding object in step 8, the expression for judging whether the target reappears is as follows: If the LCWMD values of the targets detected in three consecutive frames are between 30% and 80% of the benchmark, it is considered that the position of the target in the complex background area is detected, and the expression is as follows: ; Among them, represents the LCWMD value of the -th frame image, represents the LCWMD value of the -th frame image, represents the LCWMD value of the -th frame image; if the LCWMD value of the target is below 30% of the benchmark, the system determines that the target tracking fails and triggers a temporary loss flag, and the expression is as follows: .
[0013] Preferably, the process of extracting target features to train the short-term convolutional neural network in step 9 is as follows: S91. Calculate the size of the convolutional layer; the input channel of the convolutional layer is 1, the output channel is 8, the convolutional kernel size is 3×3, the stride is 1, and the padding is 1. The calculation expression is as follows: ; Among them, represents the feature value at the position on the output channel ; represents the corresponding pixel value in the input image; represents the weight parameter of the convolution kernel; represents the bias of the S92. Use a max pooling layer to downsample the feature map after convolution. The pooling kernel size is 2×2, the stride is 2, and the pooling calculation expression is as follows: ; where, represents the value at the position in the channel of the feature map after pooling; S93. Use the feature map obtained after convolution and pooling as the input of the first fully connected layer; the calculation expression of the fully connected layer is as follows: ; where, is the input vector; is the weight matrix of the first fully connected layer; is the bias term; is the output of the first layer; S94. Then transfer the vector output by the first fully connected layer to the second fully connected layer. This layer outputs a node for binary classification, and the calculation expression is as follows: ; where, is the weight matrix of the second fully connected layer; is the bias term; is the output of the second layer; S95. The network uses the Sigmoid activation function for binary classification on the output . The calculation expression of the Sigmoid function is as follows: ; where, is the final output of the model, and its value range is between [0,1].
[0014] Preferably, the criterion for determining whether the target has left the occlusion area in step 10 is as follows: If the LCWMD value of the target sample returns to more than 80% of the benchmark for three consecutive frames, it is considered that the target is gradually leaving the complex background area. The expression is as follows: .
[0015] Therefore, the above infrared target detection method using local contrast combined with long short-term memory of the present invention has the following beneficial effects: (1) By using the method of combining the LCWMD model with the long short-term network, small targets can be accurately located, interference in high-intensity backgrounds can be effectively suppressed, the false alarm rate can be reduced, and the detection accuracy can be significantly improved; (2) The long short-term memory network is used to process the temporal relationship of multi-frame images, making the detection of targets in consecutive frames more accurate, thereby reducing the problems of large computational complexity and slow response in the detection of consecutive frames by traditional methods; (3) The present invention introduces a short-term convolutional neural network (CNN) to learn the characteristics of occluded targets. It can still accurately track when the target is occluded by a complex background (such as high-intensity regions, target-like interference, etc.), ensuring the robustness of the target under different backgrounds. As the number of multi-frame images gradually increases, the training effect of the convolutional neural network gradually enhances, improving the detection accuracy and efficiency. Moreover, the SAM segmentation model is used to obtain the precise size of the target, thus avoiding the influence of changes in target size on detection and network learning; (4) The method proposed by the present invention can not only adapt to the detection tasks of small targets of different sizes under different backgrounds, but also ensure the real-time and efficient target tracking through continuous training and optimization.
[0016] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0017] Figure 1 It is the overall flowchart of an infrared target detection method combining local contrast with long short-term memory according to the present invention; Figure 2 It is the first-frame target capture image of the embodiment of the present invention; Figure 3 It is the long-term network target tracking image of the embodiment of the present invention; Figure 4 It is the tracking image of the target before occlusion in the embodiment of the present invention; Figure 5 It is the tracking image of the target after occlusion in the embodiment of the present invention; Figure 6 It is the tracking image of the target after reappearance in the embodiment of the present invention. Specific Embodiments
[0018] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0019] Please refer to Figures 1-6, a local contrast combined with long short-term memory infrared target detection method, comprising the following steps: Step 1: Input grayscale image; Step 2: Calculate the directional derivative of the grayscale image based on the Facet model; the specific process is as follows: S21, the pixels in the neighborhood are processed by a bivariate cubic function To express it, the expression is as follows: ; ; In the formula, and is the set of row and column indices of a 5×5 neighborhood centered at (0,0), is a known discrete orthogonal polynomial basis; Represents the fitting coefficient, which is calculated by minimizing the distance between the fitting grayscale value and the actual grayscale value To solve, the expression is as follows: ; In the formula, Represents the original grayscale value; S22. Calculate the fitting coefficients based on the orthogonal properties of the polynomials , the calculation expression is as follows: ; ; In the formula, represents a fixed convolution kernel; S23. Calculate the directional derivative of the grayscale image. The expression is as follows: ; ; In the formula, express The first directional derivative of the direction, represents the fitting coefficient, , Represents vector and The angle of the axis, express The second-order directional derivative of the direction.
[0020] Step 3: Calculate the penalty factor using directional derivatives; the specific process is as follows: The multi-derivative features of the design improvement are: ; in, is the penalty factor corresponding to each pair of different directions, Represents The second-order directional derivative in the direction, Represents The second-order directional derivative in the direction; it is calculated based on the following formula: ; ; ; Wherein, Represents the corresponding maximum value after the derivative maps in two different directions are divided point by point, Represents the absolute value of the difference between the derivative maps in two different directions, Represents a very small offset , And Respectively represent the results after filtering the derivative maps in two different directions; the reason for filtering is to prevent the target information from shrinking after differentiation. Designing four simple filters corresponding to each direction for filtering can expand the overlapping area without destroying the information of the directional derivative map; the reason why the penalty factor can suppress the background is that: infrared small and weak targets generally have approximate isotropy, while complex backgrounds have weaker isotropy. The designed penalty factor Is very strong when the isotropy is weak, so the background can be relatively suppressed; In summary, the expression of the multi-directional derivative with a penalty factor is as follows: .
[0021] Step 4, construct a three-layer double local contrast model; the specific process is as follows: S41. According to the characteristics that the target gray intensity remains stable at the center and decays towards the periphery, use the local contrast model with a three-layer sliding window structure to represent the relationship between the target and the background, namely the TC core layer, the TA decay layer, and the LB background layer, and obtain the average gray value relationship. The expression is as follows: ; ; ; Wherein, Represents the average gray value of the corresponding layer, Represents the maximum value among the average gray values of 8 background blocks, Represents the number of pixels in the corresponding layer, Represents The average gray value of the layer, Represents The average gray value of the layer, Represents In the layer at the The average gray value of a background block; S42. Calculate the local contrast using a method that combines ratio and difference. The calculation expression is as follows: ; ; ; In the formula, represents the local contrast between the TC layer and the TA layer, represents the local contrast between the TC layer and the LB layer, represents the local contrast between the TB layer and the LB layer; S43. Remove background clutter through non - negative constraints. The expression is as follows: ; ; In the formula, represents the difference between the maximum and average gray values in the th background block of the layer, represents the standard deviation between represents calculating the standard deviation, represents the maximum gray value in the th background block of the layer; obviously, the smoother the background layer, the smaller the value of S44. Construct the final three - layer double - local - contrast model , and the expression is as follows: ; In the formula, is an offset set to 1, whose function is to avoid the denominator from being too small, thereby eliminating the influence of the very small value of in the smooth background area.
[0022] Step 5. Construct the LCWMD model based on the penalty factor and the three - layer double - local - contrast model to capture the initial position of the target; among them, constructing the LCWMD model based on the penalty factor and the three - layer double - local - contrast model, and using the operation of adaptive threshold segmentation to achieve detection and capture the initial position of the target, The expressions of the model and the adaptive threshold are as follows: ; ; In the formula, and represent The mean value and variance, are given constants, and the optimal range is 0.6 to 0.8; after the above calculation, the target initial position can be captured, such as Figure 2 .
[0023] Step 6: Extract the grayscale features of the target and the background, and train a long-term convolutional neural network; the specific process is as follows: S61. First, process the first frame of the image. If a suspected target is found, confirm the target position in the corresponding areas of the second and third frames of the image, and generate a border adapted to the target size through the SAM segmentation model; interpolate and scale the target area to a size of 5×5 as a positive example sample (label +1); at the same time, collect background samples at equal intervals from the background around the target, scale them to the same size, and use them as negative example samples (label 0); after mixing the positive and negative samples, form the initial sample space; if the positive example samples are insufficient, use the SMOTE oversampling method to generate additional positive examples; S62. Calculate the size of the convolutional layer; the input channel of the convolutional layer is 1, the output channel is 16, the convolutional kernel size is 3×3, the stride is 1, and the padding is 1. The calculation expression is as follows: ; Among them, represents the position on the output channel of the eigenvalue; represents the corresponding pixel value in the input image; represents the weight parameter of the convolutional kernel; represents the th bias of the convolutional kernel; S63. Use the Max Pooling Layer to downsample the convolutional feature map. The pooling kernel size is 2×2, the stride is 2, and the pooling calculation expression is as follows: ; Among them, represents the value of the pooled feature map at the position in the th channel; S64. Use the feature map obtained after convolution and pooling as the input of the first fully connected layer; the calculation expression of the fully connected layer is as follows: ; Among them, is the input vector; is the weight matrix of the first fully connected layer; is the bias term; is the output of the first layer; S65. Then, the vector output by the first fully-connected layer is passed to the second fully-connected layer, and this layer outputs a single node for binary classification. The calculation expression is as follows: ; where is the weight matrix of the second fully-connected layer; is the bias term; is the output of the second layer; S66. The network uses the Sigmoid activation function for binary classification on the output . The calculation expression of the Sigmoid function is as follows: ; where is the final output of the model, and its value range is between [0, 1]. It can be understood as the probability that a sample belongs to a certain category. According to the value of , the category of the sample can be determined. Here, samples with a value greater than 0.5 are selected as suspected positive examples, and the output value represents the confidence that the sample is the target.
[0024] After simple training, the training loss of the convolutional neural network is as follows: Epoch [1 / 20], Loss: 1.1101 Epoch [2 / 20], Loss: 1.2572 Epoch [3 / 20], Loss: 1.1941 Epoch [4 / 20], Loss: 1.1430 Epoch [5 / 20], Loss: 1.3813 Epoch [6 / 20], Loss: 1.0870 Epoch [7 / 20], Loss: 0.9129 Epoch [8 / 20], Loss: 0.8929 Epoch [9 / 20], Loss: 0.7645 Epoch [10 / 20], Loss: 0.7187 Epoch [11 / 20], Loss: 0.6485 Epoch [12 / 20], Loss: 0.6005 Epoch [13 / 20], Loss: 0.6838 Epoch [14 / 20], Loss: 0.5226 Epoch [15 / 20], Loss: 0.4486 Epoch [16 / 20], Loss: 0.3538 Epoch [17 / 20], Loss: 0.3388 Epoch [18 / 20], Loss: 0.2283 Epoch [19 / 20], Loss: 0.1226 Epoch [20 / 20], Loss: 0.0739 Step 7: Combine the constructed LCWMD model with the long-term network for long-term infrared target tracking and detection, and determine whether the target is occluded. If it is occluded, continue with long-term tracking and detection; otherwise, proceed to the next step. The specific process is as follows: After performing the above operations, the confirmation of the target's existence and feature learning have been completed. Next, the images after the third frame will be processed.
[0025] First, according to the relationship between consecutive frames, within the local area of the next frame's image, use the trained long-term convolutional neural network to scan to obtain multiple samples of suspected targets. Select the sample with the highest confidence as the detection result of the neural network. Then, calculate the LCWMD values of all samples with high confidence. Based on this value, determine whether the target is occluded. Since the LCWMD value of the target sample will rapidly decrease when the target is occluded, the maximum LCWMD value when the target sample is not occluded is used as the benchmark to calculate the benchmark when detecting the frame , and the expression is as follows: ; Among them, represents the LCWMD value of the frame image; when the LCWMD value of the target sample drops below 80% of the benchmark, when is satisfied, it is considered that the target begins to enter a complex background area and is occluded. At this time, detect the local area and handle it according to the detection results; select the target with the largest LCWMD result and use its position as the final detection result; then, use the suspected target samples with incorrect judgments and a distance greater than 3 from the true target position as negative examples, add the corresponding labels, and use them as training samples to perform reinforcement training on this network; the same applies to the images of subsequent frames. First, use the long-term convolutional neural network to scan and repeat the above steps. The subsequent tracking results are as Figure 3 .
[0026] Step 8: Determine again whether the target is detected. If not, detect the local boundary of the occluding object. If the target is detected, proceed to the next step. When detecting the local boundary of the occluding object, it is necessary to determine whether the target reappears. If it reappears, perform long-term tracking detection; otherwise, continue with local detection. When detecting the local boundary of the occluding object, the expression for determining whether the target reappears is as follows: If the LCWMD values of the targets detected in three consecutive frames are between 30% and 80% of the benchmark, it is considered that the position of the target in the complex background area is detected. At this time, sampling is performed using a method similar to that of the initial frame, and a new sample space is formed as the sample space of the short-term convolutional neural network. Then, a short-term neural network with a simpler structure is used for training. The size of this network is only half of that of the long-term neural network, and the expression is as follows: ; Among them, represents the LCWMD value of the -th frame image, represents the LCWMD value of the -th frame image, represents the LCWMD value of the -th frame image; if the LCWMD value of the target is below 30% of the benchmark, the system determines that the target tracking fails and triggers a temporary loss flag. Next, in the next frame, the boundary of the occluding object will be obtained using the SAM segmentation model, and the long-term network will be scanned and recognized in the local area around the boundary, and verification will be performed based on the LCWMD value of the detected target until the target reappears. The expression is as follows: .
[0027] Step 9: Extract target features to train the short-term convolutional neural network; the specific process is as follows: S91: Calculate the size of the convolutional layer; the input channel of the convolutional layer is 1, the output channel is 8, the convolutional kernel size is 3×3, the stride is 1, and the padding is 1. The calculation expression is as follows: ; Among them, represents the eigenvalue at the position on the output channel ; represents the corresponding pixel value in the input image; represents the weight parameter of the convolutional kernel; represents the -th bias of the convolutional kernel; S92. Downsample the feature map after convolution using a max pooling layer with a pooling kernel size of 2×2 and a stride of 2. The pooling calculation expression is as follows: ; Among them, represents the value at position in the th channel of the pooled feature map; S93. Use the feature map obtained after convolution and pooling as the input to the first fully connected layer. The calculation expression of the fully connected layer is as follows: ; Among them, is the input vector; is the weight matrix of the first fully connected layer; is the bias term; is the output of the first layer; S94. Then pass the vector output by the first fully connected layer to the second fully connected layer. This layer outputs a single node for binary classification. The calculation expression is as follows: ; Among them, is the weight matrix of the second fully connected layer; is the bias term; is the output of the second layer; S95. The network uses the Sigmoid activation function for binary classification on the output . The calculation expression of the Sigmoid function is as follows: ; Among them, is the final output of the model, and its value range is between [0, 1].
[0028] Step 10: Combine the constructed LCWMD model with the short-term network for short-term infrared target tracking and detection, and determine whether the target has left the occlusion area. If it has not left, continue with short-term tracking and detection; otherwise, return to Step 7 for sequential execution until the target is lost. Specifically: First, based on the relationship between consecutive frames, in the local area of the next frame of the image, use the trained short-term convolutional neural network for scanning to obtain multiple samples of suspected targets, and select the sample with the highest confidence as the detection result of the neural network. Then, use traditional algorithms to calculate the LCWMD values of all samples with relatively high confidence. Select the target with the largest LCWMD result and use its position as the final detection result. Then, use the suspected target samples with incorrect judgments and a distance greater than 3 from the true target position as negative examples, add the corresponding label 0, and use them as training samples for reinforcement training with the short-term network. When processing subsequent frames of images, first use the short-term convolutional neural network for scanning and repeat the above steps until the target leaves the occlusion area. The recognition and detection results before and after the target is occluded are as shown in Figure 4 , Figure 5 . The criterion for determining whether the target has left the occlusion area is as follows: If the LCWMD value of the target sample returns to more than 80% of the benchmark for three consecutive frames, it is considered that the target is gradually leaving the complex background area. The expression is as follows: .
[0029] At this time, the same detection and tracking method is still used, but the network used is changed to a long-term convolutional neural network. The recognition and tracking results after the target reappears are as shown in Figure 6 .
[0030] Therefore, the present invention adopts the above infrared target detection method combining local contrast with long and short-term memory, which can achieve high-precision and efficient infrared small target detection and tracking in complex environments, has strong robustness and adaptability, and can meet the requirements in practical applications.
[0031] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. An infrared target detection method combining local contrast with long short-term memory, characterized in that, It includes the following steps: Step 1, input a grayscale image; Step 2, calculate the directional derivative of the grayscale image based on the Facet model; Step 3, calculate the penalty factor using the directional derivative; Step 4, construct a three-layer dual local contrast model; Step 5, construct an LCWMD model according to the penalty factor and the three-layer dual local contrast model to capture the initial position of the target; Step 6, extract the grayscale features of the target and the background, and train a long-term convolutional neural network; Step 7, combine the constructed LCWMD model with the long-term network for long-term infrared target tracking detection, and judge whether the target is occluded. If it is occluded, continue with long-term tracking detection, otherwise execute the next step; Step 8, judge again whether the target is detected. If not, detect the local boundary of the occluding object; if the target is detected, execute the next step; among them, when detecting the local boundary of the occluding object, it is necessary to judge whether the target reappears. If it reappears, perform long-term tracking detection, otherwise continue with local detection; Step 9, extract target features and train a short-term convolutional neural network; Step 10, combine the constructed LCWMD model with the short-term network for short-term infrared target tracking detection, and judge whether the target has left the occlusion area. If not, continue with short-term tracking detection, otherwise return to Step 7 for sequential execution until the target is lost.
2. The infrared target detection method combining local contrast and long short-term memory according to claim 1, wherein The process of calculating the directional derivative of the grayscale image based on the Facet model in Step 2 is as follows: S21. Represent the pixels in the neighborhood by a bi - cubic function The expression is as follows: ; ; In the formula, and are the row index set and column index set of a 5×5 neighborhood centered at (0, 0), is a known discrete orthogonal polynomial basis; represents the fitting coefficient, which is solved by minimizing the distance between the fitted gray value and the actual gray value The solution is obtained as follows: ; In the formula, represents the original grayscale value; S22. Calculate the fitting coefficients based on the orthogonality property of polynomials , and the calculation expression is as follows: ; ; In the formula, represents a fixed convolution kernel; S23, calculate the directional derivative of the grayscale image, and the expression is as follows: ; ; In the formula, express The first directional derivative of the direction, represents the fitting coefficient, , Represents vector and The angle of the axis, express The second directional derivative of the direction.
3. The infrared target detection method combining local contrast with long short-term memory according to claim 2, characterized in that The process of calculating the penalty factor using the directional derivative in Step 3 is as follows: Design an improved multi-directional derivative feature as: ; Among them, is the penalty factor corresponding to each pair of different directions, denotes the second-order directional derivative in the direction, and denotes the second-order directional derivative in the Calculate based on the following formula: ; ; ; Among them, represents the corresponding maximum value after pointwise division of the derivative graphs in two different directions, represents the absolute value of the difference between the derivative graphs in two different directions, represents the offset , and respectively represent the results after filtering of the derivative graphs in two different directions; In summary, the expression of the multi-directional derivative with a penalty factor is as follows: 。 4. The infrared target detection method combining local contrast with long short-term memory according to claim 3, wherein The process of constructing a three-layer dual local contrast model in Step 4 is as follows: S41, according to the feature that the target grayscale intensity remains stable in the center and decays towards the periphery, use a local contrast model with a three-layer sliding window structure to represent the relationship between the target and the background, namely the TC core layer, the TA decay layer, and the LB background layer, to obtain the average grayscale value relationship, and the expression is as follows: ; ; ; Among them, represents the average gray value of the corresponding layer, represents the maximum value among the average gray values of 8 background blocks in the layer, represents the number of pixels in the corresponding layer, represents the average gray value of the layer, represents the average gray value of the layer, represents in the layer, the average gray value of the nth background block; S42, use the method of combining ratio and difference to calculate the local contrast, and the calculation expression is as follows: ; ; ; In the formula, represents the local contrast between the TC layer and the TA layer, represents the local contrast between the TC layer and the LB layer, represents the local contrast between the TB layer and the LB layer; S43, remove background clutter through non-negative constraints, and the expression is as follows: ; ; In the formula, represents the difference between the maximum value and the average value of the gray values in the th background block of the represents the standard deviation between represents calculating the standard deviation, represents the maximum value of the gray values in the th background block of the S44. Construct the final three-layer double local contrast model , and the expression is as follows: ; In the formula, is an offset set to 1.
5. The infrared target detection method combining local contrast with long short-term memory according to claim 4, characterized in that: In step 5, the LCWMD model is constructed based on the penalty factor and the three-layer double local contrast model, and the detection is realized by using the operation of adaptive threshold segmentation to capture the initial position of the target. Model and adaptive threshold The expressions are as follows: ; ; In the formula, and represent the mean value and variance of, is a given constant, and the optimal range is from 0.6 to 0.
8.
6. The infrared target detection method combining local contrast with long short-term memory according to claim 5, characterized in that: The process of extracting the grayscale features of the target and the background and training a long-term convolutional neural network in Step 6 is as follows: S61, first process the first frame image. If a suspected target is found, confirm the target position in the corresponding areas of the second and third frame images, and generate a border adapted to the target size through the SAM segmentation model; interpolate and scale the target area to a size of 5×5 as a positive example sample; at the same time, collect background samples equidistantly from the background around the target, scale them to the same size, and use them as negative example samples; after mixing the positive and negative samples, form an initial sample space; if the positive example samples are insufficient, use the SMOTE oversampling method to generate additional positive examples; S62, calculate the size of the convolutional layer; the input channel of the convolutional layer is 1, the output channel is 16, the convolutional kernel size is 3×3, the stride is 1, and the padding is 1. The calculation expression is as follows: ; Among them, represents the position on the output channel ; is the eigenvalue; represents the corresponding pixel value in the input image; represents the weight parameter of the convolutional kernel; represents the bias of the th convolutional kernel. S63. Downsample the feature map after convolution using a max pooling layer with a pooling kernel size of 2×2, a stride of 2, and the pooling calculation expression is as follows: ; Among them, represents the value at the position in the th channel of the pooled feature map; S64. Use the feature map obtained after convolution and pooling as the input to the first fully connected layer; the calculation expression of the fully connected layer is as follows: ; Among them, is the input vector; is the weight matrix of the first fully connected layer; is the bias term; is the output of the first layer; S65. Then, the vector output by the first fully-connected layer is passed to the second fully-connected layer, which outputs a single node for binary classification. The calculation expression is as follows: ; Among them, is the weight matrix of the second fully connected layer; is the bias term; is the output of the second layer; S66. Network output For binary classification, the Sigmoid activation function is used on the output, and the calculation expression of the Sigmoid function is as follows: ; Among them, is the final output of the model, with a value range between [0, 1]. The output value represents the confidence that the sample is the target.
7. A method for infrared target detection combining local contrast with long short-term memory according to claim 6, characterized in that In step 7, the constructed LCWMD model is combined with a long-term network for long-term infrared target tracking detection, and it is judged whether the target is occluded. If it is occluded, continue with long-term tracking detection; otherwise, the process of the next step is as follows: First, according to the relationship between consecutive frames, in the local area of the image of the next frame, use the trained long-term convolutional neural network for scanning to obtain multiple samples of suspected targets, and select the sample with the highest confidence as the detection result of the neural network. Then, calculate the LCWMD values of all samples with high confidence. Take the maximum LCWMD value when the target sample is not occluded as the benchmark, and calculate the benchmark when detecting the frame . The expression is as follows: ; Among them, represents the LCWMD value of the th frame image; when the LCWMD value of the target sample drops below 80% of the benchmark, it is considered that the target starts to enter the complex background area and is occluded. At this time, the local area is detected and processed according to the detection results.
8. The infrared target detection method combining local contrast with long short-term memory according to claim 7, characterized in that, In step 8, when detecting the local boundary of the occluded object, the expression for judging whether the target reappears is as follows: If the LCWMD values of the targets detected in three consecutive frames are between 30% and 80% of the benchmark, it is considered that the position of the target in the complex background area is detected, and the expression is as follows: ; Among them, represents the LCWMD value of the th frame image, represents the LCWMD value of the th frame image, represents the LCWMD value of the th frame image; if the LCWMD value of the target is below 30% of the benchmark, the system determines that the target tracking fails and triggers a temporary loss flag; the expression is as follows: 。 9. The infrared target detection method combining local contrast with long short-term memory according to claim 8, characterized in that, The process of extracting target features to train a short-term convolutional neural network in step 9 is as follows: S91. Calculate the size of the convolutional layer; the input channels of the convolutional layer are 1, the output channels are 8, the convolutional kernel size is 3×3, the stride is 1, and the padding is 1. The calculation expression is as follows: ; Among them, represents the position on the output channel ; eigenvalue; represents the corresponding pixel value in the input image; represents the weight parameter of the convolution kernel; represents the bias of the S92. Downsample the feature map after convolution using a max pooling layer with a pooling kernel size of 2×2, a stride of 2, and the pooling calculation expression is as follows: ; Among them, represents the value at the position in the th channel of the pooled feature map; S93. Use the feature map obtained after convolution and pooling as the input to the first fully connected layer; the calculation expression of the fully connected layer is as follows: ; Among them, is the input vector; is the weight matrix of the first fully connected layer; is the bias term; is the output of the first layer; S94. Then, the vector output by the first fully connected layer is passed to the second fully connected layer, and this layer outputs a node for binary classification. The calculation expression is as follows: ; Among them, is the weight matrix of the second fully connected layer; is the bias term; is the output of the second layer. S95. The network uses the Sigmoid activation function for binary classification at the output The calculation expression of the Sigmoid function is as follows: ; Among them, is the final output of the model, and its value range is between [0, 1].
10. The infrared target detection method combining local contrast with long short-term memory according to claim 9, characterized in that: The criterion for judging whether the target leaves the occluded area in step 10 is as follows: If the LCWMD value of the target sample returns above 80% of the benchmark for three consecutive frames, it is considered that the target gradually leaves the complex background area, and the expression is as follows: 。
Citation Information
Patent Citations
Target tracking anti-shielding method
CN117115206A
Target tracking method of sonar image, target tracking system and computing device
CN118644526A
Visual optimization method for camera in rainy days
CN118898562A
Infrared weak and small target multi-frame detection method and system based on space-time integrated network
CN119169263A
Noise-resilient vasculature localization method with regularized segmentation
WO2022005336A1
Cited By
Target tracking and identifying system and method based on Raycore microprocessor
CN120298459A