Threshold-adaptive three-frame difference detection algorithm
By designing an adaptive threshold mechanism in the three-frame differential detection algorithm, the problem of insufficient detection accuracy and robustness caused by the uncertainty of threshold setting in the prior art is solved, and more efficient and sensitive motion object detection is achieved.
Patent Information
- Application Number
- CN202510142947.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-06
AI Technical Summary
The existing three-frame differential detection algorithm has uncertainty in the threshold setting, which is difficult to adapt to environmental changes and complex backgrounds, resulting in insufficient accuracy and robustness of detection.
An inter-frame differential detection algorithm with adaptive threshold is designed. By aligning and differential operations on the three-frame images, the threshold is dynamically adjusted to adapt to environmental changes and ensure the accuracy and robustness of the detection.
This algorithm can find the best detection threshold under the maximum number limit, improves the detection possibility of real targets, enhances the real-time and robustness of detection, and is suitable for infrared detection and other complex object detection fields.
Smart Images

Figure CN120107310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of moving target detection, and in particular to an adaptive threshold algorithm for three-frame difference detection. Background Art
[0002] The three-frame difference method is a motion target detection technology that identifies moving objects in a scene by comparing three consecutive frames of images. Compared with the two-frame difference method, the three-frame difference method has a higher computational complexity and is better at processing targets with higher moving speeds, which can reduce the "double image" phenomenon. Its application scenarios are relatively wide, including but not limited to: (1) Infrared detection: used to detect moving targets at long distances and in complex backgrounds; (2) High-speed moving object detection: For example, in traffic monitoring, vehicles usually travel at a higher speed, and the three-frame difference method can more accurately capture these fast-moving vehicles; (3) Target tracking in dynamic environments: such as athlete tracking in sports events. Since athletes' movements are fast and complex, the use of the three-frame difference method can help maintain continuous tracking of them.
[0003] Threshold setting in the three-frame difference algorithm is a very critical step. The choice of threshold directly affects the accuracy and robustness of moving target detection. In practical applications, the choice of threshold needs to weigh multiple factors, including noise suppression, adaptability to illumination changes, and sensitivity to moving targets. The contradictions in threshold selection include but are not limited to: (1) Noise suppression and target information retention: If the threshold is set too low, the algorithm may misjudge the noise in the image as a moving target, resulting in a large number of false detections. Conversely, if the threshold is set too high, some real moving areas may be filtered out, especially when the grayscale of the target does not change much, which will lead to missed detections. (2) Adaptability to illumination changes: Changes in illumination conditions in the scene can affect changes in pixel values. A fixed threshold may not be able to adapt to such changes, resulting in erroneous detection results. For example, when the illumination suddenly becomes brighter or darker, using a fixed threshold may cause a large amount of background to be mistaken for a moving target. (3) Environmental complexity: In complex environments, such as those with multiple textures and colors, it is more difficult to determine a suitable threshold. Different background features may produce different levels of noise, which affects the choice of threshold.
[0004] Currently, there is no effective adaptive threshold design algorithm. When using the three-frame detection algorithm, users usually make qualitative estimates based on experience, and then set a fixed threshold, repeatedly trial and error until the optimal threshold is found, which affects the real-time, robustness and accuracy of the application. Once the environment changes, the threshold needs to be readjusted, resulting in blindness and inaccuracy in practical use. Summary of the invention
[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and design a three-frame difference detection algorithm with an adaptive threshold, which is dynamic and adaptive and greatly improves the possibility of detecting the real target.
[0006] The technical solution adopted by the present invention is:
[0007] An adaptive threshold three-frame difference detection algorithm, the specific steps are:
[0008] S1: Align the three frames of images to be processed, and use image alignment to eliminate position deviation caused by jitter;
[0009] S2: performing a differential operation on the three aligned image frames to obtain two differential images;
[0010] Specifically: Assume that the nth frame, n-1th frame, and n-2th frame images in the video sequence are f n 、f n-1 、f n-2 , the grayscale value of the corresponding pixel in the three frames is recorded as f n (x, y), f n-1 (x, y), f n-2 (x, y), calculate the difference of three frames according to the following formula, and get two frames of difference images, which are expressed as: n (x, y) and d n-1 (x, y);
[0011] d n (x, y) = |f n (x, y)-f n-1 (x, y)|
[0012] d n-1 (x, y) = |f n-1 (x, y)-f n-2 (x, y)|;
[0013] S3: Select the maximum value of the two differential images, record them as T1 and T2 respectively, take the minimum value of T1 and T2 as the threshold T, and perform d n (x, y) and d n-1 (x, y) is compared with the threshold T pixel by pixel. The process is expressed as:
[0014]
[0015] The binary difference image b is calculated n (x, y) and b n-1 (x, y);
[0016] Then for b n (x, y) and b n-1(x, y) performs a logical AND operation, which can be expressed as:
[0017] m n (x, y) = b n (x,y)&b n-1 (x, y)
[0018] Get the motion representation frame image m n (x,y),m n The position where the (x, y) pixel is 1 is preliminarily determined to be a moving target, otherwise it is considered to be the background;
[0019] S4: Perform connectivity analysis based on the target location determined in S3. If the regions are connected, they are considered to be the same target. Otherwise, they are considered to be multiple targets. Finally, the number of targets N is determined.
[0020] S5: Determine whether the number of detected targets meets the maximum number limit detected by the optical device, that is, the maximum tolerance value M. If N is less than or equal to M, set T = T-1, and repeat the operations of steps S3 and S4 until N is greater than M. At this time, set T min =T+1,T min That is, it is the optimal detection threshold under the constraint of the maximum detection target, and it is also the most sensitive threshold;
[0021] S6: Using the updated optimal detection threshold T min , execute S3 and S4 to obtain the final detected target position, and mark the detected target on the image with a red rectangular frame for subsequent actual visualization applications;
[0022] S7: In practice, since frame images are usually continuous, the operations from S1 to S6 are repeated for all images to be processed in chronological order until the task is completed.
[0023] An adaptive threshold three-frame difference detection algorithm, the specific steps are:
[0024] S1: Align the three frames of images to be processed, and use image alignment to eliminate position deviation caused by jitter;
[0025] S2: performing a difference operation on the aligned double-frame images to obtain a difference image;
[0026] Specifically: Assume that the nth frame, n-1th frame, and n-2th frame images in the video sequence are f n 、f n-1 、f n-2 , the grayscale value of the corresponding pixel in the three frames is recorded as f n (x, y), f n-1 (x, y), f n-2(x, y), calculate the difference of three frames according to the following formula to obtain two frames of differential images, which are expressed as: d n (x, y) and d n-1 (x, y);
[0027] d n (x, y) = |f n (x, y)-f n-1 (x, y)|
[0028] d n-1 (x, y) = |f n-1 (x, y)-f n-2 (x, y)|;
[0029] S3: Select the maximum value of the two differential images, record them as threshold T1 and threshold T2 respectively, and set d n (x, y) and T1, d n-1 (x, y) is compared with T2 pixel by pixel. The process is expressed as:
[0030]
[0031] Get the binary difference image b1 n (x, y) and b1 n-1 (x, y);
[0032] Then for b1 n (x, y) and b1 n-1 (x, y) performs a logical AND operation, which can be expressed as:
[0033] m1 n (x, y) = b1 n (x,y)&b1 n-1 (x, y)
[0034] Get the motion representation frame image m1 n (x,y),m1 n The position where the (x,y) pixel is 1 is preliminarily determined to be a moving target, otherwise it is considered to be the background;
[0035] S4: Perform connectivity analysis based on the target location determined in S3. If the regions are connected, they are considered to be the same target. Otherwise, they are considered to be multiple targets. Finally, the number of targets N is determined.
[0036] S5: Use thresholds T1 and T2 to perform target detection operations to determine whether the number of detected targets meets the maximum number limit detected by the optical device, that is, the maximum tolerance value M; if N is less than or equal to M, calculate b1 respectively n (x,y) and b2 n-1The number of target areas with 1 pixels in (x, y) is N1 and N2. If N1 is less than N2, set T1 = T1-1; if N2 is less than N1, set T2 = T2-1; if N1 is equal to N2, set T1 = T1-1, T2 = T2-1; repeat the operations of steps S3 and S4; until N is greater than M, set T1 min =T1+1, T2 min =T2+1;
[0037] S6: Using the updated optimal detection threshold T1 min and T2 min , perform S3 and S4 operations to obtain the final number of detected targets and their corresponding positions, and mark the detected targets on the image with red rectangular frames for subsequent actual visualization applications;
[0038] S7: In practice, since frame images are usually continuous, the operations from S1 to S6 are repeated for all images to be processed in chronological order until the task is completed.
[0039] Due to the adoption of the above-mentioned technical solution, the present invention has the following advantages:
[0040] The three-frame differential detection algorithm with an adaptive threshold of the present invention solves the shortcomings of real-time, robustness and accuracy that the three-frame differential detection fixed threshold algorithm does not have, and solves the blindness and inaccuracy caused by environmental changes; the algorithm obtains the optimal detection threshold under the maximum number restriction condition, that is, the most sensitive detection threshold, which is not only dynamic and adaptive, but also greatly improves the possibility of detecting the real target; the present invention is not only suitable for target detection fields such as infrared detection, but can also be migrated to other complex target detection fields that require high-precision, high-efficiency processing and high motion speed, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic diagram of the process of Example 1 of the present invention.
[0042] Figure 2 It is a schematic diagram of the flow chart of Example 2 of the present invention.
[0043] Figure 3 It is a scene situation table of a dataset for detecting and tracking small aircraft targets in infrared images under ground / air background in an embodiment of the present invention.
[0044] Figure 4 yes Figure 3 The first three consecutive frames of images of the drone target in the data21 dataset.
[0045] Figure 5 yes Figure 4 The image after alignment.
[0046] Figure 6 This is the final test result diagram of Example 1 of the present invention.
[0047] Figure 7 This is a diagram of the optimal detection threshold results for each frame of image in Example 1 of the present invention.
[0048] Figure 8 This is the final test result diagram of Example 2 of the present invention.
[0049] Fig. 9 This is a diagram of the optimal detection threshold results for each frame of image in Example 2 of the present invention. DETAILED DESCRIPTION
[0050] The present invention will be further explained below in conjunction with the accompanying drawings and embodiments, which should not be used to limit the protection scope of the present invention. The purpose of disclosing the present invention is to protect all technical improvements within the scope of the present invention.
[0051] The implementation example is performed on a computer equipped with AMD Ryzen 5 5600K CPU, 16GB RAM and RTX 3060GPU hardware functions, and MATLAB 2024a software; the data set selects "weak aircraft target detection and tracking data set in infrared images under ground / air background", such as Figure 3 As shown in the figure, the data set is provided by the ATR Key Laboratory of the School of Electronic Science of the National University of Defense Technology. It provides 21 types of frame sequence infrared image data and has relatively complete information tags. The data acquisition images are taken by an infrared device with a frequency of 100Hz. As the frequency of infrared shooting is higher, the temporal and spatial changes of the target between frame sequences are less obvious and difficult to detect. Therefore, the frequency is usually reduced with a step size of 4 frames. In order to simulate the actual long-distance detection under complex backgrounds, the data21 data set is selected. This data set is one of the two data sets with the lowest signal-to-clutter ratio. The signal-to-clutter ratio describes the difficulty of detecting small infrared targets. Generally, the lower the signal-to-clutter ratio, the more difficult the target detection is. Figure 4 As shown, the first three consecutive frames of images, the target is very faint and can hardly be distinguished from the background.
[0052] Example 1
[0053] An adaptive threshold three-frame difference detection algorithm, assuming that the maximum detection tolerance M is 5, the specific steps are:
[0054] Execute S1 to Figure 4 Take the image in the figure as an example, and perform alignment operation. The method based on feature point alignment is used to eliminate the position deviation caused by jitter by image alignment. The processed image is as follows Figure 5As shown, (a) is the current frame image, (b) is the previous frame image after alignment, wherein the grayscale value of the image part that is not aligned is set to 0, as shown in the figure as the right side line and the upper side line are black; Figure (c) is the next frame image after alignment, wherein the grayscale value of the image part that is not aligned is set to 0, as shown in the figure as the left side line, the lower side line, and the upper side line are black. This step will also generate a mask, which is a binary image, usually used to indicate which pixels should be retained and which should be ignored. The mask is used to perform the S2 operation.
[0055] Execute S2 to perform a difference operation on the aligned portion of the current frame and the previous frame, and the aligned portion of the current frame and the next frame, to obtain two differential images.
[0056] Execute S3, select the maximum value of the two frames of differential images respectively, and obtain the threshold value T1 as 24 and T2 as 80, and take the minimum value T of T1 and T2 as 24; compare the two frames of differential images with the threshold value T respectively and then binarize them, that is, the position greater than or equal to T is set to 1, and the position less than T is set to 0, to obtain two binary differential images; then perform a logical AND operation on the two binary differential images to obtain a motion representation frame image, and the image position with a pixel of 1 is preliminarily determined as the target position.
[0057] Execute S4: Perform connectivity analysis based on the target location determined in S3. If the areas are connected, they are considered to be the same target. Otherwise, they are considered to be multiple targets. The current number of detected targets N is 1.
[0058] Execute S5, use the threshold T to perform target detection operation, and determine whether the number of detected targets meets the maximum number limit detected by the optical device, that is, the maximum tolerance value M. First, from S4, N=1, since M=5, N is less than M, let T=T-1, repeat the operations of S3 and S4 steps until N is greater than 5, then stop, and then let T min =T+1, we get T min =12, T min It is the optimal detection threshold under the constraint of the maximum detection target and also the most sensitive threshold.
[0059] Execute S6: Use the updated optimal detection threshold T min , execute S3 and S4 to obtain the final detected target position, and mark the detected target on the image with a red rectangular frame for subsequent actual visualization applications; Figure 6 As shown, a total of 4 targets are detected, meeting the requirement that the maximum number of detections is less than or equal to 5, including 1 real target and 3 false targets.
[0060] Execute S7: Repeat the operations from S1 to S6 for all images to be processed in chronological order until the task is completed. Figure 7As shown, 123 frames of images are processed, and it can be seen that under the condition of meeting the maximum tolerance M limit, the target detection threshold of each frame of image is different, and changes adaptively according to changes in information such as background and target position.
[0061] Example 2
[0062] An adaptive threshold three-frame difference detection algorithm, assuming that the maximum detection tolerance M is 5, the specific steps are:
[0063] Execute S1 to Figure 3 Take the image in the figure as an example, and perform alignment operation. The method based on feature point alignment is used to eliminate the position deviation caused by jitter by image alignment. The processed image is as follows Figure 5 As shown, (a) is the current frame image, Figure 5 (b) is the previous frame image after alignment, where the grayscale value of the image part that is not aligned is set to 0, which is shown as the right side line and the upper side line are black in the figure; Figure (c) is the next frame image after alignment, where the grayscale value of the image part that is not aligned is set to 0, which is shown as the left side line, the lower side line, and the upper side line are black in the figure. This step also generates a mask, which is a binary image that is usually used to indicate which pixels should be retained and which should be ignored. The mask is used to perform the S2 operation.
[0064] Execute S2 to perform a difference operation on the aligned portion of the current frame and the previous frame, and the aligned portion of the current frame and the next frame, to obtain two differential images.
[0065] Execute S3, select the maximum value of the two frames of differential images respectively, and obtain the threshold value T1 of 20 and T2 of 84. Compare the two frames of differential images with their respective maximum threshold values T1 and T2 and then binarize them to obtain two frames of binary differential images. Then perform a logical AND operation on the two frames of binary differential images to obtain a motion representation frame image, and the image position with a pixel of 1 is preliminarily determined as the target position.
[0066] Execute S4: Perform connectivity analysis based on the target location determined in S3. If the areas are connected, they are considered to be the same target. Otherwise, they are considered to be multiple targets. The current number of detected targets N is 0.
[0067] Execute S5, use thresholds T1 and T2 to perform target detection operations, and determine whether the number of detected targets meets the maximum number limit detected by the optical device, that is, the maximum tolerance value M. From S4, N = 0, since M = 5, N is less than M, and b1 is calculated respectively. n (x,y) and b2 n-1The number of regions with 1 pixels in (x, y) is N1 and N2. If N1 is less than N2, let T1 = T1-1; if N2 is less than N1, let T2 = T2-1; if N1 is equal to N2, let T1 = T1-1, T2 = T2-1; then repeat the operations of steps S3 and S4 until N is greater than 5, then let T1 min =T1+1, T2 min =T2+1, finally get T1 min =8, T2 min =23, which is the most sensitive detection threshold of the two-frame differential image.
[0068] Execute S6 and use the updated optimal detection threshold T1 min and T2 min , perform S3 and S4 operations to obtain the final detected target position, and mark the detected target on the image with a red rectangular frame for subsequent actual visualization applications; Figure 8 As shown, a total of 5 targets are detected, satisfying the requirement that the maximum number of detections is less than or equal to 5, including 1 real target and 4 false targets.
[0069] Execute S7: Repeat the operations from S1 to S6 for all images to be processed in chronological order until the task is completed. Fig. 9 As shown in the figure, 123 frames of images are processed, and it can be seen that under the conditions of meeting the maximum number and maximum tolerance M, the corresponding optimal detection threshold T1 in each frame detection min and T2 min They are all different and change adaptively according to changes in information such as background and target position.
Claims
1. An adaptive threshold three-frame difference detection algorithm, characterized in that: The specific steps are: S1: Align the three frames of images to be processed, and use image alignment to eliminate position deviation caused by jitter; S2: Perform a differential operation on the three aligned image frames to obtain two differential images; Specifically: Assume that the nth frame, n-1th frame, and n-2th frame images in the video sequence are f n 、f n-1 、f n-2 , the grayscale value of the corresponding pixel in the three frames is recorded as f n (x, y), f n-1 (x, y), f n-2 (x, y), calculate the difference of three frames according to the following formula, and get two frames of difference images, which are expressed as: n (x, y) and d n-1 (x, y); d n (x,y)=|f n (x,y)-f n-1 (x,y)| d n-1 (x,y)=|f n-1 (x,y)-f n-2 (x,y)|; S3: Select the maximum value of the two differential images, record them as T1 and T2 respectively, take the minimum value of T1 and T2 as the threshold T, and perform d n (x, y) and d n-1 (x, y) is compared with the threshold T pixel by pixel. The process is expressed as: The binary difference image b is calculated n (x, y) and b n-1 (x, y); Then for b n (x, y) and b n-1 (x, y) performs a logical AND operation, which can be expressed as: m n (x,y)=b n (x,y)&b n-1 (x,y) Get the motion representation frame image m n (x,y),m n The position where the (x, y) pixel is 1 is preliminarily determined to be a moving target, otherwise it is considered to be the background; S4: Perform connectivity analysis based on the target location determined in S3. If the regions are connected, they are considered to be the same target. Otherwise, they are considered to be multiple targets. Finally, the number of targets N is determined. S5: Determine whether the number of detected targets meets the maximum number limit detected by the optical device, that is, the maximum tolerance value M. If N is less than or equal to M, set T = T-1, and repeat the operations of steps S3 and S4 until N is greater than M. At this time, set T min =T+1,T min That is, it is the optimal detection threshold under the constraint of the maximum detection target, and it is also the most sensitive threshold; S6: Using the updated optimal detection threshold T min , execute S3 and S4 to obtain the final detected target position, and mark the detected target on the image with a red rectangular frame for subsequent actual visualization applications; S7: In practice, since frame images are usually continuous, the operations from S1 to S6 are repeated for all images to be processed in chronological order until the task is completed.
2. An adaptive threshold three-frame difference detection algorithm, characterized in that: The specific steps are: S1: Align the three frames of images to be processed, and use image alignment to eliminate position deviations caused by jitter; S2: performing a difference operation on the aligned double-frame images to obtain a difference image; Specifically: Assume that the nth frame, n-1th frame, and n-2th frame images in the video sequence are f n 、f n-1 、f n-2 , the grayscale value of the corresponding pixel in the three frames is recorded as f n (x, y), f n-1 (x, y), f n-2 (x, y), calculate the difference of three frames according to the following formula to obtain two frames of differential images, which are expressed as: d n (x, y) and d n-1 (x, y); d n (x,y)=|f n (x,y)-f n-1 (x,y)| d n-1 (x,y)=|f n-1 (x,y)-f n-2 (x,y)|; S3: Select the maximum value of the two differential images, record them as threshold T1 and threshold T2 respectively, and set d n (x, y) and T1, d n-1 (x, y) is compared with T2 pixel by pixel. The process is expressed as: Get the binary difference image b1 n (x, y) and b1 n-1 (x, y); Then for b1 n (x, y) and b1 n-1 (x, y) performs a logical AND operation, which can be expressed as: m1 n (x, y)=b1 n (x, y)&b1 n-1 (x, y) Get the motion representation frame image m1 n (x,y),m1 n The position where the (x, y) pixel is 1 is preliminarily determined to be a moving target, otherwise it is considered to be the background; S4: Perform connectivity analysis based on the target location determined in S3. If the regions are connected, they are considered to be the same target. Otherwise, they are considered to be multiple targets. Finally, the number of targets N is determined. S5: Perform target detection operation using thresholds T1 and T2 to determine whether the number of detected targets meets the maximum number limit detected by the optical device, i.e., the maximum tolerance value M; If N is less than or equal to M, calculate b1 respectively n (x, y) and b2 n-1 The number of target areas with 1 pixels in (x, y) is N1 and N2. If N1 is less than N2, set T1 = T1-1; if N2 is less than N1, set T2 = T2-1; if N1 is equal to N2, set T1 = T1-1, T2 = T2-1; repeat the operations of steps S3 and S4; until N is greater than M, set T1 min =T1+1, T2 min =T2+1; S6: Using the updated optimal detection threshold T1 min and T2 min , perform S3 and S4 operations to obtain the final number of detected targets and their corresponding positions, and mark the detected targets on the image with red rectangular frames for subsequent actual visualization applications; S7: In practice, since frame images are usually continuous, the operations from S1 to S6 are repeated for all images to be processed in chronological order until the task is completed.