Time-space domain joint point target extraction method
Through the combined point target extraction method of time-domain combined point targets, the time-domain difference and air-domain grayscale significance measurements are used to adaptively adjust the threshold, and the extraction problem of weak targets under complex backgrounds and low signal-to-noise ratios is solved, accurately tracking and correlation of targets is achieved, and the extraction accuracy and robustness of targets in the image sequence are improved.
Patent Information
- Application Number
- CN202510424839.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-29
AI Technical Summary
In image sequences with complex backgrounds and low signal-to-noise ratios, it is difficult for the prior art to accurately extract weak targets, and traditional methods show limitations when facing complex backgrounds and dynamic changes, especially in infrared imaging and medical imaging, the target signal-to-noise ratio is low and the background clutter is complex, resulting in difficulty in extracting targets.
The combined point target extraction method of time and space domain is adopted to remove background clutter through time domain differential measurement, weak targets are enhanced by the air domain grayscale significance measurement, and thresholds are adaptively adjusted to finally form a target trajectory to adapt to different application scenarios.
Effectively suppress background clutter, enhance weak target characteristics, improve the accuracy and robustness of target extraction, ensure the continuity and stability of trajectory, and is suitable for image sequence processing of weak infrared targets.
Smart Images

Figure CN120388043A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method for extracting point targets by jointly considering the spatial and temporal domains. Background Art
[0002] In the field of modern image processing, accurately extracting small and weak targets from image sequences with complex backgrounds and low signal-to-noise ratios is a highly challenging task, which is widely applied in multiple fields such as military, civilian, and medical fields.
[0003] Background clutter is one of the main obstacles to target extraction. The background in natural scenes may contain various complex textures, colors, and shapes, and these features are intertwined with the target features, making it difficult to distinguish the contours and features of the target. For example, in infrared imaging, the thermal radiation of the background may cover the thermal signal of the target, resulting in a decrease in the signal-to-noise ratio of the target. In addition, dynamic changes in the background, such as the shaking of leaves caused by wind and the ripples on the water surface, also increase the difficulty of target extraction.
[0004] Small and weak targets usually have small sizes and low contrasts, which makes them difficult to detect in images. For example, the image of a small satellite in the cosmic background may be just a faint light spot, and its brightness may be comparable to the background noise of the surrounding starry sky. In medical imaging, early tiny tumors may only show minor gray-scale changes in the image and are easily overlooked.
[0005] In a multi-frame image sequence, the movement and changes of the target need to be accurately tracked and associated to form a complete target trajectory. However, since the movement of the target may be affected by various factors, such as speed changes and acceleration changes, traditional tracking algorithms often have difficulty adapting to these changes, resulting in breaks or incorrect associations in the target trajectory.
[0006] Existing target extraction methods, such as threshold-based methods, template matching-based methods, etc., although they can achieve certain effects under certain specific conditions, often show obvious limitations when faced with image sequences with complex backgrounds and low signal-to-noise ratios. Threshold-based methods are very sensitive to the selection of thresholds. Different image sequences may require different thresholds, and the selection of thresholds often requires manual intervention and lacks self-adaptability. Template matching-based methods require prior knowledge of the approximate shape and features of the target. For unknown targets or cases where the target shape changes greatly, the matching effect will be greatly reduced. Summary of the Invention
[0007] The present invention aims to provide a spatio-temporal domain joint point target extraction method, which removes background clutter through time-domain difference measurement, enhances weak targets using spatial domain gray-scale saliency measurement, adaptively adjusts thresholds for different application scenarios, and finally forms associated trajectories in time series, having excellent complex background clutter suppression ability, and is particularly suitable for processing image sequences containing infrared small targets.
[0008] To achieve the above object, the present invention provides the following technical solutions: A spatio-temporal domain joint point target extraction method, comprising the following steps:
[0009] S1. Process the image sequence containing point targets, sequentially add the images into an image queue with a constant length, and determine parameters before processing;
[0010] S2. For the images entering the processing flow, perform preprocessing based on optical flow features in combination with the previous frame image, estimate the pixel motion amount, and screen out the pixel points with large motion amounts to form an initial target point set;
[0011] S3. For the pixel positions corresponding to each point in the initial target point set, according to the time-domain difference measurement module, calculate the fluctuation of the pixel gray-scale value in the past frames, and determine whether it is a potential target point through the time-domain adaptive threshold, and form a time-domain screened target point set;
[0012] S4. For each point in the time-domain screened target point set, according to the spatial domain gray-scale saliency measurement module, calculate its gray-scale saliency characteristics relative to the neighborhood where it is located, and determine whether it is a potential target point according to the spatial domain adaptive threshold, and form a spatial domain screened target point set;
[0013] S5. Considering the point spread and cross-pixel effects caused by the diffraction of the optical system and non-ideal imaging characteristics, cluster the point sets obtained by time-domain and spatial domain screening to obtain a candidate target set;
[0014] S6. Through a multi-frame association strategy, obtain the final target set and target trajectory.
[0015] Preferably, in the preprocessing of optical flow features in step S2, when performing optical flow estimation, set an appropriate noise threshold to avoid the interference of noise on optical flow calculation. The threshold setting needs to balance filtering out effective noise and retaining the tiny motion of the target, so it needs to be adaptively set. Introduce the image gradient as the threshold setting benchmark. For the current frame image, obtain the gradient vectors G x 、G y in the horizontal and vertical directions, and calculate the amplitude vector G of the gradients in the two directions:
[0016]
[0017] Since the gradient is affected by the data amplitude and dimension, the coefficient of variation is used to measure the degree of gradient jitter:
[0018]
[0019] where α is an adjustment factor that can be set according to different data.
[0020] Preferably, in step S3, the time-domain difference measurement module characterizes the time-domain features by the product of the standard deviation and the curvature of the vector vec. The specific formula is:
[0021] T vec =std(vec)×f curv (vec)
[0022] where std(vec) is to solve the standard deviation, and f curv is to solve the curvature. The formula for solving the curvature is:
[0023]
[0024] where x(t) and y(t) form the parametric equations of the plane curve, x'(t) and x”(t) are the first-order and second-order derivatives of x(t) respectively; y'(t) and y”(t) are the first-order and second-order derivatives of y(t) respectively.
[0025] Preferably, in step S3, for the time-domain difference measurement module, the threshold is set by an adaptive strategy. For each new image with a size of (m, n), the temporal standard deviation of each pixel position is calculated and denoted as std(t i ), i ∈ [1, m×n], the temporal length is N, and the temporal standard deviations of all pixels form a vector sequence V T . The threshold is set as:
[0026] th T =log(β·(max(V T ) - mean(V T )) + 1)
[0027] where max(·) and mean(·) are the maximum value and the average value of the sequence respectively, and β is an adjustment factor that can be set according to different data. When the numerical value of the time-domain characteristic calculated in claim 3 is higher than the threshold automatically obtained for the current frame, the pixel position is retained; otherwise, it is deleted.
[0028] Preferably, in step S4, for the spatial gray-scale saliency measurement module, a square center window is set with the candidate target point after time-domain processing as the center, and 8 square windows with the same neighborhood size are set around it. The local gray-scale levels of the 9 windows are calculated:
[0029]
[0030] where (2×r + 1) 2 -1 represents the total number of pixels excluding the central pixel in the sub-window, a and b are not both 0 at the same time, and X(p, q) is the gray value at the corresponding position in the image, and further measures the regional saliency of the target point:
[0031]
[0032] In the formula, min(·) represents the minimum value.
[0033] Principle and beneficial effects of the present technical solution:
[0034] (1) By calculating the temporal fluctuations of pixel gray values and screening with an adaptive threshold, background clutter is effectively removed, and target points are accurately identified. Calculate the gray saliency difference between the candidate target point and its neighborhood to enhance the weak target features and make them more prominent in complex backgrounds.
[0035] (2) Use optical flow estimation to screen moving pixel points, quickly locate the target area, and reduce ineffective calculations. Combine the processing results in the time domain and spatial domain, and obtain an accurate target set and trajectory through multi-frame association to ensure the continuity and stability of the trajectory.
[0036] (3) Based on the image gradient and coefficient of variation, adaptively set the noise threshold to balance noise filtering and target motion retention, and improve applicability and robustness. Set the threshold adaptively according to the standard deviation of the image time series to accurately screen potential target points and enhance adaptability and generality.
[0037] (4) Considering the point spread and cross-pixel effects caused by the optical system, cluster the screened points to restore the true shape of the target and improve the extraction accuracy. Description of the Drawings
[0038] Figure 1 It is a flowchart of a spatio-temporal domain joint point target extraction method provided by an embodiment of the present invention. Detailed Embodiment
[0039] The present invention will be further described in detail below in conjunction with the drawings and embodiments:
[0040] A spatio-temporal domain joint point target extraction method includes the following steps:
[0041] S1. For an image sequence containing point targets, start processing sequentially from the first frame. Each time an image enters the processing flow, the image is added to the image queue, and one frame of image is popped from the head of the image queue. The image queue maintains a constant length N, and the parameters are determined before processing;
[0042] S2. For each frame of image entering the processing flow, combine the previous frame of infrared image and perform preprocessing to obtain an initial set of target points;
[0043] The preprocessing method is preprocessing based on optical flow features. The motion between two adjacent frames of images is obtained according to the optical flow method. The motion amount of each pixel in the image is estimated, and those pixel points with larger motion amounts are screened out through an adaptive threshold.
[0044] Optical flow preprocessing strategy. When performing optical flow estimation, set an appropriate noise threshold to avoid the interference of noise on optical flow calculation. The threshold setting needs to balance filtering out effective noise and retaining the tiny motion of the target, so it needs to be adaptively set. Introduce the image gradient as the benchmark for threshold setting. For the current frame of image, obtain the gradient vectors G x 、G y in the horizontal and vertical directions, and calculate the magnitude vector G of the gradients in the two directions:
[0045]
[0046] Since the gradient is affected by the data magnitude and dimension, the coefficient of variation is used to measure the degree of gradient fluctuation:
[0047]
[0048] where α is an adjustment factor that can be set according to different data.
[0049] Time-domain difference measurement module. The threshold is set through an adaptive strategy. For each new image with a size of (m, n), calculate the temporal standard deviation at each pixel position, denoted as std(t i ), i ∈ [1, m×n], and the temporal length is N. The temporal standard deviations of all pixels form a vector sequence V T , and the threshold is set as:
[0050] th T =log(β·(max(V T ) - mean(V T )) + 1)
[0051] where max(·) and mean(·) are the maximum value and average value of the sequence respectively, and β is an adjustment factor that can be set according to different data. When the numerical value of the time-domain characteristic calculated in claim 3 is higher than the threshold automatically obtained for the current frame, the pixel position is retained (considered to be a possible target point), otherwise it is deleted.
[0052] S3. For the pixel positions corresponding to each point in the initial target point set, according to the time-domain difference metric module, calculate the fluctuation of the pixel gray values in the past N frames, and determine whether this point is a potential target point according to the time-domain adaptive threshold. The initial candidate target set forms a time-domain filtered target point set after being measured by the time-domain difference metric;
[0053] The time-domain difference metric module mainly characterizes its time-domain features by the product of the standard deviation of the vector vec and its curvature. The specific formula is as follows:
[0054] T vec = std(vec) × f curv (vec)
[0055] Where std(vec) refers to solving the standard deviation, and f curv is to solve the curvature, according to the following formula:
[0056]
[0057] Among them, x(t) and y(t) form the parametric equation of the plane curve, x'(t) and x''(t) are the first-order and second-order derivatives of x(t) respectively; y'(t) and y''(t) are the first-order and second-order derivatives of y(t) respectively. The curvature of the vector vec is obtained by further calculating the offline integral area according to the above formula.
[0058] S4. For each point in the time-domain filtered target point set, according to the spatial-domain gray-scale saliency metric module, calculate its gray-scale saliency characteristics relative to the surrounding neighborhood. Determine whether this point is a potential target point according to the spatial-domain adaptive threshold. The secondary filtered target set forms a spatial-domain filtered target point set after being measured by the spatial-domain gray-scale saliency metric;
[0059] The spatial-domain gray-scale saliency metric module mainly calculates the gray-scale saliency difference between the candidate target and its surrounding neighborhood. Set a square window as the central window with a certain candidate target point after time-domain processing as the center, and set 8 square windows with the same neighborhood size around it. Calculate the local gray levels of the 9 windows:
[0060]
[0061] Among them, (2×r + 1) 2 -1 represents the total number of pixels excluding the central pixel in the sub-window, not all being 0, and X(p,q) is the gray value at the corresponding position in the image. Then further measure the regional saliency of the target point:
[0062]
[0063] In the formula, min(·) represents the minimum value.
[0064] S5. Due to the diffraction and non-ideal imaging characteristics of the optical system, there will be point spread and cross-pixel effects in the photographed target. Therefore, the point target is not an isolated target point but presents the form of a point light spot. Therefore, it is necessary to cluster the point set obtained by time-domain and spatial-domain screening to obtain a candidate target set;
[0065] S6. Through a multi-frame association strategy, the final target set and target trajectory are obtained.
[0066] The above are only embodiments of the present invention, and common general technical solutions or characteristics in the solutions are not described in detail here. For those skilled in the art, without departing from the technical solution of the present invention, several deformations and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application shall be subject to the content of its claims, and the specific implementation manners described in the specification can be used to interpret the content of the claims.
Claims
1. A spatio-temporal domain joint point target extraction method, characterized in that, Including the following steps: S1. Process the image sequence containing point targets, sequentially add the images into an image queue with a constant length, and determine the parameters before processing; S2. For the images entering the processing flow, perform preprocessing based on optical flow features in combination with the previous frame image, estimate the pixel movement amount, and screen out the pixel points with large movement amounts to form an initial target point set; S3. For the pixel positions corresponding to each point in the initial target point set, according to the time-domain difference measurement module, calculate the fluctuation of the pixel gray value in the past frames, and determine whether it is a potential target point through the time-domain adaptive threshold to form a time-domain screened target point set; S4. For each point in the time-domain screened target point set, according to the spatial-domain gray-scale saliency measurement module, calculate its gray-scale saliency characteristics relative to the neighborhood where it is located, and determine whether it is a potential target point according to the spatial-domain adaptive threshold to form a spatial-domain screened target point set; S5. Considering the point spread and cross-pixel effects caused by the diffraction of the optical system and the non-ideal imaging characteristics, cluster the point sets obtained by time-domain and spatial-domain screening to obtain a candidate target set; S6. Through a multi-frame association strategy, obtain the final target set and target trajectory.
2. The spatio-temporal domain joint point target extraction method according to claim 1, wherein: Preprocessing of the optical flow features in step S2. When performing optical flow estimation, an appropriate noise threshold is set to avoid the interference of noise on the optical flow calculation. The threshold setting needs to balance filtering out effective noise and retaining the tiny movements of the target, so it needs to be adaptively set. The image gradient is introduced as the benchmark for threshold setting. For the current frame image, the gradient vectors G x and G y in the horizontal and vertical directions are obtained, and the magnitude vector G of the gradients in the two directions is calculated as follows: Since the gradient is affected by the data amplitude and dimension, the coefficient of variation is used to measure the degree of gradient jitter: where α is an adjustment factor that can be set according to different data.
3. A method for extracting point targets by spatio-temporal domain joint, according to claim 1, characterized in that: In step S3, the time-domain difference measurement module characterizes the time-domain feature by the product of the standard deviation and curvature of the vector vec, and the specific formula is: T vec = std(vec) × f curv (vec) Among them, std(vec) is to solve the standard deviation, and f curv is to solve the curvature. The formula for solving the curvature is: where x(t) and y(t) form the parametric equations of the plane curve, x'(t) and x”(t) are the first-order and second-order derivatives of x(t) respectively; y'(t) and y”(t) are the first-order and second-order derivatives of y(t) respectively.
4. A spatio-temporal domain joint point target extraction method according to claim 3, characterized in that: In step S3, for the time-domain difference metric module, the threshold is set by an adaptive strategy. For each new image with a size of (m, n) per frame, calculate the temporal standard deviation at each pixel position, denoted as std(t i ), i ∈ [1, m×n], the temporal length is N, and the temporal standard deviations of all pixels form a vector sequence V T , and the threshold is set as: th T = log(β·(max(V T ) - mean(V T )) + 1) where max(·) and mean(·) are the maximum value and average value of the sequence respectively, β is an adjustment factor that can be set according to different data. When the time-domain characteristic value calculated in claim 3 is higher than the threshold automatically obtained for the current frame, the pixel position is retained, otherwise it is deleted.
5. A method for extracting point targets by spatio-temporal domain joint, according to claim 1, characterized in that: In step S4, the spatial-domain gray-scale saliency measurement module sets a square center window centered on the candidate target points after time-domain processing, and sets 8 square windows with the same neighborhood size around it to calculate the local gray-scale levels of the 9 windows: Among them, (2×r + 1) 2 -1 represents the total number of pixels excluding the central pixel in the sub-window, a and b are not both 0 at the same time, and X(p, q) is the gray value at the corresponding position in the image, and further measures the regional saliency of the target point: In the formula, min(·) represents the minimum value.